Quantization of language models via differentiable neural architecture search
Received: 26 Apr 2026 Accepted: 29 Apr 2026
Published: 2026, vol. 30, issue 3, pp. 88–112
Abstract
Neural architecture search is a method for automatic selection of an optimal neural network architecture for a given task based on a provided dataset. One of the approaches to neural architecture search is Differentiable architecture search (DARTS). DARTS transforms the discrete search space into a continuous one, enabling the use of gradient-based optimization to learn the architectural parameters of the network. In this paper, we investigate the applicability of DARTS for selecting quantization schemes and bit-widths for different components of generative language models, and present the results of the conducted experiments. The code for training and evaluation of the quantized model is available at: https://github.com/daria1d/Darts-QAT.
Keywords: natural language processing, neural architecture search, quantization, DARTS.
BibTeX
@article{IS-Davydova2026,
author = {Davydova, Daria Nikolaevna},
title = {{Quantization of language models via differentiable neural architecture search}},
journal = {Intelligent Systems. Theory and Applications},
year = {2026},
volume = {30},
number = {3},
pages = {88--112},
}
AMSBIB
\Bibitem{IS-Davydova2026}
\by D.\,N.~Davydova
\paper Quantization of language models via differentiable neural architecture search
\jour Intelligent Systems. Theory and Applications
\yr 2026
\vol 30
\issue 3
\pages 88--112
\lang In Russian
RU
