Intelligent Systems.
Theory and Applications

(Intellektual'nye Sistemy. Teoriya i Prilozheniya)

Quantization of language models via differentiable neural architecture search

Abstract

Neural architecture search is a method for automatic selection of an optimal neural network architecture for a given task based on a provided dataset. One of the approaches to neural architecture search is Differentiable architecture search (DARTS). DARTS transforms the discrete search space into a continuous one, enabling the use of gradient-based optimization to learn the architectural parameters of the network. In this paper, we investigate the applicability of DARTS for selecting quantization schemes and bit-widths for different components of generative language models, and present the results of the conducted experiments. The code for training and evaluation of the quantized model is available at: https://github.com/daria1d/Darts-QAT.

Keywords: natural language processing, neural architecture search, quantization, DARTS.

BibTeX
@article{IS-Davydova2026,
  author  = {Davydova, Daria Nikolaevna},
  title   = {{Quantization of language models via differentiable neural architecture search}},
  journal = {Intelligent Systems. Theory and Applications},
  year    = {2026},
  volume  = {30},
  number  = {3},
  pages   = {88--112},
}
AMSBIB
\Bibitem{IS-Davydova2026}
\by D.\,N.~Davydova
\paper Quantization of language models via differentiable neural architecture search
\jour Intelligent Systems. Theory and Applications
\yr 2026
\vol 30
\issue 3
\pages 88--112
\lang In Russian
Published under Creative Commons Attribution 4.0 International (CC BY 4.0)

← Back to issue