Energy-Efficient NLP Through Tiny-Model Distillation, Pruning, and Quantized Inference
DOI:
https://doi.org/10.71146/kjmr965Keywords:
Responsible AI (RAI), efficient AI, distillation, model compression, BERT, TinyBERT, Artificial Neural Network (ANN), neural network prunning, energy-efficient AIAbstract
Transformer models achieve state-of-the-art NLP performance but face deployment constraints due to high inference energy costs. Numerous techniques have been proposed to optimize model operations. Due to limitations of hardware and time, we analyzed a few of them, which are knowledge distillation, unstructured pruning, and quantization within a unified pipeline for BERT and TinyBERT. Results demonstrate that the Distilled Student serves as the dominant baseline, offering the lowest energy per sample among all variants. Conversely, while post-training quantization and pruning reduce model size, they introduce significant kernel overheads on GPUs. Notably, combining INT8 quantization with aggressive pruning results in severe latency penalties, causing up to a 4× increase in energy consumption compared to the standard Student. These findings underscore the critical disparity between theoretical compression metrics and hardware-realized energy efficiency.
Downloads
References
J. Devlin et al., “BERT: Pre-training of deep bidirectional transformers for language understanding,” Proc. NAACL, 2019.
Strubell et al., “Energy and policy considerations for deep learning in NLP,” Proc. ACL, 2019.
R. Schwartz et al., “Green AI,” CACM, vol. 63, no. 12, 2020.
Patterson et al., “Carbon emissions and large neural network training,” arXiv:2104.10350, 2021.
S. Williams et al., “Roofline: an insightful visual performance model for multicore architectures,” CACM, vol. 52, 2009.
Yoon, J. Mun, and K.-S. Min, “Comparative study on energy con-sumption of neural networks by scaling of weight-memory energy ver-sus computing energy for implementing low-power edge intelligence,” Electronics, vol. 14, no. 13, p. 2718, 2025.
E. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
X. Jiao et al., “TinyBERT: Distilling BERT for natural language under-standing,” Findings of EMNLP, 2020.
V. Sanh et al., “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” arXiv preprint arXiv:1910.01108, 2019.
S. Han et al., “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” ICLR, 2016.
J. Frankle and M. Carbin, “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks,” ICLR, 2019.
Jacob et al., “Quantization and training of neural networks for efficient integer-arithmetic-only inference,” CVPR, 2018.
T. Dettmers et al., “LLM.int8(): 8-bit matrix multiplication for trans-formers at scale,” NeurIPS, 2022.
J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” ICLR, 2022.
R. Socher et al., “Recursive deep models for semantic compositionality over a sentiment treebank,” EMNLP, 2013.
K. T. Chitty-Venkata et al., “A Survey of Techniques for Optimizing Transformer Inference,” J. Syst. Archit., 2023.
Adepu et al., “HAP-E: Hessian-Aware Structured Pruning of LLMs for Efficient Inference,” arXiv preprint arXiv:2405.14737, 2024.
Lagunas et al., “Block pruning for faster transformers,” EMNLP, 2021.
Xiao et al., “SmoothQuant: Accurate and efficient post-training quantization for large language models,” ICML, 2023.
Frantar et al., “GPTQ: Accurate post-training quantization for gener-ative pre-trained transformers,” ICLR, 2023.
Downloads
Published
License
Copyright (c) 2026 Hamadullah Bijarani (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
