Return to Article Details Energy-Efficient NLP Through Tiny-Model Distillation, Pruning, and Quantized Inference Download Download PDF