Fine-Tuning Lightweight Transformer Models for Text Classification: A Comparative Analysis of DistilBERT, Bert, and LSTM

Authors

  • Hadiya Ali Department of Computer Science, Abbottabad University of Science and Technology, Abbottabad, Pakistan Author
  • Maleeha Riaz Department of Computer Science, University of Lahore, Sargodha Campus, Pakistan Author
  • Imran Ullah Department of Computer Science, Hazara University, Dhodial, Mansehra, Pakistan Author
  • Muhammad Saad Hussain Department of Computer Science, University of Lahore, Sargodha Campus, Pakistan Author
  • Muhammad Abu Bakar Sultan Department of Computer Science, University of Lahore, Sargodha Campus, Pakistan Author

DOI:

https://doi.org/10.63544/ggeppj73

Keywords:

Text Classification, Sentiment Analysis, Transformer, BERT, DistilBERT, LSTM, Fine-Tuning, Natural Language Processing

Abstract

Text classification is a fundamental natural language processing task, and transformer-based language models have become the dominant approach. However, large models such as BERT impose substantial computational costs that limit their deployment in resource-constrained settings, motivating interest in lightweight alternatives. This paper presents a comparative analysis of a lightweight transformer (DistilBERT), a full-scale transformer (BERT), a recurrent neural network (bidirectional LSTM), and a classical machine learning baseline (TF-IDF with logistic regression) for binary sentiment classification on the IMDB movie review dataset. All models are evaluated under an identical protocol using a subset of 8000 training and 2000 test reviews, and are compared across accuracy, precision, recall, F1-score, area under the curve, training time, and model size. BERT achieves the highest accuracy of 87.35 percent, followed by the TF-IDF baseline at 86.30 percent and DistilBERT at 85.75 percent, while the bidirectional LSTM reaches 74.50 percent. DistilBERT maintains 98.2 percent of the accuracy of BERT while using 40 percent fewer parameters and training twice as fast. A notable finding is that the classical TF-IDF baseline remains highly competitive under limited data and limited fine-tuning, exceeding DistilBERT in accuracy at a fraction of the computational cost. The results highlight the accuracy-efficiency trade-offs among modern and classical text classification methods and offer practical guidance for model selection.

REFERENCES

[1] C. C. Aggarwal and C. Zhai, “A survey of text classification algorithms,” in Mining Text Data, Springer, 2012, pp. 163–222.

[2] K. Kowsari, K. J. Meimandi, M. Heidarysafa, S. Mendu, L. Barnes, and D. Brown, “Text classification algorithms: A survey,” Information, vol. 10, no. 4, p. 150, 2019.

[3] B. Pang and L. Lee, “Opinion mining and sentiment analysis,” Foundations and Trends in Information Retrieval, vol. 2, no. 1–2, pp. 1–135, 2008.

[4] B. Liu, Sentiment Analysis: Mining Opinions, Sentiments, and Emotions. Cambridge University Press, 2015.

[5] G. Salton and C. Buckley, “Term-weighting approaches in automatic text retrieval,” Information Processing and Management, vol. 24, no. 5, pp. 513–523, 1988.

[6] T. Joachims, “Text categorization with support vector machines: Learning with many relevant features,” in Proc. ECML, 1998, pp. 137–142.

[7] S. Wang and C. D. Manning, “Baselines and bigrams: Simple, good sentiment and topic classification,” in Proc. ACL, 2012, pp. 90–94.

[8] M. Chaudhari, “Spray-Deposited Fluorine-Free Polymer Coatings via Fatty-Acid/Oxide Interphases for Extreme Water Repellency,” in Proceedings of the International Conference on Sustainable Micro‑Nano Materials & Innovative Technology (ICSUMMIT 2026), Vadodara, India, Feb. 13–14, 2026, Atlantis Highlights in Engineering, vol. 44, Atlantis Press, Jul. 2026. doi: 10.2991/978-94-6239-727-9_26.

[9] J. Pennington, R. Socher, and C. D. Manning, “GloVe: Global vectors for word representation,” in Proc. EMNLP, 2014, pp. 1532–1543.

[10] F. Ahmed, “Cloud Security Posture Management (CSPM): Automating Security Policy Enforcement in Cloud Environments,” ESP International Journal of Advancements in Computational Technology, vol. 1, no. 3, pp. 157–166, 2023. https://philarchive.org/rec/AHMCSP

[11] A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional LSTM and other neural network architectures,” Neural Networks, vol. 18, no. 5–6, pp. 602–610, 2005.

[12] S. T. Hasan, “Graph neural network-based optimization for home health aide-patient matching,” Spanish Journal of Innovation and Integrity, vol. 54, pp. 268–281, 2026.

[13] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. NAACL-HLT, 2019, pp. 4171–4186.

[14] C. Sun, X. Qiu, Y. Xu, and X. Huang, “How to fine-tune BERT for text classification?,” in Proc. China National Conf. on Chinese Computational Linguistics, 2019, pp. 194–206.

[15] S. T. Hasan, “AI-driven equity analytics in New York City home care workforce management: Identifying bias in scheduling, case assignment, workload distribution, and career advancement,” Journal of Business Insight and Innovation, vol. 4, no. 2, pp. 119–131, 2025.

[16] M. Chaudhari, “Vapor-Phase Thin-Film Polymerization for Microbial-Resistant Surface Functionalization,” in Proceedings of the International Conference on Sustainable Micro-Nano Materials & Innovative Technology (ICSUMMIT 2026), Vadodara, India, Feb. 2026, doi: 10.2991/978-94-6239-727-9_27.

[17] Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut, “ALBERT: A lite BERT for self-supervised learning of language representations,” in Proc. ICLR, 2020.

[18] X. Jiao et al., “TinyBERT: Distilling BERT for natural language understanding,” in Findings of EMNLP, 2020, pp. 4163–4174.

[19] M. A. Jawed, "Managed aquifer recharge with advanced treated municipal effluent: Modelling the impact on groundwater chemistry and native microbial ecology," International Journal of Research & Technology, vol. 12, no. 3, pp. 128–138, 2024.

[20] A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts, “Learning word vectors for sentiment analysis,” in Proc. ACL, 2011, pp. 142–150.

[21] S. Wang and C. D. Manning, “Fast dropout training,” in Proc. ICML, 2013, pp. 118–126.

[22] M. A. Jawed, "Uncertainty-aware geostatistical reconstruction of regional groundwater hydraulics under sparse monitoring conditions: A case study of the Evergreen Underground Water Conservation District (EUWCD), Texas, USA," Bishop International Journal of Mathematics and Computer Science, vol. 1, no. 1, pp. 57–71, 2025.

[23] U. Iqbal, "AI-enhanced network optimization for electric vehicle charging infrastructure expansion in the United States using graph theory and demand analytics," *Journal of Engineering and Computational Intelligence Review*, vol. 2, no. 2, pp. 112–129, 2024.

[24] Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le, “XLNet: Generalized autoregressive pretraining for language understanding,” in Proc. NeurIPS, 2019, pp. 5753–5763.

[25] T. Brown et al., “Language models are few-shot learners,” in Proc. NeurIPS, 2020, pp. 1877–1901.

[26] C. Sun, L. Huang, and X. Qiu, “Utilizing BERT for aspect-based sentiment analysis via constructing auxiliary sentence,” in Proc. NAACL-HLT, 2019, pp. 380–385.

[27] U. Iqbal, "AI-powered supplier risk intelligence: Predicting financial and geopolitical supply chain disruptions in U.S. critical industries," *Journal of Engineering and Computational Intelligence Review*, vol. 3, no. 2, pp. 173–193, 2025.

[28] Z. Sun, H. Yu, X. Song, R. Liu, Y. Yang, and D. Zhou, “MobileBERT: A compact task-agnostic BERT for resource-limited devices,” in Proc. ACL, 2020, pp. 2158–2170.

[29] A. Adhikari, A. Ram, R. Tang, and J. Lin, “Rethinking complex neural network architectures for document classification,” in Proc. NAACL-HLT, 2019, pp. 4046–4051.

[30] M. Munikar, S. Shakya, and A. Shrestha, “Fine-grained sentiment classification using BERT,” in Proc. Int. Conf. on Artificial Intelligence for Transforming Business and Society, 2019.

[31] M. Asif and M. Rafiq‑uz‑Zaman, “The silent disengagement: A quantitative analysis of leadership, recognition, and workload as predictors of quiet quitting among knowledge workers,” Al‑AASAR Journal, vol. 3, no. 1, pp. 271–303, 2026, doi: 10.63878/aaj1525.

[32] M. Asif, “Financial performance of startups linked to universities: Evidence from developing economies,” Journal of Applied Linguistics and TESOL, vol. 8, no. 3, pp. 2736–2763, 2025, doi: 10.63878/jalt2003.

[33] M. Asif, A. Ali, and F. A. Shaheen, “Assessing the effects of artificial intelligence in revolutionizing human resource management: A systematic review,” Social Science Review Archives, vol. 3, no. 4, pp. 2887–2908, 2025, doi: 10.70670/sra.v3i3.1055.

Downloads

Published

25-03-2026

How to Cite

Fine-Tuning Lightweight Transformer Models for Text Classification: A Comparative Analysis of DistilBERT, Bert, and LSTM. (2026). Journal of Engineering and Computational Intelligence Review, 4(1), 185-195. https://doi.org/10.63544/ggeppj73

Share

Similar Articles

21-30 of 66

You may also start an advanced similarity search for this article.