Open Access Peer Reviewed Monthly Est. 2014

British International Journal of Education and Social Sciences

(BIJESS)
ISSN (Print): 4519-6511 | ISSN (Online): 3342-543X
9.82
Impact Factor
13
H-Index
893+
Articles
HomeBIJESS Vol. 13, No. 4 BOOSTING TEXT CLASSIFIER PERFORMANCE WITH LLM-DRI…
📄 Research Article BIJESS Vol. 13, No. 4 (2025)

BOOSTING TEXT CLASSIFIER PERFORMANCE WITH LLM-DRIVEN DATA AUGMENTATION

Chinedu Michael Adebayo
Department of Computer Science, University of Lagos, Nigeria
British International Journal of Education and Social Sciences, Vol. 13, No. 4 (2025), pp. 13-20 | DOI: https://doi.org/10.5281/zenodo.20159212
Open Access Peer Reviewed Research Article

Abstract

This research considers the impact of data augmentation on multi-class text classification. A diverse news dataset comprising four categories was utilized for training and evaluation. Various transformer models, including BERT, DistilBERT, ALBERT, and RoBERTa, were employed to classify text across multiple categories. Based on the previous research on data augmentation, synonym replacement, antonym replacement, contextual word embedding, and the lambada method for data augmentation were chosen. Three mainstream LLMs were selected to investigate the capabilities of LLMs: LLaMA 3, GPT-4, and MistralAI. These models represent a diverse range of architectures and training data, allowing to assess the impact of different LLM capabilities on data augmentation performance. The performance of the aforementioned transformer models was evaluated using metrics such as accuracy, recall, and precision, F1-score, training time, validation, and training loss. Experiments revealed that data augmentation significantly improved the performance of transformer models in text classification tasks, with lambada augmentation consistently outperforming other methods. However, model architecture and hyperparameter tuning also played a crucial role in achieving optimal results. Obtained results have practical implications for developing NLP applications in low-resource languages, as data augmentation can help address the limitations of small datasets.
Keywords: ["augmentation","multi-class text classification","large language models","transformers","BERT","ALBERT","DistilBERT","XLM-RoBERTa"]
📑 How to Cite This Article
APA 7th Edition:
Chinedu Michael Adebayo (2025). BOOSTING TEXT CLASSIFIER PERFORMANCE WITH LLM-DRIVEN DATA AUGMENTATION. British International Journal of Education and Social Sciences, 13(4), 13-20. https://doi.org/https://doi.org/10.5281/zenodo.20159212
Vancouver Style:
Chinedu Michael Adebayo. BOOSTING TEXT CLASSIFIER PERFORMANCE WITH LLM-DRIVEN DATA AUGMENTATION. Br. Int. J. Educ. Soc. Sci.. 2025;13(4):13-20. DOI: https://doi.org/10.5281/zenodo.20159212
🔗 Other Articles in This Issue