Developing a Vietnamese Text Summarization Large Language Model on Limited Hardware

Publication Date

1-1-2026

Document Type

Conference Proceeding

Publication Title

Conference Proceedings IEEE SOUTHEASTCON

DOI

10.1109/SoutheastCon63549.2026.11476572

Abstract

Text summarization models have achieved significant growth during the last few years because of major Large Language Model (LLM) technological advancements, and are now are widely used in news distribution (TL;DR news), translation tools (DeepL Translate), or virtual assistants. However, the progress has not yet reached all languages equally. The Vietnamese language is used by more than 90 million people, but for LLM purposes, it is still considered a low-resource language and therefore faces a certain gap when. This can lead to many disadvantages for the Vietnamese monolingual population around the world, as they cannot catch up with the current informational society. This paper addresses this gap by performing a comparative study to fine-tune and evaluate two state-of-the-art Vietnamese-specific models-BartPho-syllable (BARTbased) and ViT5 (T5-based)-for the task of abstractive news summarization. The models were trained using a parameterefficient (QLoRA) approach on a curated corpus combining the public nam194/vietnews dataset with freshly scraped articles from major Vietnamese news outlets. The key finding of this research is that both sophisticated fine-tuned models were ultimately outperformed on standard quantitative metrics (ROUGE and BERTScore) by simpler extractive baselines, particularly the Lead-3 heuristic. The evaluation scores and final validation loss demonstrated that ViT5 achieved better results than BartPhosyllable in the two fine-tuned models.

Keywords

Bart, BERT, LLM, LoRA, QLoRA, T5, text summarization, Vietnamese

Department

Computer Engineering

Share

COinS