Paper deep dive
A BART-based approach with hierarchical strategy for Vietnamese abstractive multi-document summarization
Vu Nguyen Nguyen Xuan, Huy Ngo Quang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/21/2026, 4:09:43 AM
Summary
This technical report proposes a hierarchical BART-based approach for Vietnamese abstractive multi-document summarization, specifically targeting the VLSP 2022 challenge. The method employs a two-step strategy: first, shortening documents by selecting important sentences driven by the golden summary to minimize information loss, and second, generating the final summary from the aggregated sentences. The authors also contribute to the community by translating the Multinews dataset into Vietnamese, significantly expanding the available training data for this task. The approach achieves a ROUGE2-F1 score of 0.2468 on the VLSP public test set.
Entities (8)
Relation Signals (4)
Nguyen Nguyen Xuan Vu → affiliatedwith → Aimesoft JSC
confidence 100% · Nguyen Nguyen Xuan Vu Aimesoft JSC Hanoi, Vietnam
VLSP 2022 → hostedchallenge → Vietnamese Multi-document Abstractive Summarization
confidence 100% · the International Workshop on Vietnamese Language and Speech Processing (VLSP) 2022 has been holding a competition
Multinews → translatedto → Vietnamese Multinews
confidence 100% · We devote the translated Multinews dataset for the community
BARTPho → usedfor → Vietnamese Multi-document Abstractive Summarization
confidence 100% · we choose to experiment with BARTPho (Tran et al., 2021)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the International Workshop on Vietnamese Language and Speech Processing (VLSP) 2022. We choose to follow the popular hierarchical approach, i.e. condensing each document followed by aggregation and summarization. We propose a novel yet simple strategy to shorten documents that is driven by the golden summary, thus ensuring high correlation between stages of the hierarchical approach. Our method achieves a ROUGE2-F1 score of 0.2468 on the VLSP's public test set, and can produce fluent and concise summaries. Additionally, we utilize external sources for extra data, which greatly enhances the quantity of data for Vietnamese multi-document summarization. The additional data is made available for the community.
Tags
Links
- Source: https://arxiv.org/abs/2606.19591v1
- Canonical: https://arxiv.org/abs/2606.19591v1
Trouble viewing inline? Open PDF directly →
Full Text
17,620 characters extracted from source content.
Expand or collapse full text
A BART-based approach with hierarchical strategy for Vietnamese abstractive multi-document summarization Nguyen Nguyen Xuan Vu Aimesoft JSC Hanoi, Vietnam vunx@aimesoft.com Ngo Quang Huy Aimesoft JSC Hanoi, Vietnam huynq@aimesoft.com Abstract In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the International Workshop on Vietnamese Lan- guage and Speech Processing (VLSP) 2022. We choose to follow the popular hierarchical approach, i.e. condensing each document fol- lowed by aggregation and summarization. We propose a novel yet simple strategy to shorten documents that is driven by the golden sum- mary, thus ensuring high correlation between stages of the hierarchical approach. Our method achieves a ROUGE2-F1 score of 0.2468 on the VLSP’s public test set, and can produce fluent and concise summaries. Additionally, we utilize external sources for extra data, which greatly enhances the quantity of data for Vietnamese multi-document summarization. The additional data is made available for the community. 1 Introduction Living in the age of rapid technological advance- ment in communication and media platforms, hu- mans often find ourselves flooded in an excessive amount of data. In order to not only capture the essential information but also envision the big pic- ture about a topic or an event, one often has to consume numerous documents before coming up to a final conclusion. To automate this process, multi-document summarization has emerged as an attractive research topic amongst Natural Language Processing (NLP) community. In the context of this work, we will only discuss the abstractive branch of the task, which aims to generate summaries that are fluent, succinct yet able to cover principal informa- tion of a cluster of related documents. Preferably, the summaries should refrain from exactly copying text fragments from input, and instead, be para- phrased from the input contents. The task of multi-document summarization has been extensively researched for English, yet little to no work has been done for Vietnamese. There- fore, to promote research in multi-document sum- mary for Vietnamese, the International Workshop on Vietnamese Language and Speech Processing (VLSP) 2022 has been holding a competition ded- icated to Vietnamese multi-document abstractive summarization. For English, predominant strategies to this task include graph-based approaches (Liao et al., 2018; Li et al., 2020; Pasunuru et al., 2021), which are ef- ficient in capturing semantic relationships between document segments, but often have to utilize ad- ditional information, such as discourse correlation (Christensen et al., 2013). Another approach fol- lows a hierarchical scheme (Liu and Lapata, 2019; Fabbri et al., 2019; Jin et al., 2020), where each document in a cluster is mapped to a shorter in- termediate representation, before being joined to- gether to produce the final summary. While sounds natural, these works only generate the intermedi- ate representation of documents based solely on the input, which may cause information shortage compared with the final summary, undesirably en- couraging supervised models to make up content when generating summarization. Simultaneously, pretrained seq-to-seq models, which are either multipurpose (Lewis et al., 2019; Roberts et al., 2019) or summarization-specific (Zhang et al., 2020), have demonstrated the bene- fits of pretraining on the abstractive summarization task. Nonetheless, these models are not feasible for end-to-end multi-document summary due to the long input length. For instance, BART (Lewis et al., 2019) pretrained model can only take a token array of maximum length of 1024 as input, while an av- erage cluster in the VLSP2022 dataset (Tran et al., 2022) is approximately 2220-token long. Although recently there has been pretrained language models that are designed to tackle long input sequences (Guo et al., 2021; Xiao et al., 2021), to the best of arXiv:2606.19591v1 [cs.CL] 17 Jun 2026 our knowledge, there has not been any Vietnamese pretrained language model that can take a suffi- cient number of tokens as input for multi-document tasks. In this work, we aim to tackle the challenge of multi-document summarization for Vietnamese news using the hierarchical approach. However, dif- fers from previous works following this scheme, we incorporate the golden summary into the process of generating the intermediate representations of documents to ensure that the aggregated interme- diate representation has equal or more information compared with the final summary. Furthermore, despite the effort, the amount of data available for Vietnamese multi-document sum- marization (Tran et al., 2020, 2022) is still small with only a few hundred document clusters. Hence, to ease the training process of this task, we try to enlarge the dataset, by combining all available datasets for Vietnamese multi-document summa- rization along with a Vietnamese-translated version of Multinews (Fabbri et al., 2019) into one large dataset. We devote the translated Multinews dataset for the community for further research and devel- opment. Our contributions are twofold: •We approach the challenge using the well es- tablished and effective hierarchical strategy, in which we propose a simple strategy to sam- ple sentences from documents in a cluster that can minimize information loss for subsequent summary generation. Furthermore, our ap- proach can be integrated well with the avail- able seq-to-seq Vietnamese pretrained lan- guage models, such as BARTPho (Tran et al., 2021) or ViT5 (Phan et al., 2022). •We translate the Multinews dataset (Fabbri et al., 2019) into Vietnamese, yielding 50286 additional document clusters and make it pub- licly available to community. 1 2 Proposed Methods 2.1 Data Preprocessing Our approach towards the multi-document summa- rization problem follows the two-step regime: The first step involves shortening every document in a cluster, and in step two all shortened documents are used to infer the final summary for the whole 1 https://github.com/ngoquanghuy99/Abmusu2022 cluster. In step one, for each cluster, we choose the top most important sentences across all documents until a desired length threshold is reached (Figure 1). As a measurement of importance, we compute ROUGE1 score (Lin, 2004) between each sentence and the golden summary. Following the Ind-Uniq strategy as described in (Zhang et al., 2020), we score each sentence independently and only count each n-gram once. Since this sentence selecting procedure is driven by the ground truth summary, a supervised model is trained to infer the selected sentences given a single document as input. In the second step, the chosen sentences from the first step are concatenated and fed to a seq-to-seq model to generate the final summary. To make the output summary robust to the order of the initial documents, we augment the data for the second step by permuting the order of input sentences in corre- spondence with the permutation of the documents’ order, while keeping the summary untouched. We argue that by selecting sentences based on the desired summary in step 1 and use these sen- tences for summarization in step 2, we can tackle the information mismatch between the two steps and thus prevent the model from hallucination (Maynez et al., 2020) when generating summaries. 2.2 Model Architecture We train two separate seq-to-seq models, corre- sponding with the two steps described above. In step 1, the first model is trained to generate a se- quence of sentences that are selected from an input document. while in step 2, the other model uses the aggregated sentences from step 1 to finally generate the summary for a cluster of documents (Figure 2). It is noted that any seq-to-seq language model can be used for both steps, including pretrained ones. As our first attempt, we choose to experiment with BARTPho (Tran et al., 2021), a BART-like model pretrained on Vietnamese corpus, due to the fact that BART was trained on a maximum input length of 1024 tokens. The long input length of BART allows us to select more sentences in step 1, thus reducing the risk of information loss when transferring between two steps. 3 Experiments and results 3.1 Datasets The official dataset, including training, validation and test subsets, is provided by the VLSP 2022 organizer. The training set consists of 200 clusters, Seq-to-seq Step 1's target Có nguồn gốc từ tàn dư của sao chổi Marsden và Kracht để lại, mưa sao băng Delta Aquarids xuất hiện trên bầu trời từ 12/7-23/8 hàng năm, đạt cực đại vào đêm 28, rạng sáng 29/7. Delta Aquarids là một trận mưa sao băng trung bình, có thể đạt cực đại khoảng 20 vệt sao băng/giờ. Các sao băng sẽ bắt nguồn từ chòm sao Aquarius (Bảo Bình), nhưng cũng có thể xuất hiện bất cứ đâu trên bầu trời. Ngay sau mưa sao băng Delta Aquatids, vào đêm 12 rạng sáng ngày 13/8, người yên thiên văn sẽ có cơ hội chiêm ngưỡng mưa sao băng Perseids, một trong hai trận mưa sao băng đẹp nhất năm với tần suất thời điểm cực đai có thể đạt 80 vệt sao băng/giờ. Có nguồn gốc từ tàn dư của sao chổi Marsden và Kracht để lại, mưa sao băng Delta Aquarids xuất hiện trên bầu trời từ 12/7-23/8 hàng năm, đạt cực đại vào đêm 28, rạng sáng 29/7. Delta Aquarids là một trận mưa sao băng trung bình, có thể đạt cực đại khoảng 20 vệt sao băng/giờ. Điều kiện quan sát mưa sao băng Delta Aquarids năm nay khá thuận lợi vì ít bị trăng non đầu tháng ảnh hưởng. Thời gian quan sát lý tưởng nhất là sau nửa đêm, chọn nơi ít ánh sáng đèn và ô nhiễm không khí. Các sao băng sẽ bắt nguồn từ chòm sao Aquarius (Bảo Bình), nhưng cũng có thể xuất hiện bất cứ đâu trên bầu trời. Lưu ý xem dự báo thời tiết nếu có ý định quan sát. Ngay sau mưa sao băng Delta Aquatids, vào đêm 12 rạng sáng ngày 13/8, người yên thiên văn sẽ có cơ hội chiêm ngưỡng mưa sao băng Perseids, một trong hai trận mưa sao băng đẹp nhất năm với tần suất thời điểm cực đai có thể đạt 80 vệt sao băng/giờ. Source document Figure 1: An example of the training objective of the first step. Important sentences are colored in red. each cluster has the number of documents varied from 2 to 5. All documents were crawled from Viet- namese news data on a wide range of topics. To enlarge the training set, we combine the provided dataset with two other similar datasets on multi- document summarization for Vietnamese news, namely ViMs (Tran et al., 2020) and Vietname- seMDS 2 , raising the total number of clusters up to 700 clusters. However, this number of data points is still insignificant compared with the English counterpart, Multinews (Fabbri et al., 2019), which 2 https://github.com/lupanh/VietnameseMDS Doc 1Doc 2Doc n Step 1 - Extract important sentences Concatenate Inter 1Inter 2Inter n Concat Inter ... ... Step 2 - Summary Cluster summary Figure 2: Flow of inference. contains ~56,000 clusters. Therefore, we decide to translate the Multinews dataset to Vietnamese and combine this translated set with the Vietnamese ones. In the end, we have a total ~51,000 clusters in our final training set. 3.2 Hyperparameters As mentioned, for two steps of sentence selection and summary generation, we fine-tune two distinct BARTPho (Tran et al., 2021) models. For both models, we set the learning rate at 5e-5, with a linear decay rate of 0.1. Due to the nature of two steps, the maximum input and output lengths of step 1’s model both are set to be 512, while those of step 2’s are 1024 and 256, respectively. Due to hardware limitation, the step 1’s model was trained for 2,500 steps with a batch size of 2, and the other one was trained for 17,000 steps with a batch size of 4. During the decoding processes of both steps, we use a beam width of 3 with a length-penalty of 2.0 and a repetition penalty of 2.5. We also ensure that no 3-gram is repeated during generation. Table 1: Our method’s performance on the public test set, compared with baselines provided by VLSP 2022 organizer. modelR1-F1 R2-F1 RL-F1 Ours0.4550 0.2468 0.4129 extractive baseline 0.4836 0.2650 0.4421 rule baseline0.4640 0.2582 0.4284 anchor baseline0.4381 0.1931 0.3928 abstractive baseline 0.3129 0.1457 0.2797 3.3 Experimental Results 3.3.1 System evaluation Our models achieve cross-entropy losses of 1.95 and 1.97 on the validation sets of step 1 and step 2, respectively. Aside from the training and validation set, the or- ganization of VLSP also provides a test set, which is divided into a public set and a private set, along with four baselines. On the public test set, our method demon- strates a relatively good performance (ROUGE2- F1=0.2468) compared with the anchor baseline (ROUGE2-F1=0.1931) and the abstractive baseline (ROUGE2-F1=0.1457) implemented by the orga- nizer. However, we are unable to outperform the extractive (ROUGE2-F1=0.2650) and rule-based baselines (ROUGE2-F1=0.2582). 3.3.2 Human evaluation Despite having a humble performance on the public test set, based on human observation, our method can produce grammatically correct and comprehen- sible summaries. The summaries can also mostly cover the principal content of its corresponding documents. 4 Conclusion In this report, we have presented our effort on tack- ling the Vietnamese multi-document abstractive summarization challenge of VLSP 2022. Our pro- posed method shows prominent outcome and can be further tuned to produce better results. Some improvements we are thinking of include using smaller text units rather than sentences for step 1, introducing boundaries between documents when concatenating intermediate representations, and re- moving duplicated contents in the aggregated rep- resentation before feeding it to the step 2’s model. References Janara Christensen, Stephen Soderland, Oren Etzioni, et al. 2013. Towards coherent multi-document sum- marization. In Proceedings of the 2013 conference of the North American chapter of the association for computational linguistics: Human language tech- nologies, pages 1163–1173. Alexander R Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R Radev. 2019. Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model. arXiv preprint arXiv:1906.01749. Mandy Guo, Joshua Ainslie, David Uthus, Santiago On- tanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang. 2021. Longt5: Efficient text-to-text transformer for long sequences. arXiv preprint arXiv:2112.07916. Hanqi Jin, Tianming Wang, and Xiaojun Wan. 2020. Multi-granularity interaction network for extractive and abstractive multi-document summarization. In Proceedings of the 58th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 6244– 6254, Online. Association for Computational Lin- guistics. Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: De- noising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461. Wei Li, Xinyan Xiao, Jiachen Liu, Hua Wu, Haifeng Wang, and Junping Du. 2020. Leveraging graph to improve abstractive multi-document summarization. arXiv preprint arXiv:2005.10043. Kexin Liao, Logan Lebanoff, and Fei Liu. 2018. Ab- stract meaning representation for multi-document summarization. arXiv preprint arXiv:1806.05655. Chin-Yew Lin. 2004. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pages 74–81, Barcelona, Spain. Asso- ciation for Computational Linguistics. Yang Liu and Mirella Lapata. 2019. Hierarchical trans- formers for multi-document summarization. arXiv preprint arXiv:1905.13164. Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factu- ality in abstractive summarization. arXiv preprint arXiv:2005.00661. Ramakanth Pasunuru, Mengwen Liu, Mohit Bansal, Su- jith Ravi, and Markus Dreyer. 2021. Efficiently sum- marizing text and graph encodings of multi-document clusters. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4768–4779, Online. Association for Computational Linguistics. Long Phan, Hieu Tran, Hieu Nguyen, and Trieu H Trinh. 2022. Vit5: Pretrained text-to-text transformer for vietnamese language generation. arXiv preprint arXiv:2205.06457. Adam Roberts, Colin Raffel, Katherine Lee, Michael Matena, Noam Shazeer, Peter J Liu, Sharan Narang, Wei Li, and Yanqi Zhou. 2019. Exploring the limits of transfer learning with a unified text-to-text trans- former. Mai-Vu Tran, Hoang-Quynh Le, Duy-Cat Can, and Quoc-An Nguyen. 2022. Vlsp 2022 – abmusu chal- lenge: Vietnamese abstractive multi-document sum- marization. Nguyen Luong Tran, Duong Minh Le, and Dat Quoc Nguyen. 2021. Bartpho: Pre-trained sequence-to- sequence models for vietnamese. arXiv preprint arXiv:2109.09701. Nhi-Thao Tran, Minh-Quoc Nghiem, Nhung TH Nguyen, Ngan Luu-Thuy Nguyen, Nam Van Chi, and Dien Dinh. 2020. Vims: a high-quality viet- namese dataset for abstractive multi-document sum- marization. Language Resources and Evaluation, 54(4):893–920. Wen Xiao, Iz Beltagy, Giuseppe Carenini, and Arman Cohan. 2021. Primer: Pyramid-based masked sen- tence pre-training for multi-document summariza- tion. arXiv preprint arXiv:2110.08499. Jingqing Zhang, Yao Zhao, Mohammad Saleh, and Pe- ter Liu. 2020. Pegasus: Pre-training with extracted gap-sentences for abstractive summarization. In In- ternational Conference on Machine Learning, pages 11328–11339. PMLR.