Paper deep dive
Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric
Sami Shames El Deen, Mariette Awad
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/24/2026, 5:46:24 AM
Summary
This paper introduces SAraBERT, an enhanced version of the AraBERT model designed for extractive summarization of Arabic documents. The model incorporates inter-sentence transformer layers to better capture sentence boundaries and importance. Additionally, the authors propose Semantic Siamese Similarity (SSS), a novel evaluation metric that combines semantic embedding similarity, exact match (ROUGE), and frequency distribution to assess summary quality. Experiments on the Kalimat dataset (translated from CNN/Daily Mail) demonstrate that SAraBERT, particularly when paired with an RNN encoder, outperforms existing models in BLEU, ROUGE, and SSS metrics.
Entities (10)
Relation Signals (7)
SAraBERT → isenhancedversionof → AraBERT
confidence 95% · SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers
SAraBERT → uses → inter-sentence transformer layers
confidence 95% · proposes inter-sentence transformer layers for extractive summarization tasks
SAraBERT → achievesbestperformancewith → RNN
confidence 92% · SAraBert + RNN was the best performing model based on the 3 evaluation metrics.
SAraBERT → isevaluatedon → Kalimat Dataset
confidence 92% · Testing the performance of Sarabert in comparison with other published models was evaluated using BLEU, ROUGE, and Semantic Siamese similarity on Kalimat Dataset.
Semantic Siamese Similarity → combines → ROUGE
confidence 90% · propose a novel hybrid similarity metric that measures similarities between 2 text based on the semantic and syntactic features... An average of ROUGE 1, ROUGE 2, and ROUGE L is used.
Kalimat Dataset → isderivedfrom → CNN/Daily Mail
confidence 90% · We use an English benchmark dataset (CNN/Daily Mail) and translate it to Modern Standard Arabic language to train SAraBERT.
SAraBERT → outperforms → AraBert Embeddings + K-Means
confidence 85% · Simulation results showed the effectiveness of our proposed model... SAraBert + RNN... outmatches its predecessors
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document's main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs. We validated using BLEU, ROUGE, and Semantic Siamese similarity on Sarabert and published related models. Simulation results showed the effectiveness of our proposed model and motivate follow on research.
Tags
Links
- Source: https://arxiv.org/abs/2608.20964v1
- Canonical: https://arxiv.org/abs/2608.20964v1
Trouble viewing inline? Open PDF directly →
Full Text
35,155 characters extracted from source content.
Expand or collapse full text
1 EXTRACTIVE SUMMARIZATION FOR ARABIC DOCUMENTS USING SARABERT WITH A SEMANTIC SIAMESE SIMILARITY EVALUATION METRIC Sami Shames El Deen , Mariette Awad American University of Beirut Abstract—In this research, we introduce SAraBERT, an en- hanced version of AraBERT which proposes inter-sentence trans- former layers for extractive summarization tasks. To ensure that the summaries generated by SAraBERT achieve a high coverage of the document’s main ideas, we propose Semantic Siamese Similarity, a novel evaluation metric that measures the level of similarity between two text inputs. We validated using BLEU, ROUGE, and Semantic Siamese similarity on Sarabert and published related models. Simulation results showed the effectiveness of our proposed model and motivate follow on research. Impact Statement—This paper introduces SAraBERT, an im- provement on Arabert, that incorporates inter-sentence trans- former layers to perform extractive summarization on Arabic text. Furthermore, a novel called Semantic Siamese Similarity (S) is proposed to combine embedding similarity and exact match. Experiments using various evaluation metrics show that SAraBert outmatches its predecessors on Arabic extractive sum- marizing reaching new state of the art levels of performance. Index Terms—Arabic NLP , Text Summarization, Extractive Summarization, Transformers, Evaluation Metric, Sequence-to- Sequence Framework I. INTRODUCTION A S the amount of textual information available online grows rapidly, it becomes difficult for readers to read large amounts of text and find out which of these texts are useful. As a result, researchers in the field of automatic text summarization need tools that can make multiple-document reading more efficient. The task of text summarization is considered one of the most important and challenging NLP tasks. Summarization generates short text from long text, so the short text contains the most important information from the original text. The task is often divided into two paradigms known as extractive summarization and abstractive summarization. The first methodology determines the essential sections of the text using statistical tool. The summary is represented by truncating and connecting these sections. The second methodology emulates human activity in summarizing, which is based on presenting the text’s core idea in a new linguistic style and using different terms. It incorporates more complicated procedures including paraphrase, generalization, and reordering [1].This work focuses on extractive summa- rization. While there is a wealth of research in the area of extractive summarization, most of this research is based on English texts, and there is a lack of research on summarizing Arabic texts. Arabic natural language processing (NLP) [2] is considered more complex than English and other European languages. The main reason for this complexity is the highly derived and rich form of Arabic morphology. This creates a set of challenges, which include: 1) Morphological richness [3][4]: Arabic is heavily derived, inflected and has a significant impact on NLP tasks such as stemming and lemming. 2) Absence of diacritics: diacritics play an important role in determining the meaning of words and facilitating the task of tokenizing and parsing text. 3) No capitalization in Arabic language: without the usage of uppercase letters in Arabic, it will be difficult to identify proper nouns, titles, and abbreviations. In this work, we introduce SAraBERT, a novel and en- hanced version of AraBERT that adds inter-sentence trans- former layers for extractive summarization tasks. To ensure that the summaries generated achieve a high coverage of the document’s main ideas, we also propose a novel evaluation metric, Semantic Siamese Similarity that measures the level of similarity between two text inputs using contextualized embeddings in addition to exact match. We use an English benchmark dataset (CNN/Daily Mail) and translate it to Mod- ern Standard Arabic language to train SAraBERT. Testing the performance of Sarabert in comparison with other published models was evaluated using BLEU, ROUGE, and Semantic Siamese similarity on Kalimat Dataset. Results have shown the effectiveness of our proposed model. In summary, the main contributions of this work are as follows: • A novel Arabic language model for extractive summa- rization that builds on AraBERT model. • A novel hybrid similarity metric that measures similarities between 2 text based on the semantic and syntactic features. • An arabic corpus translated from English that targets the extractive summarization and Question Answering tasks, available upon request. The structure of the paper is such that, in section 2, we survey the related work and background information required for this work. Section 3 describes the model’s architecture arXiv:2608.20964v1 [cs.CL] 21 Aug 2026 2 and the proposed metric. Section 4 shows the experimental research and highlights the obtained results along with some insights. Finally Section 5 concludes the paper with follow on research directions. I. LITERATURE REVIEW To provide the necessary context to the proposed solution, it is worth investigating previous research. A. Classical Approaches Until 2017, statistical methods were predominantly used in text summarization which are based on the concept of relevance score and Bayesian classifiers [5]. Word Frequency approach is the most used methodology for sentence scoring where the sentence score is calculated by the sum of the fre- quencies while avoiding all stop words. The proposed solution in [6] generates news titles by combining Word-Frequency, sentence position and similarity measures methodologies. Term Frequency-Inverse Document Frequency (TF-IDF) is a numerical methodology that represents the importance of a word to document in a collection of documents (corps) [7]. TF- IDF is an improvement on the Word Frequency and shows how the weights are distributed on the words of a document. TF- IDF is used frequently in auto-summarization systems [8] [9] [10] [11]. An Arabic summarization system called AATSS [12] adopted an extractive text summarization approach that was mainly based on sentence weighting and scoring. However, it highly depends on the size of the document in terms of number of sentences, which leads to generation of improper grammar context, and it lacks automatic Arabic entity recog- nition system and Arabic pronoun resolution system. The classical statistical approaches are used for both single and multi-document summarization and can be used to enhance the selection of important sentences or the elimination of redundant sentences but fail to understand the text, since they only depend on statistical measures [13]. [14] proposed a solution for English language by tokenizing the text into clean sentences, then passing the sentences into a BERT model to generate embeddings. Then using K-means they generated clusters and selected summary sentences based on the closest sentence embedding to each cluster centroid. B. Machine Learning Approach Text summarization is considered as a binary classification problem, where a set of documents and their extractive sum- maries are used as a training set, and each sentence is classified as a summary sentence or non-summary based on statistical, semantic features or a combination of them [15] [16] [17] [18]. Machine learning methods can be suitable for document summarization as several methods have been introduced and proven to be effective in performing summarization tasks. We will discuss sequence to sequence models and encoder decoder architectures in what follows. Sequence-to-Sequence Model The neural abstractive summarization with sequence-to- sequence models emerged in [19] [20]. This approach has been applied to tasks such as headline generation [21] and article summarization [22]. Chopra et al. [23] show that attention approaches that are more specific to summarization can further improve the performance of models. Molham et al. [24] re-implemented the sequence-to-sequence framework on the Arabic language, which had not witnessed the employment of this model in the text summarization before. However, the authors state that the work still requires expanding the dataset to cover more articles, and to train new models that are beneficial with the Arabic language since it has a unique grammatical language written from right to left. Encoder-Decoder Model Google AI’s pre-trained language model called Bidirectional Encoder Representations from Transformers (BERT), proved highly efficient at language understanding and achieved con- vincing results in most NLP tasks. The Arabic language model ARABERT based on BERT for Arabic language was evaluated on Named Entity Recognition, Sentiment Analysis and Question Answering [25]. Abdulla et al. [26] proposed an extractive Arabic text summarizer based on ARABERT to summarize the Arabic documents by evaluating and ex- tracting the most important sentences at a document. This proposed approach generated a good summary by extracting the most important sentences from paragraphs. However, the model highly depends on sentence boundaries, its coverage accuracy decreases when the text is too long, and extracted sentences sometimes contain linguistic expressions that creates ambiguous summaries. C. Other Approaches Liu et al. of [27] managed to solve the problem of sentence boundary detection by adding an additional layer in the embed- ding that specifies the beginning and end of each consecutive sentence inserted into the encoder. Go et al. [28] explored the effects of language variants, data sizes, and fine-tuning task types in Arabic pre-trained language models. The results suggested that the variant proximity of pre-training data to fine-tuning data is more important than the pre-training data size. Moreover, other design decisions can be explored that may contribute to the fine-tuning per- formance, including vocabulary size, tokenization techniques, and additional data mixtures. Nasser et al. [29] presented an LSTM-based morphological disambiguation system for Arabic which has significantly outperformed the state-of-the-art systems. The paper suggested exploring additional deep learning architectures for morpho- logical modeling and disambiguation, especially sequence-to- sequence models. It also suggested to further investigate the role of syntax features in morphological disambiguation, and explore additional techniques for more accurate tagging. Reda et al. [30] developed a hybrid transformer based approach for Arabic text summarisation. First, they performed extractive summarization using a AraBert based sentence embedding approach to determine which sentences of an article are most suited to be part of the summary. Then, they performed abstractive summarization on the extracted summary by feeding it into the MT5 Arabic transformer. They 3 tested their model on the ESAC dataset and got a precision of 53%, a recall of 55%, and an f1 score of 49%. The main shortcoming of this work are as follows. First, pretrained models were used meaning that the models were not fine tuned specifically to perform summarization. Second, the ESAC dataset is limited to only 153 articles, and 700 summarizations. Using specialized models and a larger dataset may further improve the results. I. PROPOSED METHODOLOGY A. SAraBERT Let d denote a document containing several sentences [sent 1 , sent 2 , . . . , sent m ], where sent i is the i th sentence in the docu- ment. Extractive summarization can be defined as the task of assigning a label y i ∈ 0, 1 to each sent i , indicating whether the sentence should be included in the summary or not. Thus, if y i = 1 then the sentence is considered in the summary due to its importance in terms of document content. BERT as shown in Figure[1] is considered the model of choice because it showed better performance compared to other NLP statement embedding algorithms. To keep the original BERT pre-training objective, AraBERT uses Masked Language Modeling (MLM) to improve the pre-training pro- cess by letting the model predict the entire word rather than receiving clues from part of the word. In addition, it also uses Next Sentence Prediction (NSP) to help the model understand the relationships between sentences [25]. To achieve sentence ranking in the document based on content importance, we will apply modifications to the embedding layers as shown in Figure[2] to highlight sentence endpoints for the model to rank each sentence independently. Encoding Multiple Sentences: We insert [CLS] token be- fore each sentence and a [SEP] token after each sentence. The [CLS] is used as a symbol to represent the start of a sequence and the [SEP] indicated the end of that sequence. Inserting multiple [CLS] tokens help in specifying the boundaries of each sentence. Interval Segment Embedding: We add another layer of embedding to distinguish multiple sentences within a doc- ument. So, for sent i we assign a segment embedding E A or E B based on whether it’s even or odd. For example, For [sent 1 , sent 2 , sent 3 , sent 4 , sent 5 ] the interval segment layer assign [E A , E B , E A , E B , E A ]. The vector T i that corresponds to the i th [CLS] token, will be used as the representation for sent i . Fig. 1: Overview for the architecture of the original BERT model The sequence on top is the input document, followed by the summation of three kinds of embeddings for each token. The summed vectors are used as input embeddings to the transformer layers, generating contextual vectors for each token T i . The difference between SAraBERT and the model Fig. 2: Overview for the architecture of the SAraBERT model in Figure[1] is the insertion of [CLS] tokens before each sentence, and alternating the sentence segmenting embedding values along with the addition of a summarization layer that derives a sentence score from the [CLS] token embeddings. 4 Algorithm 1 Summarization Model Input: Text Output : Sentence Scores tokens← tokenize(Text) sentences← sentence segmentation(tokens) tokenID ← array [|sentences|] [max sent∈sentences (|sent|)] segmentID ← array [|sentences|] [max sent∈sentences (|sent|)] for i = 1 to |sentences| do sentences[i] ←”[CLS]”, sentences[i], ”[SEP]” for j = 1 to |sentences[i]| do tokenID[i][j]← getTokenID(sentences[i][j]) if i is Even then segmentID[i][j]← 0 else segmentID[i][j]← 1 end if end for end for bertEmbedding ← AraBERT(tokenID,segmentID) sentence embeddings← getCLSembeddings(bertEmbedding) sentencescores← encode(sentenceembeddings) After obtaining the embeddings of CLS tokens of each sen- tence, the tokens will get fed into an encoder ˆ Y = H(T) where H is the encoding model, T is a vector of CLS embeddings, and ˆ Y is a scoring vector with ˆy i representing the score of the i th sentence. The model H has been interchanged between Multi-Layer Perceptrons (MLP), Recurent Neural Network (RNN), and Transformers to compare the different scoring criteria each encoder have. Results are shown in section V B. Siamese Semantic Similarity (S) ROUGE [31] has been used as a metric for determining the quality of a summary by comparing a candidate summary (generated by machine) to a reference summary (generated by humans). The way it does this projects only the level of similarity on the syntactic level without the coverage of the context. Summaries may cover a large portion of the original context but with fewer words (and more verbose) which Rouge can’t keep track of. Language models have the ability to map the context of a passage into a fixed size vector. We can compute the cosine similarity of the 2 embeddings to obtain a similarity measure at the semantic level. TABLE I: An example comparing the ROUGE score and Cosine Similarity Reference...J À @ ̇ Ø – A™¢À @ ...ø AK ̇ Œ ́ROUGECosine Similarity CandidateZ AÇ” – A™¢À @ » A J K ÒÎ0.00.9074 Since cosine similarity treats all dimensions equally in the semantic space, and 2 vector points might overlap yet stay very distant from each other, the distance should be added into the evaluation, since the smaller the value is the higher the score should be, we insert the distance difference in the denominator. Moreover, we will wrap it with a square root to slow the growth speed of the value as the difference increase (this will later become helpful in interpreting the computed values). cossim(cand embedding , ref embedding ) p ||cand embedding − ref embedding || 2 + 1 (1) A value of 1 was added because the distance difference can be 0, thus to eliminate the possibility of a division by zero error. TABLE I: An example showing the cosine similarity and L2 norm when comparing a reference to predictions Reference Èk A Æ K YÀÒÀ @ » A J KCosSimL2 norm Candidate 1 È¢ Ø ... Æ¢À @ ...ø @ 0.8735.989 Candidate 2 Èͪ A Ø ... Æ¢À @ ...ø @ 0.8675.686 Moreover, ROUGE is still an important component of the formula to keep an eye on the grammar. An average of ROUGE 1, ROUGE 2, and ROUGE L is used. cos sim(cand embedding , ref embedding )· rouge(cand, ref) p ||cand embedding − ref embedding || 2 + 1 (2) Finally, frequency should be considered. Two sentences can be similar in context yet one being redundant more than the other. TABLE I: An example showing the cosine similarity, ROUGE score, L2 norm, and frequency difference when comparing a reference to predictions Èk A Æ K YÀÒÀ @ » A J K CosSimRougeL2 normFreqDiff Èk A Æ K ... Æ¢À @ ...ø @ 0.9570.2223.6371.000 » A J K » A J K Èk A Æ K Èk A Æ K YÀÒÀ @ 0.9300.8895.0051.667 Thus a penalty should be added on the difference in fre- quency distribution among words where we divide the the number of unique words over the number of total words. This step was added because an embedding does not take redundancy and coverage into consideration which can affect level of similarity for tasks such as summarization. The final Formula can be seen in Algorithm[2] TABLE IV: interpretation of S scores S scoreInterpretation 0-0.1no clear resemblance or no similarity at all 0.1-0.2the gist is clear, but not all context is covered 0.2-0.5context is significantly covered using different words ≥0.5exact replica with slight modifications 5 Algorithm 2 Semantic Textual Similarity Input: Candidate, Reference Output : SimilarityScore rouge 1 ← rouge 1 (Candidate, Reference) rouge 2 ← rouge 2 (Candidate, Reference) rouge L ← rouge L (Candidate, Reference) rouge ← rouge 1 +rouge 2 +rouge L 3 embedding 1 ← AraBERT.encode(Candidate) embedding 2 ← AraBERT.encode(Reference) cosine similarity ← embedding 1 ·embedding 1 ∥embedding 1 ∥ 2 ×∥embedding 2 ∥ 2 DistDif f ←∥embedding 1 − embedding 2 ∥ 2 F reqDif f c ← uniquewords(Candidate) totalwords(Candidate) F reqDif f r ← uniquewords(Reference) totalwords(Reference) F reqDif f ← min(F reqDiff r ,F reqDiff c ) max(F reqDiff r ,F reqDiff c ) SimilarityScore ← cosinesimilarity×rouge×F reqDiff √ DistDiff+1 IV. EXPERIMENT A. Dataset 1) CNN/ Daily Mail: We used the CNN / DailyMail dataset to train SAraBERT, it is an English dataset containing over 300,000 unique news articles written by CNN and DailyMail journalists. The dataset can be used to train a model for the task of extractive and abstractive summarization. Each document instance from the dataset consists of 3 components: 1) id: A string containing the hexadecimal SHA1 hash of the URL from which the story was obtained. 2) Article: A string containing the body of a news article. 3) Highlights: A string containing the highlights of the article written by the author of the article. Since our model should be trained on an Arabic dataset, we examined state-of-the-art machine translation model mbart [32] along with Google-Trans 1 service. We evaluated each of the models on the Opus dataset [32]. Results are shown in Table[V]. We can see competitive results from both translators, so the tie breaker was the duration required to translate a document. TABLE V: Translation Experiment Results BLEUMETEOR Duration (Docs/minute) GoogleTrans0.413± 0.3860.2430 mBart0.413± 0.3820.2520 We also generated diacriticisation for the Arabic documents [33] to see whether the diacritics have an effect on the readability level or not. Table[VI] and Figure[4] shows that the documents containing diacritics have closer readability estimations to English than documents without diacritics but it will not be a problem since [34] explained that this behavior is normal and what is important is that the correlation between Flesch and OSMAN (without diacritics) is high. We investigated if translation will affect the readability of a document and whether large difference between the English 1 pypi.org/googletrans Fig. 3: Sample English passage translated to Arabic and Diactritized and Arabic readabilities of the original and translated docu- ments will influence our model. After experimenting on 1000 translated samples, and computing the ROUGE score of the resulting summaries along with the readability of the original and translated samples, results have shown that the level of readability does not affect the model’s output and there is no correlation between readability difference and ROUGE. More details are provided in Appendix[VII-C]. TABLE VI: Average Readability Measures for Arabic MetricValue (Mean± std) OSMAN - Diacritics78.45±5.18 OSMAN + Diacritics70.28±8.95 Fig. 4: Histogram of readability comparisons between english and arabic with and without diacritics V. RESULTS Results are reported in Table[IX]. The 3 metrics used for evaluation are BLEU (Average of the 4 BLEU evaluations over uni-,bi-,tri-, and quad-grams), ROUGE (Average between F1-Scores of ROUGE-1,ROUGE-2, and ROUGE-L), and the Siamese Similarity metric. Oracle Score represents the evalu- ation of human summaries that will be used for comparison with machine summaries. ”SAraBert+RNN” was the best performing model based on the 3 evaluation metrics. Sample summaries given by the different models can be found in the Appendix[VII-B]. 6 TABLE VII: BLEU scores over Kalimat Dataset Modelbleu1 bleu2 bleu3 bleu4avgmin max SAraBert + Classifier0.4080.3850.380.3760.3870.00.99 SAraBert + RNN0.4740.450.4450.4410.4520.00.99 SAraBert + Transformer 0.3910.3680.3640.360.3710.00.99 AraBert Embeddings + K-Means 0.2840.2540.2490.2460.2580.00.991 AraVec + K-Means0.3450.3340.3310.3280.3350.00.973 Bag of Words0.3610.3510.3480.3450.3520.00.98 TABLE VIII: ROUGE scores over Kalimat Dataset Modelrouge1 rouge2 rougeLavgminmax SAraBert + Classifier0.5290.4660.510.5020.0080.99 SAraBert + RNN0.5880.5240.5680.560.0080.99 SAraBert + Transformer 0.5120.4440.4910.4820.0170.99 AraBert Embeddings + K-Means 0.4450.3660.4270.4140.0081.0 AraVec + K-Means0.5450.490.5350.5230.01.1 Bag of Words0.470.4150.460.4480.01.1 TABLE IX: Results of different models under several evaluation metrics ModelBLEUROUGE Siamese Similarity BSTS-S Oracle Score0.532 ± 0.270.699 ± 0.230.326 ± 0.30.830.003 SAraBert + Classifier0.387 ± 0.30.502 ± 0.310.211 ± 0.1580.760.018 SAraBert + RNN0.452 ± 0.20.560 ± 0.260.282 ± 0.180.790.017 SAraBert + Transformer 0.371 ± 0.30.482 ± 0.310.240 ± 0.170.770.019 AraBert Embeddings + K-Means 0.258 ± 0.210.414 ± 0.250.216 ± 0.10.760.032 AraVec + K-Means0.335 ± 0.280.523 ± 0.280.191 ± 0.180.740.039 Bag of Words0.352 ± 0.350.448 ± 0.360.217 ± 0.280.760.111 Fig. 5: Pearson Correlation between the similarity metrics VI. DISCUSSION It appears that the usage of BiLSTM encoder did better than the transformer encoder, this could be due to the need of transformers for huge amount of data to train all its attention heads and layers. Another interesting thing to notice is that Bag of Words method got a high score on the Siamese Similarity distinguishably higher than the other classical ap- proaches. One problem remains is the incapability of feeding very large documents into the models to obtain a single global lookup on the document instead of segmenting the document and loosing linked contexts between the trimmed passages. Future work will focus on the S metric, where the cosine similarity should give more importance/weight/attention to specific context features. VII. CONCLUSION Text summarization is one of the important areas of research in NLP, as textual data keep increasing each day. The proposed model ”SAraBert” works on the summarization task for MSA NLP. We applied experiments on a large-scale translated dataset and found that SAraBert with RNN encoding layers can achieve the best performance. We also created a similarity metric that evaluated the similarity of 2 documents on the semantic level and not just the syntactic level. 7 APPENDIX A. Siamese Semantic Similarity Experiments The metric that was discussed is Section[I-B] had to be tested for different cases to visualise and evaluate the overall performance of this metric. We applied S on a set of candidates shown in Table[X] to the reference found at the first row of that table. The results of S and BertScore can be found in Table[XI] along with the values required to compute S. The correlation between BertScore and S is 0.97 showing that both metrics behave very similar but S tends to be more strict with its evaluations. TABLE XI: Results of computing S and BertScore (BS) along with other required computations for the candidates in Table[X] linked by ID ID CosineROUGEnorm 2 FDSSSBS 10.5980.09.4060.9230.00.631 20.6380.1119.10.9230.0650.682 31.01.01.01.01.01.0 40.9950.9092.0121.00.9010.971 50.9800.8572.9290.9930.8290.935 60.8770.415.7190.9230.3290.772 70.7340.2058.860.9230.1370.697 80.8160.2866.8940.9230.2130.728 90.76300.197.7890.9230.1330.693 100.80000.197.5260.9230.1390.718 110.7630.197.9260.9230.1330.702 120.8240.06.7560.9230.00.717 130.8910.2385.3960.9230.1940.74 140.8120.3147.4520.9230.2330.722 150.9040.3145.2050.9230.260.741 160.7010.08.1360.9230.00.662 170.8110.0586.7770.9230.0430.669 180.7220.0788.00.9230.0520.623 190.8450.2576.2470.9230.1990.743 200.6770.2869.2430.2170.0410.616 210.8690.8746.7780.4330.3260.778 B. Sample Summaries form SaraBert This section will present sample summaries extracted via SaraBert, highlights are used to visualize the selected top 3 sentences selected to be considered as the most informative sentences. Fig. 6: Sample SaraBert summarization. Yellow highlighting represents summary extraction using MLP as encoder, Green represents RNN encoder, and Blue represents Transformer Fig. 7: Sample SaraBert+BiLSTM of a good summarization Fig. 8: Sample SaraBert+BiLSTM of a bad summarization C. Translation Quality In this section we will present sample translations and discuss the quality and performance effect on the results. Fig. 9: Sample English passage with Fleach score of 11.38 (Hard to read) Fig. 10: Sample Arabic passage (translation of Figure[9]) with Osman score of 66.73 (Slight hard to read) 8 TABLE X: List of a reference sentence (1) and set of candidate sentences(2 to 19) IDText 0 Èk A Æ K ‡ Aø È” A™£ . Z A K C JÀ @ – ÒK Ag AJ . ì ÈÉP Y÷ œ @ ̇ Ø È” A™£ ̇ ◊ AÉ YJ “ JÀ @ ...ø @ 1H . QÎ 2...ø @ 3 Èk A Æ K ‡ Aø È” A™£ . Z A K C JÀ @ – ÒK Ag AJ . ì ÈÉP Y÷ œ @ ̇ Ø È” A™£ ̇ ◊ AÉ YJ “ JÀ @ ...ø @ 4 Èk A Æ K ‡ Aø È” A™£ . Z A K C JÀ @ – ÒK Ag AJ . ì ÈÉP Y÷ œ @ ̇ Ø È” A™£ ̇ ◊ AÉ YJ “ JÀ @ » A J K 5 Èk A Æ K ‡ Aø È” A™£ . Z A K C JÀ @ – ÒK Ag AJ . ì ÈÉP Y÷ œ @ ̇ Ø È” A™£ ̇ ◊ AÉ » A J K 6 Èk A Æ K ‡ Aø È” A™£ . Z A K C JÀ @ – ÒK Ag AJ . ì ÈÉP Y÷ œ @ ̇ Ø È” A™£ ̇ ◊ AÉ YJ “ JÀ @ ...ø @ 7 ̇ ◊ AÉ ...ø @ 8 Èk A Æ K ̇ ◊ AÉ ...ø @ 9 Èk . Ag . X ̇ ◊ AÉ ...ø @ 10A” A™£ ̇ ◊ AÉ ...ø @ 11 ̇ ◊ AÉ h A Æ JÀ @ ...ø @ 12AÍ” A™£ ÈJ ” AÉ I ø @ 13h A Æ K ÈJ . k . ̇ ◊ AÉ » A J K ÈÉP Y÷ œ @ ̇ Ø 14 ̇ ◊ AÉ ...ø @ ÈÉP Y÷ œ @ ̇ Ø 15 Èk A Æ K ̇ ◊ AÉ ...ø @ Ag AJ . ì Z A K C JÀ @ P AÓ E 16P AæÉ @ Ë Q K Am . ' . AæJ Ç k . È J“÷ œ @ H P A Ø 17...ø B @ ©J ¢ Ç B È K @ ̇ Ê™K ΩÀ X ‡ @ —Í Ø ̇ Êk AÍ™” – Òí ‡ Aø 18Ω J™ J” @ ̇ Ø º Y ́ AÉ @ ‡ @ ̇ Õ i÷ fi Ö @ 19º A JÎ Èk A Æ K » A J K ‡ @ YK QK È K B ÈÉP Y÷ œ @ ̇Õ @ H . AÎ Y À AÇ“j J” Z A K C JÀ @ – ÒK h AJ . ì ̇ ◊ AÉ ° ÆJ É @ 20 Èk A Æ K ̇ ◊ AÉ ...ø @ Èk A Æ K ̇ ◊ AÉ ...ø @ Èk A Æ K ̇ ◊ AÉ ...ø @ Èk A Æ K ̇ ◊ AÉ ...ø @ Èk A Æ K ̇ ◊ AÉ ...ø @ 21 ÈÉP Y÷ œ @ ÈÉP Y÷ œ @ ̇ Ø ÈÉP Y÷ œ @ ÈÉP Y÷ œ @ È” A™£ ̇ ◊ AÉ ̇ ◊ AÉ ̇ ◊ AÉ YJ “ JÀ @ ...ø @ ...ø @ ...ø @ ...ø @ ...ø @ ...ø @ Èk A Æ K Èk A Æ K Èk A Æ K Èk A Æ K Èk A Æ K ‡ Aø È” A™£ . – ÒK – ÒK – ÒK Z A K C JÀ @ – ÒK Ag AJ . ì Fig. 11: Sample English passage with Fleach score of 83.69 (Easy to read) Fig. 12: Sample Arabic passage (translation of Figure[11]) with Osman score of 88.01 (Easy to read) The translations have shown no effect on the trained model, where the readability variation had no link with ROUGE results. We sampled 1000 documents and computed their respective readability metrics in Arabic and English (Osman and Flesch) along with the ROUGE value of the generated summary. By looking at table[XIII], it can be observed that no readability metric highly affects the ROUGE evaluation. TABLE XIII: Correlation between the different features IndexFleschOsmanRouge-ARReadability-Diff Flesch1.00.76-0.12-0.64 Osman0.761.0-0.040.01 rouge-AR-0.12-0.041.00.14 Readability-Diff-0.640.010.141.0 TABLE XII: Average value for documents computed over 1000 sample documents Flesch59.93± 10.13 Osman74.88± 7.79 Rouge-AR0.41± 0.14 Readability-Diff14.95± 6.55 9 REFERENCES [1] H. Jing, “Using hidden markov modeling to decompose human-written summaries,” Computational linguistics, vol. 28, no. 4, p. 527–543, 2002. [2] N. Hajj and M. Awad, “Weighted entropy cortical algorithms for isolated arabic speech recognition,” IEEE, 2013. [3] A. B. Al-Saleh and M. E. B. Menai, “Automatic arabic text summariza- tion: a survey,” Artificial Intelligence Review, vol. 45, no. 2, p. 203–234, 2016. [4] K. S. Al Harazin, “Multi-document arabic text summarization,” 2015. [5] V. Patil, M. Krishnamoorthy, P. Oke, and M. Kiruthika, “A statistical approach for document summarization,” Department of Computer Engi- neering Fr. C. Rodrigues Institute of Technology, Vashi, Navi Mumbai, Maharashtra, India, 2004. [6] F. Alotaiby, S. Foda, and I. Alkharashi, “New approaches to automatic headline generation for arabic documents,” Journal of Engineering and Computer Innovations, vol. 3, no. 1, p. 11–25, 2012. [7] R. Ferreira, L. de Souza Cabral, R. D. Lins, G. P. e Silva, F. Freitas, G. D. Cavalcanti, R. Lima, S. J. Simske, and L. Favaro, “Assessing sentence scoring techniques for extractive text summarization,” Expert systems with applications, vol. 40, no. 14, p. 5755–5764, 2013. [8] Q. Al-Radaideh and M. Afif, “Arabic text summarization using aggregate similarity,” in International Arab conference on information technology (ACIT2009), Yemen, 2009. [9] A. Haboush, M. Al-Zoubi, A. Momani, and M. Tarazi, “Arabic text summarization model using clustering techniques,” World of Computer Science and Information Technology Journal (WCSIT) ISSN, p. 2221– 0741, 2012. [10] F. El-Ghannam and T. El-Shishtawy, “Multi-topic multi-document sum- marizer,” arXiv preprint arXiv:1401.0640, 2014. [11] H. N. Fejer and N. Omar, “Automatic arabic text summarization using clustering and keyphrase extraction,” in Proceedings of the 6th Interna- tional Conference on Information Technology and Multimedia, p. 293– 298, IEEE, 2014. [12] N. M. Hewahi and K. A. Kwaik, “Automatic arabic text summarization system (aatss) based on semantic features extraction,” International Journal of Technology Diffusion (IJTD), vol. 3, no. 2, p. 12–27, 2012. [13] M. El-Haj, U. Kruschwitz, and C. Fox, “Multi-document arabic text summarisation,” in 2011 3rd Computer Science and Electronic Engi- neering Conference (CEEC), p. 40–44, IEEE, 2011. [14] D. Miller, “Leveraging bert for extractive text summarization on lec- tures,” arXiv preprint arXiv:1906.04165, 2019. [15] M. A. Fattah and F. Ren, “Ga, mr, ffnn, pnn and gmm based models for automatic text summarization,” Computer Speech & Language, vol. 23, no. 1, p. 126–144, 2009. [16] R. Belkebir and A. Guessoum, “A supervised approach to arabic text summarization using adaboost,” in New contributions in information systems and technologies, p. 227–236, Springer, 2015. [17] Q. A. Al-Radaideh and D. Q. Bataineh, “A hybrid approach for arabic text summarization using domain knowledge and genetic algorithms,” Cognitive Computation, vol. 10, no. 4, p. 651–669, 2018. [18] A. Nenkova and K. McKeown, “A survey of text summarization tech- niques,” in Mining text data, p. 43–76, Springer, 2012. [19] D. Bahdanau, K. Cho, and Y. Bengio, “Neural machine translation by jointly learning to align and translate,” arXiv preprint arXiv:1409.0473, 2014. [20] I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in Advances in neural information processing systems, p. 3104–3112, 2014. [21] A. M. Rush, S. Chopra, and J. Weston, “A neural attention model for abstractive sentence summarization,” arXiv preprint arXiv:1509.00685, 2015. [22] R. Nallapati, B. Zhou, and M. Ma, “Classify or select: Neural ar- chitectures for extractive document summarization,” arXiv preprint arXiv:1611.04244, 2016. [23] S. Chopra, M. Auli, and A. M. Rush, “Abstractive sentence summa- rization with attentive recurrent neural networks,” in Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p. 93–98, 2016. [24] M. Al-Maleh and S. Desouki, “Arabic text summarization using deep learning approach,” Journal of Big Data, vol. 7, no. 1, p. 1–17, 2020. [25] W. Antoun, F. Baly, and H. Hajj, “Arabert: Transformer-based model for arabic language understanding,” arXiv preprint arXiv:2003.00104, 2020. [26] A. M. Abu Nada, E. Alajrami, A. A. Al-Saqqa, and S. S. Abu-Naser, “Arabic text summarization using arabert model using extractive text summarization approach,” 2020. [27] Y. Liu, “Fine-tune bert for extractive summarization,” arXiv preprint arXiv:1903.10318, 2019. [28] G. Inoue, B. Alhafni, N. Baimukan, H. Bouamor, and N. Habash, “The interplay of variant, size, and task type in arabic pre-trained language models,” 2021. [29] N. Zalmout and N. Habash, “Don’t throw those morphological analyzers away just yet: Neural morphological disambiguation for arabic,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, p. 704–713, 2017. [30] A. Reda, N. Salah, J. Adel, M. Ehab, I. Ahmed, M. Magdy, G. Khoriba, and E. H. Mohamed, “A hybrid arabic text summarization approach based on transformers,” in 2022 2nd International Mobile, Intelligent, and Ubiquitous Computing Conference (MIUCC), p. 56–62, 2022. [31] C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out, p. 74–81, 2004. [32] Y. Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V. Chaudhary, J. Gu, and A. Fan, “Multilingual translation with extensible multilingual pretraining and finetuning,” arXiv preprint arXiv:2008.00401, 2020. [33] Z. Barqawi, “Shakkala, arabic text vocalization,” 2017. [34] M. El-Haj and P. E. Rayson, “Osman: A novel arabic readability metric,” 2016.