Paper deep dive
AraDetox: A Multi-Dialect Arabic Detoxification Dataset
Mo El-Haj
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxification dataset comprising 10,500 harmful social-media posts and 84,000 detoxified rewrites generated using GPT-5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. The generated outputs were assessed through human evaluation and automatic analyses of lexical change, semantic preservation, sentiment, and dialectal style. Results show that detoxification is primarily a meaning-preserving rewriting task: substantial lexical and structural reformulation is accompanied by consistently high semantic similarity. Human evaluation confirms successful harmful-language removal while largely preserving the original meaning. Dialectal analyses further indicate that the generated variants exhibit measurable stylistic alignment with reference Arabic dialect corpora. Comparison with existing resources highlights two complementary approaches to detoxification: minimal-edit lexical substitution and meaning-preserving reformulation. Our findings demonstrate that large-scale Arabic detoxification resources can be constructed through LLM-assisted generation and human verification. The dataset is publicly available at this https URL to support future research on Arabic detoxification, safe text generation, and multi-dialect Arabic NLP.
Tags
Links
- Source: https://arxiv.org/abs/2608.22894v1
- Canonical: https://arxiv.org/abs/2608.22894v1
Trouble viewing inline? Open PDF directly →
Full Text
61,239 characters extracted from source content.
Expand or collapse full text
AraDetox: A Multi-Dialect Arabic Detoxification Dataset Mo El-Haj VinUniversity, Vietnam Lancaster University, UK https://elhaj.uk Abstract Arabic harmful-language detection has received considerable attention, yet Arabic text detoxifi- cation remains underexplored. We introduce AraDetox, a multi-dialect Arabic detoxifica- tion dataset comprising 10,500 harmful social- media posts and 84,000 detoxified rewrites gen- erated using GPT-5 and Gemini 2.5 Flash across Modern Standard Arabic, Gulf, Levantine, and Egyptian Arabic. The generated outputs were assessed through human evaluation and au- tomatic analyses of lexical change, semantic preservation, sentiment, and dialectal style. Re- sults show that detoxification is primarily a meaning-preserving rewriting task: substantial lexical and structural reformulation is accom- panied by consistently high semantic similarity. Human evaluation confirms successful harmful- language removal while largely preserving the original meaning. Dialectal analyses further indicate that the generated variants exhibit mea- surable stylistic alignment with reference Ara- bic dialect corpora. Comparison with existing resources highlights two complementary ap- proaches to detoxification: minimal-edit lex- ical substitution and meaning-preserving re- formulation. Our findings demonstrate that large-scale Arabic detoxification resources can be constructed through LLM-assisted genera- tion and human verification. The dataset is publicly available at https://github.com/ ArabicNLP-UK/AraDetox to support future research on Arabic detoxification, safe text gen- eration, and multi-dialect Arabic NLP. 1 Introduction Research on harmful language in Arabic has largely focused on detection, resulting in numerous datasets and models for hate speech, offensive language, and toxicity classification (Al Mandhari et al., 2024). Comparatively little attention has been paid to text detoxification, where harmful content is rewritten into a safer form while preserving its meaning. Cre- ating detoxification datasets is particularly challeng- ing because it requires annotators to rewrite offen- sive, politically sensitive, and often dialectal con- tent while preserving meaning, stance, target, and communicative intent. As a result, manual dataset creation is both costly and time-consuming ( Lo- gacheva et al. , 2022; Dementieva et al. , 2024). Re- cent advances in large language models (LLMs) pro- vide a promising alternative, enabling large-scale generation of fluent rewrites that remove harmful language while retaining the underlying message. Unlike harmful-language detection, detoxification requires balancing two competing objectives: re- moving harmful expressions while preserving the original claim, stance, target, and communicative intent. This challenge is particularly pronounced in Arabic social-media discourse, where criticism, sar- casm, political disagreement, and identity-related language are often intertwined with offensive ex- pressions. As a result, successful detoxification cannot be reduced to simple lexical substitution, but instead requires broader meaning-preserving reformulation. We introduce AraDetox, a large-scale multi- dialect Arabic detoxification dataset containing 10,500 harmful social-media posts and 84,000 detoxified rewrites generated using GPT-5 and Gemini 2.5 Flash across four Arabic varieties: Mod- ern Standard Arabic (MSA), Gulf Arabic, Levan- tine Arabic, and Egyptian Arabic. The dataset was created using an LLM-assisted synthesis pipeline followed by human quality control by native Ara- bic speakers. We evaluate AraDetox using lex- ical change, semantic similarity, sentiment anal- ysis, dialectal style analysis, and human assess- ment. Our findings show that Arabic detoxifica- tion is best characterised as a meaning-preserving rewriting task. Across models and dialects, sub- stantial lexical and structural reformulation is ac- companied by consistently high semantic similarity. Human evaluation further confirms that harmful language can be removed while largely preserving the original meaning. Beyond introducing a new re- source, our work demonstrates that existing Arabic harmful-language datasets can be transformed into large-scale multi-dialect detoxification resources through LLM-assisted generation and human veri- fication. The dataset is publicly available through the AraDetox repository 1 . The main contributions of this work are threefold. First, we introduce a large-scale Arabic detoxification resource contain- ing eight rewrites for each source post across two LLM families and four Arabic varieties. Second, we provide a multi-perspective evaluation cover- ing lexical reformulation, semantic preservation, sentiment, dialectal style, and independent human assessment. Third, we release the dataset to support reproducible research on Arabic detoxification and dialect-aware safe generation. 2 Related Work Research on harmful language in Arabic has fo- cused primarily on detection, resulting in datasets and models for hate speech, offensive language, and abusive-language classification (Mubarak et al., 2021, 2022; Alghamdi et al. , 2024), as well as Ara- bic language models such as AraBERT (Antoun et al., 2020) and ARBERT/MARBERT (Abdul- Mageed et al. , 2021). Comparatively little attention has been paid to text detoxification, which aims to rewrite harmful content into safer forms while pre- serving meaning. ParaDetox introduced a parallel- data setting based on toxic and non-toxic paraphrase pairs (Logacheva et al., 2022), while subsequent work extended detoxification to multilingual set- tings through MultiParaDetox (Dementieva et al., 2024, 2025). These resources established detoxifi- cation as a generation task distinct from harmful- language detection, but existing Arabic resources remain relatively limited in scale and provide only a single detoxified rewrite per source text. For Arabic, the original MultiParaDetox dataset was derived from approximately 2,100 candidate posts and re- sulted in 1,181 manually annotated source–detoxi- fied pairs after filtering and quality control. For our empirical comparison, we use the publicly avail- able Arabic evaluation split released through Hug- ging Face 2 , which contains 400 source–detoxified pairs. Recent work has also shown that LLMs can generate high-quality synthetic datasets and annota- 1 https://github.com/ArabicNLP-UK/AraDetox 2 https://huggingface.co/datasets/textdetox/ multilingual_paradetox tions, making them increasingly useful for data cre- ation and augmentation (Gilardi et al. , 2023; Long et al. , 2024; Nadăș et al. , 2025; Yong et al. , 2024). AraDetox builds on these developments by provid- ing detoxified rewrites for Arabic social-media con- tent across Modern Standard Arabic, Gulf, Levan- tine, and Egyptian Arabic using GPT-5 and Gem- ini. Unlike existing Arabic detoxification resources, AraDetox emphasises semantic preservation while retaining the original intent, stance, and target of the source text. AraDetox substantially expands the scale of Arabic detoxification resources and, to the best of our knowledge, is the first Arabic detoxifi- cation dataset to provide rewrites generated across multiple dialects and multiple LLMs within a single framework. 3 Methodology 3.1 Dataset Construction AraDetox is derived from the Arabic Hate Speech Superset (Tonneau et al. , 2024), which inte- grates ten Arabic hate-speech and abusive-language datasets: T-HSAB (Haddad et al. , 2019), JHSC (Ah- mad et al. , 2024), MLMA (Ousidhoum et al. , 2019), L-HSAB (Mulki et al., 2019), OSACT (Mubarak et al., 2020), Let-Mi (Mulki and Ghanem, 2021), Saudi Tweets (Alshaalan and Al-Khalifa, 2020), AraCOVID19-MFH (Ameur and Aliane, 2021), Brothers (Albadi et al., 2018), and Alsafari et al. (Alsafari et al., 2020). From the 16,000 harmful posts available in the superset, we selected 10,500 posts while retaining coverage across all ten con- stituent datasets. No class balancing was applied, and the source datasets were not equally represented because they varied substantially in size. The se- lection was intended to provide broad coverage of the available harmful-language data rather than a balanced classification sample. Table 1 summarises the resulting resource. Using GPT-5 and Gemini 2.5 Flash, we gener- ated detoxified rewrites in MSA, Gulf, Levantine, and Egyptian Arabic. Models were instructed to preserve the original meaning, target, stance, sen- timent, and communicative intent while removing profanity, insults, slurs, dehumanising language, discriminatory expressions, and threats. For the di- alectal variants, outputs were additionally required to adapt the source text into Gulf, Levantine, or Egyptian Arabic. These outputs therefore combine two transformations: harmful-language detoxifica- tion and adaptation to the requested Arabic variety. Statistic Value Source posts 10,500 Detoxified outputs 84,000 Arabic varieties MSA, Gulf, Levantine, fied variants were evaluated by three native Arabic- speaking annotators, yielding 2,400 generated out- puts and 7,200 annotation records. The annotators involved in this evaluation were distinct from the Generation models Egyptian GPT-5, Gemini 2.5 Flash native Arabic speaker who performed quality con- Outputs per source post 8 Outputs generated by GPT-5 42,000 Outputs generated by Gemini 42,000 Total words 1,318,257 Average words per output 15.69 Median words per output 13 Human-evaluation sample 300 source posts Human-evaluated outputs 2,400 Annotators 3 Annotation records 7,200 Table 1: Summary statistics for AraDetox. The MSA outputs primarily represent detoxifica- tion within the same language variety, whereas the Gulf, Levantine, and Egyptian outputs jointly in- volve detoxification and dialect adaptation. Conse- quently, lexical and structural differences between the source posts and the dialectal outputs cannot be attributed exclusively to detoxification. We there- fore evaluate harmful-language removal, meaning preservation, and dialectal style as separate dimen- sions and interpret the dialectal results as evidence of joint detoxification and dialect adaptation. This produced eight detoxified variants per source post, yielding 84,000 detoxified instances and more than 1.3 million words. The full prompts are provided in Appendix A. Dataset construction incorporated au- tomatic validation, regeneration of invalid outputs, and human quality control. Automatic checks iden- tified missing fields, malformed JSON, null values, duplicated outputs, incomplete generations, and source–output alignment errors. Safety refusals, generic moderation responses, summaries, external- observer descriptions, and outputs that substan- tially altered or added to the source meaning were treated as invalid and regenerated. A native Arabic- speaking reviewer subsequently verified that the regenerated outputs performed the requested detox- ification task, preserved the source meaning, target, stance, and communicative intent, and reflected the requested Arabic variety. Further details of the gen- eration procedure and quality-control workflow are provided in Appendix B. As an additional quality-assurance step, 300 source posts were randomly selected from the 10,500-post collection. The sample was stratified to preserve the distribution of the ten source datasets across the ten source datasets and available harm categories. For each selected post, all eight detoxi- trol during dataset construction. Appendix E re- ports the distribution of the full dataset and the human-evaluation sample across the constituent source datasets. 3.2 Evaluation Design We evaluate AraDetox from five complementary perspectives: lexical change, semantic preservation, sentiment shift, dialectal style, and human eval- uation. We additionally compare AraDetox with MultiParaDetox-Ar (Dementieva et al., 2024, 2025) to contextualise its detoxification behaviour relative to an existing Arabic detoxification resource. A key challenge in detoxification is preserving meaning while removing harmful content. Offensive expres- sions often convey stance, criticism, and emotion rather than functioning as removable noise. For example, the insult خنزیر یا الكلب انت لا (“No, you are the dog, you pig”) can be rewritten as إنك بل لا، are you rather, (“No, أن ت من یتصرف بطر یقة غیر مقبولة the one behaving in an unacceptable manner”), pre- serving the criticism while removing the offensive language. Appendix F presents additional quali- tative examples from AraDetox. F1 score is not used as a generation-quality metric because detoxi- fication does not have a single correct reference sequence from which token-level true positives, false positives, and false negatives can be meaning- fully derived. Multiple lexically different rewrites may all be valid. We therefore combine surface- change measures, semantic similarity, sentiment analysis, dialectal-style analysis, and human judge- ments, with each evaluation component measuring a distinct requirement of successful detoxification. 3.2.1 Lexical Change We measure lexical change between each source post and its detoxified counterpart using six surface- level metrics: lexical overlap, Jaccard similarity, character-level edit distance, percentage of changed characters, word-level edit distance, and output length. Lexical overlap and Jaccard similarity mea- sure how much of the original vocabulary is re- tained, with higher values indicating more conser- vative rewriting. Character- and word-level edit distances quantify the amount of rewriting required to transform the source post into the detoxified out- put, while the percentage of changed characters normalises this change by source length. Output length is included to capture whether detoxifica- tion tends to shorten, preserve, or expand the orig- inal text. These measures allow us to distinguish minimal-edit detoxification from broader sentence- level reformulation. 3.2.2 Semantic Preservation To quantify semantic fidelity, we compute cosine similarity between each source post and its detoxi- fied counterpart using multilingual-e5-large (Wang et al., 2024). This embedding-based mea- sure assesses whether the detoxified text remains se- mantically close to the original despite lexical and structural changes. To complement the pairwise similarity analysis, we examine the global organ- isation of the embedding space using UMAP. By visualising source posts, detoxified outputs, and ex- ternal neutral Arabic corpora in a shared semantic space, we assess whether detoxified texts remain close to their source content while exhibiting charac- teristics associated with naturally occurring neutral language. 3.2.3 Sentiment and Neutrality Analysis Detoxification should remove harmful language without eliminating criticism, disagreement, or neg- ative opinions. We therefore examine whether detoxification systematically shifts sentiment by in- creasing neutral or positive predictions at the ex- pense of negative ones. This follows the broader use of sentiment analysis for examining evalua- tive language and affect across different domains and languages, including financial discourse, so- cial media, and multilingual settings (El-Haj et al. , 2016; Alwakid et al., 2022; Hunter et al., 2023; Huy et al. , 2026). Sentiment distributions are com- puted using three CAMeLBERT sentiment classi- fiers: CAMeLBERT-DA, CAMeLBERT-MSA, and CAMeLBERT-Mix (Inoue et al., 2021). Changes in positive, neutral, and negative predictions are analysed before and after detoxification to assess whether rewriting reduces hostile affect while pre- serving the original stance. 3.2.4 Dialectal Adaptation AraDetox includes detoxified variants in Gulf, Lev- antine, and Egyptian Arabic. To examine whether these rewrites exhibit stylistic alignment with their intended varieties, we compare them against refer- ence dialect corpora ( Zaidan and Callison-Burch, 2014) using TF–IDF cosine similarity over word and character n-grams. This analysis provides corpus-level evidence of lexical and orthographic alignment, but it is sensitive to topic, register, cor- pus composition, and vocabulary shared across Ara- bic varieties. It is therefore treated as an exploratory measure of dialectal style rather than definitive val- idation of dialect authenticity. For example, the MSA phrase الشارع أغلقوا (“they blocked the street”) may be realised as الشارع سادین in Gulf Arabic, سكروا الشارع in Levantine Arabic, and الشارع قفلوا in Egyp-tian Arabic, illustrating meaning-preserving adap- tation across Arabic varieties. 3.2.5 Human Evaluation Protocol To assess the quality of the generated outputs, three native Arabic-speaking annotators independently reviewed a random sample of 300 source posts. For each post, all eight detoxified variants were evalu- ated, yielding 2,400 generated outputs and 7,200 annotation records. Each output was assessed for of- fence removal, meaning preservation, and introduc- tion of new content. Offence removal was annotated using a three-way scheme: offence not removed, of- fence removed, or not applicable when the source text did not contain offensive content. Meaning preservation and new content were annotated using binary judgements. The full annotation guidelines are provided in Appendix D. Dialect authenticity was not included as a separate human-evaluation criterion. The dialectal findings therefore rely on corpus-level stylistic analysis and should not be in- terpreted as direct human validation of native-like dialect use. 3.2.6 Comparison with Existing Resources To contextualise AraDetox, we compare it descrip- tively with MultiParaDetox-Ar using the same lexi- cal, semantic, and sentiment-based measures. The resources differ in source-domain distribution, size, annotation procedure, and the initial sentiment and severity of their source texts. The comparison is therefore intended to characterise their rewriting behaviour rather than establish that one resource or construction method is superior. It also contrasts the relatively minimal-edit strategy represented by MultiParaDetox-Ar with the broader meaning- preserving reformulation and dialect adaptation rep- resented by AraDetox. 4 Results 4.1 Lexical Change Table 2 summarises the lexical transformations in- troduced by all eight detoxification systems. Over- lap and Jaccard measure vocabulary retention be- tween the original and detoxified texts, while CharEdit, Char%, WordEdit, and Word% capture character- and word-level modifications. Across all variants, detoxification results in substantial modifi- cation of the original text, indicating that the gener- ated outputs frequently involve sentence-level refor- mulation rather than simple replacement of offen- sive expressions. Among all systems, GPT-MSA is the most conservative. It achieves the lowest character-level edit distance (42.19), the lowest pro- portion of modified characters (61.58%), the lowest word-level edit distance (11.77), the lowest word- level modification rate (94.15%), and produces the shortest outputs (11.28 words on average). GPT- MSA therefore preserves more of the original sur- face form than the other systems. In contrast, Gemini-MSA performs more exten- sive rewriting than GPT-MSA, exhibiting larger character- and word-level modifications, higher pro- portions of changed characters, and longer out- puts. The dialectal variants generally involve even greater reformulation. Across both models, output lengths increase from an average of 13.21 words in the original posts to 15.66–16.79 words, while character-level modification rates exceed 78%. Nor- malised word-level edit rates range from 115.79% to 122.94%. Values above 100% occur because word-level edit distance is normalised by the orig- inal text length and includes insertions, deletions, and substitutions, allowing the number of edit oper- ations to exceed the number of words in the source text. GPT-Gulf exhibits the highest lexical overlap (0.2942) and Jaccard similarity (0.1719), indicating greater retention of source vocabulary. In contrast, Gemini-MSA and Gemini-Egyptian show the low- est overlap scores, suggesting a greater degree of lexical reformulation. The lexical analyses suggest that Arabic detoxification is best characterised as a meaning-preserving rewriting task rather than sim- ple lexical substitution. Whether these extensive modifications preserve the underlying message is examined through the semantic and human evalua- tions that follow. 4.2 Semantic Preservation Table 3 reports semantic similarity between the orig- inal posts and the AraDetox variants using multi- lingual E5 embeddings. Across all variants, seman- tic similarity remains consistently high, with mean scores ranging from 0.9234 to 0.9308 and all exam- ples exceeding the 0.70 similarity threshold. These results suggest that the generated rewrites preserve the core meaning of the source posts despite the substantial lexical modifications reported in Sec- tion 4.1. The dialect-aware variants achieve slightly higher similarity scores than the MSA variants, with Levantine obtaining the highest mean similarity (0.9308), followed by Egyptian and Gulf Arabic. GPT-5 and Gemini also exhibit comparable levels of semantic preservation, with mean similarities of 0.9234 and 0.9249, respectively. The relatively small differences across variants suggest that se- mantic fidelity is largely unaffected by the choice of dialect or generation model. The lexical and semantic results show that sub- stantial reformulation does not necessarily come at the expense of semantic fidelity. The generated outputs often differ considerably from the source posts at the surface level, while remaining close in meaning. This supports treating detoxification as a meaning-preserving rewriting task rather than simple lexical substitution. Figures 1 and 2 visualise the embedding space using UMAP. Across all variants, substantial over- lap is observed between the original posts and their detoxified counterparts, with no clear separation be- tween source texts and generated rewrites. Similar patterns are observed for the dialect-aware variants, whose outputs occupy largely the same regions of the embedding space as the original posts. These visualisations are consistent with the high semantic- similarity scores reported in Table 3 and provide additional evidence that detoxification and dialect adaptation preserve the semantic structure of the source content. 4.3 Sentiment Shift We investigate whether detoxification influences negative affect using three Arabic sentiment mod- els: CAMeLBERT-DA, CAMeLBERT-MSA, and CAMeLBERT-Mix. Sentiment is not treated as a direct measure of toxicity, since a post may remain critical, negative, or emotionally charged without being abusive, hateful, or discriminatory. Rather, System Overlap Jaccard CharEdit Char% WordEdit Word% Len GPT-MSA 0.2062 0.1339 42.19 61.58 11.77 94.15 11.28 GPT-Gulf 0.2942 0.1719 45.52 81.24 13.08 115.79 15.99 GPT-Levantine 0.2737 0.1610 45.90 80.66 13.24 116.14 15.66 GPT-Egyptian 0.2537 0.1454 48.07 84.72 13.67 120.51 16.04 Gemini-MSA 0.2007 0.1201 52.74 83.61 14.59 119.38 16.76 Gemini-Gulf 0.2342 0.1353 49.90 83.90 14.39 122.94 16.79 Gemini-Levantine 0.2440 0.1471 47.69 78.88 13.94 117.08 16.27 Gemini-Egyptian 0.2188 0.1263 49.93 81.26 14.50 120.85 16.75 Table 2: Lexical change statistics. Variant Mean Median Min. >0.90 Levantine 0.9308 0.9349 0.7883 80.31% Egyptian 0.9274 0.9307 0.7789 77.93% Gulf 0.9256 0.9293 0.7700 76.15% Gemini 0.9249 0.9278 0.7691 73.30% GPT-5 0.9234 0.9278 0.7504 71.05% Table 3: Semantic similarity using multilingual E5 embed- dings. Embedding-space visualisation of original and dialectal detoxified outputs Original 10 GPT-Gulf GPT-Levantine GPT-Egyptian Gemini-Gulf Gemini-Levantine 5 Gemini-Egyptian 0 5 Embedding-space visualisation of original and MSA detoxified outputs 8 7 6 5 4 3 2 1 UMAP-1 10 7.5 5.0 2.5 0.0 2.5 5.0 7.5 10.0 UMAP-1 Figure 2: UMAP visualisation of original posts and GPT/Gemini dialectal detoxified variants using multilingual E5 embeddings. Model Orig. GPT-MSA GPT-Gulf GPT-Lev. GPT-Egy. DA 74.29 53.77 70.32 71.14 69.38 MSA 75.47 53.48 71.23 73.67 70.44 Mix 72.41 49.20 68.62 69.56 66.36 Model Orig. Gem-MSA Gem-Gulf Gem-Lev. Gem-Egy. DA 74.29 55.79 68.07 70.13 70.07 Figure 1: UMAP visualisation of original posts, GPT-MSA detoxifications, and Gemini-MSA detoxifications using multi- lingual E5 embeddings. sentiment provides an additional perspective on how detoxification alters the tone of the text while preserving its meaning and intent. Table 4 and Figure 3 report the percentage of posts classified as negative before and after detoxifi- cation. The original posts are classified as negative at high rates across all three sentiment models, rang- ing from 72.41% to 75.47%. After detoxification, the strongest reduction is observed for the MSA variants. GPT-MSA reduces negative sentiment to 49.20%–53.77%, while Gemini-MSA reduces it to 53.19%–60.26%. MSA detoxification substantially softens the affective tone of harmful posts. The dialectal variants also reduce negative senti- ment, but to a more limited extent. Across both GPT and Gemini, Gulf, Levantine, and Egyptian outputs remain closer to the original sentiment distribution, MSA 75.47 60.26 69.69 72.24 70.21 Mix 72.41 53.19 65.45 67.42 65.18 Table 4: Negative sentiment before and after detoxification using CAMeLBERT-DA, MSA and Mix models. with negative rates generally between 65% and 74%. Dialectal rewriting preserves more of the original emotional tone, even after harmful expressions have been removed. The pattern is also visible in Fig- ure 4, where the MSA variants show the clearest drop in average negative sentiment, while the di- alectal variants remain nearer to the original posts. These results indicate that part of the negative af- fect detected by sentiment models is associated with abusive wording, insults, profanity, and other harm- ful linguistic constructions. Once such language is removed, the resulting texts are generally perceived as less negative, particularly in MSA. At the same time, the dialectal variants show that detoxification does not necessarily eliminate criticism, disagree- ment, or emotional stance. Rather, the generated Original MSA 2 0 2 4 6 UMAP - 2 UMAP - 2 70 60 Negative sentiment before and after detoxification GPT-MSA Word n-gram dialect-style similarity 0.40 50 GPT-Gulf 40 30 GPT-Levantine 20 0.35 10 GPT-Egyptian 0 0.30 Variant Figure 3: Negative sentiment across GPT and Gemini variants using three CAMeLBERT models. outputs preserve much of the communicative intent while reducing harmful expression. Average negative sentiment across CAMeLBERT models Gemini-MSA Gemini-Gulf Gemini-Levantine Gemini-Egyptian Reference dialect corpus 0.25 0.20 60 Figure 5: Word n-gram similarity to reference dialect corpora. 50 40 Character n-gram dialect-style similarity 30 GPT-MSA 20 10 GPT-Gulf 0 GPT-Levantine 0.76 0.74 0.72 Variant Figure 4: Average negative-sentiment rate across the three CAMeLBERT models. 4.4 Dialectal Style Analysis AraDetox includes detoxified rewrites in MSA, Gulf, Levantine, and Egyptian Arabic. To examine GPT-Egyptian Gemini-MSA Gemini-Gulf Gemini-Levantine Gemini-Egyptian Reference dialect corpus 0.70 0.68 0.66 0.64 0.62 0.60 corpus-level dialectal style, we compare the gener- ated variants with reference texts from the Arabic- Dialects corpus using TF–IDF cosine similarity over word and character n-grams. Figures 5 and 6 show the resulting similarity ma- trices. The word n-gram results provide the clearest evidence of dialectal alignment. Both MSA systems align most strongly with the MSA reference cor- pus, with Gemini-MSA achieving the highest MSA similarity (0.373), followed by GPT-MSA (0.350). The Egyptian variants likewise show strong align- ment with the Egyptian reference corpus, with GPT- Egyptian obtaining the highest similarity (0.406), followed by Gemini-Egyptian (0.383). The Gulf and Levantine results are less distinct. Gemini-Gulf aligns most strongly with the Gulf reference corpus (0.285), and Gemini-Levantine aligns most strongly with the Levantine reference corpus (0.297), suggesting clearer dialect-specific stylistic patterns. In contrast, GPT-Gulf and GPT- Levantine show their highest similarity with the Egyptian reference corpus, likely reflecting lexical overlap among informal Arabic varieties and the strong signal of the Egyptian reference set. The character n- gram analysis broadly confirms the Figure 6: Character n-gram similarity to reference dialect corpora. MSA and Egyptian findings but provides weaker separation between Gulf and Levantine. Character- level features capture orthographic and morphologi- cal properties shared across Arabic dialects, making them less discriminative than word n-grams. The results provide measurable evidence of corpus-level dialectal style, particularly for MSA and Egyptian Arabic. Gemini also shows stronger intended-variety alignment for Gulf and Levantine than GPT-5. However, the reference corpora con- sist largely of general social-media discussions and forum-style interactions, whereas AraDetox is de- rived from harmful, highly negative, politically sen- sitive, and identity-related discourse. Topic, regis- ter, and corpus composition may therefore influence the reported similarities. The weaker distinction between Gulf and Levantine further reflects the sub- stantial lexical and orthographic overlap between informal Arabic varieties. These results should consequently be interpreted as evidence of stylistic alignment rather than definitive evidence of native- like dialect authenticity. 7 CAMeLBERT-DA CAMeLBERT-MSA CAMeLBERT-Mix 0.350 0.327 0.310 0.223 0.199 0.261 0.257 0.288 0.164 0.214 0.268 0.275 0.169 0.217 0.227 0.406 0.373 0.335 0.314 0.224 0.219 0.285 0.271 0.247 0.180 0.237 0.297 0.257 0.180 0.230 0.245 0.383 0.714 0.712 0.705 0.656 0.634 0.683 0.687 0.715 0.601 0.651 0.686 0.709 0.598 0.644 0.661 0.776 0.708 0.703 0.695 0.652 0.640 0.686 0.687 0.686 0.608 0.658 0.693 0.693 0.600 0.645 0.665 0.757 Negative sentiment (%) Average negative sentiment (%) Generated variant Generated variant Cosine similarity Cosine similarity 4.5 Human Evaluation Automatic metrics provide useful insights into lex- ical change, semantic similarity, and sentiment shifts, but they cannot directly determine whether harmful content has been removed while preserv- ing the intended meaning of the original post. To address this limitation, we conducted a human eval- uation following the protocol described in Sec- tion 3.2.5. Three native Arabic-speaking annotators independently evaluated all outputs in the evalua- tion sample. The annotators represented Saudi Gulf Arabic, Jordanian Levantine Arabic, and Egyptian Arabic and were regular users of Arabic social me- dia. They were independent of the LLM-generation process and distinct from the native Arabic speaker who performed quality control during dataset con- struction. The evaluation focused on three criteria central to detoxification: offence removal, meaning preserva- tion, and introduction of new content. These dimen- sions assess whether generated rewrites success- fully remove harmful language, retain the original communicative intent, and avoid introducing infor- mation not present in the source text. To minimise potential bias, annotators were not informed that the texts had been generated by LLMs. Inter-annotator agreement was substantial across all evaluation dimensions. Fleiss’ κ values ranged from 0.715 to 0.795, indicating consistent judge- ments among annotators and supporting the relia- bility of the evaluation results (Table 5). Criterion Agreement (%) Fleiss’ κ Offence Removal 96.44 0.795 Meaning Preservation 95.30 0.715 New Content 92.00 0.716 Table 5: Inter-annotator agreement for human evalua- tion. Variant Offence Meaning New Gemini MSA 92.90 92.57 12.42 Gemini Gulf 92.35 91.46 15.52 Gemini Levantine 92.24 91.69 13.75 Gemini Egyptian 91.91 89.36 17.52 GPT MSA 87.25 86.25 13.08 GPT Gulf 87.25 91.57 24.83 GPT Levantine 87.47 91.80 24.28 GPT Egyptian 87.58 91.69 24.17 Table 6: Human evaluation results (%). Higher is better except for New Content. Table 6 shows strong performance across all three evaluation dimensions. Offence-removal rates range from 87.25% to 92.90%, showing that harm- ful language can generally be removed without sub- stantially altering the underlying message. Gemini achieves the highest offence-removal scores across all language varieties. Meaning-preservation rates are also high, rang- ing from 86.25% to 92.57%. The generated rewrites therefore largely retain the original claims, tar- gets, stances, and communicative intent of the source posts despite modifications introduced dur- ing detoxification. The clearest differences between the model families emerge in the introduction of new con- tent. Gemini introduces additional information in 12.42%–17.52% of cases, whereas GPT’s dialec- tal variants do so in 24.17%–24.83% of outputs. Although GPT tends to produce more expansive rewrites, its meaning-preservation scores remain comparable to those of Gemini, suggesting that the additional content generally supplements rather than alters the core meaning of the source text. The human evaluation findings are consistent with the automatic analyses. Both model families are effective at removing harmful language while preserving the intended meaning of the original posts. The results also point to different rewriting strategies: Gemini tends to produce more conserva- tive reformulations, whereas GPT more frequently elaborates on the source content while maintaining similar levels of meaning preservation. 4.6 Comparison with Existing Arabic Detoxification Resources To better understand the characteristics of AraDetox, we compare it descriptively with MultiParaDetox-Ar (Dementieva et al., 2024, 2025), one of the few publicly available Arabic detoxification resources. The comparison covers lexical change, semantic similarity, and sentiment shift. Because the two resources differ substantially in source distribution, size, annotation procedure, and source-text negativity, the results should be in- terpreted as a comparison of dataset characteristics rather than as a controlled benchmark. Table 7 reveals clear differences in rewriting be- haviour. MultiParaDetox-Ar follows a relatively conservative editing strategy, exhibiting higher word overlap (0.559), lower character-level modifi- cation (32.96%), and lower word-level modification (46.72%). In contrast, the AraDetox variants ex- hibit substantially greater reformulation, with word- overlap scores ranging from 0.201 to 0.294 and con- siderably higher character- and word-level modifi- cation rates than those observed in MultiParaDetox- Ar. Despite these larger lexical and structural changes, semantic similarity remains consistently high across all AraDetox variants. MultiParaDetox- Ar achieves the highest E5 similarity (0.957), consistent with its minimal-edit approach. The AraDetox variants obtain E5 similarities between 0.919 and 0.930, indicating that extensive rewriting can occur while remaining semantically close to the original content. Among the AraDetox variants, the dialect-aware rewrites achieve the strongest seman- tic preservation, with Gulf and Levantine Arabic producing the highest similarity scores. These find- ings suggest that the two resources operationalise detoxification through different rewriting strategies. MultiParaDetox-Ar primarily reflects minimal-edit detoxification, characterised by conservative lexical modifications and high surface-level similarity to the source text. In contrast, AraDetox emphasises meaning-preserving reformulation, retaining the original claim, target, and stance while permitting substantially broader lexical and structural rewrit- ing. This distinction is particularly evident in the dialectal variants, where detoxification is combined with adaptation to dialect-specific linguistic con- ventions. Dataset Pairs Overlap Char. % Word % E5 MultiParaDetox-Ar 400 0.559 32.96 46.72 0.957 AraDetox GPT-MSA 10,500 0.206 61.58 94.15 0.924 AraDetox GPT-Gulf 10,500 0.294 81.24 115.79 0.930 AraDetox GPT-Lev. 10,500 0.274 80.66 116.14 0.930 AraDetox GPT-Egy. 10,500 0.254 84.72 120.51 0.927 AraDetox Gemini-MSA 10,500 0.201 83.61 119.38 0.919 AraDetox Gemini-Gulf 10,500 0.234 83.90 122.94 0.924 AraDetox Gemini-Lev. 10,500 0.244 78.88 117.08 0.928 AraDetox Gemini-Egy. 10,500 0.219 81.26 120.85 0.925 Table 7: Comparison of MultiParaDetox-Ar and AraDetox. We further compare the datasets using negative- sentiment predictions from three CAMeLBERT sentiment models. Table 8 should be interpreted with caution because the source datasets differ sub- stantially. Only around 10–11% of MultiParaDetox- Ar source texts are classified as negative, compared with 72–75% for AraDetox, reflecting the more strongly negative nature of the harmful-language corpora from which AraDetox was derived. Even with these differences, distinct patterns emerge. MultiParaDetox-Ar shows relatively small reductions in negative sentiment (1.22–2.39 per- centage points), whereas AraDetox MSA variants produce much larger reductions. GPT-MSA re- duces negative predictions by 20.51–23.21 percent- age points across the three sentiment models, while Gemini-MSA achieves reductions of 15.21–19.22 points. The dialectal variants show smaller reduc- tions, typically between 1.80 and 7.23 points. These results reinforce the distinction between detoxifi- cation and sentiment modification. AraDetox is not designed to convert negative opinions into pos- itive ones, but to express criticism, disagreement, or opposition without insults, slurs, threats, or abu- sive language. As a result, many detoxified outputs remain negative in sentiment while no longer con- taining harmful expressions. Model Dataset Orig. Detox Red. MultiParaDetox-Ar 9.92 7.53 2.39 GPT-MSA 74.29 53.77 20.51 GPT-Gulf 74.29 70.32 3.96 GPT-Lev. 74.29 71.14 3.14 DA GPT-Egy. 74.29 69.38 4.90 Gemini-MSA 74.29 55.79 18.50 Gemini-Gulf 74.29 68.07 6.22 Gemini-Lev. 74.29 70.13 4.15 Gemini-Egy. 74.29 70.07 4.22 MultiParaDetox-Ar 11.00 9.78 1.22 GPT-MSA 75.47 53.48 21.99 GPT-Gulf 75.47 71.23 4.24 GPT-Lev. 75.47 73.67 1.80 MSA GPT-Egy. 75.47 70.44 5.03 Gemini-MSA 75.47 60.26 15.21 Gemini-Gulf 75.47 69.69 5.78 Gemini-Lev. 75.47 72.24 3.23 Gemini-Egy. 75.47 70.21 5.26 MultiParaDetox-Ar 10.14 7.78 2.36 GPT-MSA 72.41 49.20 23.21 GPT-Gulf 72.41 68.62 3.79 GPT-Lev. 72.41 69.56 2.85 Mix GPT-Egy. 72.41 66.36 6.05 Gemini-MSA 72.41 53.19 19.22 Gemini-Gulf 72.41 65.45 6.96 Gemini-Lev. 72.41 67.42 4.99 Gemini-Egy. 72.41 65.18 7.23 Table 8: Negative sentiment before and after detoxification using CAMeLBERT-DA, MSA and Mix. 5 Conclusion We introduce AraDetox, a large-scale Arabic detox- ification dataset containing 10,500 harmful social- media posts and 84,000 detoxified rewrites across MSA, Gulf, Levantine, and Egyptian Arabic. The resource combines LLM-assisted generation with automatic validation, native-speaker quality con- trol, and independent human evaluation. The find- ings show that Arabic detoxification frequently re- quires substantial reformulation while preserving the source claim, target, stance, and communica- tive intent. Human evaluation indicates high rates of offence removal and meaning preservation, al- though some outputs introduce additional content. The dialectal variants also show measurable corpus- level alignment with their intended Arabic varieties. AraDetox provides a substantial resource for devel- oping and evaluating Arabic detoxification systems, while highlighting the need for continued work on dialect validation, pragmatic meaning preservation, and downstream evaluation. 6 Limitations This study has several limitations. First, AraDetox was constructed primarily through GPT-5 and Gem- ini 2.5 Flash rather than large-scale human rewrit- ing. Although automatic validation and native- speaker quality control were applied, the outputs may retain model-specific stylistic patterns and should not be treated as naturally occurring human rewrites. Second, detailed human evaluation covered 300 source posts, corresponding to 2,400 generated out- puts, and therefore represents only a subset of the full 84,000-output resource. Rare source types or harm categories may be underrepresented. Third, semantic-preservation analyses rely partly on multilingual embedding models. High cosine similarity does not guarantee preservation of prag- matic meaning, sarcasm, humour, presupposition, target, or subtle stance. Human judgements partly address this limitation but cannot eliminate it. Fourth, dialectal evaluation relies on corpus-level word and character n-gram similarity, which may be influenced by topic, register, corpus composi- tion, and shared vocabulary across Arabic varieties. As dialect authenticity was not separately evaluated by human annotators, these results should be inter- preted as evidence of stylistic alignment rather than native-like dialect use. Fifth, comparison with MultiParaDetox-Ar is de- scriptive because the resources differ in size, source distributions, and annotation procedures. We also do not yet assess whether models trained or fine- tuned on AraDetox outperform those trained on ex- isting Arabic detoxification datasets. A controlled downstream benchmark remains future work. Finally, the dataset was generated using two pro- prietary frontier model families. Evaluating smaller open-weight models would help determine sensitiv- ity to model scale, architecture, and training data. AraDetox is also derived from Arabic social-media content and may not generalise to long-form, news, or formal political discourse. The AraDetox dataset is publicly available at https://github.com/ArabicNLP-UK/ AraDetox to support transparency and repro- ducibility. 7 Ethical Considerations This work involves harmful Arabic social-media content, including offensive, abusive, political, sec- tarian, and identity-related language. The purpose of the study is to develop safer Arabic NLP re- sources and evaluation methods for neutralising harmful language. We minimise the reproduction of harmful examples in the paper and report results in aggregate form. The study included a human evaluation compo- nent involving up to ten native Arabic-speaking adult annotators. Ethical approval was granted by an Institutional Ethical Review Board for Biomed- ical Research in Vietnam. Annotators provided informed consent and were given detailed annota- tion guidance prior to participation. No personal, sensitive, or health-related data were collected, and all procedures were conducted in accordance with institutional policy, Vietnamese regulations, and internationally recognised ethical standards. The detoxified outputs should not be treated as au- thoritative or universally acceptable rewrites, since judgements of harmfulness and acceptable neutrali- sation are shaped by social, cultural, dialectal, and political context. We also recognise that automatic detoxification can introduce risks, including over- sanitising legitimate political expression, weaken- ing the speaker’s intended stance, or removing im- portant evidence of abuse. Any practical deploy- ment of Arabic detoxification systems should there- fore include human oversight and clear moderation guidelines. Outputs that produced refusals, malformed re- sponses, summaries, external-observer descrip- tions, or substantial meaning changes were treated as invalid and regenerated during dataset construc- tion. The released resource distinguishes the com- plete LLM-generated collection from the indepen- dently human-evaluated subset. Human-evaluation labels for offence removal, meaning preservation, and new-content introduction are provided for the evaluated outputs, allowing users to identify and filter unsuccessful or uncertain generations. The complete synthetic collection should not be treated as a gold-standard dataset of human-authored detox- ifications. References Muhammad Abdul-Mageed, AbdelRahim Elmadany, and El Moatez Billah Nagoudi. 2021. ARBERT & MARBERT: Deep bidirectional transformers for Ara- bic. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Lan- guage Processing (Volume 1: Long Papers), pages 7088–7105, Online. Association for Computational Linguistics. Ashraf Ahmad, Mohammad Azzeh, Eman Alnagi, Qasem Abu Al-Haija, Dana Halabi, Abdullah Aref, and Yousef AbuHour. 2024. Hate speech detection in the arabic language: corpus design, construction, and evaluation. Frontiers in Artificial Intelligence, 7:1345445. Salim Al Mandhari, Mo El -Haj, and Paul Rayson. 2024. Is it offensive or abusive? an empirical study of hate- ful language detection of arabic social media texts. In Proceedings of the First International Conference on Natural Language Processing and Artificial Intel- ligence for Cyber Security, pages 137–146. Nuha Albadi, Mohamed Kurdi, and Swapna Mishra. 2018. Are they our brothers? analysis and detection of religious hate speech in the arabic twittersphere. In Proceedings of AICS. Seham Alghamdi, Youcef Benkhedda, Basma Alharbi, and Riza Batista-Navarro. 2024. AraTar: A corpus to support the fine-grained detection of hate speech targets in the Arabic language. In Proceedings of the 6th Workshop on Open-Source Arabic Corpora and Processing Tools (OSACT) with Shared Tasks on Arabic LLMs Hallucination and Dialect to MSA Machine Translation @ LREC-COLING 2024, pages 1–12, Torino, Italia. ELRA and ICCL. Safa Alsafari, Samira Sadaoui, and Malek Mouhoub. 2020. Hate and offensive speech detection on arabic social media. Online Social Networks and Media, 19:100096. Raghad Alshaalan and Hend Al-Khalifa. 2020. Hate speech detection in saudi twittersphere: A deep learn- ing approach. In Proceedings of the fifth Arabic natu- ral language processing workshop, pages 12–23. Ghadah Alwakid, Taha Osman, Mahmoud El Haj, Saad Alanazi, Mamoona Humayun, and Najm Us Sama. 2022. Muldasa: Multifactor lexical sentiment anal- ysis of social-media content in nonstandard arabic social media. Applied Sciences, 12(8):3806. Mohamed Seghir Hadj Ameur and Hassina Aliane. 2021. Aracovid19-mfh: Arabic covid-19 multi-label fake news & hate speech detection dataset. Procedia Com- puter Science, 189:232–241. Wissam Antoun, Fady Baly, and Hazem Hajj. 2020. AraBERT: Transformer-based model for Arabic lan- guage understanding. In Proceedings of the 4th Work- shop on Open-Source Arabic Corpora and Process- ing Tools, with a Shared Task on Offensive Language Detection, pages 9–15, Marseille, France. European Language Resource Association. Daryna Dementieva, Nikolay Babakov, and Alexander Panchenko. 2024. MultiParaDetox: Extending text detoxification with parallel data to new languages. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 124–140, Mexico City, Mexico. Association for Computational Linguis- tics. Daryna Dementieva, Nikolay Babakov, Amit Ro- nen, Abinew Ali Ayele, Naquee Rizwan, Florian Schneider, Xintong Wang, Seid Muhie Yimam, Daniil Moskovskiy, Elisei Stakovskii, Eran Kaufman, Ashraf Elnagar, Animesh Mukherjee, and Alexan- der Panchenko. 2025. Multilingual and explainable text detoxification with parallel corpora. In Proceed- ings of the 31st International Conference on Compu- tational Linguistics, pages 7998–8025, Abu Dhabi, UAE. Association for Computational Linguistics. Mahmoud El -Haj, Paul Rayson, Steve Young, Andrew Moore, Martin Walker, Thomas Schleicher, and Vasi- liki Athanasakou. 2016. Learning tone and attribution for financial text mining. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), pages 1820–1825. Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. Chatgpt outperforms crowd workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30):e2305016120. Hatem Haddad, Hala Mulki, and Asma Oueslati. 2019. T-hsab: A tunisian hate speech and abusive dataset. In International conference on Arabic language pro- cessing, pages 251–263. Springer. Myra S Hunter, Mahmoud El-Haj, Eleanor Thorne, Amanda Griffiths, and Claire Hardy. 2023. # menopause: Examining the frequency of communica- tions about menopause on twitter between 2014 and 2022. Maturitas, 177:107806. Hung Nguyen Huy, Mo El -Haj, Dawn Knight, and Paul Rayson. 2026. Freetxt-vi: A bench- marked vietnamese-english toolkit for segmenta- tion, sentiment, and summarisation. arXiv preprint arXiv:2603.05690. Go Inoue, Bashar Alhafni, Nurpeiis Baimukan, Houda Bouamor, and Nizar Habash. 2021. The interplay of variant, size, and task type in Arabic pre-trained language models. In Proceedings of the Sixth Ara- bic Natural Language Processing Workshop, Kyiv, Ukraine (Online). Association for Computational Lin- guistics. Varvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy, David Dale, Irina Krotova, Nikita Semenov, and Alexander Panchenko. 2022. ParaDetox: Detoxification with parallel data. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6804–6818, Dublin, Ireland. Association for Computational Linguistics. Lin Long, Rui Wang, Ruixuan Xiao, Junbo Zhao, Xiao Ding, Gang Chen, and Haobo Wang. 2024. On llms- driven synthetic data generation, curation, and eval- uation: A survey. In Findings of the Association for Computational Linguistics: ACL 2024, pages 11065–11082. Hamdy Mubarak, Hend Al-Khalifa, and Abdulmohsen Al-Thubaity. 2022. Overview of OSACT5 shared task on Arabic offensive language and hate speech detection. In Proceedinsg of the 5th Workshop on Open-Source Arabic Corpora and Processing Tools with Shared Tasks on Qur’an QA and Fine-Grained Hate Speech Detection, pages 162–166, Marseille, France. European Language Resources Association. Hamdy Mubarak, Kareem Darwish, Walid Magdy, Tamer Elsayed, and Hend Al-Khalifa. 2020. Overview of osact4 arabic offensive language detection shared task. In Proceedings of the 4th Workshop on open-source arabic corpora and processing tools, with a shared task on offensive language detection, pages 48–52. Hamdy Mubarak, Ammar Rashed, Kareem Darwish, Younes Samih, and Ahmed Abdelali. 2021. Arabic offensive language on Twitter: Analysis and exper- iments. In Proceedings of the Sixth Arabic Natu- ral Language Processing Workshop, pages 126–135, Kyiv, Ukraine (Virtual). Association for Computa- tional Linguistics. Hala Mulki and Bilal Ghanem. 2021. Let-mi: An arabic levantine twitter dataset for misogynistic language. arXiv preprint arXiv:2103.10195. Hala Mulki, Hatem Haddad, Chedi Bechikh Ali, and Halima Alshabani. 2019. L-hsab: A levantine twitter dataset for hate speech and abusive language. In Pro- ceedings of the third workshop on abusive language online, pages 111–118. Mihai Nadăș, Laura Dioșan, and Andreea Tomescu. 2025. Synthetic data generation using large language models: Advances in text and code. IEEE Access. Nedjma Ousidhoum, Zizheng Lin, Hongming Zhang, Yangqiu Song, and Dit-Yan Yeung. 2019. Multilin- gual and multi-aspect hate speech analysis. arXiv preprint arXiv:1908.11049. Manuel Tonneau, Diyi Liu, Samuel Fraiberger, Ralph Schroeder, Scott A. Hale, and Paul Röttger. 2024. From languages to geographies: Towards evaluating cultural bias in hate speech datasets. In Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), pages 283–311, Mexico City, Mexico. Association for Computational Linguistics. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Multilin- gual e5 text embeddings: A technical report. arXiv preprint arXiv:2402.05672. Zheng-Xin Yong, Cristina Menghini, and Stephen Bach. 2024. Lexc-gen: Generating data for extremely low- resource languages with large language models and bilingual lexicons. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 13990–14009. Omar F. Zaidan and Chris Callison-Burch. 2014. Ara- bic dialect identification. Computational Linguistics, 40(1):171–202. Appendix A Detoxification Prompt For each source post, GPT-5 and Gemini 2.5 Flash were instructed to generate detoxified rewrites in Modern Standard Arabic (MSA), Gulf Arabic, Lev- antine Arabic, and Egyptian Arabic. The prompt emphasised meaning preservation while removing harmful language. Rewrite the following Arabic social-media post into four civil and non-offensive Arabic versions: 1. Modern Standard Arabic (MSA) 2. Gulf Arabic 3. Levantine Arabic 4. Egyptian Arabic Preserve the original meaning, target, stance, crit- icism, sentiment, and communicative intent. Remove profanity, insults, slurs, dehumanising language, discriminatory expressions, threats, calls for violence, and unnecessarily inflamma- tory wording. Do not defend, refute, explain, summarise, fact- check, or respond to the statement. Rewrite di- rectly from the perspective of the original author. Produce natural social-media language in the re- quested variety and return the output as JSON: "msa_detox": "...", "gulf_detox": "...", "levantine_detox": "...", "egyptian_detox": "..." B Generation Procedure and Quality Control Detoxified rewrites were generated using GPT-5 and Gemini 2.5 Flash through their respective APIs. Generation was performed in batches using struc- tured JSON output to ensure consistent formatting across all generated variants. The final resource contains 42,000 GPT-generated rewrites and 42,000 Gemini-generated rewrites, corresponding to eight detoxified variants for each of the 10,500 source posts. Additional generations were performed dur- ing prompt development, quality control, and regen- eration of invalid outputs. GPT-5 and Gemini 2.5 Flash were selected as representatives of two con- temporary frontier LLM families with strong mul- tilingual generation capabilities. Gemini 2.5 Flash provides a favourable trade-off between generation Parameter GPT-5 Gemini 2.5 Flash API access OpenAI API Gemini API Temperature API default API default Top-p API default API default Maximum output tokens API default API default Structured JSON output Yes Yes Prompt development Iterative pilot testing Iterative pilot testing Automatic regeneration Yes Yes Human quality control Yes Yes Table 9: Generation settings used during dataset con- struction. Parameters not explicitly specified in the API requests used the provider’s default values. quality, latency, and cost for large-scale dataset con- struction, while GPT-5 was included to provide a complementary model family and enable cross- model comparisons. Our objective was not to com- pare LLM performance, but to investigate whether meaning-preserving detoxification remained con- sistent across different model families and Arabic language varieties. Prompt development. The final prompt was de- veloped through iterative pilot testing on 100 source posts. Earlier prompt versions were revised to re- duce safety refusals, summaries, external-observer responses, malformed JSON, unsupported addi- tions, and changes to the original target or stance. The final prompt was selected because it produced the most consistent structured outputs while pre- serving the source perspective and communicative intent. Table 9 summarises the generation settings used for both models. The generation scripts implemented automatic validation checks to detect missing fields, mal- formed JSON, duplicated outputs, null values, and incomplete generations. Outputs that failed these checks were automatically flagged for regeneration. Retry mechanisms were also implemented to handle API failures, empty responses, malformed outputs, and safety-related refusals. Following generation, each batch underwent a quality-control stage per- formed by a native Arabic speaker. The objective of this stage was not to rewrite or annotate the outputs, but to verify that the generation process had been completed correctly and that the resulting texts were suitable for inclusion in the dataset. Quality control focused on several aspects. The reviewer checked that each generated output corre- sponded to the correct source post and preserved its intended meaning, target, stance, and commu- nicative intent. They also verified that the MSA, Gulf, Levantine, and Egyptian variants reflected the intended language variety and exhibited natural linguistic usage. The reviewer inspected the out- puts for technical issues, including alignment errors between source posts and generated rewrites, incon- sistencies across dialectal variants, and generation failures that were not detected automatically. A particular challenge arose from the safety mechanisms of contemporary LLMs. When pre- sented with highly offensive or sensitive content, models occasionally refused the task and produced safety-oriented responses rather than detoxified rewrites (e.g. I cannot fulfil this request. I am pro- grammed to be a helpful and harmless AI assistant). In other cases, the models adopted an external ob- server perspective by describing the content of a post rather than rewriting it from the perspective of the original author (e.g. The speaker appears unhappy with the government’s stance). Such out- puts were considered invalid because they failed to perform the intended detoxification task. When these issues occurred, the affected in- stances were regenerated and, where necessary, additional clarification was provided that the task formed part of a research experiment focused on meaning-preserving detoxification. The reviewer also checked for hallucinated content, omissions of important information, and cases where the gener- ated text substantially altered the meaning of the original post. This quality-control stage served as a human ver- ification process rather than an annotation exercise. The reviewer did not manually rewrite the posts. Their role was to ensure that the LLMs correctly executed the detoxification task and that the gener- ated outputs were complete, aligned with the source posts, linguistically appropriate, and suitable for subsequent analysis. C Experimental Pipeline The experimental pipeline used in this paper con- sists of the following stages: 1. Collect harmful Arabic social-media posts from the Arabic Hate Speech Superset. 2. Generate detoxified Modern Standard Arabic (MSA) rewrites using GPT-5 and Gemini 2.5 Flash. 3. Generate detoxified Gulf Arabic rewrites using GPT-5 and Gemini 2.5 Flash. 4. Generate detoxified Levantine Arabic rewrites using GPT-5 and Gemini 2.5 Flash. 5. Generate detoxified Egyptian Arabic rewrites using GPT-5 and Gemini 2.5 Flash. 6. Compute lexical-change metrics between orig- inal and detoxified texts. 7. Compute semantic similarity using multilin- gual E5 embeddings. 8. Visualise embedding spaces using UMAP. 9. Measure sentiment shift using three CAMeL- BERT sentiment models. 10. Evaluate dialectal style similarity using au- thentic Arabic dialect corpora. 11. Conduct human evaluation using three native Arabic-speaking annotators. 12. Compute inter-rater agreement and aggregate evaluation statistics. D Human Annotation Guidelines Human annotators assessed each generated output according to the following criteria: 1. Offence Removal: whether harmful language present in the source text was successfully re- moved. 2. Meaning Preservation: whether the main meaning and communicative intent of the source text were preserved. 3. New Content: whether the generated out- put introduced information not present in the source text. Offence removal was annotated using a three-way scheme: • 0 = offence not removed, • 1 = offence removed, • 2 = not applicable (no offence present in the source text). All remaining criteria were annotated using bi- nary judgements, where 1 indicates that the crite- rion was satisfied and 0 indicates that it was not. E Human-Evaluation Sample Distribution Table 10 compares the distribution of the complete AraDetox source collection with that of the 300- post human-evaluation sample. Source dataset Evaluation sample Evaluation sample (%) JHSC 30 10.00 MLMA 30 10.00 T-HSAB 30 10.00 L-HSAB 30 10.00 AraCOVID19-MFH 30 10.00 Brothers 30 10.00 Saudi Tweets 30 10.00 Alsafari et al. 30 10.00 Let-Mi 30 10.00 OSACT 30 10.00 Total 300 100.00 Table 10: Distribution of the 300 source posts included in the human-evaluation sample, with 30 posts sampled from each dataset. F Qualitative Examples Tables 11 and 12 present representative examples from AraDetox. The examples illustrate a range of detoxification behaviours, including insult miti- gation, meaning-preserving reformulation, stance preservation, and dialect-specific adaptation. The examples also highlight differences between GPT-5 and Gemini, with some outputs favouring direct lexical substitution and others employing broader paraphrastic rewriting.