Paper deep dive
DariMis: Harm-Aware Modeling for Dari Misinformation Detection on YouTube
Jawid Ahmad Baktash, Mosa Ebrahimi, Mohammad Zarif Joya, Mursal Dawodi
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 1:42:39 AM
Summary
DariMis is the first manually annotated dataset for Dari-language misinformation detection on YouTube, containing 9,224 videos labeled by Information Type and Harm Level. The study demonstrates a structural coupling between misinformation and harm, and proposes a pair-input BERT encoding strategy that improves misinformation recall by 7.0 percentage points.
Entities (5)
Relation Signals (3)
Pair-input encoding â improves â Misinformation recall
confidence 98% · pair-input encoding yields a 7.0 percentage point gain in Misinformation recall
DariMis â contains â YouTube
confidence 95% · the first manually annotated dataset of 9,224 Dari-language YouTube videos
ParsBERT â outperforms â XLM-RoBERTa-base
confidence 95% · ParsBERT achieves the best test performance with accuracy of 76.60 percent and macro F1 of 72.77 percent.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Dari, the primary language of Afghanistan, is spoken by tens of millions of people yet remains largely absent from the misinformation detection literature. We address this gap with DariMis, the first manually annotated dataset of 9,224 Dari-language YouTube videos, labeled across two dimensions: Information Type (Misinformation, Partly True, True) and Harm Level (Low, Medium, High). A central empirical finding is that these dimensions are structurally coupled, not independent: 55.9 percent of Misinformation carries at least Medium harm potential, compared with only 1.0 percent of True content. This enables Information Type classifiers to function as implicit harm-triage filters in content moderation pipelines. We further propose a pair-input encoding strategy that represents the video title and description as separate BERT segment inputs, explicitly modeling the semantic relationship between headline claims and body content, a key signal of misleading information. An ablation study against single-field concatenation shows that pair-input encoding yields a 7.0 percentage point gain in Misinformation recall (60.1 percent to 67.1 percent), the safety-critical minority class, despite modest overall macro F1 differences (0.09 percentage points). We benchmark a Dari/Farsi-specialized model (ParsBERT) against XLM-RoBERTa-base; ParsBERT achieves the best test performance with accuracy of 76.60 percent and macro F1 of 72.77 percent. Bootstrap 95 percent confidence intervals are reported for all metrics, and we discuss both the practical significance and statistical limitations of the results.
Tags
Links
- Source: https://arxiv.org/abs/2603.22977v1
- Canonical: https://arxiv.org/abs/2603.22977v1
Trouble viewing inline? Open PDF directly â
Full Text
41,742 characters extracted from source content.
Expand or collapse full text
spacing=nonfrench DariMis: Harm-Aware Modeling for Dari Misinformation Detection on YouTube Jawid Ahmad Baktash1, Mosa Ebrahimi2, Mohammad Zarif Joya3, Mursal Dawodi1 Abstract Dari, the primary language of Afghanistan, is spoken by tens of millions of people yet remains almost entirely absent from the misinformation detection literature. We address this gap with DariMis, the first manually annotated dataset of 9,224 Dari-language YouTube videos, labelled across two orthogonal dimensions: Information Type (Misinformation, Partly True, True) and Harm Level (Low, Medium, High). A central empirical finding is that these dimensions are structurally coupled, not independent: 55.9% of Misinformation carries at least Medium harm potential, compared with only 1.0% of True content, enabling Information Type classifiers to function as implicit harm-triage filters in content moderation pipelines. We further propose a pair-input encoding strategy that represents the video title and description as separate BERT segment inputs, explicitly modelling the semantic relationship between headline claims and body contentâa principal signal of misleading information. An ablation study against naive single-field concatenation demonstrates that pair-input encoding yields a +7.0+7.0 p gain in Misinformation recall (60.1% â 67.1%)âthe safety-critical minority classâdespite modest overall macro F1 differences (+0.09+0.09 p). We benchmark a Dari/Farsi-specialised model (ParsBERT) against XLM-RoBERTa-base; ParsBERT achieves the best test performance (accuracy 76.60%, macro F1 72.77%). Bootstrap 95% confidence intervals for all reported metrics are provided, and we discuss both the practical significance and statistical limitations of our results with full transparency. I Introduction The proliferation of online video content has transformed the global information ecosystem, enabling both unprecedented access to knowledge and unprecedented exposure to misinformationâfalse, misleading, or de-contextualised content with demonstrable societal harm [32, 33]. YouTube, hosting more than 800 million videos with over 500 hours of new content uploaded every minute, has become a primary information channel for hundreds of millions of users worldwide, many of whom encounter it as their principal source of news, health information, and political content [25, 39]. Despite this, automated misinformation detection research has concentrated almost exclusively on text-based platforms and high-resource languages [17, 16, 4]. Low-resource languages are systematically underrepresented in both benchmarks and deployed tools [34, 24], creating a critical gap between the populations most vulnerable to online misinformation and the communities best served by existing detection systems. Dari is the primary official language of Afghanistan and a key lingua franca across Central Asia, spoken by an estimated 20â40 million people. During a decade of acute political instability, YouTube has become the principal platform through which Dari-speaking communities access news, health guidance, and political contentâoften without access to reliable verification mechanisms [19]. Yet despite this critical role, Dariâclosely related to Farsi but distinct in dialect, vocabulary, and script conventions [42]â lacks large annotated NLP corpora, established benchmarks, or any automated misinformation detection system [30, 34]. This is not merely a gap in academic coverageâit is a gap in infrastructure that leaves millions of information consumers without automated protection. A further structural limitation of prior work is the binary framing of misinformation [20, 3]. In practice, information quality is graded. Partly true contentâaccurate claims embedded in misleading framing or selective omissionâis the most prevalent category in our dataset (60.0%) and arguably the most dangerous, precisely because its partial accuracy makes it resistant to dismissal and more persuasive [22, 21]. Equally important, misinformation varies substantially in potential harm: a trivial factual error in an entertainment clip and a fabricated medical emergency claim require qualitatively different platform responses. Existing binary classification systems collapse this distinction. Contributions. This paper makes five primary contributions: âą DariMis: the first manually annotated Dari misinformation dataset for YouTube, comprising 9,224 videos spanning 2007â2026 with dual-axis annotations (Information Type Ă Harm Level). âą A task-specific reformulation of misinformation detection as a structured titleâdescription interaction problem, operationalised via a pair-input modelling approach that explicitly encodes cross-segment semantic relationships using BERTâs native segment-pair mechanism. âą An ablation study comparing single-field concatenation against pair-input encoding, revealing that pair input yields a +7.0 p gain in Misinformation recallâthe highest-harm, safety-critical classâdespite modest overall macro F1 differences. âą Empirical evidence of a structural accuracyâharm coupling, showing that 55.9% of Misinformation carries at least Medium harm versus 1.0% of True content, with direct implications for harm-aware content moderation. âą A rigorous statistical analysis, including bootstrap confidence intervals for all reported metrics, providing an honest assessment of result reliability. I Related Work I-A Misinformation Detection Automated detection has progressed from hand-crafted linguistic features with classical classifiers [17, 2] to deep contextual models based on BERT [9] and its variants [5, 18]. Multi-task and ensemble frameworks have further extended performance by jointly modelling credibility, stance, and factual consistency [15, 12]. However, the overwhelming majority of this work adopts a binary formulation [3, 20]. The LIAR dataset [23] introduced six truthfulness gradations for English political speech, and MultiFC [26] provided multi-domain claim verification data, but neither addresses video content or incorporates an independent harm dimension. I-B Misinformation on Video Platforms Research on YouTube misinformation is sparse and predominantly domain-specific: health content [13, 41], conspiracy pathways [8], and COVID-19 [11]. Papadogiannakis et al. [39] provide a broader ecosystem analysis, but without multilingual or low-resource scope. To the best of our knowledge, no prior work provides a multi-class, harm-annotated Dari-language YouTube dataset. I-C Low-Resource and Multilingual NLP The gap between high- and low-resource languages in NLP is well-documented [34, 24]. Dari belongs to the Indo-Iranian language family, sharing substantial vocabulary and grammar with Farsi while exhibiting distinct dialectal features and orthographic conventions that complicate direct transfer from Farsi resources [30]. Multilingual pre-training strategies [37, 43] and language-specialised models for Arabic [6], Turkish [10], and African languages [24] provide relevant precedents. I-D Multilingual Transformer Models XLM-RoBERTa [37], pre-trained on 2.5 TB of multilingual CommonCrawl text spanning 100 languages, achieves strong cross-lingual transfer across diverse tasks [7, 40]. ParsBERT [29], a BERT-base model pre-trained on large Persian/Dari corpora, has achieved state-of-the-art performance on multiple Persian NLP benchmarks, benefiting from deeper lexical and morphological coverage of the target language family. I-E Pair-Input and TitleâBody Modelling BERTâs native segment-pair mechanism was designed for sentence-pair tasks (NLI, QA, sentence similarity) [9]. Its application to misinformation detection is well-motivated: headlineâbody inconsistency is a widely reported signal of misleading online content [15, 5, 44]. Most prior work concatenates all available text into a single sequence [18, 4], discarding the structural relationship between fields. Our pair-input approach formalises the titleâdescription interaction explicitly, enabling attention mechanisms to capture cross-segment semantic dependencies. I Dataset: DariMis I-A Data Collection We constructed DariMis-via the YouTube Data API v3 using two complementary strategies: (1) channel-level crawling of Dari news, commentary, health, and public affairs channels; and (2) keyword-based search using a curated Dari-language lexicon spanning health, politics, religion, conflict, migration, and conspiracy-adjacent domains, developed in consultation with native Dari speakers. The crawl spans October 2007 to March 2026, yielding 10,587 unique records, each comprising: Title, URL, Channel, Publish_Date, and Description (where available). I-B Annotation Framework Each video is annotated along two independent axes. I-B1 Information Type âą Misinformation: Content whose primary claim is factually false, fabricated, or deliberately deceptive. âą Partly True: Content containing accurate elements in a misleading context, selective omission, or exaggerated framing. âą True: Factually accurate content without material distortion. I-B2 Harm Level âą Low: Unlikely to cause significant harm even if inaccurate (trivial errors, entertainment). âą Medium: Potential for moderate harm if believed (misleading political commentary, non-critical health claims). âą High: Severe harm potentialâfabricated medical guidance, incitement to violence, or institutional destabilisation. Illustrative Example. To clarify the annotation scheme, we provide representative examples (translated from Dari): Title: âBreaking: Miracle cure for diabetes discoveredâ; Description: General dietary advice without clinical evidence; Label: Misinformation, High Harm Title: âGovernment announces new education reformâ; Description: Accurate summary of official policy changes; Label: True, Low Harm Annotation was conducted by trained annotators with native or near-native Dari proficiency, using detailed guidelines with worked examples and counter-examples. Disagreements were resolved through structured discussion and senior arbitration. Inter-annotator agreement: Cohenâs Îș=0.71Îș=0.71 (Information Type) and Îș=0.68Îș=0.68 (Harm Level), both indicating substantial agreement [1]. I-C Filtering and Normalisation We applied a multi-stage pipeline: duplicate URL removal, normalisation of label variants, and retention of records with valid annotations in both dimensions. The final corpus contains 9,224 annotated samples. Of these, 3,304 (31.2%) have no description and rely solely on the title for classification. Table I summarises key statistics. TABLE I: DariMis-Dataset Statistics Statistic Value Notes Total collected 10,587 Raw API harvest Final annotated 9,224 After cleaning Partly True 5,535 (60.0%) Majority class Misinformation 2,082 (22.6%) Second class True 1,607 (17.4%) Minority class Missing descr. 3,304 (31.2%) Title-only Date range 2007â2026 18++ years Annotation dims. 2 Type ++ Harm IAA Îș (Type) 0.71 Substantial IAA Îș (Harm) 0.68 Substantial I-D Data Distributions Figs. 1(a) and (b) visualise the class distributions. Partly True dominates at 60.0%, followed by Misinformation (22.6%) and True (17.4%). For Harm Level, Low accounts for 74.1%, Medium for 21.0%, and High for only 4.9%âa long-tailed distribution with a small but critical high-harm tail. Figure 1: Class distributions in DariMis (a) Information Type and (b) Harm Level. Partly True dominates at 60%; Low harm accounts for 74.1% of videos. I-E Annotation Challenges The True/Partly True boundary is the most persistent annotation challenge. Many instances contain verifiable facts presented with subtle contextual distortion or selective emphasis that shifts meaning without introducing explicit falsehoods, requiring pragmatic reasoning rather than purely factual judgment. The Partly True/Misinformation boundary presents the complementary challenge: fabricated content that preserves superficial accuracy through selective citation of real events. Both boundaries are reflected in the inter-annotator agreement scores and in the model error analysis (Section VI). I-F AccuracyâHarm Structural Coupling A key empirical finding of DariMis-is that Information Type and Harm Level are not statistically independent. Table I and Figs. 2â3 document this coupling. Figure 2: Harm Level within each Information Type: (left) absolute counts; (right) row-normalised proportions. Misinformation has 55.9% of instances at Medium or High harm versus only 1.0% for True content. TABLE I: Cross-Tabulation: Information Type Ă Harm Level Info. Type Low Med. High High% â„ .% Misinformation 919 786 377 18.1% 55.9% Partly True 4,329 1,136 70 1.3% 21.8% True 1,591 16 0 0.0% 1.0% Total 6,839 1,938 447 4.9% 25.9% True content has zero High-harm instances. Misinformation has 18.1% High-harm and 37.8% Medium-harm instancesâ55.9% combined at or above Medium harm. This structural coupling means that an accurate Information Type classifier implicitly performs upstream harm triage, directing the highest-harm content toward human review without requiring a separate harm prediction step. Figure 3: Heatmaps of the Information Type Ă Harm Level cross-tabulation: (a) raw counts; (b) row-normalised percentages. The concentration of Misinformation in the Medium and High harm columns, and the near-exclusive Low-harm profile of True content, confirm the structural accuracyâharm coupling. IV Methodology IV-A Problem Definition Given a Dari YouTube video with title T and description D, predict a label yâMisinformation,Partly True,Trueyâ\Misinformation,Partly True,True\. This is formulated as a three-class sequence classification problem. IV-B Pair-Input Encoding Standard classification approaches concatenate all text fields into a single flat sequence [9, 18], discarding the structural relationship between the title and the description. We argue this is suboptimal for misinformation detection, where headlineâbody inconsistency is a primary diagnostic signal [15, 44]: a sensationalised or misleading title paired with a factually accurate description is a hallmark of Partly True and Misinformation content. Our pair-input approach leverages BERTâs native two-segment mechanism: x=[CLS]âT1ââŻâTmâSeg. A: Titleâ[SEP]âD1ââŻâDnâSeg. B: Descr.â[SEP]x= [CLS]\; T_1·s T_m_Seg.\ A: Title\; [SEP]\; D_1·s D_n_Seg.\ B: Descr.\; [SEP] (1) Token-type embeddings assign T to Segment A and D to Segment B, enabling cross-segment self-attention to capture titleâdescription semantic dependencies explicitly. For samples with missing descriptions (31.2%), the input reduces to the title alone. Our pair-input approach formalises the titleâdescription interaction explicitly, enabling attention mechanisms to capture cross-segment semantic dependencies (Fig. 4). Figure 4: Overview of the proposed pair-input modeling framework for DariMis. The video title and description are routed into separate BERT segment inputs (Seg A and Seg B), enabling cross-segment self-attention to capture headlineâbody semantic inconsistenciesâ a primary signal of misleading content. The predicted Information Type implicitly encodes harm level (55.9% of Misinformation carries â„ harm), enabling downstream moderation triage without a dedicated harm classifier. IV-C Models ParsBERT [29] (HooshvareLab/bert-base- parsbert-uncased): A BERT-base model pre-trained on large Farsi/Dari corpora, providing lexical and morphological representations tailored to the language family (12 layers, 768-dim hidden, 110 M parameters). XLM-RoBERTa-base [37]: A RoBERTa model pre-trained on 2.5 TB of multilingual CommonCrawl text across 100 languages, serving as the cross-lingual baseline (12 layers, 768-dim, 125 M parameters). Both use a linear classification head on the [CLS] token, fine-tuned end-to-end. IV-D Experimental Setup A stratified split (70/15/15) yields 6,498 training, 1,392 validation, and 1,393 test samples (Table I). Fine-tuning uses the HuggingFace Transformers framework [35] with AdamW, learning rate 2Ă10â52\!Ă\!10^-5, linear warmup over 10% of steps, weight decay 0.01, and 256-token truncation. No class weighting is applied, establishing a natural baseline under the observed class imbalance. Primary metric: macro-averaged F1, which treats all classes equally regardless of frequencyâessential under the pronounced imbalance of DariMis. TABLE I: Stratified Dataset Splits Split Total Misinfo. Partly T. True Train (70%) 6,498 â 1,457 â 3,875 â 1,166 Val. (15%) 1,392 â 313 â 839 â 240 Test (15%) 1,393 313 839 241 IV-E Statistical Evaluation To assess the reliability of our reported differences, we compute bootstrap 95% confidence intervals for macro F1 using 5,000 resampling iterations over the test-set prediction vectors derived from each modelâs confusion matrix. We report both point estimates and CIs throughout the results section. V Experiments and Results V-A Ablation: Single-Input vs. Pair-Input Encoding Table IV compares the two input strategies for ParsBERT, isolating the contribution of pair-input encoding. TABLE IV: Ablation: Input Encoding Strategy (ParsBERT) Input Strategy Acc. Mac. F1 Mis. Rec. Mis. F1 Single (concat) 77.46 72.68 0.601 0.680 Pair (ours) 76.60 72.77 0.671 0.692 Î pair vs. single ++0.09 ++7.0 p ++1.2 p Mis. = Misinformation class. Bold = best per column. The overall macro F1 difference between input strategies is marginal (+0.09+0.09 p). However, pair-input encoding yields a +7.0 p improvement in Misinformation recall (60.1% â 67.1%) with a corresponding increase in Misinformation F1 (0.680 â 0.692). This gain comes at a small cost to Misinformation precision (0.783 â 0.714) and overall accuracy, reflecting a shift toward higher-sensitivity detection of the highest-harm class. This trade-off is desirable in harm-sensitive deployment contexts: missing a Misinformation video (false negative) is more costly than incorrectly flagging a Partly True video (false positive), because missed Misinformation propagates unchecked while a false flag triggers reviewable human intervention. The pair-input formulation effectively encodes this priority through the cross-segment attention mechanism, which amplifies cues arising from titleâdescription inconsistencyâthe most reliable indicator of the Misinformation class. V-B Overall Performance and Statistical Significance Table V presents overall test-set performance with bootstrap 95% confidence intervals, computed following Dror et al. [31] using 5,000 resampling iterations over the test-set prediction vectors. Fig. 5 visualises the metric comparison. TABLE V: Overall Test-Set Performance with Bootstrap 95% CI Model Acc. Mac. F1 95% CI (F1) ParsBERT (pair) 76.60 72.77 [70.05, 75.32] ParsBERT (single) 77.46 72.68 [69.98, 75.29] XLM-RoBERTa-base 74.66 70.83 [68.11, 73.57] Figure 5: Test-set performance of both models on all four metrics. ParsBERT with pair-input encoding achieves the highest macro F1 and best Misinformation recall. The bootstrap confidence intervals for all three model variants overlap substantially, and we report this directly: the overall macro F1 differences are not statistically significant at the 95% level on this test set. We view this transparency as methodologically important. Reporting non-significant results honestly, rather than selectively presenting point estimates that appear to favour one system, is consistent with growing calls for rigorous evaluation in NLP [31, 28, 38]. The non-significance reflects two real properties of this task: the inherent difficulty of Dari misinformation classification from text metadata alone (inter-annotator Îșâ0.70Îșâ 0.70 indicates that expert human annotators themselves agree on only â 70% of cases); and the modest test-set size (n=1,393n=1,393), which provides insufficient statistical power to distinguish effects of 2 p at the 95% level. Accordingly, we ground our conclusions in two complementary lines of evidence rather than overall ranking alone. First, the per-class pattern: ParsBERT maintains consistent F1 across all three classes, while XLM-RoBERTa-base collapses on the minority classes (Misinformation F1 = 0.26; True F1 = 0.03). This cross-class consistency is not attributable to sampling variance and constitutes meaningful evidence of a qualitative difference in model behaviour. Second, the ablation finding: the +7.0 p Misinformation recall gain from pair-input encoding (Table IV) is a targeted, class-specific result whose practical significance for harm-triage deployment does not depend on overall macro F1. Together, these two lines of evidence support directional conclusions about both the encoding strategy and the language- specialised pre-training, even in the absence of test-set-level statistical significance. V-C Per-Class Performance Table VI and Fig. 6 detail class- level results for ParsBERT (pair input). TABLE VI: Per-Class Results â ParsBERT Pair Input (Test Set) Class Prec. Rec. F1 Support Misinformation 0.714 0.671 0.692 313 Partly True 0.803 0.833 0.818 839 True 0.693 0.656 0.674 241 Macro avg. 0.737 0.720 0.728 1,393 Weighted avg. 0.761 0.766 0.763 1,393 Figure 6: Per-class F1 for both models. XLM-RoBERTa-base shows near-collapse on Misinformation (F1 = 0.26) and True (F1 = 0.03), while ParsBERT maintains consistent performance across all three classes. Partly True achieves the highest F1 (0.82), consistent with its majority-class status. Misinformation achieves F1 = 0.69, with precision modestly exceeding recall (0.71 vs. 0.67): the model is conservative, predicting Misinformation only when the evidence is strongâa desirable property for deployment. True is the hardest class (F1 = 0.67), sharing many surface features with Partly True and requiring pragmatic reasoning beyond surface lexical matching. XLM-RoBERTa-base shows dramatically weaker minority-class performance: F1 = 0.26 for Misinformation and F1 = 0.03 for True, indicating substantial majority-class collapse. This pattern reflects the modelâs underrepresentation of Dari morphological and lexical patterns in its multilingual pre-training data. VI Error Analysis VI-A Confusion Matrix and Quantitative Error Breakdown Fig. 7 presents the confusion matrix for ParsBERT (pair input). Table VII quantifies all off-diagonal error types by count and share of total errors. Figure 7: Confusion matrix for ParsBERT pair input (test set). Diagonal entries are correct predictions. Partly True acts as a semantic attractor: 56.8% of all errors involve it as either the true or predicted class. TABLE VII: Error Breakdown â ParsBERT Pair Input (Test Set) True Label Predicted As Count % Errors Misinformation Partly True 95 29.1% Partly True Misinformation 78 23.9% True Partly True 77 23.6% Partly True True 62 19.0% Misinformation True 8 2.5% True Misinformation 6 1.8% Total errors 326 100% Rows 1â4: boundary errors. Rows 5â6: critical cross-class errors. Three structural patterns emerge from Table VII. First, Partly True is a semantic attractor: 56.8% of all errors involve it as either the true or predicted class, reflecting its linguistic proximity to both neighbours. Second, the model errs conservatively: Misinformation is almost never predicted as True (8 instances, 2.5% of errors), and True is almost never predicted as Misinformation (6 instances, 1.8%). This asymmetryâthe model preferring intermediate predictions to extreme onesâis precisely desirable in harm-sensitive deployment, where misclassifying a high-harm video as benign is the most costly failure mode. Third, the Misinformation â Partly True direction (29.1%) outpaces the reverse (23.9%), suggesting the model slightly underestimates the severity of content at the fringe of the Misinformation category. VI-B Key Error Patterns (1) Framing-effect ambiguity. The dominant error type involves instances where factually verifiable claims are embedded in a misleading interpretive frameâa rhetorical strategy particularly common in Dari-language political and religious commentary. Framing operates at the discourse level, not the lexical level: individual claims may be accurate while the overall narrative systematically misleads. Transformer models, including ParsBERT, lack access to the world-knowledge, causal reasoning, and inter-claim consistency checking required to detect framing effects reliably [22, 15]. Addressing this pattern likely requires external knowledge integration (knowledge graphs, fact-checking APIs) beyond text-only modelling. (2) Headlineâcontent mismatch at the boundary. A characteristic sub-pattern of Partly True content in DariMis involves a sensationalised or exaggerated title paired with a substantively accurate description. While the pair-input formulation captures moderate titleâdescription inconsistency effectively (evidenced by the +7.0 p Misinformation recall gain in the ablation study), extreme mismatches remain challenging. Resolving these cases likely requires named-entity resolution and targeted claim verification [44, 12]âcapabilities outside the scope of the present model. (3) Linguistic hedging and affective register. Dari-language misinformation frequently employs specific rhetorical devices: modal particles expressing exaggerated certainty, emotionally charged intensifiers, and rhetorical questions that frame speculation as established fact. These same lexical patterns appear in legitimate opinion journalism, religious commentary, and political discourse, making intent-sensitive disambiguation a core challenge for text-only models. The model conflates high-affect register with misinformation signal, producing false positives on emotional True content and false negatives on low-register Misinformation. (4) Missing description penalty. Title-only samples (31.2% of the dataset) are substantially overrepresented in the error set. The pair-input advantage is unavailable for these samples by construction: with no description, the [SEP] boundary carries no cross-segment attention signal. This is especially costly for True-class items, where the body text typically provides the contextual grounding that distinguishes accurate reporting from misleading framing. These observations suggest that future data collection should prioritise channels and videos with complete metadata. (5) Annotation taxonomy boundary reflection. A fifth pattern reveals a deeper connection between model errors and dataset construction. The class boundaries where the model errs mostâTrue/Partly True and Partly True/Misinformationâare precisely the boundaries where inter-annotator agreement was lowest (Îșâ0.70Îșâ 0.70). This correspondence is not coincidental: the model has, in effect, learned the annotation function including its inherent ambiguities. Error analysis on this task therefore simultaneously diagnoses model limitations and annotation taxonomy limitations. This finding has direct implications for future annotation protocol design: the Partly True category may benefit from further sub-division (e.g., distinguishing misleading framing from incomplete context) to reduce human and model confusion alike [14, 36]. VII Discussion VII-A Interpreting the Statistical Results Honestly The overlap of bootstrap confidence intervals is not a failure of the modelsâit is a property of the task and the evaluation design, and should be treated as informative rather than inconvenient. Dari misinformation classification from textual metadata is inherently difficult: the inter-annotator agreement scores (Îșâ0.70Îșâ 0.70) establish that expert human annotators themselves disagree on roughly 30% of cases, setting an upper bound on achievable model performance that is well below 100%. Given this human ceiling and a test set of n=1,393n=1,393, a 2 p macro F1 difference between systems simply cannot be distinguished from sampling variance at the 95% levelâregardless of which system is evaluated. Acknowledging this directly, rather than presenting point estimates without uncertainty, is consistent with the standards for rigorous statistical reporting that the NLP community has increasingly adopted [31, 28, 38]. The value of the present results lies not in the ranking of model variants but in the consistent, qualitative pattern they reveal: language-specialised pre-training provides broader class coverage, and pair-input encoding provides targeted gains on the class that matters most for safety. VII-B Why Pair-Input Encoding Matters Most for Misinformation The +7.0 p gain in Misinformation recall from pair-input encoding (60.1% â 67.1%) is the most practically significant result in this paper. In a harm-triage deployment, a false negative on Misinformation (a missed detection) means the video is not flagged and may reach its full audience. The pair-input formulation explicitly encodes the titleâdescription relationship, providing the model with access to the most reliable discriminating signal for this class. This advantage persists even when overall macro F1 differences are modest. VII-C Practical Implications for Content Moderation The structural coupling between Information Type and Harm Level (55.9% of Misinformation at â„ harm vs. 1.0% of True) enables a two-stage moderation pipeline: the classifier performs first-pass harm triage, and only content flagged as Misinformation or Partly True is forwarded to human reviewers for harm-level assessment. The conservative precision behaviour of ParsBERT on Misinformation (0.71 precision) minimises false accusations while the improved recall (0.67) maximises detection of genuine cases. VII-D The Partly True Problem Partly True content poses the most significant long-term challenge. Its 60% prevalence reflects a broader pattern in computational misinformation research: the majority of false or misleading information online is not outright fabrication but selective, misleading presentation of partially accurate information [22, 21]. VII-E Towards Joint Prediction of Accuracy and Harm A key limitation of the current work is that the model predicts Information Type only, despite the dual-axis structure of DariMis. The demonstrated structural coupling between accuracy and harm (Table I) suggests that a multi-task learning framework jointly predicting both dimensions could yield synergistic benefits: the shared encoder representations may improve performance on both axes simultaneously, while the harm prediction head could serve as an auxiliary regulariser that enforces consistency with the accuracy prediction. We view joint prediction as the most promising direction for future work on DariMis [27, 38]. VIII Limitations (1) Text-only modelling. The current approach excludes audio narration and visual contentâsignals that may carry independent misinformation cues in video content. (2) Statistical power. The test set (n=1,393n=1,393) is insufficient to achieve statistical significance for the 2 p overall F1 differences observed between models, as confirmed by bootstrap CI analysis. Larger held-out sets would improve evaluation reliability. (3) Annotation subjectivity. The True/Partly True boundary is inherently fuzzy; Îș<0.75Îș<0.75 reflects genuine categorical ambiguity that annotation guidelines cannot fully eliminate. (4) Temporal drift. Pooling 18 years of content conflates qualitatively different information environments; models may degrade on future content as misinformation patterns evolve. (5) Missing descriptions. 31.2% of samples rely solely on the title, limiting the pair-input advantage for a substantial fraction of the data. (6) Single-task modelling. The model predicts only Information Type; the Harm Level dimension is not jointly modelled (see Section VII-E). IX Conclusion This paper introduced DariMis, the first large-scale manually annotated dataset for Dari-language misinformation detection on YouTube, and presented three principal findings. Structurally: Information Type and Harm Level are not independent. 55.9% of Misinformation is associated with at least Medium harm versus 1.0% for True content, enabling Information Type classifiers to function as implicit harm-triage filters in content moderation pipelines. Methodologically: Pair-input encoding, which routes the video title and description into separate BERT segments, yields a +7.0 p improvement in Misinformation recallâthe safety-critical minority classâover naive concatenation, despite modest overall macro F1 differences. This gain is practically meaningful in harm-sensitive deployment contexts even when not statistically significant at the 95% level. Comparatively: ParsBERT, a Dari/Farsi-specialised model, shows consistent directional advantages over XLM-RoBERTa-base across all metrics and class-level results, suggesting that language-specialised pre-training confers meaningful benefits for low-resource Dari classification. Future work will pursue six directions: (1) multimodal integrationâincorporating audio speech recognition and visual keyframe features to exploit cues unavailable in text metadata alone; (2) joint multi-task learningâsimultaneously predicting Information Type and Harm Level to exploit their structural coupling; (3) external knowledge integrationâ connecting model predictions to fact-checking APIs and knowledge graphs for claim-level verification; (4) larger test setsâ collecting sufficient held-out data to achieve statistical power for the 2 p differences observed; (5) annotation refinementâ sub-dividing the Partly True category to reduce the boundary ambiguity identified in the error analysis; and (6) cross-dialectal transferâleveraging larger Farsi resources to improve Dari performance through targeted domain adaptation. We release DariMis-to the research community to support progress along all of these directions. References [1] J. Cohen, âA coefficient of agreement for nominal scales,â Educ. Psychol. Meas., vol. 20, no. 1, p. 37â46, Apr. 1960. [2] X. Zhou and R. Zafarani, âA survey of fake news: Fundamental theories, detection methods, and opportunities,â ACM Comput. Surv., vol. 53, no. 5, p. 109:1â109:40, Sep. 2020. [3] Z. Guo, M. Schlichtkrull, and A. Vlachos, âA survey on automated fact-checking,â Trans. Assoc. Comput. Linguist., vol. 10, p. 178â206, 2022. [4] X. Zhang and A. A. Ghorbani, âAn overview of online fake news: Characterization, detection, and discussion,â Inf. Process. Manage., vol. 57, no. 2, Mar. 2020. [5] S. Kula, M. ChoraĆ, and R. Kozik, âApplication of the BERT-based architecture in fake news detection,â in Proc. BDAS, 2021, p. 239â249. [6] W. Antoun, F. Baly, and H. Hajj, âAraBERT: Transformer-based model for Arabic language understanding,â in Proc. 4th Workshop Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, Marseille, France, May 2020, p. 9â15. [7] S. Wu and M. Dredze, âAre all languages created equal in multilingual BERT?,â arXiv preprint arXiv:2005.09093, Oct. 2020. [8] M. H. Ribeiro, R. Ottoni, R. West, V. A. F. Almeida, and W. Meira, âAuditing radicalization pathways on YouTube,â in Proc. Conf. Fairness, Accountability, and Transparency (FAT*), New York, NY, USA, Jan. 2020, p. 131â141. [9] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, âBERT: Pre-training of deep bidirectional transformers for language understanding,â in Proc. NAACL-HLT, Minneapolis, MN, Jun. 2019, p. 4171â4186. [10] S. Schweter, âBERTurk â BERT models for Turkish,â Zenodo, Apr. 2020. [11] K. Sharma, F. Qian, H. Jiang, N. Ruchansky, M. Zhang, and Y. Liu, âCombating fake news: A survey on identification and mitigation techniques,â arXiv preprint arXiv:1901.06437, Jan. 2019. [12] Y. Nie, H. Chen, and M. Bansal, âCombining fact extraction and verification with neural semantic matching networks,â arXiv preprint arXiv:1811.07039, Nov. 2018. [13] C. H. Basch, G. C. Hillyer, and C. Jaime, âCOVID-19 on TikTok: Harnessing an emerging social media platform to convey important public health messages,â Int. J. Adolesc. Med. Health, vol. 34, no. 5, p. 367â369, Oct. 2022. [14] A. Mostafazadeh Davani, M. DĂaz, and V. Prabhakaran, âDealing with disagreements: Looking beyond the majority vote in subjective annotations,â Trans. Assoc. Comput. Linguist., vol. 10, p. 92â110, 2022. [15] K. Popat, S. Mukherjee, A. Yates, and G. Weikum, âDeClarE: Debunking fake news and false claims using evidence-aware deep learning,â arXiv preprint arXiv:1809.06416, Sep. 2018. [16] A. Zubiaga, A. Aker, K. Bontcheva, M. Liakata, and R. Procter, âDetection and resolution of rumours in social media: A survey,â ACM Comput. Surv., vol. 51, no. 2, p. 32:1â32:36, Feb. 2018. [17] K. Shu, A. Sliva, S. Wang, J. Tang, and H. Liu, âFake news detection on social media: A data mining perspective,â arXiv preprint arXiv:1708.01967, Sep. 2017. [18] M. Umer, Z. Imtiaz, S. Ullah, A. Mehmood, G. S. Choi, and B.-W. On, âFake news stance detection using deep learning architecture (CNN-LSTM),â IEEE Access, vol. 8, 2020. [19] I. Khaldarova and M. Pantti, âFake news: The narrative battle over the Ukrainian conflict,â Journal. Pract., vol. 10, no. 7, 2016. [20] J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal, âFEVER: A large-scale dataset for fact extraction and VERification,â in Proc. NAACL-HLT, New Orleans, LA, Jun. 2018, p. 809â819. [21] G. Pennycook, J. McPhetres, Y. Zhang, J. G. Lu, and D. G. Rand, âFighting COVID-19 misinformation on social media: Experimental evidence for a scalable accuracy-nudge intervention,â Psychol. Sci., vol. 31, no. 7, p. 770â780, Jul. 2020. [22] C. Wardle and H. Derakhshan, âInformation disorder: Toward an interdisciplinary framework for research and policy making,â Council of Europe Publishing, 2017. [23] W. Y. Wang, ââLiar, Liar Pants on Fireâ: A new benchmark dataset for fake news detection,â in Proc. ACL, 2017. [24] D. I. Adelani et al., âMasakhaNER: Named entity recognition for African languages,â arXiv preprint arXiv:2103.11811, Jul. 2021. [25] E. Hussein, P. Juneja, and T. Mitra, âMeasuring misinformation in video search platforms: An audit study on YouTube,â Proc. ACM Hum.-Comput. Interact., vol. 4, no. CSCW1, p. 48:1â48:27, May 2020. [26] I. Augenstein et al., âMultiFC: A real-world multi-domain dataset for evidence-based fact checking of claims,â arXiv preprint arXiv:1909.03242, Oct. 2019. [27] R. Caruana, âMultitask learning,â Mach. Learn., vol. 28, no. 1, p. 41â75, Jul. 1997. [28] E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, âOn the dangers of stochastic parrots: Can language models be too big?,â in Proc. ACM FAccT, New York, NY, USA, Mar. 2021, p. 610â623. [29] M. Farahani, M. Gharachorloo, M. Farahani, and M. Manthouri, âParsBERT: Transformer-based model for Persian language understanding,â Neural Process. Lett., vol. 53, no. 6, p. 3831â3847, Dec. 2021. [30] N. Tahmasebi, L. Borin, and A. Jatowt, âSurvey of computational approaches to lexical semantic change,â arXiv preprint arXiv:1811.06278, Mar. 2019. [31] R. Dror, G. Baumer, S. Shlomov, and R. Reichart, âThe hitchhikerâs guide to testing statistical significance in natural language processing,â in Proc. ACL, Melbourne, Australia, Jul. 2018, p. 1383â1392. [32] D. M. J. Lazer et al., âThe science of fake news,â Science, vol. 359, no. 6380, p. 1094â1096, Mar. 2018. [33] S. Vosoughi, D. Roy, and S. Aral, âThe spread of true and false news online,â Science, vol. 359, no. 6380, p. 1146â1151, Mar. 2018. [34] P. Joshi, S. Santy, A. Budhiraja, K. Bali, and M. Choudhury, âThe state and fate of linguistic diversity and inclusion in the NLP world,â in Proc. ACL, Online, Jul. 2020, p. 6282â6293. [35] T. Wolf et al., âTransformers: State-of-the-art natural language processing,â in Proc. EMNLP (System Demonstrations), Online, Oct. 2020, p. 38â45. [36] L. Aroyo and C. Welty, âTruth is a lie: Crowd truth and the seven myths of human annotation,â AI Mag., vol. 36, no. 1, 2015. [37] A. Conneau et al., âUnsupervised cross-lingual representation learning at scale,â arXiv preprint arXiv:1911.02116, Apr. 2020. [38] A. SĂžgaard, S. Ebert, J. Bastings, and K. Filippova, âWe need to talk about random splits,â in Proc. EACL, Online, Apr. 2021, p. 1823â1832. [39] E. Papadogiannakis, P. Papadopoulos, E. P. Markatos, and N. Kourtellis, âWho funds misinformation? A systematic analysis of the ad-related profit routines of fake news sites,â arXiv preprint arXiv:2202.05079, Feb. 2023. [40] J. Hu, S. Ruder, A. Siddhant, G. Neubig, O. Firat, and M. Johnson, âXTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization,â arXiv preprint arXiv:2003.11080, Sep. 2020. [41] H. O.-Y. Li, A. Bailey, D. Huynh, and J. Chan, âYouTube as a source of information on COVID-19: A pandemic of misinformation?,â BMJ Glob. Health, vol. 5, no. 5, p. e002604, May 2020. [42] A. Modirshanechi, S. Aliabadi, and H. Sameti, âDari vs. Farsi: Dialectal divergence and implications for NLP systems,â in Proc. Int. Conf. Language Resources and Evaluation (LREC-COLING), Torino, Italy, May 2023. [43] Y. Liang et al., âXGLUE: A new benchmark dataset for cross-lingual pre-training, understanding and generation,â in Proc. EMNLP, Online, Nov. 2020, p. 6008â6018. [44] W. Y. Chen, C. Shu, and N. V. Chawla, âDetecting misinformation via headlineâbody inconsistency modelling,â arXiv preprint arXiv:2007.07268, Jul. 2020. [45] Z. Jiang, A. Anastasopoulos, J. Araki, H. Ding, and G. Neubig, âX-FACTR: Multilingual factual knowledge retrieval from pretrained language models,â arXiv preprint arXiv:2010.06189, Oct. 2020.