Paper deep dive
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse
Celestine Achi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 5:19:48 AM
Summary
The paper introduces the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema designed to capture the pragmatic nuances of Nigerian public discourse. Unlike traditional sentiment benchmarks (e.g., NaijaSenti, AfriSenti) that focus on three-way polarity, MIF separates surface sentiment from true communicative intent by incorporating dimensions such as register, irony, coded subtext, risk tier, and recommended communications action. The study identifies a 'Register Gap,' where frontier models like Gemini 2.5 Flash significantly improve performance (from 33.3% to 73.3% in register classification) when provided with schema-informed prompting. The framework aims to bridge the gap between literal translation and contextual understanding for media intelligence and crisis communications.
Entities (7)
Relation Signals (4)
Gemini 2.5 Flash â evaluatedusing â Meaning Intelligence Framework
confidence 100% ¡ evaluate a frontier language model (Gemini 2.5 Flash) under zero-shot and schema-informed prompting conditions.
Meaning Intelligence Framework â includesdimension â Register
confidence 100% ¡ The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent...
Meaning Intelligence Framework â includesdimension â True Intent
confidence 100% ¡ The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent...
Meaning Intelligence Framework â addressesgapin â NaijaSenti
confidence 90% ¡ Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti... The MIF addresses this gap.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task (positive, negative, neutral). We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate a frontier language model (Gemini 2.5 Flash) under zero-shot and schema-informed prompting conditions. The headline finding is the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. The composite Meaning Intelligence Score increases by 5.4 points (73.2 to 78.6) under schema-informed prompting, with the largest practical gains in register identification, coded-subtext detection (+10 points), and strategic action recommendation (+10.3 points). We release the framework specification, annotation guidelines, and the 30-item public calibration set to support reproducibility, while retaining a private holdout corpus for contamination-protected evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2606.20255v1
- Canonical: https://arxiv.org/abs/2606.20255v1
Trouble viewing inline? Open PDF directly â
Full Text
26,462 characters extracted from source content.
Expand or collapse full text
1 The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse Celestine Achi AGENTPR⢠/ AI-Powered PR / Cihan Digital Academy celestine.achi@gmail.com Abstract We introduce the Meaning Intelligence Framework (MIFâ˘), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent. Existing benchmarks for Nigerian languages, including NaijaSenti and AfriSenti, treat sentiment classification as a three-way polarity task (positive, negative, neutral). We argue that the dominant failure mode of AI systems on Nigerian discourse is not translation failure but context failure: the same utterance carries opposite pragmatic force depending on speaker, audience, and situation. The MIF operationalises this insight across nine scored dimensions: register, surface sentiment, true intent, irony, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. We construct a 30-item calibration dataset spanning Standard English, Nigerian English, Nigerian Pidgin, and code-mixed registers, and evaluate a frontier language model (Gemini 2.5 Flash) under zero-shot and schema-informed prompting conditions. The headline finding is the Register Gap: zero-shot register classification accuracy is 33.3%, rising to 73.3% (+40 points) when the model receives the MIF schema in-context. The composite Meaning Intelligence Score increases by 5.4 points (73.2 to 78.6) under schema-informed prompting, with the largest practical gains in register identification, coded-subtext detection (+10 points), and strategic action recommendation (+10.3 points). We also identify a mobilisation blind spot: the model misclassifies a mobilisation signal disguised as humour as a routine warning in both conditions, demonstrating a failure mode with direct consequences for media monitoring and crisis communications. We release the framework specification, annotation guidelines, and the 30-item public calibration set to support reproducibility, while retaining a private holdout corpus for contamination-protected evaluation. Keywords: Nigerian Pidgin, sentiment analysis, pragmatics, context-dependent NLP, African languages, benchmark, media intelligence, sarcasm detection, cultural NLP 1. Introduction Nigerian public discourse operates across at least four registers: Standard English (SE), Nigerian English (NE, characterised by English grammar with Nigerian lexical items such as go-slow, K-leg, flashed me), Nigerian Pidgin (NP, a creole with its own grammar and enormous reach), and code-mixed speech (CM) that interleaves two or more of these within a single utterance. Register shifts carry pragmatic weight: a speaker who moves from Standard English into Pidgin mid-sentence is frequently signalling emotional escalation, sarcasm, or solidarity â information that is invisible to models that treat all Nigerian text as a single linguistic variety. Consider the utterance âYou don try well well.â At a graduation ceremony, this is genuine, effusive praise: the speaker acknowledges sustained effort and achievement. Addressed to a mechanic who has returned a car still faulty for the third time, the identical words constitute sarcastic condemnation. The surface sentiment is positive in both cases. The communicative intent is diametrically opposed. Any system that labels both instances with the same sentiment tag has not understood either one. 2 This is the core thesis of the present work: the dominant failure mode of frontier AI systems on Nigerian public discourse is not a translation problem but a context problem. Existing Nigerian NLP benchmarks, most notably NaijaSenti (Muhammad et al., 2022) and the Nigerian Pidgin component of AfriSenti (Muhammad et al., 2023), have made foundational contributions by establishing annotated sentiment corpora for low-resource Nigerian languages. However, these benchmarks operate within a three-way polarity paradigm (positive, negative, neutral) that, by design, cannot capture the pragmatic divergence illustrated above. We introduce the Meaning Intelligence Framework (MIFâ˘), a nine-dimension annotation and evaluation schema that addresses this gap. The MIF separates surface sentiment (what a literal, context- blind reading of the words would conclude) from true intent (what the speaker actually means given the full context), and annotates both as independent dimensions alongside register, irony markers, coded subtext, risk tier, annotator confidence, speaker emotion, and recommended communications action. The signature diagnostic is the divergent item: an utterance where surface sentiment and true intent point in opposite directions, which we formalise as a computed flag and use as a dedicated evaluation metric. 2. Related Work 2.1 Nigerian Language Benchmarks NaijaSenti (Muhammad et al., 2022) introduced the first large-scale human-annotated Twitter sentiment dataset for the four most widely spoken languages in Nigeria, including approximately 14,000 Nigerian Pidgin tweets labelled for three-way polarity. AfriSenti (Muhammad et al., 2023) extended this to 14 African languages with over 110,000 tweets, including Nigerian Pidgin, and was used as the basis for SemEval-2023 Task 12. Oyewusi et al. (2021) proposed semantic enrichment for Nigerian Pidgin, observing that words like âgingerâ (motivation, not a plant) and âtankâ (gratitude, not a container) carry meanings invisible to standard English sentiment models. More recently, NaijaNLP (2025) surveyed the full landscape of Nigerian low-resource NLP, cataloguing datasets including SentiLeye, a lexicon-based sentiment analysis resource derived from 346,000 Nigerian banking-related tweets. Beyond sentiment, Saeed et al. (2024) introduced Implicit Discourse Relation Classification for Nigerian Pidgin, projecting Penn Discourse Treebank annotations onto synthetic NP data and training a dedicated classifier that outperformed zero-shot English baselines by 34% in F1. This work demonstrates that discourse-level NLP for Nigerian Pidgin is viable, but it addresses a fundamentally different task: classifying logical relations between sentences (cause, contrast, conjunction) rather than the communicative intent of a single utterance. INJONGO (Yu et al., 2025), the first large-scale multicultural intent detection and slot-filling dataset for 16 African languages, takes intent classification further â but its intent taxonomy consists of conversational AI actions (transfer, book_flight, play_music, make_call) designed for task-oriented dialogue systems, not the pragmatic communicative intents (PRAISE, SARC, MOBILIZE, GRIEVANCE_CODED) that the MIF addresses. The distinction matters: INJONGO asks what a user wants the system to do; the MIF asks what a speaker actually means. These contributions collectively establish the data infrastructure for Nigerian NLP. The MIF builds on this foundation but departs from it in a fundamental way: rather than asking âis this text positive, negative, or neutral?â (NaijaSenti, AfriSenti), or âwhat discourse relation links these sentences?â (Saeed et al.), or âwhat task does the user want to accomplish?â (INJONGO), the MIF asks âwhat does the speaker actually mean, and what should a communications professional do about it?â This reframing introduces dimensions (intent, irony, subtext, risk, action) that no existing benchmark addresses. 3 2.2 Sarcasm and Irony Detection Sarcasm detection is a substantial subfield of NLP, with established benchmarks including MUStARD, SemEval-2018 Task 3, and the Reddit SARC corpus. Recent work has explored prompting strategies for LLMs: Pragmatic Metacognitive Prompting (PMP) by Lee et al. (2024) and its context- aware extension (Iskandardinata et al., 2025) demonstrate that retrieval-augmented prompting can improve sarcasm detection by up to 9.87% macro-F1. SarcasmBench found that GPT-4 underperforms supervised fine-tuned smaller models on sarcasm, and that chain-of-thought prompting can actually hurt performance because sarcasm detection is a holistic cognitive process rather than a step-by-step logical one. The MIF treats irony as one dimension (D4) within a broader analytical framework, rather than as a standalone classification task. Critically, the MIF recognises that sarcasm in Nigerian discourse is often register-encoded: the switch from Standard English into Pidgin is itself a sarcasm marker, a phenomenon not captured by sarcasm benchmarks constructed from English-only or Global North corpora. 2.3 Cultural Context in African NLP Several recent works have highlighted the cultural gap in NLP evaluation. AfroBench (Adelani et al., 2023) provides a unified evaluation across 15 NLP tasks in African languages but does not include pragmatic or intent-level tasks. AfriStereo (2025) addresses stereotypical bias in LLMs from African perspectives. TriLex (2025) proposes a retrieval-augmented framework for sentiment lexicon expansion in South African languages. Ochieng et al. (2025) study LLM sentiment in low-resource, culturally nuanced Kenyan WhatsApp messages, noting that âthe same phrase may carry positive, neutral, or negative connotations depending on the speakerâs region and cultural background.â This observation aligns directly with the MIFâs foundational principle, but the MIF goes further by operationalising it into a scored, reproducible annotation schema. 3. The Meaning Intelligence Framework 3.1 Design Principles The MIF is built on three principles: The Context Rule. Never interpret a Nigerian expression literally without first anchoring it to the speaker, the audience, and the situation. This is not a guideline; it is the frameworkâs foundational axiom, and annotators who violate it during calibration testing are not certified. Surface and intent are separate. D2 (surface sentiment) is scored as if the annotator had no context. D3 (true intent) is scored with full context. When these diverge, the item is flagged as divergent â the frameworkâs signature diagnostic. Actionability. The frameworkâs ultimate output is not a label but a recommended action for a communications professional. Dimensions D6 (risk tier) and D9 (recommended action) translate linguistic analysis into operational intelligence. 3.2 Nine Dimensions Table 1 summarises the nine scored dimensions. The full specification, including all enumerated values, escalation rules, and scoring weights, is provided in the MIF Master Specification v2.0 (released as a companion document). 4 Table 1: MIF v2.0 dimension summary. Dim. Name Description D1 Register SE (Standard English), NE (Nigerian English), NP (Nigerian Pidgin), CM (code-mixed); plus shift flag D2 Surface sentiment Context-blind polarity of the words alone (POS / NEU / NEG) D3 True intent Speakerâs actual communicative goal given context (14 classes + CONTEXT_INSUFFICIENT) D4 Irony Binary + 7 marker types (exaggeration, emoji contrast, register shift, etc.) D5 Coded subtext Underlying grievance category if surface complaint is a proxy (POL, ETH, REL, ECON, REG, SEC, NONE) D6 Risk tier PR/regulatory risk: LOW, MEDIUM, HIGH, CRITICAL; plus 12-type risk vector D7 Confidence Annotator certainty (1â5) and context-dependent flag D8 Emotion Speakerâs dominant felt emotion (16 classes, nullable for composed speech) D9 Recommended action Communications response (12 classes mapped from risk tier and context) 3.3 Computed Flags Three flags are computed from the annotated dimensions rather than directly labelled: Divergent: true when D2 (surface) and D3 (intent) point in opposite polarity directions â specifically, when D2 = POS and D3 is in SARC, COMPLAIN, WARN, MOBILIZE, LAMENT, GRIEVANCE_CODED, or when D2 = NEG and D3 is in PRAISE, HOPE, SOLIDARITY, BANTER. Deceptive positive: true when D2 = POS and the item is divergent. This identifies utterances where positive surface language masks negative intent â the most commercially dangerous category for brand monitoring, as automated sentiment tools score them as praise. Human review required: true when D7 confidence ⤠2, or D5 contains ETH/REL/SEC, or D6 is HIGH/CRITICAL, or D3 is CONTEXT_INSUFFICIENT. Items triggering this flag enter a mandatory human review queue. 3.4 The Meaning Intelligence Score For model evaluation, we define the Meaning Intelligence Score (MISâ˘) as a weighted composite of per-dimension accuracies: MIS = 10¡literal + 25¡context + 20¡sarcasm_coded + 15¡emotion + 20¡risk + 10¡action where literal = mean(D1, D2 accuracy), context = D3 accuracy, sarcasm_coded = mean(D4, D5 accuracy), emotion = D8 accuracy, risk = D6 accuracy, and action = D9 accuracy. The weights reflect the frameworkâs priorities: contextual interpretation (25%) and risk mapping (20%) receive the highest weights because they represent the greatest gap between current AI capability and human expert performance, and carry the highest consequences for operational deployment. 4. Calibration Dataset 4.1 Construction We construct a 30-item calibration dataset of authored context-utterance pairs designed to probe specific failure modes across the MIFâs dimensions. Items are stratified by difficulty: 10 easy (clear 5 register, low ambiguity), 12 medium (register shifts, moderate irony, political subtext), and 8 hard (deceptive positives, coded mobilisation, discipline traps). Each item specifies the utterance, the context (a natural-language description of the communicative situation), an optional prior conversational turn, speaker type, target audience, and sector. The dataset includes constructed context pairs that demonstrate the frameworkâs core diagnostic: the same or similar utterance appears in two contrasting contexts (e.g., CAL-001 and CAL-002, the graduation-praise versus mechanic-sarcasm pair), allowing direct measurement of a modelâs context sensitivity. 4.2 Gold Labels Gold labels were assigned by the frameworkâs designer (the first author) across all nine dimensions, reviewed against the annotation guidelines, and re-annotated to v2.0 standards including risk vectors (D6b), emotions (D8), and recommended actions (D9). Calibration item CAL-026 (âDem don start againâ with no context provided) is a deliberate discipline trap: the correct D3 label is CONTEXT_INSUFFICIENT, and any annotator or model that assigns a definite intent class has violated the Context Rule. 4.3 Dataset Statistics Of the 30 items: 10 are divergent (D2 and D3 point in opposite directions), 9 are deceptive positive (positive surface + divergent), 6 trigger the human-review flag, and 25 carry non-null emotion labels (4 are null, representing composed institutional/promotional speech plus the discipline-trap item). The items span 7 sectors (technology, banking, politics, health, media, food, general/social) and all four registers. 5. Evaluation 5.1 Method We evaluate Gemini 2.5 Flash (model string google/gemini-2.5-flash) on all 30 calibration items under two conditions: Condition A (zero-shot): The model receives only the task description and the label inventories for each dimension, without the MIFâs interpretive guidance (the Context Rule, register definitions, intent class definitions, irony markers, the subtext proxy test, or risk escalation rules). Condition B (schema-informed): The model receives the full MIF guidance as a system prompt, including the Context Rule, register definitions with examples, intent class definitions with disambiguation guidance (e.g., the banter/insult boundary, the face-saving/genuine hope distinction), irony markers, the subtext proxy test, and risk-tier criteria. Both conditions use temperature 0, single pass, with the utterance, context, and prior turn presented as the user message. The model returns a structured JSON object with predictions for each dimension. 5.2 Scoring Per-dimension scoring uses exact match for register, sentiment, intent (primary), irony, risk tier, and action. Coded subtext (D5, a multi-label field) uses Jaccard overlap ⼠0.5 as the correctness threshold. Emotion (D8) is credited if the prediction matches either the gold primary or secondary emotion; null-emotion items are scored correct only if the prediction is null. Items with null gold values 6 for a dimension (notably CAL-026 for D4, D5, D6, D8, D9) are excluded from that dimensionâs accuracy computation. 5.3 Results Table 2 presents the full results. Table 2: Gemini 2.5 Flash performance on MIF v2.0 calibration set. Dimension Cond. A Cond. B Î Note D1 Register 33.3% 73.3% +40.0 Register Gap D2 Surface sentiment 66.7% 73.3% +6.7 D3 True intent 86.7% 86.7% 0.0 See §4.4 D4 Irony 96.6% 93.1% â3.4 D5 Coded subtext 73.3% 83.3% +10.0 D6 Risk tier 72.4% 79.3% +6.9 D8 Emotion 63.3% 63.3% 0.0 D9 Strategic action 55.2% 65.5% +10.3 Action Gap Divergent-item accuracy 80.0% 90.0% +10.0 Signature metric HIGH-risk recall 100% 100% 0.0 MIS⢠(composite) 73.2 78.6 +5.4 5.4 Analysis The Register Gap. The largest single improvement under schema-informed prompting is register classification: +40 percentage points (33.3% to 73.3%). Zero-shot, the model cannot reliably distinguish Nigerian English from Pidgin from code-mixed text. Since register shifts carry pragmatic information in Nigerian discourse (a slide from SE to NP often marks emotional escalation or sarcasm), register misclassification propagates errors into downstream dimensions. The Mobilisation Blind Spot. Item CAL-021, a mobilisation signal disguised as humour referencing a known protest junction, was misclassified as WARN in both conditions. The model correctly identified the item as HIGH risk (HIGH-risk recall was 100% across conditions), but could not identify the mechanism: it detected danger without detecting organisation. In live media monitoring, this distinction separates a watch-list entry from a crisis activation. The Action Gap. Strategic action recommendation (D9) was the weakest practical dimension at 55.2% zero-shot, improving to 65.5% under the schema. The model frequently identifies what a text means without knowing what a communications team should do about it. The MIFâs tier-to-action mapping, which encodes expert judgment about escalation thresholds and response protocols, provides the largest practical uplift for operational deployment. Intent accuracy. Zero-shot intent accuracy of 86.7% is notably high, unchanged under schema- informed prompting. Two qualifications apply. First, calibration items carry deliberately diagnostic context descriptions that are richer than the thin, noisy context typically available in real-world social media monitoring. Second, the residual errors concentrate on divergent items (80% to 90% accuracy under schema) and the highest-stakes intent classes (MOBILIZE, GRIEVANCE_CODED). The 500- item pilot corpus of real-world data will test whether this accuracy holds under naturalistic conditions. 7 6. Discussion The MIF addresses a gap in the African NLP evaluation landscape that sits between several well- served areas. Below it, NaijaSenti and AfriSenti provide polarity-level sentiment annotation at scale. Alongside it, INJONGO provides conversational AI intent detection for task-oriented dialogue, and Saeed et al. provide discourse relation classification for Nigerian Pidgin. Above it, full discourse analysis and pragmatic annotation remain the domain of theoretical linguistics. A systematic survey of the literature confirms that no existing framework separates surface sentiment from communicative intent as distinct scored dimensions for Nigerian or any African language, and no prior work uses the term âMeaning Intelligenceâ as a named concept in NLP. The MIF occupies this applied middle ground: rich enough to capture the pragmatic phenomena that matter for media intelligence and crisis communications, constrained enough to be annotated reliably by trained (but not specialist-linguist) annotators, and scored in a way that produces a single composite metric (the MIS) for model comparison. The frameworkâs commercial motivation â it was designed to power a media intelligence platform â is a feature, not a limitation. The inclusion of D6 (risk) and D9 (action) grounds the annotation in real-world consequences. A model that correctly identifies sarcasm but fails to recommend the appropriate communications response has not fully understood the text in any operationally meaningful sense. By including action recommendation as a scored dimension, the MIF measures not just comprehension but judgment. The divergent/deceptive-positive flags formalise a failure mode that has received insufficient attention in sentiment analysis research. A deceptive positive â an utterance that automated sentiment tools score as praise but which actually constitutes condemnation â is the highest-risk category for brand monitoring precisely because it is the category most likely to be missed. The MIF makes this failure mode visible and measurable. 7. Limitations This work has several important limitations that scope the claims we make and define the roadmap for future work. Single model, single pass. The evaluation covers one frontier model under two conditions with a single pass at temperature 0. The formal protocol calls for three-run majority voting and evaluation of at least two additional model families (GPT-class and Claude-class) before named model comparisons are published. We report Gemini 2.5 Flash by name because this is a methodological preprint, not a model ranking. Constructed calibration items. The 30-item calibration set uses authored context-utterance pairs with rich, unambiguous context descriptions. Real-world social media data offers thinner, noisier context, and we expect zero-shot accuracy to degrade under those conditions. The calibration set is designed for annotator certification and framework validation, not as a representative sample of Nigerian discourse. Single annotator for gold labels. Gold labels were assigned by the framework designer. While this ensures internal consistency with the frameworkâs design intent, it does not establish inter-annotator agreement. The formal annotation protocol calls for triple annotation with adjudication; inter-annotator agreement statistics will be reported once the annotation team is certified and the 500-item pilot corpus is labelled. 8 No fine-tuned model comparison. This evaluation tests prompting strategies only. A fine-tuned model trained on MIF-annotated data would likely show larger improvements; this is planned as future work. 8. Ethical Considerations The calibration dataset uses constructed examples rather than real social media posts, which avoids privacy concerns associated with reproducing identifiable user-generated content. All items are designed to be illustrative of discourse patterns without targeting real individuals, brands, or communities. The annotation guidelines include explicit provisions against using the framework to target ethnic or religious groups, and items flagged with ethnic, religious, or security subtext (D5 = ETH/REL/SEC) trigger mandatory human review. The private holdout component of the benchmark is maintained under strict access controls and non-disclosure agreements to prevent contamination of evaluation results. We release the 30-item public calibration set and the complete framework specification to support reproducibility while protecting the integrity of formal evaluation. 9. Conclusion and Future Work We have introduced the Meaning Intelligence Framework, a nine-dimension annotation and evaluation schema for Nigerian public discourse that goes beyond sentiment polarity to capture register, intent, irony, coded subtext, risk, emotion, and recommended action. The frameworkâs core insight â that AI failures on Nigerian discourse are context failures, not translation failures â is validated by the Register Gap finding: a 40-point improvement in register classification accuracy when a frontier model receives the MIF schema in-context, without any fine-tuning. Future work proceeds along four tracks: (1) a 500-item pilot corpus of real-world Nigerian discourse with triple annotation and inter-annotator agreement reporting; (2) formal multi-model evaluation under the three-run majority-voting protocol; (3) fine-tuning of a small language model on MIF-annotated data to measure the gap between prompting and training; and (4) extension of the framework to related West African Pidgins and code-mixed varieties. References Adelani, D. I., et al. (2021). MasakhaNER: Named entity recognition for African languages. Transactions of the Association for Computational Linguistics, 9, 1116â1131. Adelani, D. I., et al. (2023). AfroBench: How good are large language models on African languages? arXiv preprint arXiv:2311.07978. Iskandardinata, M., Christian, W., & Suhartono, D. (2025). Context-aware pragmatic metacognitive prompting for sarcasm detection. arXiv preprint arXiv:2511.21066. Lee, J., et al. (2024). Pragmatic metacognitive prompting improves LLM performance on sarcasm detection. arXiv preprint arXiv:2412.04509. Muhammad, S. H., et al. (2022). NaijaSenti: A Nigerian Twitter sentiment corpus for multilingual sentiment analysis. arXiv preprint arXiv:2201.08277. Muhammad, S. H., et al. (2023). AfriSenti: A Twitter sentiment analysis benchmark for African languages. arXiv preprint arXiv:2302.08956. Ochieng, M., et al. (2025). Reasoning beyond labels: Measuring LLM sentiment in low-resource, culturally nuanced contexts. arXiv preprint arXiv:2508.04199. 9 Oyewusi, W., et al. (2021). Semantic enrichment of Nigerian Pidgin English for contextual sentiment classification. In Proceedings of the AfricaNLP Workshop. Saeed, M., Bourgonje, P., & Demberg, V. (2024). Implicit discourse relation classification for Nigerian Pidgin. arXiv preprint arXiv:2406.18776. Shode, I., et al. (2023). NollySenti: Leveraging transfer learning and machine translation for Nigerian movie sentiment classification. arXiv preprint arXiv:2305.10971. Yu, H., Alabi, J. O., et al. (2025). INJONGO: A multicultural intent detection and slot-filling dataset for 16 African languages. arXiv preprint arXiv:2502.09814. Supplementary materials: The MIF Master Specification v2.0, Annotation Guidelines v1.0, and the 30-item public calibration set (with gold labels) are available as companion documents. The private holdout set is not released. MIFâ˘, MISâ˘, and Meaning Intelligence⢠are marks of AGENTPR.