Paper deep dive
SAHM: A Benchmark for Arabic Financial and Shari'ah-Compliant Reasoning
Rania Elbadry, Sarfraz Ahmad, Ahmed Heakl, Dani Bouch, Momina Ahsan, Muhra AlMahri, Marwa Elsaid khalil, Yuxia Wang, Salem Lahlou, Sophia Ananiadou, Veselin Stoyanov, Jimin Huang, Xueqing Peng, Preslav Nakov, Zhuohan Xie
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/26/2026, 10:49:31 PM
Summary
SAHM is a novel benchmark and instruction-tuning dataset designed for Arabic financial and Shari'ah-compliant reasoning. It addresses a significant gap in Arabic NLP by providing 14,380 expert-verified instances across seven specialized tasks: AAOIFI standards QA, fatwa-based QA/MCQ, accounting and business exams, financial sentiment analysis, extractive summarization, and event-cause reasoning. The benchmark evaluates 20 LLMs, revealing that while models show high proficiency in recognition-style tasks, they struggle significantly with generation and causal reasoning. The research demonstrates that domain-specific fine-tuning (e.g., SAHM-ALLAM-7B) can significantly close the performance gap compared to general-purpose models.
Entities (7)
Relation Signals (4)
SAHM → includestask → AAOIFI standards QA
confidence 100% · SAHM contains 14,380 expert-verified instances spanning seven tasks: AAOIFI standards QA...
SAHM → includestask → Financial Sentiment Analysis
confidence 100% · SAHM contains 14,380 expert-verified instances spanning... financial sentiment analysis...
SAHM-ALLAM-7B → isderivedfrom → SAHM
confidence 100% · fine-tuning on SAHM yields two complementary 7–8B models SAHM-ALLAM-7B
AAOIFI → governs → Shari'ah-compliant reasoning
confidence 90% · Frameworks such as AAOIFI and local regulations specify how financial instruments are structured...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:English financial NLP has progressed rapidly through benchmarks for sentiment, document understanding, and financial question answering, while Arabic financial NLP remains comparatively under-explored despite strong practical demand for trustworthy finance and Islamic-finance assistants. We introduce SAHM, a document-grounded benchmark and instruction-tuning dataset for Arabic financial NLP and Shari'ah-compliant reasoning. SAHM contains 14,380 expert-verified instances spanning seven tasks: AAOIFI standards QA, fatwa-based QA/MCQ, accounting and business exams, financial sentiment analysis, extractive summarization, and event-cause reasoning, curated from authentic regulatory, juristic, and corporate sources. We evaluate 19 strong open and proprietary LLMs using task-specific metrics and rubric-based scoring for open-ended outputs, and find that Arabic fluency does not reliably translate to evidence-grounded financial reasoning: models are substantially stronger on recognition-style tasks than on generation and causal reasoning, with the largest gaps on event-cause reasoning. We release the benchmark, evaluation framework, and an instruction-tuned model to support future research on trustworthy Arabic financial NLP.
Tags
Links
- Source: https://arxiv.org/abs/2604.19098v1
- Canonical: https://arxiv.org/abs/2604.19098v1
Trouble viewing inline? Open PDF directly →
Full Text
121,453 characters extracted from source content.
Expand or collapse full text
SAHM: A Benchmark for Arabic Financial and Shari’ah-Compliant Reasoning Rania Elbadry 1 Sarfraz Ahmad 1 Ahmed Heakl 1 Dani Bouch 1 Momina Ahsan 1 Muhra AlMahri 1 Marwa Elsaid Khalil 1 Mohamed Anwar 1 Yuxia Wang 2 Salem Lahlou 1 Sophia Ananiadou 3 Veselin Stoyanov 1 Jimin Huang 4 Xueqing Peng 4∗ Preslav Nakov 1 Zhuohan Xie 1 1 MBZUAI 2 INSAIT 3 The University of Manchester 4 The Fin AI rania.elbadry, preslav.nakov, zhuohan.xie@mbzuai.ac.ae ƁSAHM BenchmarkaCode Abstract English financial NLP has advanced rapidly through benchmarks targeting earnings analy- sis, market sentiment, tabular reasoning, and financial question answering, yet Arabic finan- cial NLP remains virtually nonexistent, de- spite 422 million speakers, $4.9 trillion in Gulf sovereign wealth, and a $4–5 trillion Is- lamic finance industry requiring specialized Shari’ah compliance over instruments like sukuk, murabaha, and takaful. We introduce SAHM, the first Arabic financial benchmark spanning seven tasks: AAOIFI standards QA, fatwa-based QA/MCQ, accounting and busi- ness exams, financial sentiment analysis, ex- tractive summarization, and event-cause rea- soning, comprising 14,380 expert-verified in- stances from authentic regulatory, juristic, and corporate sources. Evaluating 20 LLMs, we find Arabic fluency does not imply financial reasoning: models achieving 91% on recog- nition tasks drop sharply on generation, and event-cause reasoning exposes the widest per- formance gap (1.89–9.84/10). We release the benchmark and dataset to support trustworthy Arabic financial assistants. 1 Introduction The Gulf Cooperation Council (GCC) generates large volumes of Arabic financial text, including central bank reports, regulatory filings, corporate disclosures, andfatwasthat provide jurispruden- tial rulings. Despite this, systematic evaluation of Large Language Models (LLMs) on Arabic fi- nancial content remains limited. English finan- cial NLP has advanced rapidly through dedicated benchmarks ( Maia et al.,2018a;Zhu et al.,2021; Chen et al.,2021,2022;Zhao et al.,2024;Xie et al.,2025), with multilingual extensions emerg- ing for other languages ( Nie et al.,2025;Zhang et al.,2024;Peng et al.,2025a,b). Arabic bench- ∗ Corresponding author marks remain limited in scope: ArBanking77 (Jar- rar et al. ,2023) addresses only banking intent, and Arabic-centric LLMs (Sengupta et al.,2023;Team, 2025;Heakl et al.,2025a;Abbas et al.,2025) have not been evaluated on financial domains. Is- lamic finance further illustrates this gap. Unlike conventional finance, it requires Shari’ah review guided by standards issued by AAOIFI. 1 Although resources such as Fatwaset (Alyemny et al.,2023) and Hajj FQA ( Aleid and Azmi,2025) exist, they focus on general juristic QA rather than financial reasoning. As a result, LLMs remain untested on tasks that combine legal and financial analysis. We introduce SAHM, the first Arabic financial NLP benchmark unifying modern finance and Islamic jurisprudence two high-stakes domains shaping trillions in assets yet missing from LLM evaluation. SAHM spans seven expert-verified tasks grounded in AAOIFI standards, fatwa archives from seven countries, and corporate disclosures (Figure 1). Evaluating 20 LLMs reveals that Arabic fluency does not guarantee financial reasoning base Arabic models rank in the bottom 25% despite being de- signed for Arabic. However, fine-tuning on SAHM dramatically closes this gap: domain-adapted mod- els gain up to +26 points on Accounting and +25 points on Business, enabling 7–8B models to sur- pass GPT-5 on financial reasoning tasks and match 72B open-source baselines. Our contributions: •The first Arabic finance benchmark (14,380 in- stances; 7 tasks) jointly evaluating Shari’ah- compliant reasoning (fatwa QA, Islamic fi- nance standards) and core financial competen- cies (accounting MCQ, sentiment, event-cause QA), addressing a major resource gap for Ara- bic financial NLP. •A comprehensive benchmark of 20 LLMs show- ing that Arabic fluency does not guarantee fi- 1 https://aaoifi.com arXiv:2604.19098v1 [cs.CL] 21 Apr 2026 Accounting Exams MCQ قامت إحدى الشركات بشراء قطعة أرض لبناء مبنى إدارى علیھا٢٠١٨/١/١السؤال: في وبلغ المتوسط المرجح لنفقات المتراكمة١/١وبدأت الشركة في بناء المبنى الإدارى في . ولتمویل إنشاء١٢/٣١ جنیھ. وأكتمل المبنى وأصبح جاھز لإستخدام في ٩٥٠ سنوات٣ على ورقة دفع لمدة ٢٠١٨/١/١ جنیھ في ٧٠المبنى إقترضت مبلغ من كل عام، وكان على الشركة التزام آخر متمثل في١٢/٣١٪ سنویا تسد في ١٠بمعدل ٪ سنویا محرة في١٢ سنوات بمعدل فائدة ٦ جنیھ لمدة ٤٠ورقة دفع بمبلغ من كل عام. المطلوب. تبلغ الفوائد المرسملة في١/١ وتسد الفائدة في ٢٠١٧/١/١ الخیارات ...... ٢٠١٨/١٢/٣١ ١٨٠أ. (a ١٠ب. (b ١٨٠ج. (c د. لا شئ ما سبق (d a :الإجابة Bussiness Exams MCQ السؤال: تصمیم عروض المنتجات والأسعار والتوزیع الترویجیة والخدمة بشكلوالجھود یتناسب مع السوق یطلق علیھ استراتیجیة التسویق (a إدارة السوق (b إدارة المزیج التسویقي (c a :الإجابة Islamic Financial Fatwa QA ن أنھ عبارة عن عقد تأمین تجاري تقوم العلاقة فیھ على ّ بعد دراسة ھذا النظام، تبی م في قول جمھور َّ ن وشركة التأمین مباشرة.والتأمین التجاري محر ِ ّ المعاوضة بین المؤم ا َ یالفقھاء والمجامع الفقھیة؛ لأنھ من العقود المبنیة على المقامرة والمیسر. قال ﷲ تعالى: ) ِ ان َ ط ْ ی َّ الش ِ ل َ مَ ع ْ ن ِ م ٌ س ْ ج ِ ر ُ م َ لا ْ ز َ ْ الأ َ و ُ اب َ ص ْ ن َ ْ الأ َ و ُ ر ِ س ْ ی َ م ْ ال َ و ُ ر ْ م َ خ ْ ا ال َ م َّ ن ِ وا إ ُ ن َ آم َ ین ِ ذ َّ ا ال َ ھ ُّ ی َ أ ا خاصة بالقروض الربویة؛ حیث ً .كما أنھ یتضمن بنود90( المائدة/ َ ون ُ ح ِ ل ْ ف ُ ت ْ م ُ ك َّ ل َ ع َ ل ُ وه ُ ب ِ ن َ ت ْ اج َ ف ا على الوثیقة... ویحق لشركة ً نص النظام على أنھ: "توافق الشركة على أن تمنح قرض إعادة تحدید معدل الفائدة على القروض القائمة والجدیدة".وعلى ھذا، یحرم الاشتراك في مثل ھذا التأمین حتى یتم تبدیل نظامھ بنظام التأمین التعاوني الإسلامي أرجو بیان الحكم الشرعي في أحد عقود التأمین لغایات التعلیم؟ Islamic Finance Shari’ah QA في المعاملات المالیة الإسلامیة، في موضوع التامین الإسلامي: ھل یجوز اقتطاع احتیاطیات من أموال حملة الوثائق، وما ھو مصیرھا عند تصفیة الشركة؟ یجوز تحقیقا لمصلحة حملة الوثائق أن یقتطع جزء من أموالھم، أو أرباحھا احتیاطیات، أو مخصات متعلقة بصندوق التأمین على ألا تؤول إلى المساھمین، وما یتراكم في حساب التأمین یصرف في وجوه الخیر عند التصفیة a : الإجابة Event Causal Reasoning QA 1.25 أضعاف لإصدار صكوك صندوق الاستثمارات العامة بقیمة 6.5ما الذي تعكسھ نسبة التغطیة التي تجاوزت ملیارات دولار أمریكي، في ظل تصنیفات9ملیار دولار أمریكي، وحجم طلبات الاكتاب التي بلغت أكثر من الصندوق الائتمانیة القویة واستراتیجیتھ التمویلیة المتنوعة، ودوره كمحرك لتحول الاقتصادي؟ مان - مباشر:یوزع بنك نزوى ُ التقریر: ع(BKNZ)، لریاض - مباشر: أعلن صندوق الاستثمارات العامة إتمام تسعیر طرحھ لصكوكا م بالدولار ّ (، حیث سیتم توجیھ عائدات الطرح المقو ً ملیار ﷼ سعودي تقریبا4.7 ملیار دولار أمریكي )ما یعادل 1.25بقیمة أضعاف6.5وأوضح الصندوق، في بیان صحفي صادر، الیوم الخمیس، أن نسبة التغطیة تجاوزت لأغراض الصندوق العامة. ملیار ﷼33.7 ملیارات دولار أمریكي )ما یعادل 9إجمالي الإصدار، في حین وصل المجموع الكلي لطلبات الاكتاب إلى أكثر من لاستراتیجیة الصندوق التمویلیة القویة والمتنوعة، التي تحظى بدعم ً ل الطرح استمرارا ّ وتابع: "یشك [..................سعودي تقریبا(. ] كبیر من المستثمرین الدولین."وتشمل استراتیجیة صندوق الاستثمارات العامة التمویلیة طویلة الأجل مجموعة متنوعة من الأدوات ویحمل الصندوقارة. ّ التمویلیة، بما في ذلك برنامجي الصكوك والسندات، إلى جانب تمویل بھیكلیة المرابحة وتسھیلات ائتمانیة دو عند الفئة ً ائتمانیا ً تصنیفا Aa3 مع نظرة مستقبلیة "مستقرة" من وكالة التصنیف الائتماني العالمیة مودیز (Moody’s)، كما یحمل من فئة ً تصنیفا A+ من وكالة "فیتش" مع نظرة "مستقرة" Report النسبة المرتفعة لطلب الثقة الكبیرة من قبل المستثمرین الدولین في الملاءة المالیة والقوة الائتمانیة لصندوق الاستثمارات العامة. ھذه الثقة مدعومة بتصنیفاتھ الائتمانیة المتمیزة: الفئة Aa3 مع نظرة مستقبلیة "مستقرة" من وكالة مودیز، والفئة A+ من وكالة فیتش مع نظرة "مستقرة". تشیر ھذه التصنیفات الائتمانیة العالیة إلى انخفاض مخاطر الائتمان لدى الصندوق، ما یجعلھ وجھة جاذبة لمستثمرین الباحثین عن استثمارات مستقرة وآمنة في أصول ذات جودة عالیة Report Extractive Summarization مان - مباشر:یوزع بنك نزوى ُ التقریر: ع(BKNZ)، المدرج بورصة مسقط، صكوك مضاربة مجانیة على مساھمیھ، 'دبي ـ مباشر: نجحت شركة بن غاطي لتطویر العقاري، في طرح المزید من الصكوك في اكتابفي نھایة جلسة الیوم وفق بیان صحفي صادر الیوم الأربعاء، یرتبط ھذا الطرح بإصدار المطور .2024 یولیو 8یوم الاثنین الموافق ملیون دولار أمریكي، بورصة لندن وناسداك دبي، وتجاوز حجم300 بقیمة 2024لصكوك لأول مرة في فبرایر ملیون500وبالطرح الأخیر یصل الحجم الإجمالي لصفقة صكوك بن غاطي إلى بالمائة. 200الاكتاب فیھا نسبة ا تسجیلات بلغت ً ا من المستثمرین الإقلیمین والدولین، محق ً متاز ً شھدت عملیة الاكتاب الأخیرة إقبالادولار أمریكي و بالمائة واستحقاق في عام9.625ویشتمل ھذا الإدراج على عائد بقیمة [..............] .ا عن الاكتاب السابق ً ضعف4.2 ، ما یعكس التخطیط المالي الإستراتیجي لدى شركة "بن غاطي"، والتزامھا بتقدیم فرص استثماریة2027 نقطة أساس30وفاقت طلبات الاكتاب المستوى المستھدف أكثر من مرتین ، مع تسعیرھا بشكل تنافسي بفارق جذابة. على الحملات الترویجیة التي نظمتھا "بن غاطي" لمستثمرین ً نتیجة لطلب المستثمرین بناء قم بتلخیص التقریر المالي التالي باستخدام التلخیص الاستخراجي Report ملیون دولار أمریكي. وفق500وبالطرح الأخیر یصل الحجم الإجمالي لصفقة صكوك بن غاطي إلى بیان صحفي صادر الیوم الأربعاء، یرتبط ھذا الطرح بإصدار المطور لصكوك لأول مرة في فبرایر 9.625 ملیون دولار أمریكي، بورصة ویشتمل ھذا الإدراج على عائد بقیمة 300 بقیمة 2024 ، ما یعكس التخطیط المالي الإستراتیجي لدى شركة "بن غاطي"،2027بالمائة واستحقاق في عام ا من المستثمرین ً متاز ً والتزامھا بتقدیم فرص استثماریة جذابة. وشھدت عملیة الاكتاب الأخیرة إقبالا ا عن الاكتاب السابق ً ضعف4.2ا تسجیلات بلغت ً الإقلیمین والدولین، محق Financial Report Sentiment Analysis مان - مباشر:یوزع بنك نزوى ُ التقریر: ع(BKNZ)، المدرج بورصة مسقط، صكوك مضاربة مجانیة على مساھمیھ، في نھایة جلسة الیوم سھم100 صك لكل 4.47وأوضح البنك أن الصكوك بواقع الثلاثاء. ومولة من الأرباح المحتجزة لبنك، بعد موافقة البنك المركزي العماني في مایو الجاري على الشریحة25 مایو الجاري وھیئة الخدمات المالیة في 19 ملیون9.9الثانیة من الصكوك الإلزامیة القابلة لتحویل بقیمة إجمالیة قدرھا ولفت البنك إلى أن الشریحة الثانیة . ً % سنویا6﷼، وبمعدل ربح متوقع 50ھي جزء من برنامج الصكوك المعتمد العام الماضي وبقیمة إجمالیة تبلغ في ً وسبق ذلك تحقیق البنك نموا [..........]لبنك. %100ملیون ﷼ أرباح بنك نزوى ترتفع النصف الأولوك أرباحھ السنویة خلال العام بتقلیص حجم المخصات Report :اقرأ بعنایة التقریر المالي التالي واختر التصنیف الصحیح. a) Positive b) Negative c) Neutral :الخیارات c : الإجابة Figure 1:Examples of the diverse tasks included in SAHM, covering juristic Q&A, business and accounting MCQs, financial sentiment analysis, report summarization, & event causal reasoning. nancial reasoning: models that score up to 91% on MCQ-style tasks degrade substantially on open-ended generation, with the largest gap on Event–Cause QA (1.89–9.84/10). •Evidence that targeted adaptation rivals scale for Arabic financial NLP: fine-tuning on SAHM yields two complementary 7–8B models SAHM- ALLAM-7B (peak accuracy, surpassing GPT- 5 by +21.3 points on Business MCQ, 93.99% vs. 72.68%) and SAHM-JAIS-8B (uniformly pos- itive transfer across all tasks) while matching 72B open-source baselines on average demon- strating∼10×parameter efficiency and estab- lishing domain adaptation as a practical, cost- effective route to trustworthy Arabic financial assistants where frontier API access may be lim- ited. 2 Related Work Financial NLP Benchmarks.English financial NLP has matured through progressively challeng- ing benchmarks. Early work focused on classifi- cation and extraction ( Araci,2019), while recent datasets target numerical reasoning over tables (FinQA (Chen et al.,2021), TAT-QA (Zhu et al., 2021)), multi-turn dialogue (ConvFinQA (Chen et al.,2022)), and chain-of-thought verification (FinChain (Xie et al.,2025)). Comprehensive suites such as FinBen (Xie et al.,2024) and PIXIU (Xie et al.,2023) now span 24 tasks in- cluding sentiment, NER, and argument mining. Multilingual extensions have emerged for Chinese (CFinBench ( Nie et al.,2025)), and Greek (Plu- tus (Peng et al.,2025a)), demonstrating that cul- turally grounded evaluation reveals failure modes invisible in English-only testing. Yet Arabic, spo- ken by 422M people across economies managing $4.9T in sovereign wealth (Alhajraf,2025), lacks any comparable financial benchmark. Arabic NLP and the Evaluation Gap.Arabic resources have grown substantially, but remain shallow in financial coverage. ArBanking77 (Jar- rar et al.,2023) addresses banking intent detec- tion; Fatwaset ( Alyemny et al.,2023) and Hajj- FQA (Aleid and Azmi,2025) target religious QA. These datasets support general language under- standing, but do not evaluate the reasoning re- quired for regulatory compliance, numerical anal- ysis, or Shari’ah-aligned decision making. This gap is significant as Arabic financial texts present distinct challenges: mixed numeral systems (East- ern٠١٢٣and Western 0123), code-switching with English acronyms (IFRS, AAOIFI), and domain- specific terminology from Islamic jurisprudence (riba,gharar,sukuk). Meanwhile, Arabic-centric LLMs, including Jais ( Sengupta et al.,2023), Falcon-Arabic (Team,2025), AIN (Heakl et al., 2025a), and Fanar (Abbas et al.,2025), are evalu- ated only on generic benchmarks that ignore these complexities. 3 SAHM We introduce SAHM, a comprehensive benchmark for evaluating Arabic financial reasoning across di- verse, real-world tasks spanning Islamic finance, accounting, and market analysis. The benchmark is designed to capture both rule-based reasoning grounded in Shari’ah standards and applied finan- cial understanding in authentic Arabic contexts. It Chapter Splitting Human Verification OCR Islamic Books QA Generation Text Files المال یة الإس لام یة، فيعاملات في الم موضوع الوقف: ما ھي م ھام ال ناظر المت عل قة حما یة حقوق الوقف وأداء التز اما تھ؟ب لد فاع عن حقوق الوقف و الح فاظ عل یھ ود فعـ ا على الوقفالمرفو عة وى لاء الد عاأجور وك فات توثیق أع یا نھ وحقو قھومصرو أداء دیون الوقف - أداء حقوق المستحقین - حمایت ھا منئمة والع نا یة بالأو قاف ال قاـ غصب ھالاء علی ھا أو الاستی Page Images المعایر الشرعیة ـ ا، أم سلعة، ً سواء كانت نقد أم منفعة )خدمة(، ویجوز بمبلغ ثابت، أوأن تكون متغیر قائم على طریقة ٢/٢/٥معلومة لطرفین . یجوز تحدید الأجرة على جمیع العمل بحیث تستحق ]......[ المدة التي لمكاملة یحصل الانتفاع فیھا بالخدمة(، أما أجرة الفترات المعایر الشرعیة ـ ا، أم سلعة، ً سواء كانت نقد أم منفعة )خدمة(، ویجوز بمبلغ ثابت، أوأن تكون متغیر قائم على طریقة ٢/٢/٥معلومة لطرفین . یجوز تحدید الأجرة على جمیع العمل بحیث تستحق ]......[ المدة التي لمكاملة یحصل الانتفاع فیھا بالخدمة(، أما أجرة الفترات المعایر الشرعیة ـ ا، أم سلعة، ً سواء كانت نقد أم منفعة )خدمة(، ویجوز بمبلغ ثابت، أوأن تكون متغیر قائم على طریقة ٢/٢/٥معلومة لطرفین . یجوز تحدید الأجرة على جمیع العمل بحیث تستحق ]......[ المدة التي لمكاملة یحصل الانتفاع فیھا بالخدمة(، أما أجرة الفترات Gemini Gemini Human Verification Figure 2:Pipeline for constructing the Islamic Finance Shari’ah Standards QA dataset.A hybrid LLMs- human pipeline converts AAOIFI standards into QA pairs through OCR and generation stages, each followed by expert verification to ensure linguistic accuracy and legal fidelity. TaskDatasetNAvg. Words (Input)Avg. Chars (Input)Avg. Words (Answer)Avg. Chars (Answer) MCQ Accounting Exams MCQ167111.5±91.1674.3±550.51.0±0.01.0±0.0 Business Exams MCQ18346.3±12.2298.3±71.61.0±0.01.0±0.0 Islamic Financial Fatwa MCQ2,00093.1±14.7536.7±82.61.0±0.01.0±0.0 Financial Report Sentiment Analysis MCQ 80292.3±139.31,780.7±841.91.0±0.01.0±0.0 Open-Ended Event–Cause Reasoning QA80413.6±299.92,503.7±1,752.1350.6±101.82,170.4±635.8 Islamic Fatwa QA2,00064.1±36.2377.3±200.689.9±58.4492.5±324.0 Islamic Sharī'a Standards QA811140.1±5.2287.0±39.533.2±22.0192.1±129.8 Report Extractive Summarization80355.4±165.42,144.3±972.3157.4±66.5929.1±391.7 Table 1:Dataset statistics for SAHM.Mean±standard deviation of word and character counts per instance, computed over the test split of each dataset. For MCQ tasks the answer is a single letter (A–D), hence the constant 1.0 word/char count. covers multiple task formats, including question answering, multiple-choice reasoning, sentiment analysis, and summarization, enabling holistic as- sessment of model capabilities. Table1provides an overview of dataset composition, task distribu- tion, and train–test splits. 3.1 Islamic Finance Shari’ah Standards QA Finance in the Gulf and the wider MENA region differs from Western systems: banks, insurers, and capital markets must comply with Islamic prin- ciples governed by detailed Shari’ah standards. Frameworks such asAAOIFIand local regula- tions specify how financial instruments are struc- tured, e.g., lease-to-own arrangements inIjara (إ༥؇رة) and compliance requirements for Sukuk 2 issuance (Pomeranz,1997;Islamic Financial Ser- vices Board (IFSB),2024;Saudi Central Bank, 2024). Yet, most financial benchmarks implic- itly assume Western instruments such as interest- bearing loans and conventional bonds, leaving models untested on regionally critical reasoning about contract permissibility, legal constraints, and Shari’ah compliance. To address this gap, we construct the first Islamic Shari’ah Standards QA dataset directly from the official1,264-page AAOIFI compendium spanning52standards chap- 2 ݬܝިك(sukuk) are Shari’ah-compliant financial certifi- cates representing ownership in underlying assets rather than interest-bearing debt. ters, enabling systematic evaluation of LLMs on rule-based Islamic financial reasoning. We built the dataset through a multi-step pipeline that converts the AAOIFI compendium into text via OCR withGemini-2.5-Pro( Google Cloud,2025) (AppendixAprovides details) rec- ommended by Heakl et al.(2025b). Two native Islamic finance experts manually verified the ex- tracted text to preserve diacritics, numerals, and domain-specific terminology. In a review of a25% sample, the experts mea- sured a high exact-match rate of98.7±0.7% with a 95% confidence interval and strong inter-annotator agreement (κ= 0.962), confirming the reliabil- ity of the OCR pipeline. The remaining 1.3% of mismatched characters consisted primarily of mi- nor orthographic or formatting issues (e.g., spac- ing, punctuation, and occasional diacritics), which annotators corrected in the canonical text used for QA construction; in the audited sample, we did not observe OCR errors that altered the substance of any Shari’ah ruling. After cleanup, we grouped the verified text into thematic clusters, e.g., Murabaḥa (cost-plus sale) and usedGemini-2.5-Proto draft candidate Arabic question–answer pairs. Domain experts then refined and validated these samples to ensure that each question accurately captured its intended ruling and included all mandatory con- ditions and exceptions. This human-in-the-loop pipeline transforms dense regulatory prose into high-quality, legally faithful QA pairs for bench- marking Shari’ah-compliant financial reasoning. Examples from the dataset and the full pipeline ap- pear in Figures1and2, respectively. 3.2 Islamic Financial Fatwa QA We scraped fatwā archives from13official web- sites across7Arab countries to capture the breadth of real-world financial questions Muslims ask (Ta- ble5). The initial crawl yielded20kfatwas, which we cross-checked against the public Fat- waSet (Alyemny et al.,2023) to remove duplicates and then organized into 11 finance-related cate- gories (Table7), includingزႤ၍ة(almsgiving),رً؇ (usury), and۰ොໍາاਵਦ(cost-plus financing), then trans- formed these long, formal texts into concise QA pairs viaGemini-2.5-prowhile preserving their juristic meaning. Specifically, we removed intro- ductory invocations (e.g.,“ا৵ৠڎՄ៰Ղ ، واܳݱఈఃة واܳފఈఃم আॻ༟ رݿިل ا”Մ៰Ղ) and rhetorical openers (e.g.,“أ݁؇ ًأڎ”) to expose the core inquiry and ruling. We stripped HTML artifacts and redundant navigational refer- ences while retaining key metadata such as source URLs for traceability. This pipeline removes greet- ings, honorifics, hyperlinks, and scholar names while preserving Qurʾānic citations, juristic termi- nology, and legal reasoning. Further details in Ap- pendix B. Two native Arabic speakers manually re- viewed 10% of the normalized data from each cat- egory to verify clarity, linguistic fidelity, and do- main correctness. This process resulted in exactly 9,953high-quality training samples and2,000 held-out finance-focused test cases (Figure 1). Afterwards, we converted each test QA pair into multiple-choice (MCQ) format viaGemini- 2.5-Pro, enabling both open-ended fatwā reason- ing and recognition-style testing. Each MCQ consists of one correct answer derived from the source fatwā and three plausible distractors reflect- ing common misconceptions. Two native Ara- bic annotators independently reviewed the test set to assess MCQ correctness, alignment with the source fatwā, and distractor plausibility. The an- notators achieved high agreement (Cohen’sκ= 0.89). Following this pilot phase, we conducted a calibration round in which annotators discussed disagreements, resolved ambiguous cases, and re- fined shared labeling criteria. One annotator then validated the remaining MCQs under the cali- brated guidelines, ensuring that each item precisely matched its source fatwā, preserved juristic termi- nology, and avoided misleading options. A final audit confirmed that95% of MCQs aligned exactly with their original QA pairs; we discarded the re- maining5% and excluded them from evaluation. 3.3 Business & Accounting Exams MCQ Professional accounting assessment resources re- main largely English-centric, with key certifica- tions such as the CPA exam conducted exclusively in English. To address this gap and the limited availability of Arabic training materials despite the existence of IFRS translations, we design cultur- ally and linguistically adapted MCQ samples cov- ering IFRS treatments, financial ratios, budgeting, and costing, incorporating authentic Arabic finan- cial terminology such as݁أڎل دوران اݬިل(asset turnover ratio) andزႤ၍ة اܳႤ၍དྷت(corporate almsgiv- ing) within contextually accurate scenarios rather than direct translations of Western exam questions (AppendixD). We constructed the dataset by col- lecting 10 business exams and 8 accounting exams from multiple Arabic-speaking countries. We ex- tract the text from the exam PDFs viaGemini-2.5- ProfollowingHeakl et al.(2025b), after which two native Arabic-speaking annotators reviewed by comparing the OCR output against the original questions, correcting recognition errors, and vali- dating formatting. The final dataset contains457 business questions and416accounting questions, examples in Figure 1. 3.4 Financial Report Sentiment Analysis Despite managing trillions of dollars in assets, Arabic financial markets lack sentiment bench- marks tailored to region-specific financial dis- course. Existing English datasets ( Maia et al., 2018b) focus on Western market narratives and do not capture signals central to MENA markets, including OPEC+ production decisions,ݬܝިك (sukuk) issuances, subsidy reforms, and Shari’ah- compliance rulings. These challenges are am- plified by the use of culturally grounded finan- cial terminology, e.g.,۰ොໍາاਵਦ(cost-plus financing) and stylistic variation in Arabic financial reporting, where subtle modifiers can reverse sentiment po- larity. To address this gap, we construct the first Arabic financial sentiment benchmark based on au- thentic market reports rather than translated prox- ies. We collect200Arabic financial reports: 100 Islamic finance–focused and 100 general from Ar- gaam 3 , and annotate them with three document- level sentiment labels:Positive,Negative, and Neutral. Two native Arabic annotators labeled all reports using a custom web-based annotation plat- form (Figure15) following shared guidelines that emphasize holistic document interpretation rather than sentence-level cues resolving in a high inter- annotator agreement (Cohen’sκ= 0.91). We then conducted a calibration phase where annotators re- solved disagreements and refined decision criteria. For mixed-signal reports, we assign the dominant polarity if >60% of content supports it; otherwise Neutral. A third expert adjudicates residual dis- agreements. We split the dataset into120training and80test reports. 3.5 Report Extractive Summarization Extractive summarization is critical for Arabic fi- nancial reporting, where annual reports are writ- ten in Arabic but frequently contain mixed numeral systems, embedded English financial acronyms and brand names rendered in Arabic script (e.g., اৎأ؇لଫଃ اᄴᄟوܳ٭۰ ይዧٺگ؇رߌߵ اৎ؇ܳ٭۰/ IFRS andང ሒᇀ إس ྸإ / HSBC), and specialized Islamic finance termi- nology such asݬܝިك(sukuk). Misinterpreting or omitting these elements can distort regulatory in- terpretation, compliance assessment, and financial valuation. To support this task, we compile200 Arabic financial reports,100general and100Is- lamic from Argaam and annotate them with ex- tractive summaries written in Arabic by two na- tive Arabic speakers. Rather than treating sum- marization as a subjective agreement task, we use ROUGE ( Lin,2004) to measure overlap between independently produced summaries as a consis- tency check and select the more complete sum- mary as the gold reference. We split the dataset into120training reports and80test reports. Fur- ther details in Appendix F. 3.6 Event-Cause Reasoning QA Financial event-cause reasoning is underexplored in Arabic due to the lack of datasets that re- quire models to explain why financial or regula- tory events occur and what implications they en- tail. To address this gap, we introduce an event- cause reasoning task that evaluates whether models can analyze Arabic financial reports and produce analytical explanations grounded in reported finan- cial data, including market movements andݬܝިك 3 https://w.argaam.com/ issuances. We collect200Arabic financial reports: 100Islamic finance–focused and100general from Argaam. Two native Arabic financial experts anno- tate each report by formulating one analytical ques- tion that links multiple reported data points and by writing a concise expert answer explaining the un- derlying causes and implications using only infor- mation in the article. A pilot phase on20reports ensures guideline clarity. We provide further de- tails in AppendixG. We assess annotation quality at two levels: Cohen’sκ= 0.86measures agree- ment on event-cause identification, while ROUGE overlap serves as a consistency check for indepen- dently written answers. After calibration to re- solve disagreements and align on edge cases, one expert completes the remaining annotations under the agreed criteria. 4 Experiments Evaluated Models.We evaluated20models spanning Arabic-centric models ( Bari et al.,2024; Abbas et al.,2025;silma-ai,2024) (publicly avail- able instruction-tuned systems for regional adap- tation), open-weight models (Riviere et al.,2024; Kamath et al.,2025;Grattafiori et al.,2024;Yang et al.,2024;Jiang et al.,2024) (strong multilingual and general-purpose baselines), and proprietary models ( Hurst et al.,2024;OpenAI,2025;An- thropic ,2025b,c,a;Google,2024), enabling con- trolled analysis across language, scale, and capabil- ity dimensions. To assess whether domain-specific fine-tuning can close the gap between Arabic- centric and frontier models, we fine-tune three Ara- bic LLMs (ALLAM-7B, Jais-2-8B, and SILMA- 9B) on the SAHM training split using LoRA (r=64, α=128, lr=2e-4, 3 epochs). Detailed model speci- fications are provided in Table 6. We evaluate Accounting Exams, Business Ex- ams, Fatwa MCQ, and Financial Sentiment with exact-match accuracy, normalizing free-form out- puts (e.g., option text/letters) to a single choice be- fore scoring (AppendixH). For extractive summa- rization, we report ROUGE-F1 (ROUGE-1/2/L) against gold extractive references (models are in- structed to output verbatim sentences). For Fatwa QA, Shari’ah Standards QA, and Event-Cause QA, we useGemini-2.5-Flashas an LLM-as-a- judge (blind to model identity): given the origi- nal Arabic prompt, gold reference, and model an- swer, it returns a JSON-validated, additive (sum- of-components)[0,10]score under a shared rubric Model MCQ (Accuracy %↑)Open-Ended QA (Score 0–10↑) DatasetsDatasets Accounting BusinessFatwā Sentiment Mean Event-Cause QA Islamic-Standards-QA Fatwa-QA Mean Open-source Models:≥70B Parameters Qwen2.5-72B-Instruct (Yang et al.,2024) 65.87 ±2.70 74.86 ±0.32 84.65 ±0.33 75.00 ±1.25 75.108.1000 ±0.10 5.6330 ±0.10 5.3912 ±0.06 6.3747 LLaMA-3.1-70B (Grattafiori et al.,2024) 52.10 ±2.79 77.60 ±1.14 84.90 ±0.15 80.00 ±3.31 73.656.623 ±0.15 3.7245 ±0.10 4.7607 ±0.08 5.036 Open-source Models:<70B Parameters Qwen2.5-14B-Instruct (Yang et al.,2024) 49.10 ±3.93 63.39 ±0.83 76.05 ±0.85 57.50 ±3.82 61.517.4975 ±0.10 4.8806 ±0.10 4.0576 ±0.06 5.4786 Qwen2.5-7B-Instruct (Yang et al.,2024)48.50 ±2.85 59.56 ±1.14 70.00 ±0.28 55.00 ±1.91 58.276.1038 ±0.12 3.4039 ±0.10 2.6815 ±0.08 4.0631 Gemma-2-9B-IT (Riviere et al.,2024)49.10 ±2.74 63.39 ±3.83 66.60 ±0.61 55.00 ±1.44 58.527.1438 ±0.08 4.2306 ±0.08 3.4266 ±0.06 4.9336 Gemma-3-27B-IT (Kamath et al.,2025)53.89 ±2.16 73.22 ±0.32 80.65 ±0.18 80.00 ±0.72 71.948.7188 ±0.05 6.1708 ±0.08 5.1929 ±0.05 6.6942 Gemma-3-4B-IT (Kamath et al.,2025)38.32 ±2.27 67.76 ±0.32 61.35 ±0.18 75.00 ±1.44 60.617.4075 ±0.08 2.8985 ±0.08 2.4767 ±0.06 4.2609 LLaMA-3.1-8B (Grattafiori et al.,2024)41.92 ±3.28 60.66 ±4.45 64.05 ±3.62 73.75 ±5.77 60.604.9231 ±0.18 2.5168 ±0.12 1.4025 ±0.08 2.9475 Mixtral-8x7B-Instruct (Jiang et al.,2024) 32.93 ±1.04 60.66 ±0.63 62.15 ±0.34 70.00 ±0.72 56.444.5538 ±0.08 2.4980 ±0.08 1.7896 ±0.06 2.9471 Proprietary Models: Reasoning-Enhanced GPT-5 (OpenAI,2025)65.27 ±2.27 72.68 ±1.26 90.75 ±0.45 78.75 ±1.25 76.869.6831 ±0.03 8.7965 ±0.05 8.0515 ±0.04 8.8437 GPT-4o (Hurst et al.,2024)60.48 ±2.07 78.14 ±0.32 87.70 ±0.10 77.50 ±0.00 75.968.3125 ±0.06 6.6598 ±0.08 6.5219 ±0.04 7.1647 Proprietary Models: General-Purpose Claude-Opus-4.5 (Anthropic,2025b)77.84 ±2.42 76.50 ±1.14 91.75 ±0.33 75.00 ±2.50 80.279.6818 ±0.03 8.0438 ±0.05 8.8090 ±0.03 8.8449 Claude-Sonnet-4.5 ( Anthropic,2025c)78.44 ±1.20 76.50 ±1.45 88.15 ±0.38 77.50 ±1.25 80.159.3388 ±0.04 8.2588 ±0.05 7.6049 ±0.03 8.4008 Claude-Haiku-4.5 (Anthropic,2025a)67.66 ±1.80 73.77 ±1.30 84.90 ±0.40 77.50 ±1.80 75.969.1050 ±0.05 7.0002 ±0.07 6.5341 ±0.05 7.5464 Gemini-3-Flash (preview) (Google,2024) 76.05 ±1.95 74.86 ±0.95 89.90 ±0.30 81.25 ±1.25 80.529.8369 ±0.02 9.1649 ±0.03 9.1571 ±0.02 9.0798 GPT-4o-mini (Hurst et al.,2024)58.08 ±2.10 77.60 ±0.40 81.75 ±0.20 75.00 ±0.50 73.617.9613 ±0.08 5.6094 ±0.10 5.3087 ±0.06 6.2931 Arabic Models ALLAM-7B (Bari et al.,2024)44.91 ±3.55 68.31 ±3.83 74.40 ±2.83 58.75 ±2.00 61.596.8875 ±0.10 4.9364 ±0.08 4.2185 ±0.05 5.3475 Fanar-1-9B (Abbas et al.,2025)47.31 ±2.42 66.12 ±1.67 74.45 ±0.35 58.75 ±2.60 61.667.5850 ±0.10 4.9607 ±0.08 4.4600 ±0.06 5.6686 SILMA-9B (silma-ai,2024)50.90 ±21.73 69.40 ±6.61 62.55 ±5.57 30.00 ±3.75 53.211.8969 ±0.20 3.3547 ±0.12 2.0711 ±0.08 2.4409 Jais-2-8B (Sengupta et al.,2023)35.33 ±3.00 60.30 ±2.80 66.10 ±1.80 46.25 ±2.50 52.004.6922 ±0.15 4.245 ±0.10 2.5147 ±0.08 3.8133 Table 2:Unified leaderboard comparing MCQ tasks (Accuracy %) and open-ended QA tasks (Score 0–10). Values shown as mean ±std over 3 runs; open-ended scores are judged by two independent LLM judges. Open-ended QA Mean is averaged over Event-Cause QA, Islamic-Standards-QA, and Fatwa-QA. assessing alignment with the reference ruling/con- clusion, preservation of key constraints or quanti- tative fidelity, correctness (doctrinal/factual or fi- nancial reasoning), Arabic clarity, and directness/- grounding. We validate the judge with two ex- pert Arabic annotators on 200 randomly sampled outputs across the three tasks (MSE 0.41, Pearson r=0.92; inter-annotator agreementκ=0.84 on dis- cretized scores; Appendix J). All judge and model generations use greedy decoding (temperature 0; no sampling) with fixed maximum lengths; full prompts, rubrics, schema, critical checks, and set- tings appear in Appendix J. We organize our findings around three core ques- tions: (1) How do models perform across recog- nition versus generation tasks? (2) What distin- guishes strong Arabic financial reasoning from mere language fluency? (3) Where do models sys- tematically fail, and why? 4.1 Main Results Accounting Reasoning Gap.Shown in Table 2, Claude models exhibit substantial superiority on Accounting tasks, with Claude-Sonnet-4.5 exceed- ing GPT-5 by over13% the largest proprietary-to- proprietary gap in our evaluation. Crucially, this disparity cannot be attributed to general Arabic lan- guage proficiency alone, as these models achieve 100200300400500 Number of Reasoning Tokens 0 20 40 60 80 Ruling Accuracy (%) Gemini-2.5-FlashGemma-2-9BGPT-4o-mini Figure 3:Effect of reasoning token budget on rul- ing accuracy.Green indicates improvement with in- creased budget, red indicates decline, and blue indicates no change. near-parity on Business (76.50% vs.72.68%) and Fatwa (91.75% vs.90.75%) tasks. We instead at- tribute this divergence to Claude’s stronger capac- ity for procedural numerical reasoning, the ability to apply rule-based standards (e.g., IFRS, Egyp- tian Auditing Standards) through multi-step logi- cal chains. This suggests that Arabic domain rea- soning capabilities may constitute an independent axis from general language proficiency, warranting architectural investigation in future work. Notably, Gemini-3-Flash inverts the recognition-generation tradeoff, achieving the highest Open-Ended QA score despite moderate MCQ performance, likely because generative tasks afford extended reason- ing chains. This is supported by Figure 3, where Islamic Knowledge Error لا بأس بیع الأسھم وشر ائ ھا إذا كانت في شر كات لا ت عامل بال ربا، و كانت شر كات أ ملاك كالشر كات المع ما ریة وشر كات الزراعة والم صانع ا لإنتاجیة ونحو ھا، ومن ا شترا ھا جاز لھ أن یبیع ھا ویتص رف فی ھا. Error Reason: The model's primary ruling directly contradicts the reference answer. The model concludes that the transaction is "non-permissible, possibly forbidden" ( ویجعلھا مشبوھة، وقد تصل إلى التحریم), whereas the reference states it is permissible ("لا بأس"). The second error was in the reference the model provided "النھي عن الغر" was having a correct meaning but with wrong wording because "Al-Mulamasah" and "Al-Muhaqalah" are similar to the authetic Hadith ِ َّ ُ سول َ ى ر َ ھ َ ن ، ِ ة َ ر َ اض َ خ ُ ، والم ِ ة َ ل َ اق َ ح ُ الم ِ ن َ ع َ م َّ علیھ وسل ُ ى ﷲ َّ ل َ ص ِ ة َ ن َ اب َ ز ُ ، والم ِ ة َ ذ َ اب َ ن ُ ، والم ِ ة َ س َ م َ لا ُ والم. آخرا، وبیعھا على شخص ً ما حكم قیامي بشراء أسھم إحدى الشركات بمبلغ معین نقد ا على سنة، مع علمي بأن المشتري سیقوم بعد استلام ً بمبلغ أكبر من ثمن الشراء تقسیط الأسھم بیعھا لغیر؛ لحصول على مبلغ نقدي أقل من ثمن الشراء الذي التزم بھ لي؟ QuestionGround Truth السؤال المطروح یتضمنحكم شراء وبیع الأسھم مع علم بیع المشتري لھا بثمن أقل والغر،عدة جوانب فقھیة تحتاج إلى تفصیل وتحلیل، تعلق بالبیع، والتقسیط، علم: المشكلة الجوھریة في المعاملة: والمضاربة، والتحایل على أحكام الشریعة وقد،یغیر طبیعة المعاملة ویجعلھا مشبوھة ھذا العلم [.........]البائع بنیة المشتري ا، بل قد تكون ً لرأي الراجح ھو أن ھذه المعاملة غیر جائزة شرعاتصل إلى التحریم. النھي عن الغر: ورد في صحیح مسلمالأدلة الشرعیة الداعمة: [.........]محرمة لا تبایعوا في المضطرب، ولا في الملامسة، ولا في"عن النبي صلى ﷲ علیھ وسلم: النھي عنوھذه الأحادیث تحذر من البیوع التي فیھا غر وعدم وضوح ". المحاقلة المضاربة : ورد في صحیح البخاري عن النبي صلى ﷲ علیھ وسلم أنھ نھى عن ، فقد ٍ أو شراء ٍ في بیع َّ من حاب" النھي عن التحایل : ورد في الحدیث:المضاربة. الخلاصة". في النار َّ حاب Model Answer Gemma-3-27B [.............................] Concept Confusion Error الإجابة الصحیحة ھي )أ( Error Reason: The model exhibits a concept confusion error by treating total interest incurred on all outstanding debt as capitalized interest. Instead of determining the portion of interest attributable to the qualifying asset based on the timing of construction expenditures, it sums all annual interest expenses (50,000 + 72,000 + 52,000 = 174,000). This reflects a confusion between interest incurred and capitalizable interest. Ground TruthQuestion Fabricated evidence Wrong RulingWrong wording of Evidence Concept Misunderstanding لحساب الفوائد المرسملة، یجب التركیز على الدیون التي تم استخدامھا مباشرة في تمویل مشروع البناء، ٦٠٬٠٪، وسندات بمبلغ ١٠ جنیھ بمعدل فائدة ٥٠٬٠حیث تمثل ھذه الدیون في ورقة دفع بمبلغ ٪. ویتم حساب الفائدة على ورقة١٣ جنیھ بمعدل فائدة ٤٠٬٠٪، وقرض بمبلغ ١٢جنیھ بمعدل فائدة جنیھ، كما یتم حساب الفائدة٥٠٬٠ = ٪١٠ × ٥٠٬٠الدفع على النحو التالي: فائدة ورقة الدفع = جنیھ، ویتم حساب الفائدة على٧٢٬٠ = ٪١٢ × ٦٠٬٠على السندات كالتالي: فائدة السندات = على ذلك یتم التوصل إلى أن ً وبناء جنیھ، ٥٢٬٠ = ٪١٣ × ٤٠٬٠القرض كالتالي: فائدة القرض = نیھ ج١٧٤٬٠ = ٥٢٬٠ + ٧٢٬٠ + ٥٠٬٠المرسملة یساوي مجموع الفوائد Qwen2.5-72B المبلغ )جنیھ( ٣٠٬٠ ٣٢٠٬٠٬٠ ٢٠٬٠ ١ / ١٢ / ١٢ ١٢ / ٦ ١٢ / ٣١٠ / ١ ١ ٢ ٣ الدفعة ٣٠٬٠ ٢٥٠٬٠ ٥٠٬٠ تاریخ الدفع ٧ / ١ مدة الا ست فادة المتوسط المرجح تعاقدت إحدى الشركات مع شركة مقاولات على إنشاء ، جنیھ١٬٠٬٠مبنى جدید بتكلفة إجمالیة قدرھا / ١ / ١ا من ً على أن یبدأ العمل في إنشاء المبنى اعتبار ٢٠١٩ / ١٢ / ٣١ وینتھي في ٢٠١٩. ولتمویل بناء المبنى قامت الشركة بإصدار.الجدول أعلاه وقد سدت الشركة لمقاول البناء المبالغ الموضحة في وبفحص التزامات الشركة اتضح .ا ً جنیھ٢٠، والقیمة السوقیة لسھم ا ً جنیھ١٨، القیمة الاسمیة لسھم سھم٥٠٬٠ بمعدل سنوات٣ بورقة دفع لمدة جنیھ٥٠٬٠ مبلغ ٢٠١٨ / ١ / ١( اقترضت الشركة في ١أنھا تكون ما یلي ) من١٢ / ٣١، تدفع فوائدھا في ٪١٢ بمعدل فائدة جنیھ٦٠٬٠( سندات بمبلغ ٢) لتمویل عملیة البناء. ٪١٠فائدة ١ / ١، تدفع فوائده في ٪١٣ بمعدل فائدة جنیھ٤٠( قرض بمبلغ ٣) .٢٠١٧ / ١٢ / ٣١كل عام، صادرة في ٢٠١٨ / ١ / ١من كل عام، صادر في . لا شيء ما سبقد( ١٢٨٤٠ج( ١٧٤٠ب( ٦١٦٠أ( الاختیارات: Model Answer Figure 4:Qualitative error analysis showing representative failure modes.Left:Islamic knowledge error where Gemma-3-27B incorrectly rules a permissible transaction as forbidden, citing fabricated evidence with wrong word- ing of authentic Hadith.Right:Concept confusion error where Qwen2.5-72B conflates total interest incurred with capitalizable interest in a construction loan scenario. 0100200300400500 Number of Words 0.000 0.002 0.004 0.006 0.008 Density Human Responses Model Responses Figure 5:Models Talk More, Not Better.Despite models generating4-6×more fatwas text than human, models do not achieve proportionally higher accuracy, indicating that verbosity serves as proxy for uncertainty rather than expertise. the Gemini family shows increased ruling accuracy with larger reasoning token budgets. 5 Results Arabic Fluency ≠ Domain Reasoning: Event- Cause QA Exposes the GapArabic-centric pre- training provides strong foundations for Islamic jurisprudence tasks, but fails to transfer to finan- cial reasoning (Accounting, Business). Domain- specific fine-tuning on SAHM closes this gap across all Arabic LLMs, with MCQ gains of +13.7% (SAHM-ALLAM-7B), +5.8% (SAHM-JAIS-8B), and +5.2% (SILMA-9B), enabling SAHM-ALLAM-7B to surpass GPT-5 on Accounting and Business and match 72B baselines. Event-Cause QA emerges as the “true IQ test” for Arabic financial reason- ing. The spread (1.89-9.84) is the widest in the table, nearly the full scale. Proprietary models cluster tightly at the top (9.1-9.8), then a cliff drops everyone else below 8 . 7 . This is where Arabic-specific models expose their limits as the task requires causal reasoning over Arabic finan- cial text. Language fluency does not imply do- main reasoning. The task demands compositional causal inference that neither Arabic pretraining nor raw scale can approximate. Qualitative analysis of failure cases (Figure4) reveals two distinct error patterns: models exhibit surface-level familiarity with Islamic terminology without grounding in au- thoritative sources, for instance, Gemma-3-27B in- correctly rules a permissible transaction as forbid- den while citing fabricatedḥadīthevidence, and they conflate related but distinct financial concepts, as when Qwen2.5-72B confuses total interest in- curred with capitalizable interest by summing all expenses rather than computing weighted-average expenditures. The Recognition-Generation Gap.A model that can identify correct Islamic rulings when pre- sented as options should, in principle, generate co- herent fatwās from scratch. Our results challenge this assumption. On Fatwa MCQ, Claude-Opus- 4.5 and GPT-5 achieve91.75% and90.75% ac- curacy, respectively. However, their Fatwa QA scores drop to8.81and8.05out of10, a gap sug- gesting that recognition and generation tap funda- mentally different competencies. Figure 5illumi- nates one mechanism behind this gap. Human fatwās peak at approximately 50 words; model re- sponses peak at 300 words, a4-6×inflation. De- spite this verbosity, models do not achieve propor- tionally higher scores. We interpret this pattern as verbosity as uncertainty: when models lack con- fident knowledge, they hedge with additional text rather than committing to precise rulings. This finding has practical implications for deployment, response length may serve as a useful signal for an- swer confidence in Arabic financial QA systems. It further suggests that evaluation protocols should ModelROUGE-1 ROUGE-2 ROUGE-L Proprietary Models – Reasoning-Enhanced Claude-Opus-4.578.2263.1764.14 GPT-575.1963.7064.11 Claude-Sonnet-4.579.8664.9865.13 Proprietary Models – General-Purpose Claude-Haiku-4.579.3961.4063.62 GPT-4o-mini77.7962.9064.08 GPT-4o78.9163.1663.71 Gemini-3-Flash49.3635.8343.02 Gemini-2.5-Flash39.4627.1736.81 Open-source Models:≥70B parameters Gemma-3-27B-IT79.2563.5763.42 Qwen2.5-72B-Instruct40.5229.5034.04 Meta-LLaMA-3.1-70B39.6431.4032.65 Open-source Models:<70B parameters Qwen2.5-14B-Instruct44.4230.9035.82 Gemma-3-4B-IT76.5262.0660.93 Meta-LLaMA-3.1-8B66.6747.9256.10 Mixtral-8x7B-Instruct32.7113.0723.78 Qwen2.5-7B-Instruct25.1512.0121.86 Arabic Models Jais-2-8B73.6856.5461.17 Fanar-1-9B-Instruct60.5135.9746.96 ALLaM-7B-Instruct35.9722.6128.24 SILMA-9B-Instruct27.9216.6625.99 Table 3: Extractive summarization performance on Ara- bic financial reports evaluated using ROUGE F1 (%). distinguish between recognition and generation to avoid overestimating real-world model reliability. 5.1 Extractive Summarization Table3reveals a striking inversion: Claude- Sonnet-4.5 achieves the highest ROUGE-1 (79.86), while Gemini-2.5-Flash a strong open-ended rea- soner collapses to 39.46, underperforming even GPT-4o-mini (77.79). This exposes a fundamental tension: extractive summarization rewardsverba- tim selection, not generative fluency. Consider a typical report:ຶ“ۜب ༚ દઑ ᄎცཇ؇ޗ ይዧٺޚިߌߵ اܳأگ؇ري ሒᇭ ޗݠح اৎݞࣖࢴ ݆݁ اܳݱܝިك... ًگ٭݄۰ 300 ܹ݁٭ިن دور أਵਦل௧ௌ، ਊಸިرݬ۰ ܳۇٴڎن وَ؇ݿڎاك د”ሒᇀ(Binghatti Development successfully issued additional sukuk... valued at $300M, listed on the London Stock Exchange and Nasdaq Dubai). The gold summary must preserve the entity name, Islamic instrument (sukuk), exact figure, and dual listing elements paraphrasing mod- els systematically distort. Surprisingly, Gemma-3- 4B-IT achieves 76.52 ROUGE-1, rivaling Claude- Opus-4.5 (78.22) with a fraction of the param- eters, suggesting extraction benefits from con- strained generation rather than extended reason- ing. For Arabic-centric models, domain-specific tuning proves decisive: SAHM-7B-Instruct attains 57.79 ROUGE-L, outperforming ALLaM-7B by 0510152025303540 Percentage (%) Misunderstanding Concept Wrong Ruling Fabricated Evidence Hallucination Irrelevant Evidence Irrelevant Answer Wrong Wording Evidence Question Misread Concept Confusion Wrong Assumption Fact Hallucination Calculation Mistake Formula Misapplication Terminology Error Standard Violation 36.6% 25.2% 12.1% 9.9% 9.1% 5.1% 2.0% 30.5% 26.5% 13.9% 13.2% 5.3% 4.0% 3.5% 3.0% Islamic Knowledge Error Reasoning Error Figure 6: Root cause distribution of model errors across Islamic knowledge and reasoning tasks. +29.55 points, demonstrating that Arabic pretrain- ing alone does not confer financial extraction com- petence targeted domain adaptation does. 5.2 Domain Adaptation Across Arabic LLMs We systematically evaluate domain adaptation across major Arabic-centric LLMs. Table4re- veals three distinct adaptation profiles: (1)high- gainbases like ALLAM yielding SAHM-ALLAM- 7B with substantial improvement (+13.68%), (2) stable-gainbases like Jais-2 yielding SAHM-JAIS- 8B with consistent improvement across all MCQ metrics (+5.78%, with notably strong gains in Sen- timent +11.72%), and (3)selective-gainbases like SILMA that improve on some tasks (Sentiment +23.8%) but regress on others (Accounting -7.8%). 5.3 Error Analysis To diagnose failure modes, we analyze 500 ran- domly sampled incorrect responses across all datasets, grouped by required competence:Is- lamic Knowledge Errors(Fatwa QA, Shari’ah Standards QA, Fatwa MCQ) andReasoning Er- rors(Accounting, Business, Event-Cause QA); summarization errors are treated in § 5.1and sen- timent via accuracy. The two annotators (see §3) jointly adjudicated each error against the gold ref- erence through aconsensus protocol, addition- ally verifying cited religious evidence for Islamic tasks. Consensus over independent annotation with post-hoc IAA maximizes taxonomy coverage and ensures consistent categorization across het- erogeneous errors requiring both jurisprudential and financial expertise. Error Breakdown.Figure6reveals that two er- ror types dominate:Misunderstanding Concept andWrong Ruling, together accounting for 58.5% MCQ (Accuracy %↑)Open-Ended QA (Score 0–10↑) ModelAccounting Business Fatwā Sentiment Mean Event-Cause Fatwa-QA Islamic-Std Mean Base Models ALLAM-7B44.9168.3174.4058.7561.596.894.944.225.35 Jais-2-8B35.3360.3066.1046.2552.004.692.514.243.81 SILMA-9B50.9069.4062.5530.0053.211.903.352.072.44 Fine-tuned Models SAHM-ALLAM-7B 71.40(+26.5)93.99(+25.7)74.45(+0.1)61.25(+2.5)75.27(+13.7)6.79(-0.1)6.48(+1.5)4.12(-0.1)5.80(+0.5) SAHM-JAIS-8B40.72(+5.4)62.30(+2.0)70.14(+4.0)57.97(+11.7)57.78(+5.8)5.25(+0.6)4.69(+2.2)4.97(+0.7)4.97(+1.16) SILMA-9B (fine-tuned) 43.11(-7.8)75.96(+6.6)60.60(-2.0)53.75(+23.8)58.36(+5.2)2.01(+0.1)3.67(+0.3)3.67(+1.6)3.12(+0.7) Table 4: Domain adaptation across Arabic LLMs. MCQ accuracy (%) and Open-Ended QA scores (0–10) before and after fine-tuning on SAHM.Bold model names(SAHM-ALLAM-7B, SAHM-JAIS-8B) denote the two released SAHM-family artifacts; SILMA-9B is included as a comparison case illustrating that adaptation outcomes depend on base-model properties. of all failures. Fabricated Evidence (11.4%) and Hallucination (9.3%) follow. Notably, calculation mistakes contribute only 0.3%, models rarely fail at arithmetic but frequently fail atknowing which arithmetic to perform. 012345678910111213 Number of Evidences (Quran + Hadith) 20 30 40 50 60 Ruling Accuracy (%) Trend (quadratic) Average Accuracy Figure 7: Effect of number of evidences from Hadith and Quran on Ruling Accuracy. Effect on Evidence Count on Accuracy.Fig- ure 7examines whether the presence of scriptural evidence (Qur’ānic verses andḥadīth) in reference answers correlates with model accuracy. We ob- serve a logarithmic relationship: accuracy rises from 28% with zero evidence to approximately 55% with six or more citations. This pattern ad- mits two interpretations. Optimistically, models may leverage textual evidence as grounding sig- nals. Pessimistically, questions with more evi- dence may simply be easier or more frequently rep- resented in training data. The increased variance at higher evidence counts (shaded region) suggests the relationship is not deterministic. 6 Conclusion and Future Work We introduced SAHM, the first Arabic financial NLP benchmark integrating modern finance and Shari’ah-compliant reasoning across seven tasks. Evaluating 20 LLMs reveals Arabic fluency does not imply financial reasoning, and fine-tuning on SAHM yields two complementary released mod- els: SAHM-ALLAM-7B surpasses GPT-5 on Ac- counting and Business while matching 72B base- lines, and SAHM-JAIS-8B achieves uniformly pos- itive transfer across all tasks, demonstrating that targeted domain adaptation outperforms scale. We release all resources to support trustworthy Arabic financial assistants. Several directions extend this work. First, SAHM currently focuses on formal financial text; incor- porating informal genres such as retail investor discourse, social media financial discussions, and dialectal Arabic would broaden coverage. Sec- ond, Arabic financial reports frequently contain tables, charts, and mixed-format documents; ex- tending the benchmark to multimodal reasoning over structured financial data is a natural next step. Third, our evaluation assesses answer correctness but not evidence traceability; future metrics should explicitly verify cited Qur’anic verses,ḥadīthre- ports, and AAOIFI standard references. Fourth, cross-lingual transfer from English financial bench- marks to Arabic remains unexplored; investigat- ing whether English financial reasoning capabili- ties transfer to Arabic could reduce data require- ments. Finally, regional variation in Shari’ah inter- pretation across different supervisory bodies war- rants task variants that evaluate model robustness to jurisdictional differences in Islamic finance rul- ings. Limitations Scope and coverage.SAHM is built from curated, document-grounded sources and covers as much of the available public material as feasible; however, practical access and usage constraints on some on- line sources limit the extent to which additional genres can be incorporated at this time. As a re- sult, while the benchmark provides strong prove- nance and reduces ambiguity, it does not yet cover all Arabic financial genres (e.g., informal retail- investor discourse) or fully capture regional and in- stitutional variation in Arabic financial writing. Shari’ah-related content.For Shari’ah-oriented questions, SAHM evaluates faithfulness to the ref- erenced material and the reasoning constraints re- flected in the provided sources; since interpreta- tions may differ across jurisdictions and supervi- sory bodies, the benchmark is not intended to ad- judicate between schools of thought, but rather to test source-grounded answering under the stated as- sumptions. Futureevaluationdirections.As future work, we plan to develop evaluation metrics that explicitly assess (i) the existence and correctness of cited, source-verifiable evidence including traceable sup- port from the underlying materials (e.g., fatwa text, and financial report statements) and, when answers cite religious evidence, the correctness of references such as Qur’anic verses, hadith reports, or named fiqh sources; and (i) the accuracy of book/standard citations in model outputs (e.g., cor- rect document title, section/article identifiers, and pointers that match the relevant source segment), enabling more direct measurement of citation faith- fulness and evidence-groundedness. Ethical Statement and Broad Impact Licensing.We release SAHM under a dual li- cense: (1) code and evaluation scripts under MIT License, and (2) annotation data under C BY-NC 4.0, restricting commercial use while enabling aca- demic research. Users must independently obtain source documents where applicable. Availability.The benchmark data, models, and evaluation scripts are publicly available at https://huggingface.co/SahmBenchmark, https://github.com/rania-hossam/SAHM. Acknowledgments We acknowledge The Fin AI community for its re- search support, feedback, and collaborative envi- ronment that contributed to this work. References Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur A. Chowdhury, Fahim Dalvi, Ka- reem Darwish, Nadir Durrani, Mohamed Elfeky, Ahmed K. Elmagarmid, Mohamed Y. Eltabakh, Masoomali Fatehkia, Anastasios Fragkopoulos, Maram Hasanain, Majd Hawasly, Mus’ab Husaini, Soon-Gyo Jung, Ji Kim Lucas, Walid Magdy, Safa Messaoud, Abubakr Mohamed, Tasnim Mohiuddin, Basel Mousi, Hamdy Mubarak, Ahmad Musleh, Zan Naeem, Mourad Ouzzani, Dorde Popovic, Amin Sadeghi, Husrev Taha Sencar, Mohammed Shinoy, Omar Sinan, Yifan Zhang, Ahmed Ali, Yassine El Kheir, Xiaosong Ma, and Chaoyi Ruan. 2025. Fanar: An Arabic-centric multimodal generative AI platform.ArXiv preprint, abs/2501.13944. Hayfa A Aleid and Aqil M Azmi. 2025.Hajj-FQA: A benchmark Arabic dataset for developing question- answering systems on Hajj fatwas.Journal of King Saud University Computer and Information Sciences, 37(6):135. Salem Alhajraf. 2025.Strategic role of sovereign wealth funds in the Gulf’s energy transition and eco- nomic diversification. Technical report, Rice Univer- sity’s Baker Institute for Public Policy. Ohoud Alyemny, Hend S. Al-Khalifa, and Abdulrah- man A. Mirza. 2023.A data-driven exploration of a new Islamic fatwas dataset for Arabic NLP tasks. Data, 8(10):155. Anthropic. 2025a.System Card: Claude Haiku 4.5.An- thropic. Anthropic. 2025b.System Card: Claude Opus 4.5.An- thropic. Anthropic. 2025c.System Card: Claude Sonnet 4.5. Anthropic. Dogu Araci. 2019.FinBERT: Financial sentiment analysis with pre-trained language models.ArXiv preprint, abs/1908.10063. M Saiful Bari, Yazeed Alnumay, Norah A. Alzahrani, Nouf M. Alotaibi, Hisham A. Alyahya, Sultan Al- Rashed, Faisal A. Mirza, Shaykhah Z. Alsubaie, Has- san A. Alahmed, Ghadah Alabduljabbar, Raghad Alkhathran, Yousef Almushayqih, Raneem Alna- jim, Salman Alsubaihi, Maryam Al Mansour, Ma- jed Alrubaian, Ali Alammari, Zaki Alawami, Ab- dulmohsen Al-Thubaity, Ahmed Abdelali, Jeril Kuri- akose, Abdalghani Abujabal, Nora Al-Twairesh, Areeb Alowisheq, and Haidar Khan. 2024. AL- LaM: Large language models for Arabic and English. Preprint, arXiv:2407.15390. Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. FinQA: A dataset of nu- merical reasoning over financial data . InProceed- ings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 3697–3711, Online and Punta Cana, Dominican Republic. Asso- ciation for Computational Linguistics. Zhiyu Chen, Shiyang Li, Charese Smiley, Zhiqiang Ma, Sameena Shah, and William Yang Wang. 2022. ConvFinQA: Exploring the chain of numerical rea- soning in conversational finance question answering. InProceedings of the 2022 Conference on Empiri- cal Methods in Natural Language Processing, pages 6279–6292, Abu Dhabi, United Arab Emirates. As- sociation for Computational Linguistics. Google. 2024.Gemini 3 Flash model card.Google. Google Cloud. 2025. Gemini 2.5 pro — genera- tive ai on vertex ai.https://cloud.google.com/ vertex-ai/generative-ai/docs/models/gemini/ 2-5-pro. Last accessed: 2025-10-06. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, Aiesha Letman, Akhil Mathur, Alan Schel- ten, Alex Vaughan, et al. 2024.The Llama 3 herd of models.ArXiv preprint, abs/2407.21783. Ahmed Heakl, Sara Ghaboura, Omkar Thawakar, Fa- had Shahbaz Khan, Hisham Cholakkal, Rao Muham- mad Anwer, and Salman H. Khan. 2025a.AIN: The Arabic inclusive large multimodal model.ArXiv preprint, abs/2502.00094. Ahmed Heakl, Muhammad Abdullah Sohail, Mukul Ranjan, Rania Elbadry, Ghazi Shazan Ahmad, Mo- hamed El-Geish, Omar Maher, Zhiqiang Shen, Fahad Shahbaz Khan, and Salman Khan. 2025b. KITAB-Bench: A comprehensive multi-domain benchmark for Arabic OCR and document under- standing. InFindings of the Association for Compu- tational Linguistics: ACL 2025, pages 22006–22024, Vienna, Austria. Association for Computational Lin- guistics. Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, Alek- sander Mądry, Alex Baker-Whitcomb, Alex Beutel, Alex Borzunov, Alex Carney, Alex Chow, et al. 2024. GPT-4o system card.Preprint, arXiv:2410.21276. Islamic Financial Services Board (IFSB). 2024.Islamic financial services industry stability report 2024. Mustafa Jarrar, Ahmet Birim, Mohammed Khalilia, Mustafa Erden, and Sana Ghanem. 2023.ArBank- ing77: Intent detection neural model and a new dataset in modern and dialectical Arabic . InPro- ceedings of ArabicNLP 2023, pages 276–287, Sin- gapore (Hybrid). Association for Computational Lin- guistics. Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gi- anna Lengyel, Guillaume Bour, Guillaume Lam- ple, Lélio Renard Lavaud, Lucile Saulnier, Marie- Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024.Mix- tral of experts.Preprint, arXiv:2401.04088. Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Ta- tiana Matejovicova, Alexandre Ramé, Morgane Riv- ière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean-bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lu- cas Beyer, Xiaohai Zhai, Anton Tsitsulin, et al. 2025.Gemma 3 technical report.ArXiv preprint, abs/2503.19786. Chin-Yew Lin. 2004.ROUGE: A package for auto- matic evaluation of summaries. InText Summariza- tion Branches Out, pages 74–81, Barcelona, Spain. Association for Computational Linguistics. Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018a.W’18 Open Chal- lenge: Financial opinion mining and question an- swering. InCompanion of the The Web Confer- ence 2018 on The Web Conference 2018, W’18, pages 1941–1942, Lyon , France. ACM. Macedo Maia, Siegfried Handschuh, André Freitas, Brian Davis, Ross McDermott, Manel Zarrouk, and Alexandra Balahur. 2018b. W’18 Open Chal- lenge: Financial opinion mining and question an- swering. InCompanion Proceedings of the The Web Conference 2018, W ’18, page 1941–1942, Re- public and Canton of Geneva, CHE. International World Wide Web Conferences Steering Committee. Ying Nie, Binwei Yan, Tianyu Guo, Hao Liu, Haoyu Wang, Wei He, Binfan Zheng, Weihao Wang, Qiang Li, Weijian Sun, Yunhe Wang, and Dacheng Tao. 2025.CFinBench: A comprehensive Chinese finan- cial benchmark for large language models. InPro- ceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Com- putational Linguistics: Human Language Technolo- gies (Volume 1: Long Papers), NAACL’25, pages 876–891, Albuquerque, New Mexico. Association for Computational Linguistics. OpenAI. 2025.GPT-5 system card.OpenAI. Xueqing Peng, Triantafillos Papadopoulos, Efstathia Soufleri, Polydoros Giannouris, Ruoyu Xiang, Yan Wang, Lingfei Qian, Jimin Huang, Qianqian Xie, and Sophia Ananiadou. 2025a. Plutus: Benchmarking large language models in low-resource Greek finance. ArXiv preprint, abs/2502.18772. Xueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yu- peng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shan- shan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopou- los, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Guojun Xiong, Zhiyang Deng, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Ar- man Cohan, Xiao-Yang Liu, Jimin Huang, Ale- jandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian- Yun Nie, Sophia Ananiadou, and Qianqian Xie. 2025b.MultiFinBen: A multilingual, multimodal, and difficulty-aware benchmark for financial LLM evaluation.ArXiv preprint, abs/2506.14028. Felix Pomeranz. 1997. The accounting and auditing or- ganization for Islamic financial institutions: An im- portant regulatory debut.Journal of International Accounting, Auditing and Taxation, 6(1):123–130. Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, Peter Liu, Pouya Tafti, Abe Friesen, Michelle Casbon, Sabela Ramos, Ravin Kumar, Charline Le Lan, Sammy Jerome, An- ton Tsitsulin, Nino Vieillard, Piotr Stanczyk, Sertan Girgin, Nikola Momchev, et al. 2024.Gemma 2: Improving open language models at a practical size. ArXiv preprint, abs/2408.00118. Saudi Central Bank. 2024.Saudi Central Bank (sama) regulatory framework for Islamic finance. Neha Sengupta, Sunil Kumar Sahu, Bokang Jia, Satheesh Katipomu, Haonan Li, Fajri Koto, Osama Mohammed Afzal, Samta Kamboj, Onkar Pandit, Rahul Pal, Lalit Pradhan, Zain Muham- mad Mujahid, Massa Baali, Alham Fikri Aji, Zhengzhong Liu, Andy Hock, Andrew Feldman, Jonathan Lee, Andrew Jackson, Preslav Nakov, Timothy Baldwin, and Eric P. Xing. 2023. Jais and Jais-chat: Arabic-centric foundation and instruction- tuned open generative large language models.ArXiv preprint, abs/2308.16149. silma-ai. 2024.SILMA 9B Instruct v1.0.https://huggingface.co/silma-ai/ SILMA-9B-Instruct-v1.0. Falcon-LLM Team. 2025.Falcon-Arabic: A break- through in Arabic language models. Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, et al. 2024. FinBen: A holistic financial benchmark for large language models.Advances in Neural Information Processing Systems, 37:95716–95743. Qianqian Xie, Weiguang Han, Xiao Zhang, Yanzhao Lai, Min Peng, Alejandro Lopez-Lira, and Jimin Huang. 2023. PIXIU: A large language model, instruction data and evaluation benchmark for fi- nance . InProceedings of the 37th International Conference on Neural Information Processing Sys- tems, NeurIPS’23, Red Hook, NY, USA. Curran As- sociates Inc. Zhuohan Xie, Daniil Orel, Rushil Thareja, Dhruv Sahnan, Hachem Madmoun, Fan Zhang, Debo- priyo Banerjee, Georgi Georgiev, Xueqing Peng, Lingfei Qian, Jimin Huang, Jinyan Su, Aaryamon- vikram Singh, Rui Xing, Rania Elbadry, Chen Xu, Haonan Li, Fajri Koto, Ivan Koychev, Tanmoy Chakraborty, Yuxia Wang, Salem Lahlou, Veselin Stoyanov, Sophia Ananiadou, and Preslav Nakov. 2025.FinChain: A symbolic benchmark for verifi- able chain-of-thought financial reasoning.Preprint, arXiv:2506.02515. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jian- hong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jingren Zhou, Junyang Lin, Kai Dang, Keming Lu, Keqin Bao, Kexin Yang, Le Yu, Mei Li, Mingfeng Xue, Pei Zhang, Qin Zhu, Rui Men, Runji Lin, Tian- hao Li, Tianyi Tang, Tingyu Xia, Xingzhang Ren, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yu Wan, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, and Zihan Qiu. 2024.Qwen2.5 technical report.ArXiv preprint, abs/2412.15115. Xiao Zhang, Ruoyu Xiang, Chenhan Yuan, Du- anyu Feng, Weiguang Han, Alejandro Lopez-Lira, Xiao-Yang Liu, Meikang Qiu, Sophia Ananiadou, Min Peng, Jimin Huang, and Qianqian Xie. 2024. Dólares or Dollars? Unraveling the bilingual prowess of financial LLMs between Spanish and En- glish. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Min- ing, August 25-29, 2024, KDD’24, pages 6236– 6246, Barcelona, Spain. ACM. Yilun Zhao, Hongjun Liu, Yitao Long, Rui Zhang, Chen Zhao, and Arman Cohan. 2024.Finance- MATH: Knowledge-intensive math reasoning in fi- nance domains. InProceedings of the 62nd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), ACL’24, pages 12841–12858, Bangkok, Thailand. Association for Computational Linguistics. Fengbin Zhu, Wenqiang Lei, Youcheng Huang, Chao Wang, Shuo Zhang, Jiancheng Lv, Fuli Feng, and Tat-Seng Chua. 2021.TAT-QA: A question answer- ing benchmark on a hybrid of tabular and textual con- tent in finance. InProceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 3277–3287, Online. Association for Computational Linguistics. A Islamic Finance Shari’ah Standards QA: Sources and Processing OCR Quality Evaluation.We developed a ded- icated OCR quality evaluation tool to systemati- cally assess recognition accuracy in Arabic legal– financial documents. The tool compares raw machine-extracted text against both the original scanned page and a manually corrected reference, enabling fine-grained verification of OCR fidelity at the page level. For each document, the system pairs a scanned page image (e.g.,page_001.png) with its cor- responding OCR output (page_001.txt) and presents them side by side: the original page image appears in the left panel, while the OCR-generated Arabic text is shown on the right. Annotators in- spect these pairs to identify errors such asَݧ ݁ڰگިد (missing text),۰༲٭ොේ ଫଃ༚ فරඞأ(incorrect characters), ߙߵོصగၵ၍ ؇ت ༠؇ޗ(incorrect word order), andڣگڎان اܳٺྡྷފ٭ݑ(formatting loss). When needed, they cor- rect the OCR output using an editable field while monitoring a live similarity score reflecting the edit distance between the corrected and original text. In addition to direct corrections, annotators la- bel common OCR failure modes, including dis- torted symbols (ر݁ިز ༠؇ݬ۰ ݁ލި۰۱), punctuation er- rors (أۊޚ؇ء ఈః༟݁؇ت اܳଫଐڢࡗࡲ), and inaccurate numerals (أرڢ؇م ଫଃ༚ دڢ٭گ۰), and may add targeted comments on recurring issues such as confusion between visu- ally similar Arabic characters (e.g.,بvs.ن) or misinterpretation ofاܳྥލശܭ(diacritics). The system automatically computes a quanti- tative quality score using character-level edit dis- tance and maps similarity percentages to four in- terpretable categories:Excellent(≥95%),Good (80–95%),Partial(50–80%), andPoor(<50%). All evaluation steps are logged as structured JSON records, including original and corrected text, simi- larity scores, identified error types, annotator com- ments, and timestamps, supporting reproducibility and auditability. Beyond per-page inspection, the pipeline en- ables aggregate analysis of OCR performance across document collections, allowing researchers to identify systematic error patterns, benchmark OCR quality across heterogeneous Arabic sources, and inform downstream normalization and model refinement. Overall, this human-in-the-loop methodology ensures that OCR text used down- stream such as in Arabic financial NLP bench- marks and model training is verifiably accurate and free from recognition errors that could af- fectاݿٺڎل اܳሒᇼདྷ(jurisprudential reasoning) or اܳٺ༲ܹ٭ܭ اሒᇿ؇ৎ(financial analysis). Figure 8illustrates the OCR quality evaluation interface used during annotation. PromptDesign:To ensure high-quality OCR ex- traction, we design a constrained prompt (Figure9) that enforces verbatim transcription, preservation of diacritics and formatting, and strict exclusion of non-textual artifacts. These constraints are crit- ical for maintaining fidelity in Arabic legal doc- uments, where minor textual variations can alter meaning. Similarly, we construct a controlled prompt for question–answer generation (Figure10) that restricts outputs to the explicit content of the Shari’ah standards. This design prevents hallucina- tion, preserves juridical precision, and ensures that generated QA pairs remain faithful to the source text. B Islamic Fatwa Dataset: Sources and Processing The purpose of this evaluation is to determine whether an AI-generated multiple-choice question (MCQ) accurately tests the same Islamic jurispru- dence concept as the originalڣٺިىQ&A pair. The goal is to maintain both pedagogical soundness and factual correctness. A well-formed MCQ must re- main conceptually aligned with the original ruling (اࠍુળ اܳሒᇼདྷ), preserve the main݁ڰ۳ިم ڣگ۳with- out distortion, and use appropriate݁ݱޚܹ༲؇ت ڣگ۳٭۰ to reflect the opinion of the original scholar (มฆُڰ ৎا). Evaluators must ensure that the question targets the central legal issue and does not introduce unrelated details or alter the scenario in a way that changes the ruling. This evaluation is conducted through a structured annotation dashboard (Figure 11) that presents the original fatwā alongside the generated MCQ for systematic validation. The fatwā Q&A pairs used in this evaluation are collected from a di- verse set of authoritative online sources (Table 5). For an MCQ to be marked asቕሶఈః݁(RELE- VANT), it must meet four main criteria. First, conceptual alignment (اৎިاء݁۰ اৎڰ؇۱٭݄٭۰) the question should test the same core ruling as the source fatwa and stay faithful to its reasoning and condi- tions. Second, correct answer accuracy (۰ً؇༥دڢ۰ ا اܳݱۜ٭۰༲) the indicated correct option must exactly match the original answer, remain free of con- tradictions, and use precise Islamic legal terms. Third, distractor quality (ۏިدة اࠍ٭؇رات اࠍ؇ޗ۰٪) incor- rect options should be plausible but clearly wrong according to the fatwa, reflecting common misun- derstandings rather than random or nonsensical an- swers. Finally, question clarity (وݪިح اܳފޝال) the MCQ must be clearly phrased, grammatically cor- rect inاܳأݠ۰ਃಸ, and provide enough context to be an- swerable without referencing the original text. Figure 8: OCR quality evaluation interface for the Shari’ah Standards QA dataset. The tool displays each scanned page from the AAOIFI Shari’ah Standards (left) alongside the OCR-extracted Arabic text (right) to support man- ual quality verification. Annotators compare the original page with the extracted text, flag recognition errors in diacritics, numerals, and domain-specific terminology, and add corrective notes (bottom). A progress bar tracks annotation completion and overall OCR accuracy. WebsiteLinkCountry Dar Al Ifta in Saudi Arabiahttps://w.alifta.gov.sa/Saudi Arabia Dar Al Ifta in Egypthttps://w.dar-alifta.orgEgypt Dar Al Ifta in Jordanhttps://aliftaa.joJordan Al Shaikh Abdual Aziz Ibn Baz https://binbaz.org.saSaudi Arabia Al Shaikh Mohammad Ibn Othaimin https://binothaimeen.net/siteSaudi Arabia Al Shaikh Abdual Aziz Al Ashaikhhttps://w.mufti.af.org.saSaudi Arabia Al Shaikh Saleh Al Fwzanhttps://w.alfawzan.af.org.saSaudi Arabia Al Shaikh Saleh Bin Humaid https://w.ibnhomaid.af.org.sa/Saudi Arabia Al Shaikh Abdullah Al Maneehttps://al-manee.comSaudi Arabia IslamWebhttps://w.islamweb.comQatar FatwaPediahttps://fatwapedia.comSaudi Arabia IslamQA https://islamqa.infoSyria IslamOnlinehttps://islamonline.netQatar Table 5: Primary online fatwā archives used for collecting Islamic financial question–answer pairs. These official and widely recognized sites span seven Arab countries, providing diverse juristic opinions and real-world financial scenarios. The URLs shown correspond to the original Arabic portals from which data was programmatically scraped and later cleaned for inclusion in the dataset. Prompt for Arabic OCR Text Extraction Task.You are an expert Arabic OCR system spe- cialized in legal and financial documents. Given a scanned page image from an official Islamic finance standard, your task is to extract the textverbatimin Arabic with maximum fidelity to the original source. Extraction guidelines. •Preserve the original wording exactly; donot paraphrase, summarize, or infer missing content. •Preserve diacritics (اܳྥލശܭ) whenever present in the source. •Preserve all numerals exactly as written; do not normalize or convert number formats. •Preserve punctuation, headings, lists, and para- graph boundaries as faithfully as possible. •Do not correct perceived grammatical, typo- graphical, or stylistic issues. •If text is unclear or partially illegible, extract the most faithful representation without guessing. What to ignore. •Page numbers, running headers, footers, or deco- rative elements not part of the main content. •Marginal artifacts or scanning noise that do not belong to the text. Critical rule.Donotadd explanations, comments, translations, or annotations. Output Arabic text only. Input.A single scanned page image from the AAOIFI Shari’ah Standards. Output format.Return the extracted Arabic text as plain UTF-8 text, preserving line breaks and para- graph structure. Figure 9: Prompt for Arabic OCR text extraction with strict verbatim fidelity and formatting preservation. Conversely, an MCQ should be marked asଫଃ༚ ቕሶఈః݁(NOT RELEVANT) if it fails any major re- quirement. Conceptual misalignment occurs when the question tests a different topic, oversimplifies a complex juristic issue, or changes critical con- text such as conditions (ཇوط) or scenarios. Incor- rect answer issues include a keyed option that con- tradicts the fatwa, multiple potentially correct an- swers, or misleading explanations. Poor distractor quality arises when wrong options are obviously incorrect, factually wrong aboutاݿఈఃم, or too am- biguous. Technical problems include grammar er- rors that affect meaning, vague or incomplete ques- tions, or improper mixing of different݁ڍا۱صin a way that confuses the intended ruling. Prompt for Shari’ah Standards Question– Answer Generation Task.You are an assistant supporting the creation of evaluation data for Islamic finance. Given a verified excerpt from an official Shari’ah standard, your task is to draft candidate Arabic question–answer pairs that reflect theexplicit ruling stated in the text. Question generation. •Formulate a clear, focused question that asks about the ruling, condition, or permissibility de- scribed in the excerpt. •Do not introduce hypothetical scenarios or facts not present in the source text. •Ensure the question can be answered directly and completely from the provided excerpt. Answer generation. •Base the answer strictly on the given excerpt; do not add external knowledge. •Preserve the original legal meaning, mandatory conditions, and stated exceptions. •Do not simplify, reinterpret, or generalize the rul- ing beyond what the text explicitly states. •Use formal Arabic consistent with fiqh al- muʿāmalāt terminology. Restrictions. •Do not issue personal opinions or normative judg- ments. •Do not cite sources outside the provided text. •Do not omit conditions, constraints, or qualifiers that affect the ruling. Input template. STANDARD_EXCERPT: arabic_text Output format. ”question”: ”...”, ”answer”: ”...” Figure 10: Prompt for generating Arabic QA pairs from Shari’ah standard excerpts with strict fidelity to explicit rulings. The evaluation process follows a clear four-step workflow. First, read the original Q&A carefully, identify the primaryુળ༡, anyཇوطor exceptions, and the supporting evidence such as Qur’anic verses orಱڎ༡. Second, analyze the generated MCQ to check conceptual consistency, faithful- ness of the correct answer, and plausibility of dis- Figure 11: Custom annotation interface used to validate automatically generated multiple-choice questions (MCQs) for the Islamic Finance Fatwa Q&A dataset. The interface displays each original question–answer pair on the left and the corresponding AI-generated MCQ on the right, including the question, answer options, and the automati- cally selected correct choice. Annotators review conceptual alignment between the MCQ and the original fatwā, verify the correctness and terminology of the marked answer, and assess the plausibility and pedagogical value of distractors. The bottom panel provides structured evaluation criteria and issue tagging to ensure consistent, high- quality validation. tractors. Third, look for red flags such as contra- dictions, oversimplification, missing qualifiers, or scenario changes. Finally, make a decision: label the MCQ asቕሶఈః݁if it meets all core criteria (mi- nor language or formatting issues may be tolerated) or asቕሶఈః݁ ଫଃ༚if any critical issue is present. This structured approach ensures that evaluation is con- sistent, transparent, and preserves the integrity of Islamic legal reasoning in AI-generated questions. Normalization Prompt.To standardize fatwā question–answer pairs, we design a constrained prompt (Figure 12) that removes non-essential el- ements such as greetings and formatting artifacts while preserving the original wording and legal intent. This ensures consistency across examples without altering the underlyingሒᇼཇ ુળ༡. C Evaluated Models This appendix briefly documents the rationale be- hind the selection of models evaluated in Table2, with model specifications summarized separately in Table 6. The goal is not comparative analysis, but transparency regarding model coverage across language focus, scale, and accessibility. Arabic Focused Models:ALLAM-7B- Instruct (Bari et al.,2024), Fanar-1-9B- Instruct (Abbas et al.,2025), and SILMA- 9B-Instruct (silma-ai,2024) were selected to represent publicly available instruction-tuned models explicitly adapted for Arabic. These sys- tems reflect different training strategies and base model lineages, providing coverage of current Arabic-centric development efforts. Open-Source Multilingual Models.Qwen2.5 models (Yang et al.,2024), LLaMA-3.1 mod- els (Grattafiori et al.,2024), Gemma-2 and Gemma-3 models ( Riviere et al.,2024;Kamath et al.,2025), and Mixtral-8x7B-Instruct (Jiang et al.,2024) were included as strong open-weight baselines spanning a wide range of parameter scales. These models are widely used, well- documented, and provide reference points for general-purpose multilingual performance on Ara- bic financial and jurisprudential tasks. Proprietary Models.GPT-5 (OpenAI,2025), GPT-4o ( Hurst et al.,2024), Claude-4.5 vari- ants (Anthropic,2025b,c,a), and Gemini-3- Flash (Google,2024) were evaluated as closed- source upper-bound references. Their inclusion ModelOrganizationSize Source / Notes Arabic-Focused Models ALLAM-7B-InstructSDAIA / ALLaM-AI 7B(Bari et al.,2024) Fanar-1-9B-InstructQCRI9B(Abbas et al.,2025) SILMA-9B-InstructSILMA AI9B(silma-ai,2024) Strong Multilingual / General Open-Source Models Qwen2.5-72B-InstructAlibaba72B(Yang et al.,2024) LLaMA-3.1-70B-Instruct Meta70B(Grattafiori et al.,2024) Qwen2.5-14B-InstructAlibaba14B(Yang et al.,2024) Qwen2.5-7B-InstructAlibaba7B(Yang et al.,2024) Gemma-2-9B-ITGoogle9B(Riviere et al.,2024) Gemma-3-27B-ITGoogle27B(Kamath et al.,2025) Gemma-3-4B-ITGoogle4B(Kamath et al.,2025) LLaMA-3.1-8B-InstructMeta8B(Grattafiori et al.,2024) Mixtral-8x7B-InstructMistral AI8×7B (Jiang et al.,2024) Proprietary Models (Upper-Bound References) GPT-5OpenAI–(OpenAI,2025) (API) GPT-4oOpenAI–(Hurst et al.,2024) (API) GPT-4o-miniOpenAI–(Hurst et al.,2024) (API) Claude Opus 4.5Anthropic–(Anthropic,2025b) (API) Claude Sonnet 4.5Anthropic–(Anthropic,2025c) (API) Claude Haiku 4.5Anthropic–( Anthropic,2025a) (API) Gemini-3-Flash (preview) Google DeepMind–( Google,2024) (API) Table 6: Models evaluated in this study, grouped into Arabic-focused models, strong multilingual open-source baselines, and proprietary frontier models used as upper-bound references. enables contextualization of open and Arabic- focused models against contemporary frontier systems, without implying direct comparability or deployment parity. D Business and Accounting Exam Extraction Prompts Business and accounting exams in Arabic exhibit heterogeneous layouts, ranging from narrative exercise-based formats to tabular true/false ques- tions. To reliably extract structured MCQs from these sources, we use two task-specific prompts tai- lored to the dominant document formats observed in the collected exams: an exercise-based extrac- tion prompt (Figure 13) and a table-oriented extrac- tion prompt (Figure14). E Financial Sentiment Annotation Guidelines We annotate Arabic financial reports using a document-level sentiment scheme designed to re- flect overall market impact rather than sentence- level polarity. Annotation follows a structured human-in-the-loop workflow supported by a cus- tom web-based interface (Figure 15), with clear decision rules (Figure16) to ensure consistency across Islamic and conventional financial report- ing. F Arabic Finance Extractive Summarization Annotation Guidelines We annotate Arabic financial reports for extractive summarization using a structured human-in-the- loop workflow supported by a custom web-based interface (Figure 17) and guided by explicit an- notation criteria (Figure18). Arabic financial re- ports also exhibit recurring linguistic and format- ting challenges, including specialized terminology, code-switching, and mixed numeral systems (Ta- ble 8). Annotators are instructed to select sen- tences that preserve key financial facts, numerical values, and regulatory references without introduc- ing paraphrasing or abstraction. Prompt for Fatwā Q&A Normalization Task.You are an expert Arabic copy-editor special- izing in Islamic finance Q&A. Given aQUESTION and anORIGINAL ANSWER, your goal is to pro- duce a concise, self-contained question and answer pair in Arabic by removing only non-essential ele- mentswithout paraphrasing or changing the juristic intent. Donotsummarize or rephrase; keep the orig- inal wording as much as possible. 1. Referral flag.Before editing, set IS_MAINLY_REFERRAL: •”YES”if the answer mainly redirects to another fatwā, link, or reference and does not provide a sub- stantive independent ruling. •”NO”otherwise. 2. Clean the question.Edit minimally while pre- serving wording and fiqh intent: •Remove greetings, honorifics, and personal ap- peals (e.g.,۰༡؇ᆙᆊ اܳލ٭ڹՄ៰Ղا ۬గఒݿاܳފఈఃم༟ ܹ٭ુળ). •Remove formal closings (e.g.,أرۏި ݁ۇٴુળ اܳٺ୍ଲموඹජاቕመ ً اଫଃ༠ Մ៰Ղا). •Remove the scholar’s name if it is only a form of address; keep it only if the question explicitly seeks that scholar’s specific fatwā or opinion. •Ensure the final question reads as a natural, stan- dalone query. 3. Clean the answer.Edit minimally while preserv- ing wording and reasoning: •Remove formal openings and closings so the an- swer starts with substantive content. •Remove all fatwā numbers, hyperlinks, and nav- igational phrases, editing surrounding text just enough to remain grammatical. •Convert Arabic-Indic numerals to Western numer- als. •Remove purely formulaic closings such asوڣگુળ اՄ៰Ղ andواՄ៰Ղ أ༟when they are not part of practical ad- vice. •Always preserve Qurʾānic verses and sūrah refer- ences, ḥadīth attributions, and citations of scholars and their opinions. Global rule.Always deleteallfatwa numbers from the cleaned question and cleaned answer. Input template.TITLE: title QUESTION: question ORIGINAL ANSWER: answer Output format. ”IS_MAINLY_REFERRAL”: ”YES” or ”NO”, ”cleaned_question”: ”...”, ”cleaned_answer”: ”...” Figure 12: Prompt for Arabic fatwā Q&A normaliza- tion with minimal editing and preservation of juristic intent. Prompt for Extracting Arabic Accounting Exam MCQs (Exercise-Based Format) Task.You are an expert system for extracting Arabic accounting exam questions. The input is a scanned exam page containing exercises that begin with the keywordદઊݠஓ(Exercise), followed by a number. Extraction instructions. •Identify all exercises that begin withદઊݠஓfollowed by a numeral (e.g.,١ દઊݠஓ,٢ દઊݠஓ). •For each exercise, extract the full contextual text that follows the exercise header. •Within each exercise, identify all multiple-choice questions (numbered 1, 2, 3, etc.). •For each MCQ, combine the exercise context with the specific question text to form a complete ques- tion. •Extract all answer choices labeled asأ,ب,ج, and د. •Identify the correct answer by detecting underlined text in the choices; underlining indicates the cor- rect option. Critical rules. •Preserve the original Arabic text exactly; do not paraphrase or normalize. •ExtractallMCQs appearing under each exercise. •If no underlined choice is visible, set the correct answer tonull. Output format.Return a JSON object with the fol- lowing structure: ”exercises”: [ ”exercise_number”: ”...”, ”exercise_context”: ”...”, ”questions”: [ ”question_number”: ”...”, ”full_question_text”: ”...”, ”choices”: :”” ”...”, :”” ”...”, :”” ”...”, :”” ”...” , ”correct_answer”: ”...”, ”is_underlined”: true ] ], ”page_info”: ”total_exercises”: ..., ”total_questions”: ..., ”language”: ”Arabic”, ”subject”: ”Accounting” Figure 13: Prompt for extracting MCQs from Arabic accounting exams with exercise-based layouts. G Event–Cause Reasoning Annotation Guidelines We construct event–cause reasoning instances for Arabic financial reports using a structured annota- Prompt for Extracting Arabic Business Exam MCQs (Tabular Format) Task.You are an expert system for extracting Ara- bic business and accounting exam questions from scanned images containing tabular layouts. Document characteristics. •Each row corresponds to one question. •Questions are numbered using Arabic numerals (e.g.,١٢٤,١٢٥). •Answer choices typically include༃ ࡷ (True) and ۊޚ؊(False), with optional additional choices. •The correct answer is highlighted with a yellow background. Extraction instructions. •Identify all question rows in the table. •Extract the question number and full Arabic ques- tion text for each row. •Extract all visible answer choices. •Identify the correct answer by detecting yellow highlighting. •Label questions astrue_falseor multiple_choiceaccordingly. Critical rules. •Preserve Arabic text exactly as written, including diacritics. •If yellow highlighting is ambiguous or not visible, sethas_yellow_highlighttofalse. •Extractallvisible questions on the page. Output format.Return a JSON object with the fol- lowing structure: ”questions”: [ ”question_number”: ”...”, ”question_text”: ”...”, ”question_type”: ”...”, ”choices”: ”a”: ”...”, ”b”: ”...”, ”c”: ”...” , ”correct_answer”: ”...”, ”correct_choice_text”: ”...”, ”has_yellow_highlight”: true, ”subject_area”: ”business” ], ”page_info”: ”total_questions”: ..., ”format”: ”table_with_yellow_highlighting”, ”language”: ”Arabic”, ”question_type”: ”mixed” Figure 14: Prompt for extracting MCQs from Arabic business and accounting exams with tabular layouts. tion framework with explicit quality control proce- dures (Figure19). CategoryTotal Count Zakat (زႤ၍ة)4,888 Riba (رً؇)2,454 Murabaha (۰ොໍາاਵਦ)1,389 Gharar (ਵؗر)860 Waqf (وڢژ)730 Ijara (إ༥؇رة)571 Maysir (ཏ݁)372 Musharaka (݁ލ؇رᄎც)242 Mudharaba (݁ݯ؇ر۰ً)228 Takaful (Ⴄၽّڣܭ)187 Sukuk (ݬܝިك)32 Total records11,953 Table 7: Distribution of questions across Islamic fi- nance categories in the final dataset. IssueExample from the report Islamic Finance Terminology “ّأݞߌ߳ اࡺ࢘ࢦި لܭ اৎފٺڎام وّޚިߌߵ اܳݱܝިك واܳފۇٴڎات” Code-switchingዛው૭“ڎف ّأݞߌ߳ اࡺ࢘ࢦި لܭ اৎފٺڎام وّޚިߌߵ اܳݱܝިك واܳފۇٴڎات، وز ل؇دة ނڰ؇ڣ٭۰ اܳگޚ؇ع ”...Fitch Mixed Numeral Systems “اෛຶڰݥ ܾ ا౫ౙభ દઊᄴᄟި ٢٧ ر ل؇ل ڢޚݠي )٧٫٤ ܹ݁٭؇ر دور༟ ሒᇭ (؇م ”٢٠٢٣— combines Arabic currency and West- ern digits Table 8: Key text difficulties in Arabic financial reports with real examples H MCQ Answer Normalization and Scoring To ensure fair and reproducible evaluation of multiple-choice questions, we normalize model outputs before computing accuracy. Large lan- guage models frequently generate free-form re- sponses (e.g., explanations, mixed scripts, or mul- tiple answer mentions) rather than a single option label. Normalization procedure.For each model out- put, we apply the following steps: •Normalize Unicode and Arabic script by remov- ing diacritics, collapsing repeated whitespace and punctuation, and mapping Eastern Arabic digits (e.g.,١٢٣٤) to Western digits (1234). •Extract the first explicit answer mention using a cascade of regular expressions that handle: –Latin option labels (e.g.,A,B, “Option C”), –Arabic option letters (e.g.,أ,ب,ج), –Spelled-out Arabic forms (e.g.,ً؇ء), –Numeric indices (e.g.,1–4). Figure 15: Custom annotation platform used to label Arabic financial reports for sentiment analysis. Annotators reviewed full reports, assigned sentiment classes, and flagged ambiguous cases for expert adjudication. Document-Level Financial Sentiment An- notation Guidelines Core principle.Assign a single sentiment label based on theoverall dominant sentiment of the en- tire report, not on individual sentences or isolated phrases. Annotation procedure. •Read the complete document before assigning any label. •Identify the main financial outcome, thesis, and conclusion. •Give greater weight to headlines, executive sum- maries, and concluding sections than to supporting details. Handling mixed sentiment. •Dominant sentiment rule: assignPositiveor Negative if one polarity accounts for more than 60% of the salient content. •Neutral default: assignNeutralwhen positive and negative signals are balanced or when the re- port is primarily factual. Decision criteria. •Positive: growth announcements, profit increases, successful expansions, or favorable forecasts. •Negative: losses, declining performance, regula- tory issues, or adverse outlooks. •Neutral: factual reporting, balanced analysis, or informational updates without a clear directional impact. Quality control.Annotators label reports indepen- dently, resolve disagreements during a calibration phase, and refine shared decision criteria. A third domain expert adjudicates remaining conflicts. Each report receives exactly one final sentiment label. Figure 16: Guidelines for document-level sentiment an- notation of Arabic financial reports. Scoring.We compute accuracy as an exact match between the normalized predictionˆyand the gold labely. For example, the output “2 ሒሃ ۰ً؇༥ا ྟ૭ص ݬ٭؇۰༚ اࠍુળ” is normalized toB, while “اࠍ٭؇ر )ج( ۱ި اܳݱۜ٭ں” is normalized toC. Outputs that do not contain a valid option after normalization are marked incorrect. This procedure ensures that evaluation is ro- bust to superficial variation in formatting, language mixing, and numeral systems, and that all models are assessed under a consistent and deterministic scoring protocol. I Instruction Templates for SAHM Tasks To enable a unified instruction-tuning and evalua- tion setup across heterogeneous tasks, we convert each SAHM task into a standardized instruction format. Table 9lists the canonical task instructions used in our benchmark, shown in their original Ara- bic formulation alongside an English translation for clarity. The Arabic prompts constitute the ac- tual inputs used during model evaluation, while the English versions are provided solely to document task intent and facilitate reproducibility. J LLM-as-a-Judge Protocol, Validation, and Reproducibility J.1 Judge Protocol and Reproducibility We evaluate the three open-ended tasks (Fatwa QA, Shari’ah Standards QA, and Event–Cause QA) us- ing an LLM-as-a-judge setup withGemini-2.5- Flash. For each instance, the judge receives: (i) the exact Arabic prompt shown to the model (in- cluding any report/excerpt and question), (i) the gold reference answer, and (i) the model’s can- didate answer. The judge is blind to model iden- Figure 17: Custom web-based annotation interface for extractive summarization. Annotators view Arabic financial reports, select key sentences containing figures, decisions, and disclosures, and mark them for gold-standard sum- maries. tity and always observes the inputs in fixed, explic- itly labeled fields (prompt,ground_truth,candi- date_answer) to avoid ordering or positional am- biguity. The judge returns a structured JSON ob- ject containing: (a) rubric sub-scores whose sum defines an overall score in[0,10], (b) task-specific critical error flags (e.g., contradiction with the ref- erence, omission of critical constraints, normaliza- tion of unlawful elements, or fabrication/alteration of figures), and (c) a brief explanatory note. Task- specific evaluation rubrics are defined for fatwa QA (Figure 20), Islamic finance QA (Figure21), and financial analysis tasks (Figure22). We enforce a strict JSON schema during pars- ing. If a response is invalid JSON or violates the schema, we retry once with the same inputs and an explicit JSON-only instruction; persistent fail- ures are marked invalid and excluded from aggre- gate scores (we report the invalid-rate). We run the judge deterministically (temperature= 0.0, greedy decoding, max output tokens= 4096), and there- fore do not perform repeated judging or score av- eraging. Full judge prompts, rubrics, and task- specific schemas are provided in the following sub- sections. J.2 Human Alignment Study (Judge Validation) To validate the LLM judge against expert evalua- tion, we conduct a human alignment study on200 randomly sampled open-ended outputs spanning Fatwa QA, Shari’ah Standards QA, and Event– Cause QA. Two expert Arabic annotators inde- pendently score each model response using the same[0,10]additive rubric provided to the judge (Section J). We compare the judge’s scores (from Gemini-2.5-Flash) to the mean of the two human scores, obtaining an MSE of0.41and a Pearson correlation ofr= 0.92. Inter-annotator agreement is high (κ= 0.84computed on discretized integer scores). These results indicate that the LLM-as-a- judge scores closely track expert human judgments under our rubric. J.3 Cross-Judge Validation of Open-Ended Evaluations To address concerns about potential model-family bias in our LLM-as-judge evaluation, we re-ran all open-ended tasks with two independent judges: Gemini-2.5-Flash (our primary judge) and GPT-4o. Each model was evaluated 3 times under greedy DatasetOriginal Arabic PromptEnglish Translated Prompt Islamic Sharia Standards QA আॻ༟ ݁أ؇لଫଃ وأႤၽ༡م اࡺ࢘ࢦި لܭ اݿሒᇧఈః واৎأ؇݁ఈఃت اৎ؇ܳ٭۰ ً ء؇ಸ اܳފޝال: اܳདྷ٭۰، أۏص আॻ༟ اܳފޝال اܳٺ؇ࣖࢻ ሒᇿڢ۰.Text: Question. Outputاոຖॆूּ١ وؔؠո ቘً۠ܙاּ܅ اڤ֍ໃמ١: Based on Islamic finance standards and Shari’ah- compliant rulings, answer the following question accurately.Text:Question: Question.Answer (Shari’ah-compliant):Output Islamic Fatwa QAআॻ༟ أႤၽ༡م اܳདྷ لأ۰ اݿఈః݁٭۰ واܳڰگ۬ اݿሒᇧఈః، أۏص আॻ༟ ً ء؇ಸ اܳފޝال اܳٺ؇ሒᇿ ًޚݠ لگ۰ ݁ڰݱᄭᄥ و݁ڎ؇ً ۰ᆇᅦدᄭᄟ ۇٴڎ ا݁Ⴄၽن. اܳފޝال: .QuestionText:OutputAnswer: Based on Islamic jurisprudence (fiqh) and Shari’ah rulings, answer the following question in a detailed manner, supported by evidence when possible.Text:Question: Question.Answer: Output Islamic Financial Fatwa MCQ اڢݠأ اܳފޝال اܳٺ؇ሒᇿ ًأۇٴ؇ل۰ واଫଐ༠ ا༥؇۰ً اܳݱۜ٭۰༲ وڣگ؇ ً اܳފޝال: .Question اࠍ٭؇رات: Ⴄၽ༡م اܳདྷ لأ۰.Text: أරඝجරඞ ف اࠍ٭؇ر اܳݱۜ٭ں ڣگޔ. .ChoicesAnswer: Read the following question carefully and choose the correct answer according to Shari’ah rul- ings.Text:Question: Question. Choices: Choices.Answer:Output only the correct op- tion letter. Accounting Exams MCQاڢݠأ اܳފޝال اܳٺ؇ሒᇿ ًأۇٴ؇ل۰ واଫଐ༠ ا༥؇۰ً اܳݱۜ٭۰༲.Text: اܳފޝال: .Question اࠍ٭؇رات: .ChoicesAnswer: أරඝجරඞ ف اࠍ٭؇ر اܳݱۜ٭ں ڣگޔ. Read the following question carefully and choose the correct answer.Text:Question: Question. Choices: Choices.Answer:Output only the correct option letter. Business Exams MCQاڢݠأ اܳފޝال اܳٺ؇ሒᇿ ًأۇٴ؇ل۰ واଫଐ༠ ا༥؇۰ً اܳݱۜ٭۰༲.Text: اܳފޝال: .Question اࠍ٭؇رات: .ChoicesAnswer: أරඝجරඞ ف اࠍ٭؇ر اܳݱۜ٭ں ڣگޔ. Read the following business/management ques- tion carefully and choose the correct answer.Text: Question: Question. Choices: Choices.An- swer:Output only the correct option letter. FinancialReport SentimentAnalysis MCQ اڢݠأ ًأۇٴ؇ل۰ اܳٺگݠߌߵ اሒᇿ؇ৎ اܳٺ؇ሒᇿ واଫଐ༠ اܳٺݱྡྷ٭ژ اܳݱۜ٭ں ݆݁ ݁ۇٴޙިر اܳٺگݠߌߵ: .Input اৎފྥټ݄ݠ.Text:/ ሒᇀ؇ຬإ)Answer: .(ࣖࢴ؇ො ݿܹม / Read the following financial report carefully and choose the correct label from an investor’s per- spective.Text:Report: Input.Answer:(Posi- tive / Negative / Neutral). ReportExtractive Summarization ڢܾ ਐಸܹۛ٭ݧ اܳٺگݠߌߵ اሒᇿ؇ৎ اܳٺ؇ሒᇿ ً؇ݿٺ༱ڎام اܳٺܹۛ٭ݧ اݿٺۛݠاሒᇚ Summarization). (Extractive ۰٭ᆇᆅأ ଫ܋ا ܭ৵৩ৠا ଫଐ༠ا ݁ٴ؇ཇة ݆݁ اܳۇٴݧ اݬঌॻ دون ّأڎلܭ أو إ༟؇دة ݬ٭؇۰༚، ورᚹ ዛّ؇ اۏأܭ اৎܹۛݧ ۋިا–30 ሒᇿ%40 ݆݁ ܾಸ ڰݴ ૭ܹފܹ۳؇. اܳۇٴݧ، ورআॻ༟ ّண ارڢ؇م واܳگݠارات واܳۇٴٺ؇༇ واܳٺިار༂.Text: أරඝج اৎܹۛݧ ڣگޔ دون أي اܳٺگݠߌߵ: .InputAnswer: .حཇ Summarize the following financial report using extractive summarization (select sentences verba- tim, keep original order, target 30–40% length, fo- cus on numbers/decisions/outcomes/dates).Text: Report: Input.Answer:Output the extractive summary only (no extra text). Event–Cause Reasoning QA আॻ༟ اܳٺگݠߌߵ اሒᇿ؇ৎ اܳٺ؇ሒᇿ، أۏص আॻ༟ اܳފޝال اܳٺ༲ܹ٭႟ၽ૰ ঌॻ ً ء؇ಸ ݁ڰݱܭ ودڢ٭ݑ ݁ؕ اܳଐام ً؇ৎأߺࠊ݁؇ت اܳިاردة ሒᇭ اܳۇٴݧ ڣگޔ. Question. :اܳފޝال Input. :ሒᇿ؇ৎاܳٺگݠߌߵ اText: OutputAnswer: Based on the following financial report, answer the analytical question in a detailed and accurate way, grounded only in the provided text.Text:Fi- nancial report: Input. Question: Question. Answer:Output Table 9: Instruction templates used for SAHM tasks (Arabic prompts are used in evaluation; English translations document task intent). decoding, and we report mean±std across runs for both judges. If our primary judge were systemat- ically favoring Gemini-family models, switching to GPT-4o should lower Gemini scores; instead, theyriseunder GPT-4o (Gemini-3-Flash: Islamic- Std 9.18→9.76, Fatwa 9.17→9.32), i.e., the oppo- site of what circular bias would predict. Model rankings are preserved across judges: top-tier models (Gemini-3-Flash, Claude-Opus-4.5) and bottom-tier models (SILMA-9B, LLaMA-3.1-8B) remain in the same groupings regardless of judge. Tight confidence intervals (±0.02–0.10) across 3 runs confirm reproducibility under greedy decod- ing. Full per-model, per-judge scores appear in Ta- ble10. J.4 Frontier Model Error Analysis To diagnose why frontier models fail on specific SAHM tasks, we conducted a systematic error anal- ysis of GPT-5 and Gemini-3-Flash across Account- ing, Business, and Summarization. Two native Arabic-speaking annotators with financial back- grounds jointly reviewed every error, examined Event-Cause QAIslamic-Std QAFatwa QA ModelGemini Judge GPT-4o Judge Gemini Judge GPT-4o Judge Gemini Judge GPT-4o Judge Gemini-3-Flash9.84±0.029.97±0.029.18±0.039.76±0.019.17±0.029.32±0.02 Claude-Opus-4.59.67±0.039.32±0.978.06±0.059.53±0.028.79±0.039.18±0.02 Claude-Sonnet-4.59.32±0.049.68±0.048.24±0.059.28±0.027.58±0.038.86±0.02 GPT-4o8.30±0.069.50±0.096.64±0.088.53±0.036.50±0.048.04±0.01 Gemma-3-27B8.70±0.059.74±0.026.15±0.088.41±0.035.18±0.057.35±0.02 Qwen2.5-72B8.08±0.108.45±0.175.61±0.106.96±0.075.37±0.066.30±0.01 Fanar-1-9B7.57±0.108.03±0.224.94±0.086.03±0.034.44±0.065.15±0.04 Gemma-3-4B7.39±0.089.09±0.082.88±0.085.32±0.072.46±0.064.39±0.02 ALLAM-7B6.87±0.107.79±0.164.92±0.085.95±0.034.20±0.054.49±0.03 LLaMA-3.1-70B6.60±0.157.15±0.303.70±0.104.67±0.054.74±0.082.24±0.02 Mixtral-8x7B4.53±0.085.14±0.092.48±0.083.43±0.081.78±0.062.80±0.04 SILMA-9B1.88±0.201.43±0.373.33±0.122.33±0.032.05±0.081.61±0.05 LLaMA-3.1-8B4.90±0.182.49±0.172.50±0.121.85±0.031.38±0.080.71±0.02 SAHM-ALLAM-7B6.50±0.107.02±0.126.30±0.026.59±0.044.24±0.044.51±0.03 Table 10: Cross-judge validation on open-ended tasks. All evaluations run 3 times under greedy decoding with two independent judges (Gemini-2.5-Flash and GPT-4o). Rankings are preserved across judges; Gemini-family scores rise under GPT-4o, disconfirming circular bias. ModelAccounting (%) Business (%) Fatwā MCQ (%) Sentiment (%) Proprietary Claude-Opus-4.578.04±2.4276.14±1.1491.57±0.3361.25±2.50 Claude-Sonnet-4.577.25 ± 1.2077.05 ± 1.4588.83 ± 0.38 66.25 ± 1.25 Gemini-3-Flash74.65±1.9575.41±0.9590.07±0.3070.00±1.25 GPT-563.67±2.2772.31±1.2691.15±0.4562.50±1.25 GPT-4o59.28±2.0778.32±0.3287.50±0.1061.25±0.00 Gemini-2.5-Flash55.49±2.1375.05±0.8386.02±1.2258.33±4.39 Open-source≥70B Qwen2.5-72B63.08±2.7075.23±0.3283.63±0.3364.00±1.25 LLaMA-3.1-70B49.11±2.7975.58±1.1482.90±0.1551.25±3.31 Open-source<70B Gemma-3-27B53.29±2.1674.13±0.3280.67±0.1864.17±0.72 Gemma-2-9B46.71±2.7465.31±3.8370.43±0.6154.17±1.44 Qwen2.5-14B48.49±3.9364.66±0.8375.18±0.8560.83±3.82 Qwen2.5-7B46.11±2.8563.02±1.1469.70±0.2854.17±1.91 Gemma-3-4B38.12±2.2767.58±0.3261.30±0.1862.08±1.44 Mixtral-8x7B31.74±1.0459.38±0.6362.32±0.3458.33±0.72 LLaMA-3.1-8B38.93±3.2858.64±4.4560.35±3.6252.08±5.77 Arabic Models Fanar-1-9B43.51±2.4270.13±1.6774.60±0.3560.42±2.60 SILMA-9B49.32±21.7360.11±6.6153.85±5.5725.75±3.75 ALLAM-7B42.24±3.5564.75±3.8372.25±2.8356.50±2.00 Table 11: MCQ evaluation across 3 runs using each model’s recommended temperature. Values shown as mean±std. Rankings remain consistent across runs, confirming robustness of our main findings to decoding configuration. the model’s chain-of-thought against the gold ref- erence, diagnosed the root cause, and agreed on a category before recording it (full agreement af- ter joint adjudication). We use a shared taxonomy across both models: Misunderstanding Concept (correct setup, wrong principle applied),Concept Confusion(conflates two related but distinct con- cepts),Hallucination(generates facts not in the question),Question Misread(answers a different question),Calculation Mistake(arithmetic error), andDomain Knowledge Gap(lacks the terminol- ogy entirely). J.4.1 GPT-5 on Accounting Seventy percent of GPT-5 accounting errors stem from domain reasoning failures: Misunderstand- ing Concept (39%) and Concept Confusion (31%), while Calculation Mistakes account for only 20%. Extractive Summarization Annotation Guidelines Task overview.Annotators create extractive sum- maries by selecting the most important sentencesver- batimfrom each Arabic financial document. The goal is to produce a concise summary that preserves critical financial information and reflects the docu- ment’s main message. Document-level assessment. •Read the entire document before selecting any sen- tences. •Identify the document type (e.g., earnings report, regulatory announcement, market analysis, com- pany news). •Segment the text into sentences using Arabic punc- tuation marks (، ؛ .). •Target a summary length of approximately 30– 40% of the original document. Critical content to prioritize.Annotators must in- clude sentences containing: •Financial figures:ارً؇ح/اࠍފ؇߉ߵ(profits/losses), اߌߵادات(revenues),اܳྡྷފص اৎ٪ި ل۰(percentages). •Performance indicators:ஓި/اෛຶڰ؇ض(growth/de- cline),ارّڰ؇ع/۱ٴިط(increase/decrease). •Strategic decisions:اݿٺۜިاذ(acquisition), اࣖࢾ݁؇ج(merger),اܳٺިݿؕ(expansion). •Regulatory or official actions:ڢݠارات اୖ٭۰٪(au- thority decisions),اৎިاڣگ؇ت(approvals),اܳଫଐاۊ٭ݧ(li- censes). Sentence scoring.Each sentence is scored on a 1–5 scale: •5: Critical financial data or main announcement. •4: Important context or cause–effect explanation. •3: Supporting detail required for clarity. •2: General market or background information. •1: Redundant or generic statements. Selection procedure. •Select all sentences scored5. •Add sentences scored4until the target length is reached. •Include a sentence scored3only if necessary for coherence. Final validation.Before submission, annotators ver- ify that the summary: •Includes all key financial figures and the main an- nouncement. •Is coherent and understandable on its own. •Falls within the 30–40% length target. •Avoids repetition and generic background content. Common errors to avoid. •Selecting sentences based solely on position in the document. •Omitting numerical or regulatory information. •Including repetitive or stylistic filler content. •Exceeding the target summary length without jus- tification. Figure 18: Guidelines for extractive summarization an- notation of Arabic financial reports. This confirms our Section5.3finding that models rarely fail at arithmetic but frequently fail at know- ing which arithmetic to perform. Concretely, GPT- 5 (i) reaches correct intermediate calculations but selects the wrong accounting standard at the deci- sion point, (i) confuses closely related financial concepts (e.g., treating a direct relationship as in- verse), and (i) in the most striking cases (15% of errors), arrives at the correct answer through valid reasoning and then fabricates a non-existent rule to justify switching to a wrong option. These patterns indicate that the model’s weaknesses lie primarily in conceptual grounding rather than computational ability. In particular, errors often arise at the stage of mapping problem statements to the appropriate accounting principle or standard, rather than dur- ing numerical execution. This suggests that im- proving domain-specific reasoning and conceptual alignment is more critical than enhancing raw cal- culation capabilities for such tasks. J.4.2 GPT-5 on Extractive Summarization GPT-5 understands report content but fails at task execution, selecting background sentences over key financial figures (42.5% of errors), copying entire reports instead of respecting the 30–40% compression target (16.3%), and introducing con- tent from unrelated reports (11.3%). This explains why GPT-5 collapses on extractive summarization (ROUGE-L: 33.37) despite being our strongest model on open-ended reasoning: the task rewards verbatim selection discipline, not generative flu- ency. J.4.3 Gemini-3-Flash on Accounting Concept Confusion (27.8%) and Misunderstand- ing Concept (19.4%) together account for 47% of errors, particularly in auditing standards and for- eign currency hedging. Unlike GPT-5, Gemini-3- Flash also exhibits Hallucination (8.3%) and Ques- tion Misread (8.3%), while Calculation Mistakes remain rare (5.6%). J.4.4 Gemini-3-Flash on Business Reasoning Error (39.5%) and Concept Confu- sion (37.2%) dominate at 77% combined, con- centrating in Strategic Management, Marketing, and Entrepreneurship, areas requiring culturally grounded Arabic business knowledge. Domain Knowledge Gap (9.3%) appears as a distinct cate- gory: cases where the model lacks specialized Ara- bic business terminology entirely. Event–Cause Reasoning QA Annotation and Quality Control Task objective.Annotators construct one event–cause reasoning instance per Arabic financial report. Each instance consists of (i) an analytical question that requires causal or interpretive reasoning and (i) a concise expert-written answer grounded exclusively in the information provided in the report. The task evaluates whether models can explain whyfinancial or regulatory events occurred andwhat their implications are, rather than recalling isolated facts. Question construction.The question must: •Be analytical in nature (e.g., “why did this occur?” or “what does this indicate?”). •Connect multiple data points from the report (e.g., financial figures, growth rates, market reactions). •Avoid purely descriptive prompts (e.g., “what was the profit?”). •Be answerable using only information stated or implied in the report. Answer construction.The answer must: •Be written in Arabic and remain concise. •Rely exclusively on the content of the report, without external knowledge or speculation. •Explicitly reference numerical figures and percentages when available. •Provide economic or financial interpretation (e.g., performance drivers, risk implications, or market significance). •Preserve technical and domain-specific terminology. Focus areas.Annotators prioritize questions involving: •Market trend analysis and its implications. •Performance comparison between companies or sectors. •Economic significance of observed data patterns. •Risk assessment based on reported financial indicators. Quality control procedure.We enforce quality control through a multi-stage human validation workflow: 1.Pilot annotation.Two native Arabic financial experts independently annotate a pilot subset of 20 reports, each producing an event–cause question and an analytical answer. 2.Agreement assessment.We evaluate agreement at two complementary levels: •Event–cause identification: measured using Cohen’sκ, assessing consistency in identifying salient events and their causes. •Answer consistency: measured using ROUGE overlap between independently written answers, used as a consis- tency check rather than a correctness metric. 3.Calibration.Annotators review disagreements from the pilot phase, discuss ambiguous cases (e.g., implicit causal- ity, multi-factor events, overlapping economic drivers), and refine shared annotation criteria. This calibration aligns interpretation standards and reduces annotation drift. 4.Full annotation.After calibration, one expert annotates the remaining reports under the agreed guidelines. 5.Audit and correction.A senior annotator audits a random sample of completed annotations to verify that each instance: •Identifies a plausible event and its cause(s) supported by the report. •Includes relevant numerical evidence when available. •Provides an analytical explanation rather than a descriptive summary. Annotations that fail these checks are revised or discarded. Final dataset format.Each finalized instance consists of a financial report, one analytical event–cause question, and one expert-written answer. This format supports evaluation using both exact-match and partial-match metrics and enables controlled benchmarking of causal reasoning in Arabic financial text. Figure 19: Guidelines and quality control workflow for event–cause reasoning annotation in Arabic financial re- ports. J.4.5 Cross-Model Convergence The most striking finding is that GPT-5 and Gemini-3-Flash, built by different organizations with different architectures and training data, share the same dominant failure mode: conceptual con- fusion between related domain principles (70% for GPT-5, 47–77% for Gemini-3-Flash), with arith- metic errors rare in both (20% and 5.6% respec- tively). This convergence strongly suggests that Arabic financial reasoning remains a genuine chal- lenge for frontier models, not an artifact of our eval- uation design. K Decoding Configuration and Variance Analysis Our main evaluations in the paper use greedy de- coding (temperature 0) for reproducibility. To ver- ify that our conclusions are not artifacts of this choice, particularly given that some models (e.g., GPT-5) do not support temperature 0, and oth- ers (e.g., Qwen3) recommend non-zero tempera- tures for best performance, we re-ran all MCQ LLM-as-a-Judge Rubric for Fatwa QA (Arabic) Role.You are an expert evaluator in Islamic fatwa (iftāʾ). Inputs (provided each time). •category(optional context) – may be empty (e.g., riba, zakat, takaful) •prompt– the full Arabic prompt shown to the model (instructions + question) •ground_truth– the reference fatwa answer (Arabic) •candidate_answer– the model answer to evaluate (Arabic) Task.Judge how wellcandidate_answermatchesground_truthinruling (ḥukm),justification, andoperative con- straints/qualifications. Prioritize doctrinal correctness and required conditions/exceptions. Do not penalize stylistic paraphrase if the core ruling and constraints are preserved. Be concise and deterministic. Scoring (sum to exactly 10). 1.Coverage of core ruling (0–4).The candidate must clearly state the same central hukm (e.g., permissibility/prohi- bition, validity/invalidity)andinclude the key justification present in the ground truth. One-word/minimal answers without essential justification should receive a much lower score (e.g., 0–1). 2.Conditions, exceptions, constraints (0–2).Does it retain critical restrictions, qualifiers, or carve-outs that materi- ally affect the ruling? 3.Doctrinal/factual accuracy (0–2).No misstatements that would change the fatwa; no implicit legalization of pro- hibited elements (e.g., ribā); no misleading generalizations or invented requirements. 4.Clarity & Arabic language quality (0–1).Clear Arabic, understandable structure, minimal ambiguity appropriate for a fatwa answer. 5.Directness & fatwa format (0–1).Directly answers the question; avoids long digressions; phrasing suitable for a fatwa. Critical checks (true/false). •contradicts_ground_truth: Does the candidate contradict the central ruling? •omits_critical_conditions: Does it omit key conditions/exceptions that change the ruling? •introduces_unlawful_elements: Does it introduce/normalize prohibited elements (e.g., ribā)? •hallucinated_citations: Misleading/fabricated sources claimed that distort the ruling? •non_answer_or_evasive: Does it avoid giving a clear ruling? •off_topic_or_unsafe: Off-topic or otherwise inappropriate? Output format (strict).Outputonlyvalid JSON (no prose, no code fences), following this schema: ”scores”: ”coverage_core_ruling”: <float 0-4>, ”conditions_exceptions”: <float 0-2>, ”factual_doctrinal_accuracy”: <float 0-2>, ”clarity_language”: <float 0-1>, ”directness_format”: <float 0-1> , ”overall”: <float 0-10>, ”critical_checks”: ”contradicts_ground_truth”: <true/false>, ”omits_critical_conditions”: <true/false>, ”introduces_unlawful_elements”: <true/false>, ”hallucinated_citations”: <true/false>, ”non_answer_or_evasive”: <true/false>, ”off_topic_or_unsafe”: <true/false> , ”note”: ”<short NOTE in NOTE_LANG>” Figure 20: Evaluation rubric used for LLM-based judgment of fatwa QA responses. evaluations 3 times using each model’s recom- mended temperature settings. Table 11reports mean±std across runs. Rankings remain consis- tent across runs and across temperature settings, confirming that our findings are robust to decoding configuration. The low variance observed across repeated runs further indicates that performance differences are stable rather than driven by sam- pling noise. Overall, these results suggest that model comparisons in our benchmark are reliable under both deterministic and stochastic decoding regimes. Notably, models that perform strongly under greedy decoding maintain their relative ad- vantage under higher-temperature settings, indicat- ing that gains are not dependent on sampling vari- ability. Similarly, lower-performing models do not LLM-as-a-Judge Rubric for Islamic Fi- nance QA (Arabic) Role.You are an expert evaluator in Islamic finance (Fiqh al-mu'āmalāt). Inputs (provided each time). •topic(optional context) – may be empty •question– in Arabic •ground_truth– the reference correct answer (Arabic) •candidate_answer– the model answer to evalu- ate (Arabic) Task.Judge how wellcandidate_answermatches ground_truthinmeaning,ruling,justification, and constraints. Prioritize doctrinal correctness and com- pleteness of key conditions/exceptions. Do not pe- nalize stylistic paraphrase if the core ruling and con- straints are preserved. Be concise and deterministic. Scoring (sum to exactly 10). 1.Coverage of core ruling (0–4). 2.Conditions, exceptions, constraints (0–2). 3.Doctrinal/factual accuracy (0–2). 4.Clarity & Arabic language quality (0–1). 5.Directness & on-topic (0–1). Critical checks (true/false). •contradicts_ground_truth •omits_critical_conditions •introduces_unlawful_elements •hallucinated_citations •non_answer_or_evasive •off_topic_or_unsafe Outputformat(strict).Outputonlyvalid JSON (no prose, no code fences), following this schema: ”scores”: ”coverage_core_ruling”: <float 0-4>, ”conditions_exceptions”: <float 0-2>, ”factual_doctrinal_accuracy”: <float 0-2>, ”clarity_language”: <float 0-1>, ”directness_format”: <float 0-1> , ”overall”: <float 0-10>, ”critical_checks”: ”contradicts_ground_truth”: <true/false>, ”omits_critical_conditions”: <true/false>, ”introduces_unlawful_elements”: <true/false>, ”hallucinated_citations”: <true/false>, ”non_answer_or_evasive”: <true/false>, ”off_topic_or_unsafe”: <true/false> , ”note”: ”<short NOTE in NOTE_LANG>” Figure 21: Evaluation rubric used for LLM-based judg- ment of Islamic finance QA responses. benefit significantly from increased randomness, suggesting that errors are rooted in systematic lim- itations rather than decoding strategy. This con- sistency reinforces the validity of our evaluation pipeline across diverse inference settings. It also highlights that task difficulty and domain reason- ing, rather than decoding configuration, are the primary drivers of performance differences in our benchmark. K.1 Doctrinal Variation in Shari’ah Rulings Islamic jurisprudence is not monolithic: the four major Sunnimadhahib(Hanafi, Maliki, Shafi’i, Hanbali), alongside Shia schools, may offer differ- ent rulings on the same financial question. This raises a natural concern for benchmark construc- tion: do our reference answers represent a sin- gle doctrinal position that could penalize models trained on legitimate alternative rulings? We analyzed this question during dataset con- struction and report our findings here. Category # Samples Significant Dispute Sukuk62 (33%) Takaful3810 (26%) Zakat792173 (22%) Gharar14920 (13%) Ijara10212 (12%) Murabaha23420 (9%) Maysir646 (9%) Riba40732 (8%) Total289 (14.4%) Table 12: Distribution of doctrinal variation in Islamic finance categories. Variation concentrates in zakat (cal- culation methods differ across schools), takaful (a mod- ern instrument with no classical precedent and evolv- ing rulings), and sukuk (small sample, actively debated across regulatory bodies). For all disputed cases, ref- erence answers explicitly present alternative valid posi- tions, and our rubric accepts any legitimate ruling as correct. (1) Most questions test consensus rulings. Seventy-four percent of samples involve cross- madhabagreement on established Islamic finance principles, the prohibition ofriba(usury), contract invalidation due togharar(excessive uncertainty), the impermissibility ofmaysir(gambling), rather than narrow inter-school disputes. (2) Evaluation targets the ḥukm , not the evi- dencepath.Reference fatwas and model outputs naturally vary in cited Qur’anic verses,ḥadīth,fiqh sources, and reasoning detail, making exact-match evaluation infeasible. We therefore score at the ḥukm(ruling) level: the rubric evaluates the final ruling and its operative constraints. A model cit- ing different but valid evidence while reaching the correct ruling is not penalized. LLM-as-a-Judge Rubric for Financial Analysis & Capital Markets (Arabic) Role.You are an expert evaluator in financial analysis and capital markets. Inputs (provided each time). •prompt– the full Arabic prompt (report/excerpt + question) shown to the model •ground_truth– the reference ideal analytical answer (Arabic) •candidate_answer– the model answer to evaluate (Arabic) Task.Judge how wellcandidate_answermatchesground_truthinconclusions,reasoning, anduse of provided fig- ures. Prioritize factual/quantitative fidelity, correct interpretation of financial concepts (e.g., spreads, yields, coverage, issuance, capital structure, Basel I, supply/demand dynamics), and avoidance of hallucinated data. Do not penalize stylistic paraphrase if core insights and numeric takeaways align with the reference. Scoring (sum to exactly 10). 1.Core conclusion alignment (0–4).Does the candidate capture the main thesis and key takeaways of the ground truth (what/why/so-what)? 2.Quantitative fidelity & use of figures (0–2).Correctly cites/uses the reported numbers (e.g., percentages, amounts, maturities, oversubscription) without inventing or altering figures. Any simple computations/comparisons must be consistent. 3.Financial reasoning soundness (0–2).Causality and mechanisms are plausible and consistent with standard fi- nance/econ logic (e.g., pricing vs. credit risk, duration/tenor structure, demand/oversubscription signals, capital adequacy). 4.Clarity & Arabic language quality (0–1).Clear Arabic, coherent structure, minimal ambiguity. 5.Directness & on-topic grounding (0–1).Answers what was asked; stays anchored in the provided scenario/data (no generic filler). Critical checks (true/false). •contradicts_ground_truth: contradicts the central conclusion of the reference •fabricates_or_alters_numbers: introduces numbers not present or materially distorts reported figures •hallucinates_context_or_sources: injects external context/sources not in the prompt that change the assessment •flawed_financial_logic: serious finance/econ reasoning error that would mislead the conclusion •non_answer_or_evasive: avoids providing an analytical answer •off_topic_or_unsafe: off-topic or otherwise inappropriate Output format (strict).Outputonlyvalid JSON (no prose, no code fences). Return JSON strictly in this schema (all fields required): ”scores”: ”coverage_core_conclusion”: <float 0-4>, ”quantitative_fidelity”: <float 0-2>, ”financial_reasoning”: <float 0-2>, ”clarity_language”: <float 0-1>, ”directness_grounding”: <float 0-1> , ”overall”: <float 0-10>, ”critical_checks”: ”contradicts_ground_truth”: <true/false>, ”fabricates_or_alters_numbers”: <true/false>, ”hallucinates_context_or_sources”: <true/false>, ”flawed_financial_logic”: <true/false>, ”non_answer_or_evasive”: <true/false>, ”off_topic_or_unsafe”: <true/false> , ”note”: ”<short NOTE in NOTE_LANG>” Figure 22: Evaluation rubric used for LLM-based judgment of financial analysis and event–cause reasoning tasks. (3) Quantified dispute distribution.For the 26% of samples with some degree of dispute, Ta- ble12reports how often the reference answer it- self flags a significant disagreement across recog- nized schools. “Significant Dispute” means two or more recognizedmadhahibhold materially differ- ent rulings—i.e., theḥukmitself changes (e.g., per- missible vs. impermissible), not just the supporting evidence. (4) Manual error analysis confirms failures are genuine,notdoctrinal.To verify that low scores do not stem from penalizing valid alternative po- sitions, we analyzed 500 randomly sampled er- rors (Figure 6). Model failures are unambigu- ous: wrong rulings (25.2%), fabricated evidence (12.1%), and misquotedḥadīth(2.0%)not cases where the model offered a legitimate but different school’s position.