Paper deep dive
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark
Charalampos Mastrokostas, Nikolaos Giarelis, Nikos Karacapilidis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/21/2026, 1:11:14 AM
Summary
This paper introduces DemosQA, a novel Greek Question Answering dataset derived from social media (Reddit) to capture cultural and social nuances, addressing the lack of community-driven resources. It proposes a memory-efficient evaluation framework using 4-bit quantization and empirically evaluates 11 monolingual and multilingual Large Language Models (LLMs) across six Greek QA datasets using three prompting strategies to assess performance on under-resourced languages.
Entities (11)
Relation Signals (10)
DemosQA → derivedfrom → r/greece
confidence 95% · DemosQA encompasses a wide variety of questions and answers... extracted from the “r/greece” subreddit
GPT-4o-mini → evaluatedon → Greek QA
confidence 95% · empirically evaluate 11 monolingual and multilingual LLMs ... including GPT-4o mini
4-bit model quantization → usedby → Evaluation Framework
confidence 95% · first framework to leverage 4-bit model quantization
Llama Krikri 8B → basedon → Llama-3-8B
confidence 92% · Llama Krikri 8B ... built on ... Llama 3 8B
Meltemi 7B → basedon → Mistral-7B
confidence 92% · Meltemi 7B ... built on Mistral 7B
DemosQA → comparedwith → Greek Truthful QA
confidence 90% · We conducted a comparative analysis of DemosQA and five existing Greek QA datasets
DemosQA → comparedwith → Belebele
confidence 90% · We conducted a comparative analysis of DemosQA and five existing Greek QA datasets
DemosQA → comparedwith → Greek Medical MCQA
confidence 90% · We conducted a comparative analysis of DemosQA and five existing Greek QA datasets
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent advancements in Natural Language Processing and Deep Learning have enabled the development of Large Language Models (LLMs), which have significantly advanced the state-of-the-art across a wide range of tasks, including Question Answering (QA). Despite these advancements, research on LLMs has primarily targeted high-resourced languages (e.g., English), and only recently has attention shifted toward multilingual models. However, these models demonstrate a training data bias towards a small number of popular languages or rely on transfer learning from high- to under-resourced languages; this may lead to a misrepresentation of social, cultural, and historical aspects. To address this challenge, monolingual LLMs have been developed for under-resourced languages; however, their effectiveness remains less studied when compared to multilingual counterparts on language-specific tasks. In this study, we address this research gap in Greek QA by contributing: (i) DemosQA, a novel dataset, which is constructed using social media user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist; (ii) a memory-efficient LLM evaluation framework adaptable to diverse QA datasets and languages; and (iii) an extensive evaluation of 11 monolingual and multilingual LLMs on 6 human-curated Greek QA datasets using 3 different prompting strategies. We release our code and data to facilitate reproducibility.
Tags
Links
- Source: https://arxiv.org/abs/2602.16811v1
- Canonical: https://arxiv.org/abs/2602.16811v1
Trouble viewing inline? Open PDF directly →
Full Text
53,234 characters extracted from source content.
Expand or collapse full text
Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark Charalampos Mastrokostas, Nikolaos Giarelis, Nikos Karacapilidis Industrial Management and Information Systems Lab, MEAD University of Patras, Rio Patras, Greece cmastrokostas@ac.upatras.gr, giarelis@ceid.upatras.gr, karacap@upatras.gr Abstract Recent advancements in Natural Language Processing and Deep Learning have enabled the development of Large Language Models (LLMs), which have significantly advanced the state-of-the-art across a wide range of tasks, including Question Answering (QA). Despite these advancements, research on LLMs has primarily targeted high-resourced languages (e.g., English), and only recently has attention shifted toward multilingual models. However, these models demonstrate a training data bias towards a small number of popular languages or rely on transfer learning from high- to under-resourced languages; this may lead to a misrepresentation of social, cultural, and historical aspects. To address this challenge, monolingual LLMs have been developed for under-resourced languages; however, their effectiveness remains less studied when compared to multilingual counterparts on language-specific tasks. In this study, we address this research gap in Greek QA by contributing: (i) DemosQA, a novel dataset, which is constructed using social media user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist; (i) a memory-efficient LLM evaluation framework adaptable to diverse QA datasets and languages; and (i) an extensive evaluation of 11 monolingual and multilingual LLMs on 6 human-curated Greek QA datasets using 3 different prompting strategies. We release our code and data to facilitate reproducibility. Keywords: Large Language Models, Natural Language Processing, Question Answering, Greek Language, Language Resources, Social Media, Greek NLP 1. Introduction Research on the field of Natural Language Pro- cessing (NLP) focuses on the development of meth- ods that enable machines to process and under- stand human language. Recent advances in NLP and Deep Learning have led to the emergence of Large Language Models (LLMs), which demon- strate strong natural language understanding and reasoning capabilities, while achieving state-of-the- art performance across a plethora of tasks (Minaee et al., 2025; Naveed et al., 2025). Often referred to as foundation models, LLMs are trained on mas- sive corpora using substantial computational re- sources (e.g., GPU clusters) and can subsequently be adapted to a variety of NLP tasks with compara- tively fewer resources (Bommasani et al., 2022). Earlier LLMs, such as GPT-3 (Brown et al., 2020) and Llama-2 (Touvron et al., 2023), primarily sup- ported English, due to being predominantly trained on English corpora. In contrast, more recent mod- els, such as GPT-4 (Achiam et al., 2024), Llama 3 (Grattafiori et al., 2024) and Gemma 2 (Riviere et al., 2024), demonstrate multilingual capabili- ties by training on corpora from a diverse set of languages. Despite these advances, multilingual LLMs exhibit several limitations. Recent studies have highlighted that: (i) they do not adequately address the imbalance of training resources be- tween high- and under-resourced languages (Blasi et al., 2022); (i) they often apply the same learning techniques, without considering the grammatical and syntactical differences among languages (Blasi et al., 2022); and (i) they may misrepresent social, cultural and historical aspects of underrepresented languages (Qin et al., 2025). Consequently, the per- formance of multilingual models can vary substan- tially across languages and tasks. In addition, their evaluation remains largely limited to a small num- ber of popular languages, with under-resourced ones rarely assessed in a comprehensive way. This study focuses on the Question Answering (QA) task, which has been significantly advanced by LLMs (Minaee et al., 2025), with particular em- phasis on Standard Modern Greek. The language’s unique alphabet, rich morphology, and complex syntax make building accurate NLP models espe- cially challenging. Moreover, recent reviews under- line the scarcity of models, datasets, and compar- ative evaluations for Greek QA (Bakagianni et al., 2025; Papantoniou and Tzitzikas, 2024; Giarelis et al., 2024c). Despite these challenges, only a few studies have explored Greek LLMs. Specifically, two recent works introduce the first Greek LLMs, re- porting state-of-the-art performance across several Greek NLP tasks (Voukoutis et al., 2024; Rous- sis et al., 2025); another work (Pavlopoulos et al., 2025) evaluates the strengths and weaknesses of both an open-weights and a proprietary LLM (GPT- 4o mini (Hurst et al., 2024)) on several Greek NLP tasks, but not including QA. Building on the gaps and challenges highlighted arXiv:2602.16811v1 [cs.CL] 18 Feb 2026 above, this study aims to advance Greek QA through the following contributions: •We introduce DemosQA, a novel Greek QA dataset, which is constructed using social me- dia user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist; • We propose a memory-efficient LLM evalua- tion framework that can be adapted to different QA datasets and languages. To our knowl- edge, it is the first framework to leverage 4-bit model quantization (Dettmers and Zettlemoyer, 2023), reducing hardware requirements (large and costly GPUs) with minimal loss of accu- racy; • We empirically evaluate 11 monolingual and multilingual LLMs supporting Greek on 6 human-curated Greek QA datasets using 3 different prompting strategies; • We make our code and data public to facilitate the reproducibility of our research 1 . Research questions (RQs) investigated in this study include: • RQ1: How do open-weights monolingual LLMs perform compared to open-weights multilin- gual LLMs on Greek QA? •RQ2: Can open-weights LLMs achieve the state-of-the-art performance of a proprietary LLM (GPT-4o mini) on Greek QA? •RQ3: How do different prompting strategies influence model accuracy across Greek QA datasets? •RQ4: Is it possible to construct a high-quality, human-curated QA dataset from social media content? The remainder of this paper is structured as fol- lows: LLMs and QA datasets supporting Greek are presented in Section 2. The proposed QA dataset is described in Section 3, while our evalu- ation framework and experimental results are pre- sented in detail in Section 4. Concluding remarks, future research directions, limitations and ethical considerations are discussed in Section 5. 2. Related Work In this study, we focus on LLMs with at least 7 billion parameters, as such models demonstrate substan- tially stronger natural language understanding and 1 The code will be made public after the peer-review process. The dataset is available at the following link: https://huggingface.co/datasets/IMISLab/DemosQA reasoning capabilities compared to smaller ones (Minaee et al., 2025; Naveed et al., 2025). We consider their instruction-tuned variants, which are optimized for in-context learning and can be directly prompted to perform a variety of NLP tasks, unlike their base counterparts. Throughout this paper, model sizes are abbreviated (e.g., 7B denotes 7 billion parameters). 2.1. Greek and Multilingual Large Language Models This subsection presents instruction-tuned monolin- gual (Greek) LLMs and multilingual LLMs that sup- port Greek without additional post-training. The lat- ter are predominantly post-trained for a small num- ber of popular languages, resulting in the marginal- ization of under-resourced languages. Meltemi 7B (Voukoutis et al., 2024) and Llama Krikri 8B (Roussis et al., 2025) are the first Greek LLMs, built on Mistral 7B (Jiang et al., 2023) and Llama 3 8B (Grattafiori et al., 2024), respectively. They were adapted from their base models through additional pre-training on large Greek corpora fol- lowed by instruction tuning, which enabled their conversational capabilities. Experimental results reported by the authors show that both Greek LLMs outperform their original instruction-tuned counter- parts across several Greek NLP tasks. The Mistral (Jiang et al., 2023) model family in- cludes two multilingual LLMs. Mistral Nemo 12B demonstrates strong performance on several high- resource European and Asian languages; however, its performance on under-resourced languages has not been evaluated by its authors. Ministral 8B performs well on multiple English and multilingual benchmarks; however, the exact number of sup- ported languages has not been disclosed. Llama 3.1 8B (Grattafiori et al., 2024) and Gemma 2 9B (Riviere et al., 2024) follow a similar training strategy. Both were pre-trained on large, multilingual web-scale corpora that also include mathematical and code reasoning data. Despite considering many languages during pre-training, these models officially support only a limited set of high-resource languages (e.g., eight languages for Llama 3.1 8B). Teuken 7B (Ali et al., 2024) and EuroLLM 9B (Martins et al., 2025) follow a similar multilingual training strategy. Unlike other models, both employ custom multilingual tokenizers to officially support 24 and 35 languages, respectively, including Greek. However, a key limitation of these models is their imbalanced language distribution, with most under- resourced European languages being severely un- derrepresented in the training data. Aya Expanse 8B (Dang et al., 2024) officially supports 23 languages, including Greek, and is built on the same architecture as Command R 7B (Aakanksha et al., 2025). Aya Expanse 8B em- ploys a cross-lingual transfer learning technique that trains expert models for linguistically related language groups using translated synthetic data from English. The weights from the best-performing expert models are then merged to form the final instruction-tuned model. In summary, although the availability of multi- lingual and Greek LLMs has increased, their true capabilities in Greek remain underexplored due to the scarcity of multiple human-curated Greek QA datasets across diverse domains for systematic evaluation, as confirmed by previous studies (Bak- agianni et al., 2025; Papantoniou and Tzitzikas, 2024). 2.2. Greek and Multilingual QA Datasets In our search for Greek or multilingual QA datasets supporting Greek, we focused on high-quality, human-curated resources to avoid machine transla- tion errors. This choice was motivated by research outcomes revealing that non-curated, machine- translated datasets can negatively impact the evalu- ation of text generation tasks (Graham et al., 2019). Using these criteria, we identified five QA datasets suitable for the purposes of our study. The Greek Medical MCQA dataset (Voukoutis et al., 2024) contains 2,034 QA pairs from the medical exams of the Hellenic National Academic Recognition and Information Center (DOATAP 2 ). Of these, 1,602 pairs were reserved for model training, with the remaining ones used for validation. Most QA pairs consist of a question, five possible answer options and a single correct one. The Greek Truthful QA (Voukoutis et al., 2024) is a human-curated, machine-translated version of Truthful QA (Lin et al., 2022). It contains 817 questions designed to challenge misconceptions or false beliefs held by humans. For our evaluation, we select its multiple-choice version, in the hardest difficulty setting (mc1_targets), where only one an- swer is considered correct out of a list of possible ones. Unlike other considered QA datasets, the number of possible answers per question varies in Truthful QA. BELEBELE (Bandarkar et al., 2024) is a human- curated dataset containing 900 QA pairs available in 122 languages, including Greek. Each entry consists of a short passage, a question, four candi- date answers, and a single correct one. Although framed as a QA dataset, BELEBELE primarily tar- gets reading comprehension to evaluate language understanding and transfer capabilities of LLMs. The authors make clear that the dataset is English- centric, since the QA pairs were translated from 2 https://w.doatap.gr English and do not fully capture the cultural or lin- guistic nuances of non-English languages. INCLUDE (Romanou et al., 2024) is a multiple- choice QA dataset containing 197,243 QA pairs across 44 languages collected from local exams. It spans a comprehensive range of topics, including academic exams (e.g., Humanities, STEM Fields, Law, etc.) and professional certifications, thus en- abling per-language assessment of regional and domain-specific knowledge in multilingual LLMs. The Greek part includes a test subset of 552 QA pairs, where each question is accompanied by four candidate answers and a single correct one. Greek ASEP MCQA (Kyriazi and Prokopidis, 2025) comprises 1,200 multiple-choice questions and their corresponding answers, extracted from the Greek Supreme Council for Civil Personnel Se- lection (ASEP) exams. The dataset covers several topics, including Greek law, politics, public admin- istration, e-governance, and modern Greek history. As in the previous dataset, each question is ac- companied by four candidate answers and a single correct one. Collectively, these five human-curated datasets offer a valuable foundation for evaluating Greek and multilingual LLMs across diverse domains. Never- theless, they are limited in capturing community- driven content, motivating the creation of De- mosQA, a novel dataset of Greek QA pairs sourced from social media. 3. The DemosQA Dataset Reddit is a popular social media platform structured as a collection of forums, where users engage in discussions on a broad range of topics, from every- day life to specialized domains such as economics and politics (Medvedev et al., 2019). Each forum, known as a subreddit, enables users to create posts, participate in comment-based discussions, and collectively rank content through an upvoting or downvoting mechanism that promotes the most relevant contributions. Moreover, each subreddit is moderated by a group of trusted users responsi- ble for enforcing community rules and maintaining discussion quality. Proferes et al. (2021) reviewed more than 700 research works utilizing Reddit data across vari- ous NLP applications, confirming its value as a rich and diverse source of user-generated text. For the Greek language, several subreddits exist, with “r/greece” being the largest and most active one, comprising more than 260,000 members. Posts within this community are categorized by topic, re- viewed by moderators before publication, and gov- erned by explicit guidelines discouraging offensive or irrelevant content. To the best of our knowledge, no prior work has explored QA datasets derived from Greek social media. To address this gap, we introduce De- mosQA, the first dataset of community-reviewed Greek QA pairs collected from social media. Its name derives from the Greek word "δῆμος" (mean- ing “the people”), reflecting the dataset’s demo- cratic and participatory nature. DemosQA encom- passes a wide variety of questions and answers spanning domains such as everyday life, history, science, and politics, providing a valuable resource for studying real-world discussions in Greek and advancing research on language understanding within socially grounded contexts. The DemosQA dataset comprises questions ex- tracted from the “r/greece” subreddit, each accom- panied by four candidate answers, the selected best answer and its index, the date of posting, and the corresponding Reddit post ID. Candidate an- swers are ranked based on community voting, with the highest-upvoted response designated as the reference answer. This community-driven ranking mechanism not only ensures that the dataset cap- tures genuine user preferences but also establishes a meaningful benchmark for assessing how closely large language models align with human judgments of response quality. The complete dataset collec- tion and curation process is detailed in the following subsections. 3.1. Data Collection Several tools have already been developed for col- lecting Reddit data (Proferes et al., 2021); however, their use has become increasingly restricted and costly due to recent changes to Reddit’s API ac- cess policies (Wright, 2024). Consequently, this study employs the PRAW 3 Python library, which provides controlled access to Reddit content (lim- ited to approximately 200 posts per search request), while fully adhering to the platform’s official API guidelines. To retrieve a larger volume of data (i.e., thousands of posts), we manually compiled a list of 120 Greek search keywords and combined them with multiple sorting filters based on post popularity (i.e., “top”, “hot”, “relevance”, “comments”, “new”) and time range (i.e., “all”, “year”, “month”, “week”). Our data crawling script iteratively applies these search combinations and introduces short time de- lays between requests to ensure compliance with the API’s rate limits. To directly focus on Greek QA content, we col- lected posts from the r/greece subreddit catego- rized underερωτήσεις(questions). For each post, we extracted its ID, title, main text, publication date, and responses. In addition to our primary collec- tion, we incorporated data from GreekReddit (Mas- trokostas et al., 2024), which consists exclusively 3 https://pypi.org/project/praw/ of categorized Reddit posts without user answers. We identified question posts from GreekReddit by detecting the presence of question marks and then used their IDs to retrieve the corresponding an- swers. 3.2. Data Pre-Processing We applied a series of pre-processing techniques to ensure the quality and consistency of the proposed dataset. First, we identified engaging question posts with a minimum of five upvotes and five an- swers to ensure a sufficient candidate pool. Then, we removed duplicates and posts containing only images without textual content. We also excluded posts flagged as adult content to assure the overall appropriateness of the dataset. In the resulting subset, we collected the ten highest-upvoted answers (wherever available) for each post to serve as a set for further manual fil- tering. Since answers originate from each post’s comment tree, only the top-level comments that directly respond to the question were considered, thus limiting secondary discussion responses. Fi- nally, the remaining QA pairs were cleaned by re- moving redundant whitespace characters. Follow- ing this pre-processing pipeline, more than 2,100 samples were retained for manual curation. 3.3. Data Curation To further enhance dataset quality, we manually reviewed all pre-processed data through the follow- ing steps. First, we conducted a thorough review to remove all questions and answers that contain offensive language, hate speech or "troll" content (e.g., sarcasm, misleading information). This step ensured the informative and neutral tone of the dataset (see Table 8 in Appendix A). Second, to address the fact that in many posts the question was not properly posed in the title and/or the main text, we concatenated these two fields. Third, instead of relying solely on the upvote count, we manually selected the four most relevant comments to serve as candidate answers for each question. This step limits comments that do not directly address the post question. After this step, we marked the highest upvoted answer as the best one. Finally, we randomly shuffled the order of an- swers to mitigate potential LLM selection bias to- ward the first option (Khatun and Brown, 2024). The resulting dataset comprises 600 curated ques- tions, each paired with four candidate answers and one best answer. 3.4. Comparative Analysis of Greek QA Datasets We conducted a comparative analysis of DemosQA and five existing Greek QA datasets, considering the number of documents for evaluation, dataset domain and a series of word count percentiles for questions and correct answers (Table 1). Most of the compared datasets are intended solely for evaluation, with the exception of Greek Medical MCQA, for which we used the validation subset. Collectively, these datasets cover diverse domains, providing a representative basis for cross-dataset QA evaluation. As shown in Table 1, there exists substantial vari- ation in both question and answer lengths across datasets. DemosQA features the longest ques- tions, with a median length of 84.5 words, fol- lowed by BELEBELE, which also includes addi- tional contextual passages. In contrast, the remain- ing datasets contain considerably shorter ques- tions, with a median (P50) ranging from 9 to 13 words. Regarding answer length, DemosQA in- cludes the longest answers (median: 54.5 words). Additionally, DemosQA contains answers of vary- ing length, which is evident from the numeric differ- ences across the percentiles. The other datasets are characterized by notably concise answers, mostly containing fewer than 20 words. This comparative analysis highlights the linguistic diversity and complexity of DemosQA, distinguish- ing it from other Greek QA datasets that typically contain shorter and more uniform QA pairs. These characteristics make DemosQA a valuable bench- mark for assessing LLM performance in realistic, community-driven Greek text. In the following sec- tion, we present our experimental setup, describing the models, prompting strategies, and evaluation framework employed to measure the capabilities of the selected LLMs in Greek QA tasks. 4. Experiments To assess the performance of LLMs on Greek QA, we conducted a series of experiments. This section outlines the experimental setup, the adopted eval- uation framework, and the results obtained. Our goal is to examine the effectiveness of both multi- lingual and Greek-adapted LLMs in understanding and generating accurate responses to Greek ques- tions across diverse topics. 4.1. Setup For our experiments, we used a computer equipped with an Intel Core i5 CPU, 64 GB of RAM, and an NVIDIA GPU with 12 GB of VRAM. LLM inference was developed using Huggingface Transformers (Wolf et al., 2020). Since LLMs typically require large amounts of VRAM, which are only available in high-end GPUs, we applied a 4-bit model quanti- zation technique (Dettmers and Zettlemoyer, 2023) as implemented in the bitsandbytes project 4 . This approach substantially reduces memory require- ments for LLM inference with minimal accuracy loss; for instance, the weights of a 7B-parameter model require approximately 14 GB of VRAM in 16-bit precision, but only 3.5 GB in 4-bit precision. Model performance on multiple-choice QA tasks was evaluated using the accuracy metric from the scikit-learn library (Pedregosa et al., 2011). All models were deployed locally, except for GPT-4o mini, which was accessed through the OpenAI API. 4 https://pypi.org/project/bitsandbytes/ Dataset# Docs Domain TypeP5 P25 P50 Mean P75 P95 P99 DemosQA600Social Question 26 53 84.5 103.04 132.25 243 347.7 Answer 11 31 54.5 80.22 105 222 362.04 BELEBELE (Greek) 900General Question 58 77 98 100.33 121 147 181 Answer 1 344.96711 15.01 Greek Medical MCQA 432Medical Question 3 699.96121831 Answer 1 2 3.5 4.67612 17.69 Greek Truthful QA 817General Question 5 79 11.031322 40.68 Answer 2.8 7 10 10.021318 21.84 Greek ASEP MCQA 1200 Civil Service Question 4 6 10 11.41426 37.01 Answer 2 478.38 11.25 2127 INCLUDE (Greek) 552Education Question 5 9 13 22.8 27.25 75.9 129.96 Answer 1 357.2392035 Table 1: Number of evaluation documents and domain per dataset, with a statistical overview of their question and answer word counts. 4.2. Evaluation Framework We evaluated the considered LLMs (see Table 2) on several multiple-choice QA tasks. To ensure reproducibility, we set a fixed random seed and employed greedy decoding, which corresponds to a model temperature of 0.0 (Renze and Guven, 2024). The correct answer was extracted from the model’s output using rule-based parsing and reg- ular expressions, since instruction-tuned models often include greetings or explanatory text along- side their selected answer. If a valid answer could not be extracted, it was labeled as “No match”. Furthermore, we employed three prompting strategies to identify the most effective approach. The first one, the Instruction prompt, directs the model to select the best answer. The second, the Role prompt, assigns a specific role to the model (e.g., “You are a language model for the Greek language”). Following this role assignment, the model is instructed to select the best answer. The third strategy, a zero-shot Chain-of-Thought (CoT) prompt, builds on the Role prompt and additionally instructs the model to reason step-by-step (Kojima et al., 2022) (See Table 7 in Appendix A for the ex- act prompts). Since the evaluation datasets have different formats, our framework standardizes them to ensure consistent evaluation across all models. 4.3. Experimental Results This subsection presents the experimental results for the QA tasks. Tables 3–5 summarize the exper- imental results of each model across all datasets, using the instruction, role and CoT prompting strate- gies, respectively. Finally, Table 6 reports on the average accuracy scores across all three strate- gies. Table 3 reports the experimental results for the instruction prompt. Specifically, GPT-4o mini achieves the highest accuracy across all datasets, while a clear performance gap is observed between this proprietary model and all other open-weight ones in Greek Medical MCQA, Greek ASEP MCQA, and INCLUDE (Greek). Gemma 2 9B ranks second overall, having similar performance with GPT-4o mini on Greek Truthful QA and BELEBELE (Greek). The third best performing model across all datasets is Llama Krikri 8B, which equals the performance of GPT-4o mini in DemosQA. In contrast, most multi- lingual LLMs underperform compared to the above- mentioned models, with Teuken 7B v0.4 exhibiting the lowest accuracy scores. Table 4 presents the experimental results for the role prompt. Similarly to Table 3, there is a wide performance gap between GPT-4o mini and the rest of the models for Greek Medical MCQA, Greek ASEP MCQA and INCLUDE (Greek). However, for the rest of the datasets considered in our study, Gemma 2 9B attains the best accuracy scores on Greek Truthful QA and BELEBELE (Greek), while Llama Krikri 8B achieves the best accuracy score in DemosQA and demonstrates comparable perfor- mance to Gemma 2 9B. The rest of the models un- derperform compared to the aforementioned ones, with Teuken 7B v0.4 having the worst performance. Table 5 reports on the experimental results col- lected for the CoT prompt. Similarly to the previous tables, there is a large performance gap between GPT-4o mini and the rest of the models for Greek Medical MCQA, Greek ASEP MCQA and INCLUDE (Greek). In contrast with Table 4, the best perfor- LLMFull Model Name Greek Adapted Open- Weights GPT-4o minigpt-4o-mini-2024-07-18-- Gemma 2 9Bgoogle/gemma-2-9b-it-D Llama Krikri 8Bilsp/Llama-Krikri-8B-InstructDD Meltemi 7B v1.5ilsp/Meltemi-7B-Instruct-v1.5D Llama 3.1 8Bmeta-llama/Llama-3.1-8B-Instruct-D EuroLLM 9B v1utter-project/EuroLLM-9B-Instruct-D Ministral 8Bmistralai/Ministral-8B-Instruct-2410-D Mistral NeMo 12B mistralai/Mistral-Nemo-Instruct-2407-D Aya Expanse 8BCohereLabs/aya-expanse-8b-D Command R 7BCohereLabs/c4ai-command-r7b-12-2024-D Teuken 7B v0.4openGPT-X/Teuken-7B-instruct-research-v0.4-D Table 2: Considered LLMs for the evaluation Acc (%)DemosQA Greek Truthful QA BELEBELE (Greek) Greek Medical MCQA Greek ASEP MCQA INCLUDE (Greek) GPT-4o mini57.1761.6989.4469.2176.2566.49 Gemma 2 9B54.8359.0088.8946.0665.7551.99 Llama Krikri 8B57.1737.8277.3344.4458.9249.82 Meltemi 7B v1.542.6736.3564.7836.8157.5043.84 Llama 3.1 8B47.1737.2164.8925.2343.5829.53 EuroLLM 9B v141.6735.0152.8939.8153.1738.41 Ministral 8B42.3333.4160.1024.7739.6732.43 Mistral NeMo 12B34.6741.7469.3325.6942.4237.32 Aya Expanse 8B52.3342.9682.3334.4957.5845.83 Command R 7B46.8341.3774.2230.0957.5042.93 Teuken 7B v0.423.0016.4033.8922.2222.5826.45 Table 3: Experimental results for the instruction prompt. Acc (%) denotes the macro model accuracy. The best and second-best results are highlighted in bold and underline respectively. Acc (%)DemosQA Greek Truthful QA BELEBELE (Greek) Greek Medical MCQA Greek ASEP MCQA INCLUDE (Greek) GPT-4o mini55.1754.9689.1165.7475.0064.67 Gemma 2 9B56.17 59.6189.2245.1465.4252.90 Llama Krikri 8B56.3353.9881.8946.9965.9253.08 Meltemi 7B v1.550.1737.4561.1131.7158.0841.12 Llama 3.1 8B52.5038.5669.2226.1651.5836.78 EuroLLM 9B v141.8335.8662.7838.4356.2544.38 Ministral 8B46.1730.7264.6729.4045.2535.51 Mistral NeMo 12B44.8342.8469.7832.6449.6739.31 Aya Expanse 8B53.8342.1181.7837.7357.9248.55 Command R 7B52.3344.1969.3327.7854.2538.41 Teuken 7B v0.424.3325.0942.1126.3935.9228.44 Table 4: Experimental results for the role prompt. Acc (%) denotes the macro model accuracy. The best and second-best results are highlighted in bold and underline respectively. mance on BELEBELE is achieved by GPT-4o mini. Nonetheless, for the DemosQA and Greek Truthful QA datasets, the best accuracy scores are attained by Llamma Krikri 8B and Gemma 2 9B, respec- tively. These models achieve comparable accuracy across most datasets, while the remaining models undeperform, with Teuken 7B v0.4 having again the worst performance. Table 6 summarizes the average accuracy scores across the three prompting strategies for each model and dataset combination. As shown, GPT-4o mini ranks first in terms of accuracy across most datasets, severely outperforming the open- weights models on INCLUDE, Greek Medical and ASEP MCQA. The best accuracy score on De- mosQA and Greek Truthful QA were achieved by Llama Krikri 8B and Gemma 2 9B, respectively. These models attain similar accuracy scores across most datasets, while the rest undeperform, with Teuken 7B v0.4 attaining the worst accuracy. Acc (%)DemosQA Greek Truthful QA BELEBELE (Greek) Greek Medical MCQA Greek ASEP MCQA INCLUDE (Greek) GPT-4o mini54.3354.9688.5668.0675.1764.67 Gemma 2 9B53.6756.1882.3346.358.7550.36 Llama Krikri 8B56.0054.2282.3346.0666.2551.81 Meltemi 7B v1.548.1732.8057.6727.5549.8334.96 Llama 3.1 8B55.0040.8874.8925.4652.1739.31 EuroLLM 9B v140.1736.4767.6741.6756.8348.91 Ministral 8B44.5031.4663.1123.3844.7536.05 Mistral NeMo 12B46.6738.5671.0032.6451.0035.69 Aya Expanse 8B46.8331.2168.6736.8152.5846.92 Command R 7B49.8337.0959.2229.6346.7531.88 Teuken 7B v0.423.3326.1942.6725.6935.5829.71 Table 5: Experimental results for the CoT prompt. Acc (%) denotes the macro model accuracy. The best and second-best results are highlighted in bold and underline respectively. Acc (%)DemosQA Greek Truthful QA BELEBELE (Greek) Greek Medical MCQA Greek ASEP MCQA INCLUDE (Greek) GPT-4o mini55.5657.2089.0467.6775.4765.28 Gemma 2 9B54.89 58.2686.8145.8363.3151.75 Llama Krikri 8B56.5048.6780.5245.8363.7051.57 Meltemi 7B v1.547.0035.5361.1932.0255.1439.97 Llama 3.1 8B51.5638.8869.6725.6249.1135.21 EuroLLM 9B v141.2235.7861.1139.9755.4243.90 Ministral 8B44.3331.8662.6325.8543.2234.66 Mistral NeMo 12B42.0641.0570.0430.3247.7037.44 Aya Expanse 8B51.0038.7677.5936.3456.0347.10 Command R 7B49.6640.8867.5929.1752.8337.74 Teuken 7B v0.423.5522.5639.5624.7731.3628.20 Table 6: Experimental results across all prompts (mean accuracy). Acc (%) denotes the macro model accuracy. The best and second-best results are highlighted in bold and underline respectively. Overall, the top three models across the consid- ered QA datasets were GPT-4o mini, Greek Llama Krikri 8B, and the multilingual Gemma 2 9B. Despite their smaller parameter sizes, the two open-weight models achieved results comparable to the propri- etary GPT-4o mini, with the exception of the Greek ASEP, Medical MCQA, and INCLUDE datasets, which cover the civil service, medical, and edu- cational domains, respectively. Although GPT-4o mini demonstrates state-of-the-art performance, it attains lower average scores across all prompt- ing strategies on datasets that require common- sense reasoning. Llama Krikri 8B and Gemma 2 9B achieve the highest average scores on DemosQA and Greek Truthful QA, respectively. When comparing the evaluation results from Ta- bles 3–5, we notice a performance difference be- tween the proprietary GPT-4o mini and the open- weights models across the three prompting strate- gies considered. Our evaluation reveals that GPT- 4o performs better with simple user instructions, whereas open-weights models require prompts that specify a role to improve their performance. In addition, CoT has shown that it can lead to im- proved performance for models having increased reasoning capabilities; however, it reduces the ac- curacy of most open weights models due to pos- sible hallucinations introduced during reasoning. Finally, the performance gaps of open-weights mod- els in domain-specific datasets (i.e., Greek Medical MCQA, Greek ASEP MCQA and INCLUDE) indi- cate that the lack of specialized knowledge cannot be compensated for by optimizing prompt engineer- ing. It is important to note that comparisons between GPT-4o mini, accessed via the OpenAI API, and the open-weight models that are loaded locally, are not strictly equivalent. The proprietary model’s size and architecture are undisclosed (Chen et al., 2023), so we cannot guarantee that the same model version was served throughout our experiments. Addition- ally, API responses may incorporate contributions from system-level components beyond the base LLM (Neumann et al., 2025), whereas open-weight models provide fully transparent and stable check- points. 5. Discussion 5.1. Concluding Remarks This study aims to address the existing gap in Greek QA. To this end, (i) we introduce DemosQA, a novel QA dataset that enriches the limited set of human- annotated Greek QA resources, (i) we propose an adaptable and memory-efficient LLM evaluation framework that can run on commodity hardware, and (i) we leverage a diverse set of Greek QA datasets to comprehensively evaluate the reason- ing and linguistic capabilities of several LLMs. The experimental findings yield several key insights in response to our research questions: • Among the open-weight models, the Greek Llama Krikri 8B consistently outperforms most multilingual counterparts across multiple datasets (RQ1); •Recent open-weight LLMs have substantially narrowed the performance gap with the propri- etary GPT-4o mini on several datasets (RQ2); •Llama Krikri 8B and Gemma 2 9B appear as the most competitive open-weight models, achieving comparable performance across most datasets; •The instruction prompting strategy performs best for GPT-4o mini, while the role prompt yields performance benefits for the open- weights models; at the same time, open-weight models with enhanced reasoning capabilities can benefit from zero-shot CoT (RQ3); • It is feasible to construct high-quality, human-curated QA datasets from community- reviewed knowledge sources, as highly upvoted social media content provides reliable QA pairs. The consistency of model rankings across DemosQA and existing Greek QA benchmarks further advocates the quality of our dataset (RQ4). 5.2. Future Research Directions Overall, this work establishes a solid foundation for future research on Greek QA. By releasing De- mosQA and our evaluation framework, we aim to encourage the development of more linguistically in- clusive and culturally grounded LLMs. Future work may explore instruction-tuning Greek models with domain-specific data, extending DemosQA with ad- ditional social media sources, and adopting hybrid evaluation methods that combine human and auto- matic assessment for deeper insights into model behavior. To address the imbalance between high- and under-resourced languages, we advocate de- veloping Greek LLMs trained from scratch on cor- pora that include regional dialects (Chatzikyriakidis et al., 2024) and polytonic Greek texts (Kaddas et al., 2023), enhancing their social, historical, and cultural understanding. Moreover, new high-quality NLP datasets are needed, as multilingual ones often contain trans- lation errors and fail to capture language-specific nuances (Bandarkar et al., 2024; Roussis et al., 2025). Evaluating future LLMs on a broader set of Greek QA benchmarks (Peng et al., 2025; Chla- panis et al., 2025) using the proposed framework will further generalize and validate our findings. Fi- nally, future LLMs supporting the Greek Language could be evaluated on other NLP tasks, such as Greek Text Summarization, where typically small encoder-decoder models are utilized (Giarelis et al., 2024a,b). 5.3. Limitations This study has a few limitations that outline direc- tions for future improvement. First, our evalua- tion covers only a limited number of high-quality Greek QA datasets, reflecting the current scarcity of such resources. Second, we focus exclusively on Greek QA and do not extend our analysis to other under-resourced languages, where performance differences are expected due to language-specific training disparities. Third, we do not include larger multilingual LLMs in our evaluation, as no compa- rable large-scale open-weight Greek models are currently available. 5.4. Ethical Considerations All data used in this study were collected through the official Reddit API and are publicly avail- able. Data collection strictly adhered to Reddit’s API usage policies, including rate limits, which were respected by introducing short pauses be- tween consecutive requests. The collected con- tent was processed exclusively for academic and non-commercial research purposes. Furthermore, we manually reviewed and filtered the proposed dataset to remove any inappropriate, offensive, or non-informative material, ensuring ethical handling and high-quality data curation. References Aakanksha, Arash Ahmadian, Marwan Ahmed, Jay Alammar, Milad Alizadeh, Yazeed Alnumay, Sophia Althammer, Arkady Arkhangorodsky, Vi- raat Aryabumi, Dennis Aumiller, et al. 2025. Com- mand A: An Enterprise-Ready Large Language Model. ArXiv:2504.00698. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Ale- man, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2024. GPT-4 Technical Report. ArXiv:2303.08774. Mehdi Ali, Michael Fromm, Klaudia Thellmann, Jan Ebert, Alexander Arno Weber, Richard Rutmann, Charvi Jain, Max Lübbering, Daniel Steinigen, Johannes Leveling, et al. 2024. Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs. ArXiv:2410.03730. Juli Bakagianni, Kanella Pouli, Maria Gavriilidou, and John Pavlopoulos. 2025. A systematic sur- vey of natural language processing for the Greek language. Patterns, page 101313. Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Vari- ants. In Proceedings of the 62nd Annual Meet- ing of the Association for Computational Linguis- tics (Volume 1: Long Papers), pages 749–775, Bangkok, Thailand. Association for Computa- tional Linguistics. Damian Blasi, Antonios Anastasopoulos, and Gra- ham Neubig. 2022. Systematic Inequalities in Language Technology Performance across the World’s Languages. In Proceedings of the 60th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 5486–5505, Dublin, Ireland. Association for Computational Linguistics. Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. 2022. On the Opportunities and Risks of Foundation Models. ArXiv:2108.07258. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D. Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sas- try, Amanda Askell, et al. 2020. Language Mod- els are Few-Shot Learners. Advances in Neural Information Processing Systems, 33:1877–1901. Stergios Chatzikyriakidis, Chatrine Qwaider, Ilias Kolokousis, Christina Koula, Dimitris Papadakis, and Efthymia Sakellariou. 2024. GRDD: A Dataset for Greek Dialectal NLP. ArXiv:2308.00802. Lingjiao Chen, Matei Zaharia, and James Zou. 2023. How is chatgpt’s behavior changing over time? Odysseas S. Chlapanis, Dimitris Galanis, Nikolaos Aletras, and Ion Androutsopoulos. 2025. Greek- BarBench: A challenging benchmark for free- text legal reasoning and citations. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 25099–25119, Suzhou, China. Association for Computational Linguis- tics. John Dang, Shivalika Singh, Daniel D’souza, Arash Ahmadian, Alejandro Salamanca, Made- line Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, et al. 2024. Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier. ArXiv:2412.04261. Tim Dettmers and Luke Zettlemoyer. 2023. The case for 4-bit precision: k-bit Inference Scaling Laws. In Proceedings of the 40th International Conference on Machine Learning, pages 7750– 7774. PMLR. Nikolaos Giarelis, Charalampos Mastrokostas, and Nikos Karacapilidis. 2024a. Greek wikipedia: A study on abstractive summarization. In Proceed- ings of the 13th Hellenic Conference on Artificial Intelligence, SETN ’24, New York, NY, USA. As- sociation for Computing Machinery. Nikolaos Giarelis, Charalampos Mastrokostas, and Nikos Karacapilidis. 2024b. Greekt5: Sequence- to-sequence models for greek news summariza- tion. In Artificial Intelligence Applications and Innovations, pages 60–73, Cham. Springer Na- ture Switzerland. Nikolaos Giarelis, Charalampos Mastrokostas, Ilias Siachos, and Nikos Karacapilidis. 2024c. A re- view of greek nlp technologies for chatbot devel- opment. In Proceedings of the 27th Pan-Hellenic Conference on Progress in Computing and In- formatics, PCI ’23, page 15–20, New York, NY, USA. Association for Computing Machinery. Yvette Graham, Barry Haddow, and Philipp Koehn. 2019. Translationese in Machine Translation Evaluation. ArXiv:1906.09833. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. 2024. The Llama 3 Herd of Models. ArXiv:2407.21783. Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Os- trow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bres- sand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7B. ArXiv:2310.06825. Panagiotis Kaddas, Basilis Gatos, Konstantinos Palaiologos, Katerina Christopoulou, and Kon- stantinos Kritsis. 2023. Text Line Detection and Recognition of Greek Polytonic Documents. In Document Analysis and Recognition – IC- DAR 2023 Workshops, pages 213–225, Cham. Springer Nature Switzerland. Aisha Khatun and Daniel G. Brown. 2024. A Study on Large Language Models’ Limita- tions in Multiple-Choice Question Answering. ArXiv:2401.07955. Takeshi Kojima, Shixiang (Shane) Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large Language Models are Zero-Shot Reason- ers. Advances in Neural Information Processing Systems, 35:22199–22213. Penny Kyriazi and Prokopis Prokopidis. 2025. Mul- tiple choice qa greek asep. Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. TruthfulQA: Measuring how models mimic human falsehoods. In Proceedings of the 60th Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland. Association for Computational Linguistics. Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, et al. 2025. Eu- roLLM: Multilingual Language Models for Europe. Procedia Computer Science, 255:53–62. Charalampos Mastrokostas, Nikolaos Giarelis, and Nikos Karacapilidis. 2024. Social Media Topic Classification on Greek Reddit. Information, 15(9):521. Alexey N. Medvedev, Renaud Lambiotte, and Jean- Charles Delvenne. 2019. The Anatomy of Reddit: An Overview of Academic Research. In Dynam- ics On and Of Complex Networks I, pages 183– 204, Cham. Springer International Publishing. Shervin Minaee, Tomas Mikolov, Narjes Nikzad, Meysam Chenaghlu, Richard Socher, Xavier Am- atriain, and Jianfeng Gao. 2025. Large Language Models: A Survey. ArXiv:2402.06196. Humza Naveed, Asad Ullah Khan, Shi Qiu, Muham- mad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. 2025. A Comprehensive Overview of Large Lan- guage Models. ACM Transactions on Intelligent Systems and Technology, page 3744746. Anna Neumann, Elisabeth Kirsten, Muhammad Bi- lal Zafar, and Jatinder Singh. 2025. Position is power: System prompts as a mechanism of bias in large language models (llms). In Proceedings of the 2025 ACM Conference on Fairness, Ac- countability, and Transparency, FAccT ’25, page 573–598, New York, NY, USA. Association for Computing Machinery. Katerina Papantoniou and Yannis Tzitzikas. 2024. NLP for The Greek Language: A Longer Survey. ArXiv:2408.10962. John Pavlopoulos, Juli Bakagianni, Kanella Pouli, and Maria Gavriilidou. 2025.Open or Closed LLM for Lesser-Resourced Languages? Lessons from Greek. ArXiv:2501.12826. Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Pretten- hofer, Ron Weiss, Vincent Dubourg, Jake Van- derplas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Édouard Duchesnay. 2011. Scikit-learn: Machine learn- ing in python. Journal of Machine Learning Re- search, 12(85):2825–2830. Xueqing Peng, Triantafillos Papadopoulos, Efs- tathia Soufleri, Polydoros Giannouris, Ruoyu Xi- ang, Yan Wang, Lingfei Qian, Jimin Huang, Qian- qian Xie, and Sophia Ananiadou. 2025. Plutus: Benchmarking large language models in low- resource Greek finance. In Proceedings of the 2025 Conference on Empirical Methods in Natu- ral Language Processing, pages 30176–30202, Suzhou, China. Association for Computational Linguistics. Nicholas Proferes, Naiyan Jones, Sarah Gilbert, Casey Fiesler, and Michael Zimmer. 2021. Study- ing Reddit: A Systematic Overview of Disciplines, Approaches, Methods, and Ethics. Social Media + Society, 7(2):20563051211019004. Libo Qin, Qiguang Chen, Yuhang Zhou, Zhi Chen, Yinghui Li, Lizi Liao, Min Li, Wanxiang Che, and Philip S. Yu. 2025. A survey of multilingual large language models. Patterns, 6(1):101118. Matthew Renze and Erhan Guven. 2024. The Effect of Sampling Temperature on Problem Solving in Large Language Models. In Findings of the As- sociation for Computational Linguistics: EMNLP 2024, pages 7346–7356, Miami, Florida, USA. Association for Computational Linguistics. Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, Johan Ferret, et al. 2024. Gemma 2: Improving Open Language Models at a Practical Size. ArXiv:2408.00118. Angelika Romanou, Negar Foroutan, Anna Sot- nikova, Zeming Chen, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Maheshwary, Micol Al- tomare, Mohamed A. Haggag, Snegha A, et al. 2024. INCLUDE: Evaluating Multilingual Lan- guage Understanding with Regional Knowledge. ArXiv:2411.19799. Dimitris Roussis, Leon Voukoutis, Georgios Paraskevopoulos, Sokratis Sofianopoulos, Prokopis Prokopidis, Vassilis Papavasileiou, Athanasios Katsamanis, Stelios Piperidis, and Vassilis Katsouros. 2025. Krikri: Advancing Open Large Language Models for Greek. ArXiv:2505.13772. Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Niko- lay Bashlykov, Soumya Batra, Prajjwal Bhar- gava, Shruti Bhosale, et al. 2023. Llama 2: Open Foundation and Fine-Tuned Chat Models. ArXiv:2307.09288. Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos, Sokratis Sofianopoulos, Prokopis Prokopidis, Vassilis Papavasileiou, Athanasios Katsamanis, Stelios Piperidis, and Vassilis Katsouros. 2024. Meltemi: The first open Large Language Model for Greek. ArXiv:2407.20743. Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, et al. 2020. Transformers: State-of-the-art natural language processing. In Proceedings of the 2020 Confer- ence on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38– 45, Online. Association for Computational Lin- guistics. Jennifer C Wright. 2024. Stakeholder Manage- ment in Change Initiatives: Reddit Changes Its API Pricing. SAGE Publications: SAGE Busi- ness Cases Originals, 1 Oliver’s Yard, 55 City Road, London EC1Y 1SP United Kingdom. Appendix A. Model Prompts & Dataset Curation Examples Prompt Type GreekEnglish (Translated) Instruction Διάλεξε την καλύτερη απάντηση στην παρακάτω ερώτηση και απάντησε μόνο με το γράμα (Α, Β, Γ ή Δ). Select the best answer to the following question and answer only with the letter (A, B, C or D). Role Είσαι ένα γλωσικό μοντέλο για την εληνική γλώσα. Επίλεξε μόνο την καλύτερη απάντηση από τις διαθέσιμες απαντήσεις στην παρακάτω ερώτηση. Γράψε το κείμενο της επιλεγμένης απάντησης. You are a language model for the Greek language. Select only the best answer from the available answers to the following question. Write the text of the selected answer. Chain-of- Thought (CoT) Είσαι ένα γλωσικό μοντέλο για την εληνική γλώσα. Επίλεξε μόνο την καλύτερη απάντηση από τις διαθέσιμες απαντήσεις στην παρακάτω ερώτηση. Γράψε το κείμενο της επιλεγμένης απάντησης. Παρακαλώ σκέψου βήμα προς βήμα. You are a language model for the Greek language. Select only the best answer from the available answers to the following question. Write the text of the selected answer. Please think step-by-step. Table 7: Multiple-choice QA model prompts used in our study alongside their English translations. QuestionDisregarded AnswerReasonSelected Answer Μπορω να πάω στους ολυμπιακούς; Αντικειμενικά ένας κοινός άνθρωπος αν διάλεγε ένα άθλημα που δεν απαιτεί κάποια τρομερή φυσική κατάσταση πχ σκοποβολή, θα μπορούσε να συμετέχει στους ολυμπιακούς; Can I go to the Olympics? Objectively, if a common person chose a sport that doesn’t require terrible physical condition e.g. shooting, could they participate? Πάνε Αυστραλία και πες ότι είσαι break dancer. Θα πας αμέσως. Go to Australia and say you are a break dancer. You’l go immediately. Sarcasm Δεν είναι ακατόρθωτο, αλά θέλει σκληρή προπόνηση και να περάσεις σε προκριματικούς αγώνες. It is not impossible, but it requires hard training and passing qualifying matches. Παιδιά καμία συμβουλή τι γυμναστική να κάνω στο σπίτι για να πέσει η κοιλιά; Αν είναι να χάσω τον χρόνο μου μέσα στο σπίτι ας κάνω καλό στην υγεία μου τουλάχιστον Guys any advice on what home workout to do to lose belly fat? If I’m going to waste my time at home at least let me do good for my health Για κοιλιά· Δίαιτα. For belly? Diet. Unhelpful Comment Μείωσε τις θερμίδες που τρως καθημερινά (...) Τώρα, για να απαντήσω σε αυτό που ρωτάς, η γυμναστική που σε βοηθά να χάσεις λίπος ειναι αυτό που λένε (...) Reduce your daily calories (...) Now, to answer what you ask, the workout that helps you lose fat is what they call (...) Είναι νόμιμο να σου κρατήσουν λεφτά από το μισθό σου για παράπτωμα εν ώρα εργασίας; Καλησπέρα σε όλους, για να μην τα πολυλογώ χθες στην δουλεία έκανα μια μεγάλη γκαφα. (...) Is it legal to deduct money from your salary for misconduct during work hours? Good evening everyone, to make a long story short yesterday at work I made a big blunder. (...) [deleted] Deleted Comment ́Οχι. Αν θέλουν ας σου κάνουν αγωγή ή να σε απολύσουν. Σφάλματα είναι μέρος της δουλειάς και το ρίσκο που αναλαμβάνει η επιχείρηση. (...) No. If they want, let them sue you or fire you. Mistakes are part of the job and the risk the business undertakes. (...) Είναι νόμιμο να σου κρατήσουν λεφτά από το μισθό σου για παράπτωμα εν ώρα εργασίας; Καλησπέρα σε όλους, για να μην τα πολυλογώ χθες στην δουλεία έκανα μια μεγάλη γκαφα. (...) Is it legal to deduct money from your salary for misconduct during work hours? Good evening everyone, to make a long story short yesterday at work I made a big blunder. (...) ́Ελα, πες τι π*παριά έκανες, μας έχεις ιντριγκάρει. Εξάλου μόνο αυτοί που δεν κάνουν τίποτα δεν κάνουν λάθη Come on, say what bullsh*t you did, you intrigued us. Besides, only those who do nothing make no mistakes Offensive Language ́Οχι. Αν θέλουν ας σου κάνουν αγωγή ή να σε απολύσουν. Σφάλματα είναι μέρος της δουλειάς και το ρίσκο που αναλαμβάνει η επιχείρηση. (...) No. If they want, let them sue you or fire you. Mistakes are part of the job and the risk the business undertakes. (...) Table 8: Examples of the manual curation process (English translations are provided below the original Greek text).