Paper deep dive
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
Soyeon Kim, Cheongwoong Kang, Myeongjin Lee, Eun-Chul Chang, Jaedeok Lee, Jaesik Choi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 6/21/2026, 8:51:17 AM
Summary
K-MetBench is a new multi-dimensional diagnostic benchmark designed for evaluating large language models (LLMs) and multimodal large language models (MLLMs) in the field of meteorology, specifically tailored for the Korean context. It addresses four critical gaps identified in current meteorological AI evaluation: the modality gap (visual reasoning of charts), the reasoning gap (validity of rationales), the geo-cultural gap (Korean-specific topography and regulations), and the granularity gap (performance across specific sub-domains). The benchmark consists of 1,774 questions derived from National Meteorological Engineer certification exams, featuring expert-verified rationales and multimodal elements like weather maps and Skew-T Log-P diagrams. Experimental results show that while top-tier global models excel in general accuracy, specialized Korean models and advanced reasoning models show significant performance in local contexts, highlighting that parameter scaling alone does not resolve cultural and domain-specific dependencies.
Entities (8)
Relation Signals (5)
K-MetBench โ evaluates โ Gemini-3-Pro-Preview
confidence 100% ยท We evaluated a diverse array of models... such as GPT-5.2 and Gemini-3-Pro-Preview
K-MetBench โ evaluates โ EXAONE-4.0
confidence 100% ยท including EXAONE-4.0, A.X-4.0, VARCO-Vision-2.0, and HyperCLOVA X
K-MetBench โ isgroundedin โ National Meteorological Engineer certification exam
confidence 100% ยท K-MetBench is constructed from raw data drawn from the National Meteorological Engineer certification examinations
K-MetBench โ addressesgapsin โ Meteorology
confidence 90% ยท K-MetBench serves as a roadmap for developing reliable, culturally aware expert AI agents in meteorology.
Korea Meteorological Administration โ issuesregulationsfor โ Meteorology
confidence 90% ยท regulations issued by the Korea Meteorological Administration (KMA).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The development of practical (multimodal) large language model assistants for Korean weather forecasters is hindered by the absence of a multidimensional, expert-level evaluation framework grounded in authoritative sources. To address this, we introduce K-MetBench, a diagnostic benchmark grounded in national qualification exams. It exposes critical gaps across four dimensions: expert visual reasoning of charts, logical validity via expert-verified rationales, Korean-specific geo-cultural comprehension, and fine-grained domain analysis. Our evaluation of 55 models reveals a profound modality gap in interpreting specialized diagrams and a reasoning gap where models hallucinate logic despite correct predictions. Crucially, Korean models outperform significantly larger global models in local contexts, demonstrating that parameter scaling alone cannot resolve cultural dependencies. K-MetBench serves as a roadmap for developing reliable, culturally aware expert AI agents. The dataset is available at this https URL .
Tags
Links
- Source: https://arxiv.org/abs/2604.24645v1
- Canonical: https://arxiv.org/abs/2604.24645v1
Trouble viewing inline? Open PDF directly โ
Full Text
123,129 characters extracted from source content.
Expand or collapse full text
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology Soyeon Kim KAIST, INEEJI Seongnam, Korea soyeon.k@kaist.ac.kr Cheongwoong Kang and Myeongjin Lee KAIST Seongnam, Korea cw.kang; lmjk311@kaist.ac.kr Eun-Chul Chang and Jaedeok Lee Kongju National University Gongju, Korea echang@kongju.ac.kr, ruio1084@gmail.com Jaesik Choi โ KAIST, INEEJI Seongnam, Korea jaesik.choi@kaist.ac.kr Abstract The development of practical(multimodal) large language model assistants for Korean weather forecasters is hindered by the ab- sence of a multidimensional, expert-level eval- uation framework grounded in authoritative sources. To address this, we introduce K- MetBench, a diagnostic benchmark grounded in national qualification exams. It exposes crit- ical gaps across four dimensions: expert vi- sual reasoning of charts, logical validity via expert-verified rationales, Korean-specific geo- cultural comprehension, and fine-grained do- main analysis. Our evaluation of 55 mod- els reveals a profoundmodality gapin in- terpreting specialized diagrams and areason- ing gapwhere models hallucinate logic de- spite correct predictions. Crucially, Korean models outperform significantly larger global models in local contexts, demonstrating that parameter scaling alone cannot resolve cul- tural dependencies. K-MetBench serves as a roadmap for developing reliable, cultur- ally aware expert AI agents. The dataset is available at https://huggingface.co/ datasets/soyeonbot/K-MetBench. 1 Introduction Large language models(LLMs)and multimodal large language models(MLLMs)have shown growing promise in scientific domains(Taylor et al. ,2022;Team et al.,2023;OpenAI,2025), achieving performance matching passing thresh- olds on professional certification exams(Singhal et al.,2023;Katz et al.,2024). As these models are increasingly positioned as assistants for domain- specific tasks, there is a growing need for evalua- tion frameworks that go beyond surface-level cor- rectness and more precisely characterize domain- โ Corresponding author. ID: 206|Part: 5| Korean-Specific: False Question: When a jet streak is located within an upper- level trough as shown in the figure, in which region does strong convergence occur? Choices: 1. Region d, 2. Region b, 3. Region c, 4. Region a Answer: 2 Expert-Verified Rationale: Due to the curvature effect within a jet streak located in an upper-level trough, ... Therefore, the region of strong convergence is Region b. (1) Visual Reasoning beyond OCR to Chart interpretation (2) Reasoning Evaluation Expert-aligned rationale scoring (3) Local Specificity Korea-specific topography & regulations (4) Granular Analysis Weakness diagnosis in 5 sub-fields 100 KTS 125 KTS 150 KTS a b c d Figure 1:An example of the K-MetBench dataset (translated into English).K-MetBench provides evaluation across four critical dimensions:(1)multi- modal understanding,(2)expert-level reasoning,(3) geo-cultural context sensitivity, and(4)fine-grained do- main knowledge across five meteorological sub-fields. relevant competencies(Liang et al.,2022). How- ever, existing benchmarks for vertical domains of- ten summarize performance using a single aggre- gate score, making it difficult to understand why a model succeeds or fails in practice. In complex ap- plied fields such as meteorology, this coarse eval- uation obscurescritical limitationsfor real-world deployment. We identify four recurring limitations in current evaluations of meteorological reasoning. First,the modality gap. Meteorological anal- arXiv:2604.24645v1 [cs.CL] 27 Apr 2026 ysis, inherently multimodal, requires the synthesis of numerical data, textual descriptions, and special- ized visual charts(e.g., weather maps, skew-T log- P diagrams).However, most scientific benchmarks remain predominantly text-based and provide lim- ited assessment of a modelโs ability to interpret domain-specific charts and spatial patterns. As a result, visual understanding capabilities central to operational forecasting remain under-evaluated. Second,the reasoning gap. Conventional benchmarks primarily rely on answer accuracy, without explicitly evaluating the validity or struc- ture of the underlying reasoning. In high-stakes do- mains like weather forecasting, a correct prediction reached through shallow heuristics or incomplete logic may still lead to brittle or unreliable behav- ior( Turpin et al.,2023). Without access to expert- aligned rationales, it is difficult to distinguish gen- uine understanding from shortcut learning. Third,the geo-cultural gap. Many existing datasets emphasize global or universal physical principles while abstracting away local geographic and institutional context. In meteorology, how- ever, local topography, climatological conventions, and region-specific regulations play a substantial role in interpretation and decision-making. Mod- els trained and evaluated solely on decontextual- ized data may therefore fail to generalize reliably to region-specific applications. Fourth,the granularity gap. Aggregate per- formance scores often mask uneven competence across sub-domains. A model may perform well on factual recall or chart interpretation while strug- gling with quantitative reasoning or applied dy- namics. Without fine-grained analysis, such dis- parities remain difficult to diagnose. To address these limitations, we introduce K-MetBench, a Korean meteorological bench- mark designed for multi-dimensional evaluation of LLMs and MLLMs. Rather than treating meteo- rological expertise as a monolithic capability, K- MetBench decomposes evaluation along four com- plementary axes:(1)multimodal understanding of meteorological charts and symbols,(2)reason- ing qualityassessed using expert-verified ratio- nales,(3)sensitivity to geo-cultural and regional context, and(4)fine-grained coverageacross five officially defined meteorological sub-domains. Through this structured design, K-MetBench is in- tended as a diagnostic tool that helps reveal which aspects of meteorological reasoning remain chal- lenging for current models, and why. 2 Related Work Existing benchmarks for meteorological and cli- mate reasoning reflect diverse assumptions about knowledge sources, modalities, and evaluation ob- jectives. Rather than treating them as competitors, we situate them along complementary axes that highlight different aspects of domain expertise. ClimaQA( Manivannan et al.,2024)evaluates climate question answering using textbooks as the knowledge source. By grounding questions in established instructional materials, it emphasizes conceptual understanding and theoretical reason- ing characteristic of graduate-level climate science. While this approach provides scientific rigor, it re- mains purely text-based and does not assess visual interpretation or operational reasoning grounded in real-world artifacts. ClimateIQA(Chen et al., 2025)constructs instruction-style QA data from numerical weather prediction(NWP)heatmaps and associated geospatial metadata. This enables evaluation of visual pattern recognition and struc- tured data interpretation. WeatherQA(Ma et al., 2024)further targets operational forecasting sce- narios by combining multiple meteorological im- ages with expert-written mesoscale discussions. These datasets advance multimodal evaluation, but their emphasis remains on task-level performance rather than fine-grained diagnosis across sub-fields or distinct reasoning failures. In the Korean-language evaluation landscape, KMMLU( Son et al.,2025)is derived from official national examinations, measuring expert-level lin- guistic competence across a wide range of profes- sions. Since it is based on official Korean exams, KMMLU captures linguistic and cultural aspects of the Korean language. KMMLU-Redux( Hong et al.,2025)is a reconstructed version of KMMLU that removes erroneous, ambiguous, or contami- nated items to improve reliability. While these benchmarks offer high reliability and clear passing criteria, they are primarily text-based and treat me- teorological knowledge as a small subset within a broader evaluation suite, limiting their ability to an- alyze domain-specific competencies in depth. 3 K-MetBench Construction K-MetBench is designed to complement the exist- ing benchmarks by explicitly separating and jointly examining four dimensions that are often conflated in prior benchmarks. Rather than introducing new task formats, K-MetBench focuses on providing di- Table 1:Comparison with existing benchmarks.K-MetBench distinguishes itself by covering four key axes: visual understanding,rationale reliability,geo-cultural alignment, andfine-grained diagnosis in sub-domains. DatasetLang. DomainTest Size(Source)ModalityReasoningGeo-CulturalGranularity KMMLU(Son et al.,2025)KorGeneral35k(License Exam)TextรKoreaร(45 Subjects) KMMLU-Redux(Hong et al.,2025)KorGeneral2.6k(License Exam)TextรKoreaร(14 Subjects) ClimaQA(Manivannan et al.,2024)EngClimate566(Autogenerate)TextรGlobalร(3 Tasks) ClimateIQA(Chen et al.,2025)EngClimate152k(Template)Image+TextรGlobalร(4 Tasks) WeatherQA(Ma et al.,2024)Eng Weather Forecast 600(Template)Image+TextรUnited Statesร(2 Tasks) K-MetBench(Ours)KorMeteorology1.7k(License Exam)Image+TextExpert-VerifiedKorea5 Sub-domains Note for K-MetBench.Modality:Includes multimodal questions evaluating interpretation of professional weather charts.Rea- soning:Provides rationale verified by domain experts.Geo-Cultural:Includes questions requiring knowledge of local geography and regulations that are specific to Korea(e.g., the Korea Meteorological Administration(KMA)protocols).Granularity:Sup- ports fine-grained diagnosis across the five sub-domains officially defined in Korea Engineer Meteorology certification exam. Table 2:Detailed statistics of K-MetBench.The dataset is structured into four key dimensions to en- able structured evaluation: Modality, Reasoning, Geo- cultural, and Granularity. Diagnostic AxisStatisticValue 1. OverviewTotal Questions1,774 2. ModalityImage+Text Questions82(4.62%) (Visual Understanding) (Charts, Diagrams) 3. ReasoningAvg. Rationale Length93.72 tokens (Expert Rationales)Text-Only Reasoning121(6.82%) Multimodal Reasoning20(1.13%) 4. Geo-CulturalKorean-Specific73(4.11%) ( Local Knowledge ) Questions 5. GranularityPart 1: Forecast Theory 373(21.03%) (5 Subject Areas)Part 2: Observation332(18.71%) Part 3: Atmos. Dynamics 359(20.24%) Part 4: Climatology376(21.20%) Part 5: Atmos. Physics 334(18.83%) Note:The number of tokens is calculated using the gemini-2.5-flashtokenizer.(Atmos.: Atmospheric) agnostic visibility into how and where current mod- els succeed or fail when approaching expert-level meteorological reasoning. 3.1 Data Collection and Processing K-MetBench is constructed from raw data drawn from the National Meteorological Engineer certi- fication examinations, covering 25 exam sessions between March 16, 2003 and March 5, 2022. The initial pool comprised 2,500 multiple-choice ques- tions. Because these examinations are generated from a shared question bank, substantial overlap exists across years. To construct a balanced and non-redundant benchmark, we applied a multi- stage filtering and augmentation pipeline. For deduplication( Lee et al.,2021), we first ap- plieddifflib.SequenceMatcherwith a similar- ity threshold of 0.6, removing exact duplicates as well as items with trivially permuted answer op- tions. Importantly, questions with inverted logic (e.g.,โhighestโvs.โlowestโ,โsaturatedโvs.โunsat- uratedโ)were manually reviewed and retained, as they probe distinct reasoning behaviors despite sur- face similarity. This process yielded a refined set of 1,774 questions. To reduce memorization and contamination ef- fects, we applied two transformations. First, we randomized answer option orders for all questions. Second, we paraphrased question stems using Gemini-2.5-Pro, with strict constraints to pre- serve technical terminology and domain-specific meaning. The system prompt used for paraphras- ing is provided in Appendix C.1. To maintain quality, a human researcher reviewed and refined 14.88%(264/1,774)of the paraphrased items. For multimodal questions, both text and visual elements were extracted from the original examina- tion PDFs. Three researchers reviewed and cross- checked all extracted images to correct parsing arti- facts such as missing axis labels, distorted symbols, or incomplete annotations. As a design choice to separate perceptual challenges from reasoning dif- ficulty, mathematical formulas embedded as im- ages were transcribed into LaTeX code to prevent OCR bottlenecks, while meteorological charts and diagrams were preserved in their original format. 3.2 Subset 1: Multimodal Diagnosis The multimodal subset of K-MetBench consists of 82 questions(4.62% of the dataset)that require interpretation of meteorological visuals. Unlike general-purpose multimodal benchmarks that fo- cus on object recognition or scene description, this subset targets domain-specific charts and symbolic representations. The included materials span sur- Table 3:Distribution of K-MetBench across five sub-domains.The number of questions for each subject area is reported, with the number of reasoning questions featuring expert-verified rationales in parentheses(Reas. stands for Reasoning). PartSubject Area Overall VolumeModalityGeo-Cultural Total(Reas.)Text(Reas.)Image + Text(Reas.)Korean(Reas.) 1 Weather Analysis & Forecast Theory373(28)364(24)9(4)6(0) 2 Meteorological Observation Methods332(28)318(24)14(4)0(0) 3 Atmospheric Dynamics359(29)340(25)19(4)0(0) 4 Climatology376(28)363(24)13(4)50(7) 5 Atmospheric Physics334(28)307(24)27(4)17(0) Sum Total Coverage1,774(141)1,692(121)82(20)73(7) face weather maps, upper-level charts(e.g., 200 and 500 hPa),and thermodynamic diagrams such as Skew-T Log-P plots and emagrams derived from radiosonde measurements. Solving these questions requires extracting structured informa- tion, including pressure gradients, wind vectors, and thermodynamic indicesโfrom dense visual fields that cannot be resolved through OCR alone. Consequently, this subset assesses the ability of MLLMs to integrate textual meteorological knowl- edge with the interpretation of domain-specific vi- sual cues. Representative examples are provided in Appendix Table 6. 3.3 Subset 2: Reasoning-Aware Evaluation To evaluate reasoning quality beyond final answer correctness, K-MetBench includes a reasoning- aware subset consisting of 141 questions paired with expert-verified rationales. These rationales serve as reference explanations for assessing the validity, coherence, and depth of model-generated reasoning. Rationale construction followed a two- stage process. First,GPT-5was used to gener- ate initial reasoning drafts, guided by prompts that emphasized logical flow, factual consistency, clar- ity, and completeness. Second, two meteorology professors reviewed these drafts, correcting factual errors, refining physical explanations, and resolv- ing ambiguities. We employ an LLM-as-a-Judge framework( Zheng et al.,2023)to score model- generated rationales against the expert-verified rationales as reference standard. The system prompts used for reasoning generation and evalu- ation are detailed in AppendixC.5andC.6. To validate the reliability of this framework in a spe- cialized domain, we conduct a meta-evaluation( Li et al.,2024)comparing LLM judgments with hu- man expert scores. The experimental and survey protocols are provided in AppendixD.2andC.9. 3.4 Subset 3: Geo-Cultural Sensitivity Meteorological reasoning is strongly influenced by local geography, climate patterns, and institu- tional conventions. To capture this dependency, we annotate aKorean-Specificsubset compris- ing 73 questions that involve implicit, speaker- centric, or high-context expressions specific to the Korean Peninsula. Candidate items were identi- fied using prompt-enhanced LLMs(GPT-4.1and Gemini-2.5-Pro)designed to detect references to localized phenomena, such as regional topog- raphy(e.g., the Yeongdong region)or regulations issued by the Korea Meteorological Administra- tion(KMA).These candidates were subsequently reviewed and validated by two researchers to en- sure relevance and correctness. Rather than test- ing translation ability, this subset probes whether models can appropriately ground meteorological knowledge in region-specific context. As such, it provides a controlled setting for analyzing geo- cultural alignment in domain-specific reasoning. 3.5 Subset 4: Domain Specificity To enable fine-grained analysis of meteorological expertise, K-MetBench is organized into five of- ficial subject areas defined in the Korean Meteo- rological Engineer certification exam. These in- clude: Part 1(Weather Analysis and Forecast The- ory),Part 2(Meteorological Observation Meth- ods),Part 3(Atmospheric Dynamics),Part 4 (Climatology),and Part 5(Atmospheric Physics). Each subject area targets a distinct aspect of pro- fessional competence, ranging from chart inter- pretation and numerical weather prediction prin- ciples to instrumentation, large-scale atmospheric motion, climate systems, and thermodynamic cal- culations. This structure allows model perfor- mance to be examined at a level of granularity that is not visible from aggregate scores alone. By aligning evaluation with established subject bound- aries, this design facilitates diagnosis of domain- specific strengths and weaknesses, for example, distinguishing models that perform well on de- scriptive climatology but struggle with quantitative dynamics or thermodynamics. 4 Experiments 4.1 Experimental Setup Evaluated Models.To ensure a comprehen- sive benchmark, we evaluated a diverse ar- ray of models categorized by scale, training language, and modality support. The selec- tion includes proprietary state-of-the-art mod- els renowned for superior reasoning capabil- ities, such asGPT-5.2(evaluated with and without reasoning modules enabled) (OpenAI, 2025) and Gemini-3-Pro-Preview ( Team et al., 2023). We also incorporated open-source mod- els ranging from 0.6B to 235B parameters, exemplified byInternVL3.5( Wang et al., 2025)andQwen3-VL(Yang et al.,2025), along- side large-scale foundation models such asgpt- oss-120b(Agarwal et al.,2025),command- a-reasoning-08-2025( Cohere et al.,2025), andLlama-3.2-90B-Vision-Instruct(Meta, 2024). To investigate the impact of geo-cultural knowledge, we specifically included Korean- centric models, includingEXAONE-4.0(Research et al.,2025),A.X-4.0(Lab,2025),VARCO- Vision-2.0( Cha et al.,2025), andHyperCLOVA X(Yoo et al.,2024). Finally, strictly text-based baselines were established by evaluating non- multimodal models solely on the textual compo- nents of questions to quantify text dependency. Geo-Cultural Disambiguation Protocol.To es- tablish a fair evaluation protocol for global models, we designed four experimental configurations that cross-reference question formulation with prompt- ing conditions. This setup ensures that models are assessed on their meteorological competence rather than their ability to decode localized linguis- tic ambiguities. For question formulation, we com- pared anImplicitcondition, using original speaker- centric terms likeโOur country,โagainst anEx- plicitcondition, which replaces these with proper nouns(e.g.,โSouth Koreaโ)to isolate and evalu- ate pure domain knowledge. Regarding prompt- ing conditions, beyond aStandardprompt that injects an expert persona, we introduced anAd- vancedprompt providing explicit disambiguation (e.g.,โโOur countryโrefers to South Koreaโ).This advanced protocol serves as a specialized support layer, mitigating performance degradation caused by implicit geo-cultural references and enabling global models to compete on an equal footing. Comparison with Existing Benchmarks.We evaluated models using the official test sets of all datasets, employing the Chain-of-Thought (CoT) (Wei et al.,2022)protocol for Weath- erQA. Task orthogonality was analyzed using Kendallโs Tau-b rank correlation coefficient. To align the distance-based Haversine metric of Cli- maIQA(where lower is better)with standard ac- curacy metrics, we inverted the sign of ClimaIQA scores prior to calculating correlations. Meta-Evaluation Setup: Validating LLM-as-a- Judge.Given the specialized nature of meteo- rology, validating the reliability of commercial LLMs as judges is crucial. We conducted a meta- evaluation comparing human expert judgments with LLM judgments. We selected ten represen- tative questions varying in difficulty and type, and collected reasoning outputs from ten open-source LLMs. Two human experts provided gold standard scores, whileGemini-2.5-Proserved as the AI evaluator. Both parties utilized identical expert- verified references and a scoring rubric across four axes: Factuality, Logicality, Depth, and Clarity. We calculated Kendallโs Tau-b(ฯ b )correlation be- tween human and AI scores, confirming the align- ment of the automated judge(ฯ b >0.8).We also computed Krippendorffโsฮฑ(interval)and Intr- aclass Correlation Coefficient(ICC, 2-way mixed, absolute)to assess inter-rater reliability, which in- dicated acceptable agreement(ฮฑ>0.7).To in- vestigate whether incorporating human expert ra- tionales improves the alignment between the LLM evaluator and human judgment, we compared the correlations of their scores under conditions with and without rationale availability. Implementation Details.To ensure a fair com- parison, we utilizedStandardprompts across all models. We applied a zero-shot setting to all text, multimodal, and reasoning questions to eval- uate intrinsic capabilities. We computed accu- racy by extracting final answers via regular expres- sions. To rigorously assess instruction-following Table 4:K-MetBench performance scores across diverse models.Models are sorted by accuracy. Accuracy score ranges from 0 to 100, while the reasoning score(Reas.)ranges from 4 to 20. The highest scores in each column are shown inboldfor proprietary and open-source models, respectively.(Acc.: Accuracy,K: Korean model,V: Vision language model,R: Reasoning model.) Type Model Type Acc.Reas.Geo-Cult.ModalityGranularity(P1โP5) K V RKorean Text Multi P1 P2 P3 P4 P5 gemini-3-pro-preview(Thinking)V R93.7 18.0190.4 94.6 75.6 92.5 97.9 94.2 92.8 91.6 gpt-5.2(Thinking)VR87.817.3380.890.629.386.393.488.086.285.3 Proprietary gpt-5.2V77.6 17.3975.3 79.0 50.0 77.2 81.3 71.9 81.4 76.3 Multilingual Models Qwen3-VL-235B-A22B-ThinkingV R84.4 17.2272.686.248.881.5 88.6 87.2 83.2 82.0 Qwen3-VL-32B-ThinkingVR78.616.1960.379.951.274.385.278.878.776.3 command-a-reasoning-08-2025R77.8 14.1274.6 77.8- 73.4 85.2 73.8 78.8 78.5 gpt-oss-120bR77.316.1262.077.3-72.585.876.577.474.9 Qwen3-30B-A3B-Thinking-2507R76.7 15.7667.6 76.7- 75.5 82.1 75.6 74.9 75.9 InternVL3.5-38B-InstructV57.311.3847.958.140.256.064.848.761.455.7 Llama-3.2-90B-Vision-InstructV56.9 9.7252.1 58.2 30.5 57.1 59.3 52.4 62.2 53.3 Phi-451.511.7540.851.5-52.553.850.055.145.3 Korean Models A.X-4.0K76.1 15.4678.976.1- 76.6 77.7 68.2 81.3 76.5 EXAONE-4.0-32BKR59.913.5759.259.9-58.264.852.463.161.2 VARCO-Vision-2.0-14BK V58.7 11.2457.5 59.5 42.7 59.0 62.3 54.3 61.7 56.0 A.X-4.0-LightK55.711.4560.655.7-55.854.450.961.455.7 A.X-4.0-VL-LightK V52.5 9.7654.8 53.0 42.7 51.5 50.6 50.1 58.0 52.1 Open-source HyperCLOVAX-SEED-Think-14BKR50.811.2952.150.8-51.653.841.855.651.1 capabilities, we counted any output that violated the required format as a failure case. We em- ployed the vLLM library(Kwon et al.,2023)with its default configurations, except forA.X-4.0-VL- LightandLlama-3.2-90B-Vision-Instruct, which were run using Hugging Face Transform- ers. The random seed was fixed at 42, and sam- pling temperatures were set to 0.1 by default, while a temperature of 1.0 was employed for reasoning models. All prompts and questions were provided in the original Korean to strictly evaluate localized comprehension without translation artifacts. 5 Results Beyond simple leaderboards, we dissect the perfor- mance of models across four dimensions to reveal their true capabilities and limitations. The Modality Gap: Text-Only vs. Multimodal. Figure2reveals a distinctdentedshape along the Multimodalaxis, confirming that visual reason- ing is the primary bottleneck for current MLLMs. Specifically, models exhibited a sharp accuracy decline(avg.โ18.55%)on multimodal ques- tions compared to text-only ones. This deficit is most pronounced in professional tasks involving Skew-T Log-P diagrams and surface weather maps, where models failed to extract key data despite their general vision capabilities. The Reasoning Gap: Knowledge vs. Reasoning. Table4and Figure2highlight a distinct gap be- tween answer accuracy and reasoning quality. Al- though Kendallโsฯ b (0.78)indicates a general cor- relation(Appendix Figure 7),qualitative analysis reveals that models frequently provide correct an- swers accompanied by insufficient rationales, in- cluding the use of improper or hallucinated termi- nology(Appendix Table 7).Additionally, while models achieve high accuracy on simple retrieval tasks, performance significantly degrades on calcu- lation and multi-step reasoning tasks, even when CoT promptingโexplicitly guiding the model to use a<scratchpad>( Nye et al.,2021)โis em- ployed. The Geo-Cultural Gap.Table4reveals that large multilingual models struggle with the Korean-Specificsubset(e.g., Changma, topog- raphy)despite their scale. The Korean-centric A.X-4.0(72B)scored 78.9, outperforming the largerQwen3-VL-235B-Thinking(72.6). This confirms that parameter scaling does not automatically grant proficiency in local domains. Granular Domain Analysis. Finally, decom- posing performance across the five official subject areas reveals fine-grained disparities masked by ag- gregated scores. As shown in Table 4, models gen- Figure 2:Holistic performance analysis of top-6 models across five dimensions.The radar chart visual- izes model capabilities in Accuracy,Reasoning,Geo- Cultural alignment(K-Specific),Modality(Text-only vs. Multimodal), andGranularity(Subject Parts 1โ5). While models show balanced performance across theo- retical subjects, a sharp decline is observed in theMul- timodalaxis, highlighting the modality gap. erally exhibit robust performance in Part 2(Mete- orological Observation),which focuses on instru- mentation and factual knowledge(e.g.,Gemini-3- Proreaching 97.9).However, significant perfor- mance drops are observed in calculation-intensive and abstract domains like Part 3(Atmospheric Dy- namics)and Part 5(Atmospheric Physics).A strik- ing example is the Korean modelA.X-4.0, which achieves its highest accuracy in Part 4(Climatol- ogy)(81.3)โlikely benefiting from training on lo- cal meteorological lawsโbut struggles dispropor- tionately in Part 3(68.2),where understanding syn- optic motions is required. This granular diagno- sis identifies specific domain weaknesses: while models may possess sufficient regulatory knowl- edge(Part 4),they require targeted fine-tuning to enhance quantitative reasoning in thermodynamics and dynamics(Part 3, 5). Orthogonality between Existing Baselines.As shown in Figure 3, we analyzed Kendallโsฯ b correlations to assess the independence of K- MetBench. While theText-Onlysubset corre- lates strongly with general Korean benchmarks (KMMLU-Redux,ฯ b = 0.78),we observe a dis- tinct decoupling in complex capabilities. Notably, the correlation weakens for theReasoningsubset (ฯ b = 0.66)and drops sharply for theMultimodal KMMLU Figure 3:Correlation analysis with existing bench- marks.The heatmap visualizes Kendallโsฯ b correla- tion coefficients between K-MetBench metrics and ex- isting benchmarks. subset(ฯ b = 0.29).Furthermore, correlations with external weather baselines(e.g., ClimaQA, Cli- maIQA, and WeatherQA)remain consistently low across both reasoning and multimodal dimensions (avg.ฯ b <0.14).This quantitative gap demon- strates that K-MetBench evaluates specialized do- main logic and visual interpretation skills that are orthogonal to general linguistic proficiency and ex- isting meteorological tasks. Meta-Evaluation: Human-LLM Agreement. We validated our reasoning evaluation framework by measuring inter-rater agreement on 100 sam- pled responses(Table 5).All axes surpassed the reliability threshold(ฮฑ>0.7),withReasoning To- talachieving a robustฮฑof 0.838. Additionally, Figure4illustrates a strong correlation between human and LLM scores. Thew/ rationaleset- ting yielded a Kendallโsฯ b of 0.99 with low vari- ance, slightly outperforming thew/o rationaleset- ting(ฯ b = 0.96). Table 5:Inter-rater agreement analysis.The agree- ment between the average scores of two human experts and the LLM evaluator. We report Krippendorffโsฮฑ (interval)and Intraclass Correlation Coefficient(ICC, two-way mixed, absolute agreement). Evaluation Axis KrippendorffโsฮฑICCN Factuality0.8270.829 100 Logicality0.8270.830 100 Depth0.7420.747 100 Clarity0.8250.827 100 Reasoning Total0.8380.841 100 6810121416 Human Expert Score 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 LLM-Judge Score A.X-4.0 InternVL3_5-14B-Instruct InternVL3_5-8B-Instruct Llama-3.1-8B-Instruct Qwen3-0.6B Qwen3-VL-235B-A22B-Thinking VARCO-VISION-2.0-14B command-a-reasoning-08-2025 gpt-oss-120b gpt-oss-20b Model size Performance tier 0.6B 20B 235B Top Mid Low w/ rationale w/o rationale Figure 4:Scatter plot comparing human expert vs. LLM-judge scores.Thew/ rationalecondition(ฯ b = 0.99)shows slightly higher precision and lower vari- ance than thew/o rationalecondition(ฯ b = 0.96), while both maintain a strong correlation. 6 Discussion 6.1 The Challenge of Visual Reasoning in Specialized Domains The observed modality gap in Table4and Fig- ure2underscores a fundamental limitation: cur- rent MLLMs lack thedomain-specific visual lit- eracynecessary for forecasting. Although profi- cient in general recognition, models struggle to ground specialized visual patternsโsuch as iso- bars, fronts, and wind barbsโin physical princi- ples. This indicates that training on general image- text pairs is insufficient for mastering the fine- grained visual reasoning required in specialized scientific domains. 6.2 Geo-Cultural Alignment in Meteorology Meteorology requires applying universal laws to lo- calized contexts. The observed performance gap indicates a critical lack ofgeo-cultural alignment in global models. Despite linguistic fluency, mul- tilingual models frequently hallucinate on specific Korean geographic and terminological nuances. Consequently, effective deployment in vertical do- mains demands more than mere scaling; it requires rigorous alignment with local topographic and le- gal contexts to bridge the gap between general ca- pability and expert-level application. 6.3 Superficial Reasoning vs. Causal Deduction The observation that models output correct an- swers with shallow or erroneous explanations points toshortcut learning(Geirhos et al.,2020)โ a reliance on surface-level associations rather than genuine understanding. Furthermore, the inabil- ity to reach expert-level performance on formula- based problems(e.g., calculating geostrophic wind speed)highlights a critical deficiency in applying physical laws. Addressing this requires shifting from general instruction tuning to training on high- quality reasoning trace data grounded in rigorous physical principles. 6.4 Reliability of Automated Evaluation in Specialized Domains Our results confirm thatGemini-2.5-Prois a reliable proxy for human experts in meteorology. The high agreement inFactualityandLogicality in Table5demonstrates objective evaluation of logic and evidence. WhileDepthshowed slightly more subjectivity, the overall consistency sup- ports the frameworkโs robustness. Furthermore, the tight correlation observed in the scatter plot in Figure 4indicates that expert rationales effec- tively minimize variance. However, the modelโs high intrinsic knowledge ensures reliable grading even in their absence. These findings demonstrate that, when guided by high-quality rubrics, modern LLMs are cost-effective and reliable judges even in fine-grained domains like meteorology. This vali- dates adopting the LLM-as-a-Judge framework for the reasoning evaluation in this study. 7 Conclusion We present K-MetBench, a multi-dimensional benchmark for fine-grained evaluation of large lan- guage models in meteorological reasoning. By decomposing performance across modality, rea- soning quality, geo-cultural context, and domain- specific sub-fields, K-MetBench provides diagnos- tic insights that are not observable from aggre- gate accuracy alone. Our evaluation reveals per- sistent challenges in interpreting domain-specific visual artifacts, producing coherent expert-level ra- tionales, and grounding meteorological knowledge in local context. In addition, analysis across of- ficial subject areas exposes uneven performance that is obscured by holistic scores. Overall, K- MetBench is intended as a diagnostic complement to existing benchmarks, helping identify where cur- rent models succeed and where targeted improve- ments are needed for reliable deployment in spe- cialized scientific domains. Limitations While K-MetBench serves as a rigorous diagnos- tic tool for meteorological AI, we acknowledge several limitations. First, regarding modality, the benchmark focuses on static visual reasoning(e.g., snapshot weather charts).While interpreting these charts is fundamental to forecasting, the current dataset does not evaluate the temporal reasoning re- quired to interpret atmospheric evolution, such as sequential radar imagery or satellite loops. Second, the dataset is geo-specifically rooted in the Korean context. Although this design effectively evaluates geo-cultural alignmentโa key contribution of our workโit inherently limits direct generalizability to other climatic regions without adaptation. Fi- nally, we utilized the official examination passing criteria(60%)as a proxy for human competency. While this provides a validated baseline for quali- fication, a fine-grained human expert ceiling ( e.g., the upper-bound score of top-tier meteorologists) was not explicitly measured in this study. Future work will focus on establishing this upper bound to quantify theโsuper-humanโgap precisely. Ethical Considerations We adhered to copyright laws and ethical guide- lines in constructing K-MetBench. The dataset is derived from National Meteorological Engineer ex- aminations administered from March 16, 2003 to March 5, 2022; among 43 sessions in this period, we used only the 25 that were officially released to the public. We also obtained explicit permis- sion from the Human Resources Development Ser- vice of Korea(HRDK)to use these materials for re- search and to release the refined dataset in an open repository. In addition, the dataset was reviewed to ensure that it contains no personally identifiable information or harmful content. For human annotation, we involved two domain experts from collaborating institutions in the same funded project: one university professor and one research professor. The same experts conducted both reference-rationale verification and scoring of model-generated reasoning, and these activities were compensated separately on a per-item basis in accordance with our institutionโs internal stan- dards for expert advisory and review work. We consider this compensation appropriate given the expertsโseniority, domain expertise, and expected time commitment. Licensing and Legal Compliance The K-MetBench dataset is derived from public ex- amination materials managed by the HRDK. We conducted a rigorous legal review to ensure com- pliance with theO๏ฌicial Information Disclosure Actand relevant copyright laws(Copyright Act Art. 24-2, 25)in Korea . We confirmed that the questions are not classified as restricted informa- tion. To support the research community, the cu- rated dataset is released via an open repository un- der the C BY-NC-ND license, permitting non- commercial research use while preserving the in- tegrity of the original artifacts. Acknowledgments We express our gratitude to the Human Resources Development Service of Korea(HRDK)for allow- ing the use of National Technical Qualification Ex- amination data for research purposes. We would like to thank Seongsu Bae and the anonymous re- viewers for their valuable comments. This research was supported by the High- Performance Computing Support Project, funded by the Ministry of Science and ICT(MSIT)and the National IT Industry Promotion Agency(NIPA) under grant No. RQT-25-070278(providing 40 H100 GPUs).This work was also supported by the Institute for Information & Communications Tech- nology Planning & Evaluation(IITP)grant funded by the Korea government(MSIT)(No. RS-2019- I190075, Artificial Intelligence Graduate School Program(KAIST);and No. RS-2022-I220984, Development of Artificial Intelligence Technology for Personalized Plug-and-Play Explanation and Verification of Explanation),and by the Korea Me- teorological Administration(KMA)and National Institute of Meteorological Sciences(NIMS)under grant No. KMA2021-00123(Developing Intelli- gent Assistant Technology and Its Application for Weather Forecasting Process). Data and Code Availability The dataset is hosted on HuggingFace at https://huggingface.co/datasets/ soyeonbot/K-MetBench.The evaluation toolkit is available athttps://github.com/ kmetbench/kmetbench-release .The K- MetBench leaderboard is publicly available at https://kmetbench.github.io/. References Sandhini Agarwal, Lama Ahmad, Jason Ai, Sam Alt- man, Andy Applebaum, Edwin Arbus, Rahul K Arora, Yu Bai, Bowen Baker, Haiming Bao, and 1 others. 2025. gpt-oss-120b & gpt-oss-20b model card.arXiv preprint arXiv:2508.10925. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wen- bin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, and 1 others. 2025. Qwen2. 5-vl technical report.arXiv preprint arXiv:2502.13923. Young-rok Cha, Jeongho Ju, SunYoung Park, Jong- Hyeon Lee, Younghyun Yu, and Youngjune Kim. 2025. Varco-vision-2.0 technical report.arXiv preprint arXiv:2509.10105. Jian Chen, Peilin Zhou, Yining Hua, Dading Chong, Meng Cao, Yaowei Li, Wei Chen, Bing Zhu, Junwei Liang, and Zixuan Yuan. 2025. Climateiqa: A new dataset and benchmark to advance vision-language models in meteorology anomalies analysis. InPro- ceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 5322โ5333. Team Cohere, Arash Ahmadian, Marwan Ahmed, Jay Alammar, Milad Alizadeh, Yazeed Alnumay, Sophia Althammer, Arkady Arkhangorodsky, Vi- raat Aryabumi, Dennis Aumiller, and 1 others. 2025. Command a: An enterprise-ready large language model.arXiv preprint arXiv:2504.00698. Robert Geirhos, Jรถrn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks.Nature Machine Intelligence, 2(11):665โ673. Seokhee Hong, Sunkyoung Kim, Guijin Son, Soyeon Kim, Yeonjung Hong, and Jinsik Lee. 2025.From KMMLU-redux to pro: A professional Korean benchmark suite for LLM evaluation. InFindings of the Association for Computational Linguistics: EMNLP 2025, pages 19067โ19096, Suzhou, China. Association for Computational Linguistics. Jenny Y. Huang, Yunyi Shen, Dennis Wei, and Tamara Broderick. 2026.Dropping just a handful of pref- erences can change top large language model rank- ings . InThe Fourteenth International Conference on Learning Representations. Human Resources Development Service of Ko- rea. 2024.Examination standards for me- teorological engineer(2023.1.1โ2026.12.31). https://w.q-net.or.kr/pageLink.do? link=cst/cstReport.Accessed: 2026-01-06. Available at Q-Net. Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo. 2024. Gpt-4 passes the bar exam.Philosophical Transactions of the Royal Society A, 382(2270):20230254. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serv- ing with pagedattention. InProceedings of the 29th symposium on operating systems principles, pages 611โ626. SKT AI Model Lab. 2025.A.X 4.0. Katherine Lee, Daphne Ippolito, A. Nystrom, Chiyuan Zhang, D. Eck, Chris Callison-Burch, and Nicholas Carlini. 2021.Deduplicating training data makes language models better. InAnnual Meeting of the Association for Computational Linguistics. Haitao Li, Qian Dong, Junjie Chen, Huixue Su, Yu- jia Zhou, Qingyao Ai, Ziyi Ye, and Yiqun Liu. 2024. Llms-as-judges: a comprehensive survey on llm-based evaluation methods.arXiv preprint arXiv:2412.05579. Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Ku- mar, and 1 others. 2022. Holistic evaluation of lan- guage models.arXiv preprint arXiv:2211.09110. Chengqian Ma, Zhanxiang Hua, Alexandra Anderson- Frey, Vikram Iyer, Xin Liu, and Lianhui Qin. 2024. Weatherqa: Can multimodal language mod- els reason about severe weather?arXiv preprint arXiv:2406.11217. Veeramakali Vignesh Manivannan, Yasaman Jafari, Srikar Eranky, Spencer Ho, Rose Yu, Duncan Watson-Parris, Yian Ma, Leon Bergen, and Taylor Berg-Kirkpatrick. 2024. Climaqa: An automated evaluation framework for climate question answer- ing models.arXiv preprint arXiv:2410.16701. Meta. 2024.Llama 3.2 model card. Accessed: 2024- 01-04. Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari, Henryk Michalewski, Jacob Austin, David Bieber, David Dohan, Aitor Lewkowycz, Maarten Bosma, David Luan, Charles Sutton, and Augustus Odena. 2021.Show your work: Scratchpads for interme- diate computation with language models .arXiv preprint arXiv:2112.00114. OpenAI. 2025.Update to gpt-5 system card: Gpt-5.2. Accessed: 2026-01-04. A Yang Qwen, Baosong Yang, B Zhang, B Hui, B Zheng, B Yu, Chengpeng Li, D Liu, F Huang, H Wei, and 1 others. 2024. Qwen2. 5 technical re- port.arXiv preprint. LG Research, Kyunghoon Bae, Eunbi Choi, Kibong Choi, Stanley Jungkyu Choi, Yemuk Choi, Kyubeen Han, Seokhee Hong, Junwon Hwang, Taewan Hwang, and 1 others. 2025. Exaone 4.0: Unified large language models integrating non-reasoning and reasoning modes.arXiv preprint arXiv:2507.11407. Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mah- davi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, and 1 others. 2023. Large language models encode clinical knowledge.Nature, 620(7972):172โ180. Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheon- bok Park, Kang Min Yoo, and Stella Biderman. 2025. KMMLU: Measuring massive multitask language understanding in Korean. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Hu- man Language Technologies(Volume 1: Long Pa- pers), pages 4076โ4104, Albuquerque, New Mexico. Association for Computational Linguistics. Ross Taylor, Marcin Kardas, Guillem Cucurull, Thomas Scialom, Anthony Hartshorn, Elvis Saravia, Andrew Poulton, Viktor Kerkez, and Robert Stojnic. 2022. Galactica: A large language model for science.arXiv preprint arXiv:2211.09085. Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean- Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Mil- lican, and 1 others. 2023. Gemini: a family of highly capable multimodal models.arXiv preprint arXiv:2312.11805. Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2023. Language models donโt always say what they think: Unfaithful explanations in chain-of- thought prompting.Advances in Neural Information Processing Systems, 36:74952โ74965. Weiyun Wang, Zhangwei Gao, Lixin Gu, Hengjun Pu, Long Cui, Xingguang Wei, Zhaoyang Liu, Linglin Jing, Shenglong Ye, Jie Shao, and 1 others. 2025. In- ternvl3. 5: Advancing open-source multimodal mod- els in versatility, reasoning, and efficiency.arXiv preprint arXiv:2508.18265. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elic- its reasoning in large language models.Advances in neural information processing systems, 35:24824โ 24837. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388. Kang Min Yoo, Jaegeun Han, Sookyo In, Heewon Jeon, Jisu Jeong, Jaewook Kang, Hyunwook Kim, Kyung- Min Kim, Munhyong Kim, Sungju Kim, and 1 others. 2024. Hyperclova x technical report.arXiv preprint arXiv:2404.01954. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, and 1 others. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information pro- cessing systems, 36:46595โ46623. Appendix Table of Contents A Dataset Examples15 B Case Study of Reasoning Answer15 C Prompts and Questionnaires for Benchmark Construction15 C.1 Question Paraphrasing. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .15 C.2 Identification of Korean-Specific Subset. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .15 C.3 Implicit vs. Explicit Dataset Design for Korean-Specific Subset. . . . . . . . . . . . . . . . . .15 C.4 Evaluation Prompts for Korean-Specific Subset. . . . . . . . . . . . . . . . . . . . . . . . . . .16 C.5 Prompt for Reference Rationale Generation. . . . . . . . . . . . . . . . . . . . . . . . . . . . .16 C.6 Questionnaire for Expert Verification on Reference Rationale. . . . . . . . . . . . . . . . . . . .16 C.7 Reasoning Prompt for Open-Source LLMs. . . . . . . . . . . . . . . . . . . . . . . . . . . . .17 C.8 Prompt for LLM-as-a-Judge Evaluation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .17 C.9 Questionnaire for Expert Scoring of LLM Reasoning. . . . . . . . . . . . . . . . . . . . . . . .17 D Experimental Setups17 D.1 Reasoning Model Inference. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .17 D.2 Meta-Evaluation for LLM-as-a-Judge. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .18 E Additional Results and Discussion18 E.1 Detailed Orthogonality Analysis of K-MetBench. . . . . . . . . . . . . . . . . . . . . . . . . .18 E.2 Detailed Analysis of K-MetBench Performance. . . . . . . . . . . . . . . . . . . . . . . . . . .19 E.3 Results of Meta Evaluation of LLM-as-a-Judge. . . . . . . . . . . . . . . . . . . . . . . . . . .20 E.4 MCQA Accuracy vs. Reasoning Score. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .20 E.5 Computational Cost and Efficiency Analysis. . . . . . . . . . . . . . . . . . . . . . . . . . . .21 F Compute Resources21 G Hierarchical Topic Distribution21 H Robustness of Conclusions Under Small Subsets38 H.1 Statistical Robustness Diagnostics. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .38 H.2 Robustness of Key Findings to Critical Data Perturbation. . . . . . . . . . . . . . . . . . . . . .39 Table 6:Representative examples of K-MetBench tasks.The examples are organized by modality:Text-only (Top)andMultimodal(Bottom).We showcase three task types within each modality:Standard(fundamental knowledge),K-Specific(geo-cultural context),andReasoning(complex deduction).Partdenotes the correspond- ing subject from the five official fields.Gray textindicates English translations. ModalityStandard MCQAK-Specific MCQAReasoning MCQA Text-OnlyID:1535,Part:5 ์ง๋ฌธ:์์ธต์ผ๊ธฐ๋์ํ์ฉ์๋ํด์ฌ๋ฐ๋ฅด๊ฒ ์ค๋ช ํ๊ฒ์? Question: Which of the following is a correct description regarding the utilization of upper-level weather charts? 1.500 hPa์ผ๊ธฐ๋์ํ๋ญ๊ธฐ์๊ณจ์์๋ฑ์จ์ ์ ์งํญ์ด๋ฑ๊ณ ์ ์์งํญ๋ณด๋คํด๊ฒฝ์ฐ์๋๊ทธ ๊ธฐ์๊ณจ์ํ๋ฐฉ์์ฝํ์์น๊ธฐ๋ฅ๊ฐ์๊ณ ,์ ๋ฐฉ์ ์ฝํํ๊ฐ๊ธฐ๋ฅ๊ฐ์๋ค. 2.300 hPa๋ฉด์์๋์จ๋๊ฐ์งํ,๋ณต์ฌ์์ํฅ์ ๋ฐ์ผ๋ฏ๋ก์ ์ ๋ถ์์ด์ฉ์ดํ๋ค. 3.300 hPa์ ํธ๊ธฐ๋ฅ์ถ๊ตฌ์์ข์ธก์ํ๊ฐ๊ธฐ๋ฅ, ์ฐ์ธก์์์น๊ธฐ๋ฅ๊ฐ์์ผ๋ฉฐ,์ ๊ตฌ์์๋์ข์ธก์ ์์น๊ธฐ๋ฅ,์ฐ์ธก์ํ๊ฐ๊ธฐ๋ฅ๊ฐ์๋ค. 4.500 hPa๊ธฐ๋ฅ๊ฐ์ง์ํ๋ญ์ ์ ์์์ง์ผ๋ก ๋ถ๋ฉด์ด์ ์ ์ํ์ฑ์ผ๋ก์์ ์ฒ์ด๋ํ๋๋ค. 1. In a cold trough on a 500 hPa chart, if the amplitude of the isotherms is larger than the amplitude of the contours(geopotential height), there is a weak updraft behind the trough and a weak downdraft ahead of it. 2. On the 300 hPa surface, temperature is affected by topography and radiation, making frontal analysis easy. 3. At the exit of a 300 hPa jet stream, there is a downdraft on the left and an updraft on the right; at the entrance, there is an updraft on the left and a downdraft on the right. 4. If the 500 hPa airflow blows perpendicular to a surface cold front, the front becomes active and severe weather occurs. ์ ๋ต:1 Ground Truth: 1 ID:65,Part: 5 ์ง๋ฌธ:๋ค์์ํ๊ตญ์ง์ญ์์ํฅ์์ฃผ๋๊ณ ๊ธฐ์์ ํน์ฑ์์ค๋ช ํ๊ฒ์ด๋ค. ๋ด์ฉ์ด์ณ์ง์์๊ฒ์? Question: The following describes the characteristics of high-pressure systems affecting theKorean region. Which statement is incorrect? 1.์๋ฒ ๋ฆฌ์๊ณ ๊ธฐ์์๊ฒจ์ธ์ฒ ์์ถฅ๊ณ ๊ฑด์กฐํ ๋ ์จ๋ฅผ๋ง๋ ๋ค. 2.์คํธ์ธ ํฌํด๊ณ ๊ธฐ์์๋ํด์์ง๋ฐฉ์ ๊ณ ์จํ์์์ผ์ผํจ๋ค. 3.๋ถํํ์๊ณ ๊ธฐ์์๊ณ ์จ๋ค์ตํ๋ฉฐ,์ฌ๋ฆ์ฒ ์ ๋ฌด๋์ด๋ ์จ๋ฅผ๋ง๋ ๋ค. 4.์ด๋์ฑ๊ณ ๊ธฐ์์์ํฅ์๋ฐ์ผ๋ฉด๋ด์๋ ๋ฐ๋ปํ๋ ์จ,๊ฐ์์๋๋ง์๋ ์จ๊ฐ๋๋ค. 1. The Siberian High creates cold and dry weather during the winter. 2. The Okhotsk Sea High causes high-temperature phenomena in the east coastal regions. 3. The North Pacific High is hot and humid, creating sweltering weather during the summer. 4. Under the influence of migratory highs, the weather becomes warm in spring and clear in autumn. ์ ๋ต:2 Ground Truth: 2 ID:18,Part: 2 ์ง๋ฌธ:๋น์ด์์ฐจ์์์ฌ๋ฐ๋ฅด๊ฒ๋ํ๋ธ๊ฒ์ ๋ฌด์์ ๋๊น? Question: What is the correct dimensional representation of specific heat? 1.[$L^2T^2ฮธ^-1$] 2.[$L^2T^-2ฮธ^-1$] 3.[$ML^-1T^-2$] 4.[$ML^2T^-2$] ์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ:๋น์ด์๋จ์์ง๋๋น๋จ์ ์จ๋์์น์ํ์ํ์๋์ง๋ก์์ฐจ์์ (์๋์ง)/(์ง๋ยท์จ๋)= $(ML^2T^-2)/(Mฮธ)= L^2T^-2ฮธ^-1$์ด๋ฏ๋ก2๋ฒ์ด๋ง๊ณ ,4๋ฒ์ ์๋์ง์์ฒด์์ฐจ์,3๋ฒ์์๋ ฅ์์ฐจ์,1๋ฒ์ ์๊ฐ์ง์๊ฐ๋ถํธ๊ฐ๋ฐ๋๋ผํ๋ฆฝ๋๋ค. Expert-Verified Rationale: Specific heat is the energy required to raise the temperature of a unit mass by one unit. Its dimension is(Energy)/(Massยท Temperature)= $(ML^2T^-2)/(Mฮธ)= L^2T^-2ฮธ^-1$. Therefore, option 2 is correct. Option 4 represents the dimension of energy itself, option 3 represents the dimension of pressure, and option 1 is incorrect because the sign of the time exponent is reversed. ์ ๋ต:2 Ground Truth: 2 MultimodalID:460,Part: 3 ์ง๋ฌธ:๋ถ๋ฐ๊ตฌ์์๋ํ๋๋์ง๊ท ํ ($ V_g$),์ค์ ํ($ V$),์ํ๊ฐ์๋ ($( d vdt)_H$)์ฌ์ด์๊ด๊ณ๋ฅผ ์ฌ๋ฐ๋ฅด๊ฒํํํ๊ทธ๋ฆผ์์ด๋๊ฒ์ธ๊ฐ? Question: Which figure correctly represents the relationship between the geostrophic wind ($ V_g$),the actual wind($ V$),and the horizontal acceleration($( d v dt)_H$)in the Northern Hemisphere? 1. 2. 3. 4. ์ ๋ต:3 Ground Truth: 3 ID:687,Part: 4 ์ง๋ฌธ:์ ์๋๊ทธ๋ฆผ์ํ๊ตญ์์ด๋ค์ง์ ์์ฐํ๊ท ๋ฌผ์์ง๋ฅผ๋ณด์ฌ์ค๋ค. ์ด๊ทธ๋ฆผ์์D๋ถ๋ถ์ด ์๋ฏธํ๋๊ฒ์๋ฌด์์ธ๊ฐ? Question: The presented figure shows the annual average water balance of a certain location in Korea. What does section D in this figure represent? 1.ํ ์์๋ถ์๊ณผ์ 2.ํ ์์๋ถ์๋ณด์ถฉ 3.ํ ์์๋ถ์์ด์ฉ 4.ํ ์์๋ถ์๊ฒฐํ 1. Soil moisture surplus 2. Soil moisture recharge 3. Soil moisture utilization 4. Soil moisture deficit ์ ๋ต:3 Ground Truth: 3 ID:460,Part: 3 ์ง๋ฌธ:๋ค์๊ทธ๋ฆผ์ด๋ณด์ฌ์ฃผ๋์ญ์ ์ธต์์ข ๋ฅ๋ก ์ณ์๊ฒ์? Question: Which of the following is the correct type of inversion layer shown in the figure below? 1.๋ณต์ฌ์ญ์ 2.๋๋ฅ์ญ์ 3.์ ์ ์ญ์ 4.์นจ๊ฐ์ญ์ 1. Radiation inversion 2. Turbulence inversion 3. Frontal inversion 4. Subsidence inversion ์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ:๊ทธ๋ฆผ์ฒ๋ผ์งํ์์๋ฐ๋ก ์์ํ๋์์์ญ์ ์ธต์ด์๋ก๊ฐ์๋ก์ฝํ๋๋ ํํ๋์ผ๊ฐ์งํ๋ณต์ฌ๋๊ฐ์ผ๋ก์๊ธฐ๋ ๋ณต์ฌ์ญ์ ์์ ํ์ด๋ฉฐ,์นจ๊ฐ์ญ์ ์๊ณ ๊ธฐ์ํ ํ๊ฐ๋ฅ๋ก์์ธต์๋ถ๋ฆฌ๋์ด๋ํ๋๊ณ ์ ์ ์ญ์ ์ ์ ์ ๋ฉด์๋ฐ๋ผ๊ฒฝ์ฌ์ ธ์์ผ๋ฉฐ๋๋ฅ์ญ์ ์์ฃผ๊ฐ ํผํฉ์ธต๊ผญ๋๊ธฐ์ํ์ฑ๋์ด์งํ์์์์ํ์ง ์์ผ๋ฏ๋ก๊ทธ๋ฆผ๊ณผ๋ค๋ฅด๋ค. Expert-Verified Rationale: As shown in the figure, a shallow inversion layer starting directly from the surface and weakening with height is typical of a radiation inversion caused by nocturnal surface radiative cooling. In contrast, a subsidence inversion appears separated in the upper atmosphere due to downdrafts under high pressure, a frontal inversion is tilted along a frontal surface, and a turbulence inversion forms at the top of the daytime mixed layer rather than starting at the surface; thus, they differ from the figure. ์ ๋ต:1 Ground Truth: 1 Table 7:Comparison ofInternVL3.5-8B-InstructandQwen3-VL-235B-A22B-Thinkingon K-MetBench reasoning question.( X :Incorrect, โ : Correct,shaded text: Highlighted errors) Reasoning MCQAReasoning Evaluation ofInternVL3.5-8B-InstructReasoning Evaluation ofQwen3-VL-235B-A22B- Thinking ID:105,Part:3 ์ง๋ฌธ:์ง๊ท ํ๊ด๊ณ์์์๊ธฐ์๊ฒฝ๋์ ๊ณต๊ธฐ๋ฐ๋๊ฐ ์ผ์ ํ๋ค๋ฉด, ํ์์ ์ด๋ป ๊ฒ๊ฒฐ์ ๋๋๊ฐ? Question: In the geostrophic wind for- mula, if the pressure gradient and air den- sity are constant, how is the wind speed determined? 1.1.๊ณ ๋๊ฐ์ฆ๊ฐํจ์๋ฐ๋ผ์ฆ๊ฐํ๋ค. 2.2.์๋์์ฆ๊ฐ์๋ฐ๋ผ๊ฐ์ํ๋ค. 3.3.๊ธฐ์์์ฆ๊ฐ์๋ฐ๋ผ์ฆ๊ฐํ๋ค. 4.4.๊ณต๊ธฐ์์ ์ฑ์ดํฌ๋ฉด์ฆ๊ฐํ๋ค. 1. It increases as altitude increases. 2. It decreases as latitude increases. 3. It increases as pressure increases. 4. It increases if air viscosity is high. ์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ:์ง๊ท ํ์Vg = | p|/(ฯ f)์ด๋ฏ๋ก๊ธฐ์๊ฒฝ๋์๋ฐ๋๊ฐ์ผ์ ํ๋ฉด์ฝ๋ฆฌ์ฌ๋ฆฌ๋งค๊ฐ๋ณ์f=2ฮฉ sinฯ๋ง์ด ๋ณ์๊ฐ๋์ด์๋๊ฐ์ฆ๊ฐํ ์๋กf๊ฐ์ปค ์ ธํ์์๊ฐ์ํ๋ฉฐ,์ ๋๊ธฐ์์ํฌ๊ธฐ (๋ณด๊ธฐ3)๋๊ณ ๋(๋ณด๊ธฐ1), ์ ์ฑ(๋ณด๊ธฐ4) ์์ด๊ด๊ณ์์์ง์ ๋ฑ์ฅํ์ง์๋๋ค. Expert-Verified Rationale: Since the geostrophic wind is defined as Vg = | p|/(ฯ f),if the pressure gradient and den- sity are constant, the Coriolis parameter f=2ฮฉ sinฯ becomes the only variable. As latitude increases, f increases, causing the wind speed to decrease. The magnitude of absolute pressure(Option 3),altitude (Option 1),and viscosity(Option 4)do not appear directly in this equation. ์ ๋ต:2 Ground Truth: 2 Response fromInternVL3.5-8B-Instruct ์์ฑ๋_์ถ๋ก _๊ทผ๊ฑฐ:์ง์ฐํ๊ด๊ณ์์์ํ์์ด๊ธฐ์๊ฒฝ๋ก์๋ณํ์ ์ง์ ์ ์ผ๋ก์ฐ๊ด๋๋ค๋๊ฒ์์ดํดํ ์์์ต๋๋ค.๊ธฐ์๊ฒฝ๋๊ฐ ์ฆ๊ฐํ๋ฉดํ์๋์ฆ๊ฐํ๋๊ฒฝํฅ์ด์์ผ๋ฉฐ,์ด๋ํ์์ด๊ธฐ์์์ฆ๊ฐ์ ๋ฐ๋ผ์ฆ๊ฐํ๋ค๋3๋ฒ์ต์ ์ด์ ๋ต์ด๋ผ๋๊ทผ๊ฑฐ๋ฅผ์ ๊ณตํฉ๋๋ค. Expert-Verified Rationale: From thegeostrophic wind relationship, we can understand that wind speed is directly associated with changes in the pressure path. Since wind speed tends to increase when the pressure gradient increases, this provides grounds for Option 3โthat wind speed increases as pressure increasesโbeing the correct answer. ์ ๋ต:3 X Predicted Answer: 3 Reasoning Evaluation from Human Experts: ์ฌ๋์ ๋ฌธ๊ฐํ๊ท ์ ์:์ฌ์ค์ ์ ํ์ฑ(1),๋ ผ๋ฆฌ์ ์๊ฒฐ์ฑ(1),์ถ๋ก ์๊น์ด (1),ํํ์๋ช ํ์ฑ(1) Average of Human Expert Score:factuality(1),logicality(1),depth(1), clarity(1) ํ๊ฐ์๊ฒฌ1:์ค๋ต๋ฐ์ฌ์ฉ๋์ฉ์ด๋ถ์ ์ :๊ธฐ์๊ฒฝ๋ก,์ง์ฐํ๋ฑ ํ๊ฐ์๊ฒฌ2:๋ฌธ์ ์๋ง์ง์๋์ถ๋ก ๊ทผ๊ฑฐ๋ฐ๋ต๋ํ๋ฆผ. Evaluation Comment 1:Incorrect answer and inappropriate use of terminology:โpressure pathโ,โdelayed windโ,etc. Evaluation Comment 2:The reasoning is irrelevant to the question, and the answer is incorrect. Reasoning Evaluation from LLM-Evaluator(Gemini-2.5-Pro): ์ฌ์ค_์ค๋ฅ_๋ชฉ๋ก:[ ์ง๊ท ํ(geostrophic wind)์โ์ง์ฐํโ์ด๋ผ๋์กด์ฌํ์ง์๋๊ธฐ์ํ์ฉ์ด๋ก ์๋ชป์ง์นญํ์์ต๋๋ค., ๋ฌธ์ ์์โ๊ธฐ์๊ฒฝ๋๊ฐ์ผ์ ํ๋คโ๋ํต์ฌ๊ฐ์ ์๋ฌด์ํ๊ณ ,โ๊ธฐ์๊ฒฝ๋๊ฐ ์ฆ๊ฐํ๋ฉดํ์์ด์ฆ๊ฐํ๋คโ๋๊ด๊ณ๋ฅผ์ถ๋ก ์๊ทผ๊ฑฐ๋ก์ฌ์ฉํ์์ต๋๋ค. โ๊ธฐ์๊ฒฝ๋(pressure gradient)โ์โ๊ธฐ์(pressure)โ์ด๋ผ๋๋ฌผ๋ฆฌ์ ์ผ๋ก๋ค๋ฅธ ๊ฐ๋ ์ํผ๋ํ์ฌ,๊ธฐ์๊ฒฝ๋์ํ์์๊ด๊ณ๋ฅผ๊ทผ๊ฑฐ๋ก๊ธฐ์๊ณผํ์์๊ด๊ณ๋ฅผ ์ค๋ช ํ๋์ค๋ฅ๋ฅผ๋ฒํ์ต๋๋ค. ] ํ๊ฐ_์ ์:์ ํ์ฑ:1,๋ ผ๋ฆฌ์ _์๊ฒฐ์ฑ:1,์ถ๋ก ์_๊น์ด:1,ํํ์_๋ช ํ์ฑ: 1 ํ๊ฐ_์ฌ์ :์ํ์AI์๋ต๋ณ์๋ชจ๋ ํ๊ฐ์ถ์์์ตํ์ ์๊ธฐ๋กํ์ต๋๋ค. โ์ ํ์ฑโ์ธก๋ฉด์์,โ์ง๊ท ํโ์โ์ง์ฐํโ์ด๋ผ๋ํ๊ฐ์ฉ์ด๋ก์ง์นญํ์ผ๋ฉฐ, โ๊ธฐ์๊ฒฝ๋โ์โ๊ธฐ์โ์ํผ๋ํ๋์น๋ช ์ ์ธ๊ฐ๋ ์ค๋ฅ๋ฅผ๋ณด์์ต๋๋ค. ๋ํ ๋ฌธ์ ์ํต์ฌ์ ์ ์ธโ๊ธฐ์๊ฒฝ๋์ผ์ โ์๋ฌด์ํ์ฌ์ฌ์ค๊ด๊ณ๋ฅผ ์๊ณกํ์ต๋๋ค. โ๋ ผ๋ฆฌ์ ์๊ฒฐ์ฑโ์ธก๋ฉด์์,๋ฌธ์ ์์กฐ๊ฑด๊ณผ๋ฌด๊ดํ์ผ๋ฐ๋ก ์ ๊ทผ๊ฑฐ๋ก์ ์ํ๊ณ ,โ๊ธฐ์๊ฒฝ๋โ์โ๊ธฐ์โ์ํผ๋ํ์ฌ๊ฒฐ๋ก ์๋์ถํ๊ธฐ ๋๋ฌธ์๊ทผ๊ฑฐ๊ฐ๊ฒฐ๋ก ์์ ํ๋ท๋ฐ์นจํ์ง๋ชปํฉ๋๋ค. โ์ถ๋ก ์๊น์ดโ์ธก๋ฉด์์, ์ง๊ท ํ์ํต์ฌ์๋ฆฌ์ธ์ฝ๋ฆฌ์ฌ๋ฆฌํ๊ณผ์๋์์๊ด๊ณ์๋ํ์ดํด์์ด ํผ์์ ์ธ์ฌ์ค์๋์ดํ๋๋ฐ๊ทธ์ณค์ต๋๋ค. โํํ์๋ช ํ์ฑโ์ธก๋ฉด์์๋, ์กด์ฌํ์ง์๋์ฉ์ด์๋ชจํธํํํ(โ๊ธฐ์๊ฒฝ๋ก์๋ณํโ)์์ฌ์ฉํ์ฌ์๋ฏธ ์ ๋ฌ์์ฌ๊ฐํ๊ฒฐํจ์ด์์ต๋๋ค. List_of_Factual_Errors:[ It incorrectly referred toโgeostrophic windโasโdelayed windโ,a non-existent meteorological term., It disregarded the problemโs key assumption that theโpressure gradient is constantโand used the relationshipโwind speed increases as pressure gradient increasesโas the basis for reasoning., It confused the distinct physical concepts ofโpressure gradientโandโpressureโ, erroneously explaining the relationship between pressure and wind speed based on the relationship between pressure gradient and wind speed., ]Evaluation_Scores:Accuracy: 1, Logical_Completeness: 1, Depth_of_Reasoning: 1, Clarity_of_Expression: 1 Evaluation_Reasoning:The candidate AIโs response recorded the lowest scores across all evaluation axes. In terms ofโAccuracyโ,it referred to โgeostrophic windโusing the hallucinated termโdelayed windโand demonstrated a fatal conceptual error by confusingโpressure gradientโwith โpressureโ.Furthermore, it distorted factual relations by ignoring the problemโs core premise that theโpressure gradient is constantโ.In terms of โLogical_Completenessโ,the reasoning failed to support the conclusion at all because it presented generalities irrelevant to the problem conditions and derived the conclusion by confusingโpressure gradientโwithโpressureโ.In terms ofโDepth_of_Reasoningโ,it merely listed superficial facts without understanding the relationship between the Coriolis force and latitude, which is the core principle of geostrophic wind. In terms ofโClarity_of_Expressionโ, there were serious flaws in meaning delivery due to the use of non-existent terms and ambiguous expressions(e.g.,โchange in pressure pathโ). Response fromQwen3-VL-235B-A22B-Thinking ์์ฑ๋_์ถ๋ก _๊ทผ๊ฑฐ:์ง๊ท ํ์Vg = |p|/(ฯ f)์ด๋ฏ๋ก ๊ธฐ์๊ฒฝ๋์ ๋ฐ๋๊ฐ ์ผ์ ํ๋ฉด ์ฝ๋ฆฌ์ฌ๋ฆฌ๋งค๊ฐ๋ณ์ f=2ฮฉ sinฯ๋ง์ด๋ณ์๊ฐ๋์ด์๋๊ฐ์ฆ๊ฐํ ์๋กf ๊ฐ์ปค์ ธํ์์๊ฐ์ํ๋ฉฐ,์ ๋๊ธฐ์์ํฌ๊ธฐ(๋ณด๊ธฐ 3)๋๊ณ ๋(๋ณด๊ธฐ1),์ ์ฑ(๋ณด๊ธฐ4)์์ด๊ด๊ณ์์์ง ์ ๋ฑ์ฅํ์ง์๋๋ค. Expert-Verified Rationale: Since the geostrophic wind is defined as Vg = |p|/(ฯ f),if the pressure gradient and density are constant, the Coriolis parameter f=2ฮฉ sinฯ becomes the only variable. As latitude increases, f increases, causing the wind speed to decrease. The magnitude of absolute pressure(Option 3),altitude (Option 1),and viscosity(Option 4)do not appear di- rectly in this equation. ์ ๋ต:2 โ Predicted Answer: 2 Reasoning Evaluation from Human Experts: ์ฌ๋์ ๋ฌธ๊ฐํ๊ท ์ ์:์ฌ์ค์ ์ ํ์ฑ(5),๋ ผ๋ฆฌ์ ์๊ฒฐ์ฑ(5),์ถ๋ก ์๊น์ด(4.5),ํํ์๋ช ํ์ฑ(5) Average of Human Expert Score:factuality(5),log- icality(5),depth(4.5),clarity(5) Comments: Reasoning Evaluation from LLM-Evaluator (Gemini-2.5-Pro): ์ฌ์ค_์ค๋ฅ_๋ชฉ๋ก:[] ํ๊ฐ_์ ์:์ ํ์ฑ:5,๋ ผ๋ฆฌ์ _์๊ฒฐ์ฑ:5,์ถ๋ก ์ _๊น์ด:5,ํํ์_๋ช ํ์ฑ:5 ํ๊ฐ_์ฌ์ :์ํ์AI์๋ต๋ณ์๋ชจ๋ ํ๊ฐ์ถ์์ ๋งค์ฐ์ํ์์ค์๋ณด์์ต๋๋ค. โ์ ํ์ฑโ์ธก๋ฉด์ ์,์ ์๋์ง๊ท ํ์๋๊ณต์,์ฝ๋ฆฌ์ฌ๋ฆฌ๊ณ์์์ ๋์๊ด๊ณ,๊ทธ๋ฆฌ๊ณ ๊ฒฐ๋ก ๊น์ง๊ธฐ์ํ์ ์ฌ์ค์์๋ฒฝ ํ๊ฒ๋ถํฉํ๋ฉฐ์ด๋ ํ์ค๋ฅ๋๋ฐ๊ฒฌ๋์ง์์์ต๋ ๋ค. โ๋ ผ๋ฆฌ์ ์๊ฒฐ์ฑโ์ธก๋ฉด์์,๋ฌธ์ ์์กฐ๊ฑด(๊ธฐ์ ๊ฒฝ๋,๋ฐ๋์ผ์ )์ผ๋ก๋ถํฐํ์์ด์ฝ๋ฆฌ์ฌ๋ฆฌ๊ณ์์ ๋ฐ๋น๋กํ๋ค๋ํต์ฌ๊ด๊ณ๋ฅผ๋ช ํํํ๊ณ ,์ด๋ฅผ์๋ ์์๊ด๊ณ๋กํ์ฅํ์ฌ๊ฒฐ๋ก ์๋์ถํ๋๊ณผ์ ์ด๋น ์ฝ์์ด์๋ฒฝํ๊ฒ์ฐ๊ฒฐ๋์์ต๋๋ค. โ์ถ๋ก ์๊น์ดโ ์ธก๋ฉด์์,์ ๋ต์๊ทผ๊ฑฐ๋ฅผ์ ์ํ๋๊ฒ์๊ทธ์น์ง์ ๊ณ ,๋งค๋ ฅ์ ์ธ์ค๋ต์ ํ์ง(3๋ฒ,4๋ฒ)๊ฐ์์ค๋ต์ธ ์ง๋ฅผ์ง๊ท ํ์๊ธฐ๋ณธ๊ฐ์ (๋ง์ฐฐ๋ฌด์)๊ณผํต์ฌ๋ณ์ (๊ธฐ์๊ฒฝ๋)๋ฅผ ๋ค์ด ๋ช ํํ ๋ฐ๋ฐํจ์ผ๋ก์จ ๋ฌธ์ ์ ๋ํ๊น์ด์๋์ดํด๋ฅผ๋ณด์ฌ์ฃผ์์ต๋๋ค. โํํ์ ๋ช ํ์ฑโ์ธก๋ฉด์์๋๊ตฐ๋๊ธฐ์์ดํต์ฌ๋ด์ฉ์๊ฐ ๊ฒฐํ๊ณ ๋ช ํํ๊ฒ์ ๋ฌํ์ฌ์ดํด๋๋ฅผ๋์์ต๋๋ค. ์ด๋๋ชจ๋ฒ์ ์ธํด์ค์ด๋ผํ ์์์ต๋๋ค. List_of_Factual_Errors:[] Evaluation_Scores:Accuracy:5, Logi- cal_Completeness: 5, Depth_of_Reasoning: 5, Clarity_of_Expression: 5 Evaluation_Reasoning:The candidate AIโs response demonstrated a very high level of performance across all evaluation axes. In terms ofโAccuracyโ,the pre- sented geostrophic wind speed formula, the relation- ship between the Coriolis parameter and latitude, and the conclusion perfectly matched meteorologi- cal facts with no errors found. In terms ofโLogi- cal_Completenessโ,the process of clarifying the core relationship that wind speed is inversely proportional to the Coriolis parameter from the problem condi- tions(constant pressure gradient and density)and ex- tending this to the relationship with latitude to de- rive the conclusion was perfectly connected without logical leaps. In terms ofโDepth_of_Reasoningโ,it demonstrated a deep understanding of the problem not only by providing the basis for the correct answer but also by clearly refuting why the attractive distrac- tors(Options 3 and 4)were incorrect, citing the ba- sic assumption of geostrophic wind(ignoring friction) and the key variable(pressure gradient).In terms of โClarity_of_Expressionโ,it delivered the core content concisely and clearly without redundancy, enhancing comprehensibility. This can be considered an exem- plary explanation. A Dataset Examples Table6presents representative examples from K- MetBench, organized into two primary modality groups:Text-onlyandMultimodal. Within each modality, we further stratify the tasks into three dis- tinct categories to evaluate comprehensive meteo- rological capabilities. Text-only Tasksassesses linguistic reasoning and theoretical knowledge without visual interpre- tation. This group includes(a)Standard MCQA for fundamental concepts,(b)K-Specific MCQA which requires geo-cultural knowledge specific to the Korean Peninsula, and(c)Reasoning MCQA that demands multi-step logical deduction. Multimodal Tasksintroduces visual data inter- pretation, a critical skill for meteorologists. This group parallels the text-only structure with(d) Standard, (e)K-Specific,and(f)Reasoningsub- sets, but specifically evaluates the modelโs ability to analyze weather charts, satellite imagery, and at- mospheric diagrams. This structured categoriza- tion allows for a clear comparison of model per- formance across different modalities and levels of domain expertise. B Case Study of Reasoning Answer Two human experts and the LLM-Evaluator (gemini-2.5-pro)conducted evaluations using identical rubrics. As shown in Table 7, we ob- served consensus between the human and AI eval- uators for theInternVL3.5-8B-Instructand Qwen3-VL-235B-A22B-Thinkingmodels: both correctly identified incorrect answers and high- lighted inappropriate terminology in the reasoning rationale. Notably, the human expert made a specific er- ror by misreadingโ๊ธฐ์๊ฒฝ๋โ(pressure gradient) asโ๊ธฐ์๊ฒฝ๋กโ(pressure path).While the quantita- tive scores assigned by the human expert and the LLM-Evaluator were comparable, the granularity of their feedback differed significantly. The hu- man expert made an implicit judgment, providing summary comments alongside the score. In con- trast, the LLM-Evaluator generated more detailed outputs, including explicit justifications and com- prehensive lists of factual errors. C Prompts and Questionnaires for Benchmark Construction This section details the prompts utilized for data augmentation(paraphrasing)and the identification of domain-specific subsets. These processes were conducted to enhance the quality of the dataset and provide rich learning signals. C.1 Question Paraphrasing To diversify sentence structures and lexical expres- sions while preserving the original semantic mean- ing of the questions, we utilized theGemini-2.5- Promodel. Figure5presents the specific system prompt employed for this paraphrasing task. C.2 Identification of Korean-Specific Subset To identify questions containing Korean-specific geographical and cultural contexts(theKorean- Specificsubset)from the total pool of 1,774 ques- tions, we established a hybrid pipeline combining LLM-based filtering with human verification. LLM-Aided IdentificationThe screening pro- cess involved independent filtering using two dis- tinct models:Gemini-2.5-ProandGPT-4.1. The identification prompts for each model were op- timized through an iterative refinement process to maximize recall. Figures 12and14illustrate the fi- nal enhanced prompts used for identifying Korean- specific context questions, respectively. Human Selection ProcessBased on the LLM filtering,Gemini-2.5-Proextracted 135 candi- dates, whileGPT-4.1extracted 95 candidates. We consolidated these results into a union of 149 unique questions. Subsequently, two human re- searchers performed cross-validation on this candi- date set to finalize theKorean-Specificsubset. The selected questions typically contain high-context keywords such asโOur countryโ(์ฐ๋ฆฌ๋๋ผ),โKo- rean Peninsula,โ โJeju,โ โSeoul,โ โYeongdong,โ โSoutherly windโ(๋งํ๋),โTaebaek Mountains,โ andโ24 Solar Terms.โ C.3 Implicit vs. Explicit Dataset Design for Korean-Specific Subset To ensure a fair evaluation of local context under- standing regardless of the modelโs primary training language, we constructed a dual-version dataset by converting implicit questions into explicit ones. โขImplicit Questions:These refer to the original items containing high-context expressions that presuppose the speakerโs spatiotemporal and cul- tural location(e.g.,โOur country,โโMaparam,โ โEast Coastโ). โขExplicit Questions:These refer to the modi- fied items where human researchers manually replaced high-context references with objective and unambiguous terminology(e.g., changing โOur countryโtoโSouth KoreaโorโMaparamโ toโSoutherly wind, a pure Korean termโ). Table8presents comparative examples of these original implicit questions and their explicit coun- terparts. Table 8:Examples of Context Transformation from implicit to explicit forms ID Implicit(Before)Explicit(After) All์ฐ๋ฆฌ๋๋ผํ๊ตญ์ง์ญ 618์์ธํ๊ตญ์์์ธ์ง์ญ 1037 24์ ๊ธฐ๋์์์์ง์ญ์24์ ๊ธฐ 822๋ํด์ํ๊ตญ์ง์ญ์๋ํด์ 557๊ฒจ์ธ์ฒ ๋ฐํด๋ง์ ์์์๊ธฐ์๊ณจ์ด ์ ๊ทผํ๊ณ ์๋ค. ๊ฒจ์ธ์ฒ ๋ฐํด๋ง์ผ๋ก๋ถํฐ์๊ท๋ชจ ๊ธฐ์๊ณจ์ดํ๊ตญ์ง์ญ์ผ๋ก์ ๊ทผํ ๋์ํฉ์์ 271, 1744 ๋งํ๋ํ๊ตญ์ง์ญ์์ง๋ฐฉํ์ธ๋งํ๋ C.4 Evaluation Prompts for Korean-Specific Subset This section details the construction of system prompts designed to evaluate the modelโs under- standing of geo-cultural contexts. To encourage the model to effectively utilize its latent local knowledge, we designed anAdvanced Promptthat explicitly defines the speakerโs persona(i.e., a Korean meteorology expert)and clarifies that the questions are contextually situated in Korea. To quantify the prompting gainโthe extent to which this contextual cuing aids performanceโ and to ensure equitable evaluation for non-Korean models, we also established aStandard Promptas a control group. Figure 22presents the standard sys- tem prompt used for the baseline experiment, while Figure24displays the advanced system prompt used to test the activation of geo-cultural knowl- edge. C.5 Prompt for Reference Rationale Generation To secure high-quality reasoning references(ratio- nales)for the benchmark, we utilized theGPT-5 model. The prompt engineering process employed an iterative refinement technique. Specifically, we established a loop where anEnhancermodel drafted the initial prompt and aCriticmodel identi- fied weaknesses for revision, usingGPT-5for both roles to derive the optimal instruction. The final system prompt used for rationale generation is pre- sented in Figure 16. To ensure comprehensive coverage, the target questions were selected via stratified sampling to include all subject areas, modalities(text-only/mul- timodal),and Korean-specific items. Furthermore, to guarantee the validity of the reasoning paths, we enforced a strict filtering protocol: if the model generated an incorrect answer, the generation pro- cess was repeated until a rationale leading to the correct answer was produced. C.6 Questionnaire for Expert Verification on Reference Rationale To ensure the reliability of the LLM-as-a-judge pipeline, two human experts conducted a rigor- ous verification of the generated rationales from September 9 to October 19, 2025. This process was critical for establishing the integrity of the ref- erence data. Before the verification process, the experts were given written instructions describing the purpose of the study, the expected completion time, and how their judgments would be used in the research. They were asked to assess each gener- ated rationale in terms of factual accuracy, logical soundness, completeness, and conciseness, and to mark whether the rationale should be adopted as is System Prompt for Question Paraphrasing(Gemini-2.5-Pro) Task:Paraphrase the following multiple-choice question about meteorology in Korean. Rules: - Preserve the core meaning and all technical terminology. - Change the sentence structure or phrasing for a more natural flow. Output Format:Provide only the final paraphrased text. Do not include any introductory phrases or explanations. โ Original Question:โoriginal_questionโ Figure 5:System prompt used to paraphrase Korean meteorological questions or revised. When revisions were needed, they were instructed to provide either minor-fix or major-fix notes. An example of the questionnaire used for this process is shown in Table 11. Out of 142 rationales initially generated by GPT-5, experts provided feedback for revision on 19 cases(13.38%).The revisions primarily ad- dressed technical accuracy and clarity. Specifi- cally, experts corrected erroneous terminology(5 cases),such as changingโ๋น์ด์ฉ๋โ(specific heat capacity)toโ๋น์ดโ (specific heat)orโ์งํ์์ฉ ๋์ด๋โtoโํ์ฑ์์ฉ๋์ด๋โ (planetary vortic- ity).They also reinforced variable explanations and standard units(3 cases);for instance, refining the phrasingโamong the temperatures handledโ to โamong the variables handled in atmospheric sci- enceโ becauseโํผํฉ๋นโ(mixing ratio)is not a tem- perature variable. Additionally, the revisions in- cluded full sentence rewriting(7 cases),supple- mentary explanations(2 cases),and minor stylis- tic polishing(2 cases)to align with standard Ko- rean meteorological conventions(e.g., standardiz- ingโํฌํ ์ ์จ๋โtoโ์จ์โ,โ๋ฐํธ๋กํฝโtoโ์์โ, andโ๋จ์ด๊ฐ์ดโtoโ๋จ์ด์์ถโ). In addition to refining the AI-generated ratio- nales, this expert review also identified inherent de- fects in the raw exam data. One question(ID 276) was discarded from the dataset as it was deemed logically unsolvable. Furthermore, questions with syntactic errors(IDs 308, 650)or issues with op- tion configuration/double answers(IDs 14, 583, 1665)were precisely corrected based on expert consultation. Through this process, we secured the integrity of the final 141 reasoning evaluation sam- ples. Table 11presents the specific questionnaire used for this expert verification process. C.7 Reasoning Prompt for Open-Source LLMs Figure 18presents the system prompt utilized for generating reasoning paths and answers from open- source LLMs. It is important to note that this prompt serves as the standard instruction for the main inference phase of our benchmark evaluation protocol, rather than an experimental variation. C.8 Prompt for LLM-as-a-Judge Evaluation Figure20illustrates the specific system prompt em- ployed for the LLM-as-a-Judge evaluation pipeline. The prompt was meticulously designed with the following key considerations to ensure robust alignment with human expert evaluation: โขUnified Evaluation Scale:We adopted the iden- tical 1-to-5 Likert scale and four evaluation axes โFactual Accuracy, Logical Soundness, Depth of Reasoning,andClarity & Concisenessโused by human experts. This unification allows for di- rect statistical comparison and correlation analy- sis between LLM and expert scores. โขEnforced Chain of Thought(CoT):To en- hance consistency, the prompt explicitly man- dates a step-by-step thinking process. The eval- uator is required to verify facts against the pro- vided expert reference materialbeforeassigning scores, thereby minimizing hallucinations and ensuring evidence-based grading. โขExplicit Scoring Criteria:To prevent arbitrary scoring, we defined concrete rubrics for spe- cific score tiers(e.g., distinguishing between a 5-point perfect answer and a 3-point answer with minor errors). โขStructured Output:The prompt enforces a strict JSON output format that separates the โList of Factual Errorsโfrom the quantitative scores. This structural constraint compels the model to explicitly isolate factual hallucinations from qualitative reasoning flaws. C.9 Questionnaire for Expert Scoring of LLM Reasoning Table 12outlines the questionnaire and scoring rubric provided to two meteorology professors. Crucially, this rubric served as the blueprint for the LLM-as-a-Judge prompt described above, en- suring that both human and AI evaluators oper- ated under identical standards regarding accuracy and reasoning quality. The human experts were also provided with written scoring instructions that described the study purpose, the expected anno- tation time, and the four evaluation axes: factual accuracy, logical soundness, depth of reasoning, and clarity. To reduce bias, model identities were blinded in the scoring materials, and the experts were instructed to judge only the content of the gen- erated reasoning against the expert-verified refer- ence rationale. An example of the scoring ques- tionnaire is provided in Table 12. D Experimental Setups D.1 Reasoning Model Inference We activate the thinking mode for hybrid models by settingenable_thinking = True(Qwen3- *B,EXAONE-4.0- *),reasoning_effort = 'high'(gpt-5.2),andthinkingLevel = 'high'(gemini-3-pro-preview).In contrast, InternVL3.5- *- Instructis evaluated in standard instruct mode. D.2 Meta-Evaluation for LLM-as-a-Judge To validate the reliability of the LLM evaluator, we designed a meta-evaluation protocol consisting of three steps: 1)generating reasoning paths and answers using various open-source LLMs; 2)per- forming LLM-as-a-Judge evaluation using expert- verified references and a specific rubric(based on a 5-point Likert scale across four evaluation axes); and 3)obtaining scores from two human experts using the identical rubric to calculate the statistical correlation between the LLM judge and human ex- perts. The detailed prompt for the main inference, the judge prompt, and the expert scoring question- naire are provided in Appendix C.7,C.8, andC.9, respectively. The full questionnaires and written instructions provided to the human experts are in- cluded in Appendix C.6and AppendixC.9(Fig- ure 10and Figure11). Sampling Strategy of Target ModelsTo ensure that the LLM-as-a-Judge can reliably evaluate rea- soning capabilities across a broad spectrum of pro- ficiency, we employed a performance-based strat- ified sampling strategy. We categorized the pool of candidate models into three distinct tiersโTop, Mid, and Lowโbased on their normalized reason- ing scores on the 141 reasoning questions. From these strata, we selected representative models to form a final set of 10 target models for the meta- evaluation, ensuring that the judge is tested against both high-quality coherent reasoning and lower- quality outputs. The list of sampled models is de- tailed in Table 9. Stratified Sampling of Evaluation ItemsTo es- tablish a robust gold standard for scoring, we se- lected 10 representative reasoning questions. In- stead of random selection, we applied a stratified sampling to ensure both comprehensiveness and discriminatory power. The selection process in- volved the following criteria: โขItem Di๏ฌiculty:We classified the 141 reasoning questions into three difficulty tiers based on the average normalized reasoning scores of 10 open- source LLMs: Hard(Top 30%),Mid(40%),and Easy(Bottom 30%).We sampled 3, 4, and 3 Table 9:List of Sampled Models for Meta- Evaluation of LLM-as-a-Judge.Models were se- lected via stratified sampling based on their normalized reasoning score tiers to ensure diverse evaluation tar- gets. Reas. denotes the normalized reasoning score on theReasoningsubset questions. Tier Model NameFamilySize(B)Reas. Top Qwen3-VL-235B-A22B-Thinking Qwen235.0 4.31 gpt-oss-120bOpenAI120.0 4.03 A.X-4.0SKT72.0 3.87 Mid command-a-reasoning-08-2025 Cohere111.0 3.53 gpt-oss-20bOpenAI20.0 3.39 VARCO-VISION-2.0-14BNCSoft14.0 2.81 InternVL3.5-14B-InstructOpenGVLab15.0 2.36 Low Llama-3.1-8B-InstructMeta8.0 1.91 InternVL3.5-8B-InstructOpenGVLab8.0 1.77 Qwen3-0.6BQwen0.6 1.15 questions from each respective group to balance the difficulty distribution. โขDiscriminatory Power:Within each difficulty tier, we prioritized questions with a high stan- dard deviation in accuracy across the 10 models. A high standard deviation indicates that the ques- tion effectively discriminates between high- and low-performing models. โขCategory Coverage:The selection was fur- ther constrained to ensure a balanced inclusion of text-only, multimodal, and Korean-specific questions, as well as coverage across the official exam subject areas(Parts 1, 3, 4, and 5). Based on these criteria, the final 10 questions selected for meta-evaluation are: IDs 105, 1618, 1590, 14, 456, 1694, 963, 1745, 131, and 1224. Ta- ble 10details the characteristics of these sampled items. E Additional Results and Discussion E.1 Detailed Orthogonality Analysis of K-MetBench The Uniqueness of Visual Reasoning.As illus- trated in Figure 7, theMultimodalsubset of K- MetBench displays consistently low correlations (avg.ฯ b <0.30)1across all external bench- marks, including text-based baselines(KMMLU- Pro, KMMLU-Redux, ClimaQA)and weather- domain vision benchmarks(ClimaIQA, Weath- erQA).This disconnect quantitatively confirms the modality gap, demonstrating that the ability to in- Table 10:Statistics of Selected Evaluation Items. Mean and Std. Dev. represent the item-wise normal- ized reasoning scores(1-5)across the selected models. TierID Mean Std. Dev. Part Note Hard 1052.231.723- 1618 2.351.655- 1590 2.551.913- Mid 142.701.891- 4562.731.883- 1694 3.101.935- 9633.131.814Korean Easy 1745 3.351.894- 1313.401.745- 1224 3.531.764- terpret meteorological charts and symbols is a dis- tinct skill set not linearly correlated with general linguistic or textual reasoning capabilities. To investigate the orthogonality of our bench- mark, we further analyzed correlations with KMMLU-Pro and -Redux. For KMMLU-Redux, where only the test set is publicly available, we specifically partitioned the data into the 39 ques- tions derived from the 2022 Meteorological Engi- neer exam versus the remaining 2,547 general ques- tions. The sample pool for this analysis consisted of 25 open-source VLLMs forMultimodalsubset comparisons and 52 open-source models for other subsets(excluding the proprietary models listed in Table 4). As shown in Figure6, KMMLU-Pro exhib- ited weaker correlation due to domain divergence. Within KMMLU-Redux, the isolated 39-question meteorological subset showed lower correlation (ฯ b = 0.70)than the full dataset(ฯ b = 0.78), suggesting that this small subset is insufficient to capture comprehensive meteorological capabil- ity. Crucially, a significant drop in correlation is observed for K-MetBenchโs multimodal and rea- soning subsets, highlighting the structural gap be- tween our multimodal evaluation and existing text- only licensing exams. E.2 Detailed Analysis of K-MetBench Performance Table4presents the comprehensive leaderboard of K-MetBench, evaluating a diverse range of propri- etary and open-source models. The results are cate- gorized by model type, capabilities(Korean-native, Multimodal, Reasoning),and granular domain per- Figure 6:Heatmap of Kendallโsฯ b rank correla- tions between K-MetBench and KMMLU, KMMLU- Redux).In KMMLU-Redux,โ39โdenotes the subset of 39 Meteorological Engineer Exam questions, while โAllโ39โrefers to the remaining subset excluding these meteorological questions. Acc.: Accuracy. formance. SOTA Performance and the Impact of Rea- soning.Proprietary models dominate the upper echelon of the leaderboard.gemini-3-pro- preview(Thinking)achieves state-of-the-art per- formance with a total accuracy of 93.7%, signif- icantly outperforming other contenders. A no- table trend is the efficacy ofThinking(reason- ing)models; for instance,gpt-5.2(Thinking) scores 87.8%, showing a substantial improvement (+10.2%p)over its standard counterpart,gpt-5.2 (77.6%).This pattern reinforces that chain-of- thought capabilities are crucial for solving complex meteorological problems. Open-Source Landscape.In the open-source domain, theQwen( Yang et al.,2025;Bai et al., 2025;Qwen et al.,2024)series exhibits ex- ceptional performance.Qwen3-VL-235B-A22B- Thinkingleads this category with 84.4%. Even smaller models likeQwen3-VL-32B-Thinking (78.6%)surpass much larger non-reasoning mod- els(e.g.,gpt-oss-120b, 77.3%),highlighting the efficiency of reasoning-enhanced architectures in specialized scientific domains. The Modality Gap.A critical disparity exists between textual and visual reasoning. While top models achieve near-perfect scores on theTextsub- set(e.g., Gemini: 94.6%),their performance drops significantly on the Multimodal subset ( Gemini: 75.6%).This modality gap is even more pro- nounced in other models;gpt-5.2 (Thinking) sees a drastic decline from 90.6%(Text)to 29.3% (Multi).This indicates that while current LLMs excel at theoretical knowledge retrieval, they still struggle with interpreting professional meteorolog- ical charts and diagrams. Geo-Cultural Alignment and Granularity. Korean-native models demonstrate distinct advan- tages in localized contexts.A.X-4.0achieves a highK-Specificscore of 78.9%, outperforming several larger global models in this specific subset, despite a lower overall accuracy. In terms of domain granularity(P1โP5),models generally perform best inMeteorological Observation(P2), likely due to the descriptive nature of the questions, while struggling more inAtmospheric Dynamics (P3)andAtmospheric Physics(P5),which require deeper calculation and physical conceptualization. E.3 Results of Meta Evaluation of LLM-as-a-Judge Rank Preservation Analysis.In benchmark evaluation, the accuracy of relative ranking is often more critical than absolute scores. The slope graph in Figure 8compares the rankings assigned by hu- man experts and the LLM. Although minor rank fluctuations exist, the overall trend distinguish- ing high-performing models from low-performing ones is preserved. Human Rank LLM Rank A.X-4.0 (#2)(#2) A.X-4.0 InternVL3_5-14B-Instruct (#8) (#7) InternVL3_5-14B-Instruct InternVL3_5-8B-Instruct (#9)(#9) InternVL3_5-8B-Instruct Llama-3.1-8B-Instruct (#7) (#7) Llama-3.1-8B-Instruct Qwen3-0.6B (#10)(#10) Qwen3-0.6B Qwen3-VL-235B-A22B-Thinking (#1)(#1) Qwen3-VL-235B-A22B-Thinking VARCO-VISION-2.0-14B (#6)(#6) VARCO-VISION-2.0-14B command-a-reasoning-08-2025 (#4)(#4) command-a-reasoning-08-2025 gpt-oss-120b (#3)(#3) gpt-oss-120b gpt-oss-20b (#5)(#5) gpt-oss-20b Figure 8: Slope graph of rank changes of reasoning evaluation scores of two human experts vs. LLM evaluator.The crossing lines indicate minor discrep- ancies, but the overall performance tiers remain largely consistent. E.4 MCQA Accuracy vs. Reasoning Score As illustrated in Figure9, we analyze the relationship between answer accuracy and qualitative reasoning capabilities. The color gradient represents theReasoning Score Gap, defined as the disparity between the reason- ing score of correctly answered items and the overall average(i.e., Reasoning Score Gap= Reasoning Score |A=correct โReasoning Score total ). We observe a strong positive correlation(r= 0.959)between QA accuracy and reasoning scores, Figure 7:Heatmap of Kendallโsฯ b rank correlations between K-MetBench and existing benchmarks.Red denotes high positive correlation, while blue indicates negative correlation. indicating that models that derive correct answers also tend to generate higher-quality reasoning traces. A distinct scaling law is also evident; larger models(shown by marker size)consistently pop- ulate the upper-right quadrant, achieving superior performance in both metrics. Two notable trends appear among specific models: High Reasoning but Low Accuracy:Qwen3- VL-8B-Thinkingemerges as an outlier. Despite its relatively low accuracy, it maintains a high rea- soning score. This suggests that while the model generates detailed โthinkingโ processes, its limited capacity(8B)often leads to hallucinations or logi- cal fallacies in the final deduction. Impact of Reasoning Optimization:The benefit of reasoning-specific training is high- lighted by the Command family.command-a- reasoning-08-2025significantly outperforms its predecessor,c4ai-command-a-03-2025, in both accuracy and reasoning quality, validating the efficacy of reasoning-enhanced fine-tuning. Figure 9:Scatter plot of MCQA Accuracy vs. Rea- soning Score.The x-axis represents the answer accu- racy, while the y-axis denotes the qualitative reasoning score evaluated by the judge. Marker sizes are propor- tional to the model parameter count. The strong correla- tion(r= 0.959)confirms that high-performing models generally provide more reliable reasoning traces. E.5 Computational Cost and E๏ฌiciency Analysis We evaluated the normalized total GPU compute time for 100 questions against the 150-minute exam limit(โ2.50 GPU-hours). Standard instruction-tuned models(e.g., Qwen2.5-VL-Instruct)demonstrated negligi- ble cost(<0.01GPU-hours),operating orders of magnitude faster than the human time constraint. Reasoning models exhibited significant com- putational overhead.WhileQwen3- VL-8B- Thinking(2.4 GPU-hours)remained within the limit, larger models likeQwen3-VL-32B- Thinking(3.8 GPU-hours)andcommand-a- reasoning(20.8 GPU-hours)exceeded the threshold, highlighting the substantial resource trade-off required for deep reasoning. While this heavy computational overhead may yield deeper reasoning traces, it poses challenges for time-sensitive forecasting applications where rapid decision-making is critical. However, em- ploying tensor parallelism can effectively reduce wall-clock inference time. F Compute Resources We evaluated all open-source models on an inter- nal cluster equipped with 40 NVIDIA H100 80GB PCIe GPUs. To maximize inference efficiency, we utilized the vLLM library for all benchmark evaluations. The evaluation covered 52 text-only and multimodal models across all subsets of K- MetBench, totaling approximately 192.14 H100 GPU hours(153.01 and 39.13 GPU hours for stan- dard MCQA and reasoning MCQA, respectively). G Hierarchical Topic Distribution Figures29through31illustrate the comprehensive hierarchical taxonomy of the K-MetBench dataset, aligned with the official evaluation cri- teria of the National Meteorological Engineer writ- ten examination( Human Resources Development Service of Korea,2024). The dataset spans five major subject areas:Weather Analysis and Fore- casting Theory,Meteorological Observation Meth- ods,Atmospheric Dynamics,Climatology, andAt- mospheric Physics. As depicted in Figure29-31, each subject area and hierarchy demonstrates the benchmarkโs fine- grained granularity and comprehensive coverage of meteorological domain knowledge. The numer- ical values in parentheses represent the estimated count of questions belonging to each specific cate- gory. To map the 1,774 questions to this detailed hierarchy, we employedGemini-2.5-Profor au- tomated classification. These counts serve as an indicative reference, highlighting the datasetโs bal- anced coverage across the theoretical and practical spectrums of meteorology. Table 11:Expert verification questionnaire for reference rationales ID QuestionChoicesExact Answer Generated Rationale Adopt? Note 1 (Minor Fix) Note 2 (Major Fix) 18 ๋น์ด์... 1. ...2 ๋น์ด์... yes-- Note:Thegray-shaded cellsindicate the items to be answered by the expert. Adoption Criteria(Accuracy, Logical Soundness, Completeness, and Conciseness)are provided separately. (a)Text-only Reasoning Questions (b)Multimodal Reasoning Questions Figure 10:Examples of the expert verification questionnaire for reference rationales(in Korean) Table 12:Questionnaire for expert scoring of open-source LLM reasoning results ID Question Choices Exact Answer GT-R Target RFact (1-5) Sound (1-5) Depth (1-5) Clear (1-5) Total (4-20) Note 18 ๋น์ด์... 1. ...2 ๋น์ด์...๋น์ด์ฉ๋์... 12159- Note:GT-R: Expert-verified Gold Rationale, Target R: LLM Generated Reasoning. Scoring: 1(Poor)โ 5(Excellent).The gray-shaded cellsindicate the items to be answered by the expert. Figure 11:Examples of the questionnaire for expert scoring of open-source LLM reasoning results(in Korean) System Prompt for Identifying Regional Questions(Gemini-2.5-Pro) ๋๋์ด์ ๋ถํฐ๊ธฐ์ํ์ ๋ฌธ๊ฐ์ผ. ์ฃผ์ด์งโ๊ธฐ์๊ธฐ์ฌ๋ฌธ์ โ๋ฅผ๋ถ์ํด์ค. ๊ฐ๋ฌธ์ ์๋ด์ฉ๊ณผ์ ํ์ง๋ฅผ์ ์คํ๊ฒ๊ฒํ ํด์,์ ์ธ๊ณ์ ์ผ๋ก์ ์ฉ๋๋ ์ผ๋ฐ์ ์ธ๊ธฐ์์ง์์ด์๋๋ผ์ค์งํ๊ตญ์์ง๋ฆฌ,๊ธฐํ,๊ธฐ์์์คํ ,๊ด๋ จ๊ธฐ๊ด์๋งํด๋น๋๋โ์ง์ญ์ฑ์ด๊ฐํโ ๋ฌธ์ ๋ฅผ์๋ณํด์ค. ๋ค์๊ณผ๊ฐ์๊ธฐ์ค์์ฌ์ฉํด์๋ฌธ์ ๋ฅผ์ ๋ณํด: -ํ๊ตญ์ํน์ ์ง์ญ(์: ์๋,์์,์ธ๋ฆ๋,์ํด์)์๊ธฐ์ํ์์๋ค๋ฃจ๋๋ฌธ์ -ํ๊ตญ์์๋ง์ฌ์ฉํ๋ํน์ ๊ธฐ์์ฉ์ด๋์๋ณด์์คํ ์๊ดํ๋ฌธ์ (์: ๋๋ค์๋ณด) -ํ๊ตญ๊ธฐ์์ฒญ๋๋๊ด๋ จ๊ธฐ๊ด์์ญํ ์ด๋์ ๋ฌด์์ง์ ์ ์ผ๋ก๊ด๋ จ๋๋ฌธ์ -ํฉ์ฌ,์ฅ๋ง,ํํ๋ฑํ๊ตญ์ํฐ์ํฅ์๋ฏธ์น๋๊ธฐ์ํ์์ด๋ผ๋,๊ทธ๋ด์ฉ์ดํ๊ตญ์ํน์ํ์ํฉ(์:ํน์ ์ง์ญ์ ์ํฅ,ํ๊ตญ์์๋ณด์ฒด๊ณ)๊ณผ๊ฒฐ๋ถ๋์ด์์๊ฒฝ์ฐ์๋ง์ ํ ๊ฒฐ๊ณผ๋๋ค๋ฅธ์ค๋ช ์์ด,์ค์งํด๋น๋ฌธ์ ์ID๋ฒํธ๋งํฌํจ๋์ซ์(int)ํ์์ผ๋ก์ถ๋ ฅํด์ค. Figure 12:The system prompt used by Gemini-2.5-Pro to filter Korea-specific meteorological questions English Translation of the System Prompt for Identifying Regional Questions(Gemini-2.5- Pro) You are an expert in meteorology. Analyze the providedโMeteorological Engineer Exam Questionsโ.Carefully review the content and options of each question to identifyโstrongly regionalโquestions that pertain solely to South Koreaโs geography, climate, weather systems, and relevant institutions, rather than general meteorological knowledge applicable globally. Use the following criteria to select the questions: - Questions dealing with meteorological phenomena in specific regions of Korea(e.g., Yeongdong, Yeongseo, Ulleungdo, West Coast) - Questions regarding specific meteorological terms or forecasting systems used exclusively in Korea(e.g., Neighbor- hood Forecast) - Questions directly related to the roles or operations of the Korea Meteorological Administration(KMA)or relevant institutions - Even for weather phenomena that significantly impact Korea, such as Asian Dust, Changma(rainy season),or Typhoons, select them only if the content is tied to Koreaโs specific context(e.g., impact on a specific region, Koreaโs forecasting system) Output the result strictly as integers representing the ID numbers of the relevant questions, without any further explanation. Figure 13:English translation of the system prompt used by Gemini-2.5-Pro to filter Korea-specific meteoro- logical questions System Prompt for Identifying Regional Questions(ChatGPT-4.1) #์ญํ ๋น์ ์โ๋ํ๋ฏผ๊ตญ๊ธฐ์๊ธฐ์ฌ์๊ฒฉ์ํโ์๊ณผ๋ชฉยท์ถ์ ๊ฒฝํฅ์์์๋๊ธฐ์๊ต์ก์ ๋ฌธ๊ฐ์ ๋๋ค.์ ์ธ๊ณ์ ์ผ๋กํต์ฉ ๋๋์ผ๋ฐ๊ธฐ์์ง์๊ณผ,ํ๊ตญ๊ธฐ์ํ์ ยท์ ๋ยท์งํยทํ์ ๋ฑ์ง์ญํนํ์ง์์์ฐจ์ด๋ฅผ๋ช ํํ๊ตฌ๋ถํ ์์์ต๋๋ค. #๋ชฉํ ์๋์์ ์๋๊ธฐ์๊ธฐ์ฌ์ํ๋ฌธ์ ์คโ์ ์ธ๊ณ์ด๋์๋๋์ผํ๊ฒ์ ์ฉ๋๋์ผ๋ฐ๊ธฐ์์ง์โ์ด์๋,โํ๊ตญ์ ํนํ๋์ง์ญ์ฑยท์ ๋์ฑยทํ์ ์ ๋ฐฐ๊ฒฝ์ด๊ฐํ๊ฒ๋ฐ์๋๋ฌธ์ ๋งโ์๊ณจ๋ผ๋ฌธ์ ID(๋ฒํธ)๋ง๋ฐํํ์ธ์. #ํ๋จ๊ธฐ์ค -ํ๊ตญ์ ์ฉ๊ด์ธก์ฅ๋นยท์ฉ์ด(์: ์ฅํ๊ณ ์ฃผํ๊ทผ์ ๋ฐฐ์ด๊ด์ธก๋ง,ASOS-K) -ํ๊ตญ๊ธฐ์์ฒญ(KMA)ํ์ ์ ์ฐจยท๋ฒ๊ทยท๊ณ ์(์: ํน๋ณด๋ฐํจ๊ธฐ์ค,๊ด์ธก๋ณด๊ณ ์๊ฐ๊ท์ ) -ํ๊ตญ์งํยท๊ธฐํํน์์ฑ(์: ์ํด์๋๊ตฌ๋ฆํน์ฑ,ํํ์๋ฅ๊ฒฝ๋กํต๊ณ) -์ฐ๋ฆฌ๋๋ผ์์๋ง์ฌ์ฉํ๊ฑฐ๋์ํ๋ฒ์์ํฌํจ๋๋๊ณ ์ ์ดํยท๋จ์ยท์ถ์ฝ์ด -์์๊ฐ์์์๊ฐ์ ํ์๊ณ ,ICAOยทWMO๋ฑ๊ตญ์ ๊ณตํต๊ท๊ฒฉ๋ง๋ค๋ฃจ๋ฉดโ์ ์ธ๊ณ์ผ๋ฐโ์ผ๋ก๊ฐ์ฃผ #์ถ๋ ฅํ์ -์๋ฌด์ค๋ช ์์ด์ซ์๋ง์ถ๋ ฅ ์: โ4โ -๋ฆฌ์คํธ๊ฐ๋น์ด์์ผ๋ฉดโ์์โ์ด๋ผ๊ณ ๋ง์์ฑ #์ฌ๊ณ ๊ณผ์ -๋ฌธ์ ํ๋์ฉ์ฝ๊ณ ,์โํ๋จ๊ธฐ์คโ์๋ฐ๋ผํ๊ตญํนํ์ฌ๋ถ๋ฅผ๋จผ์ ๋จธ๋ฆฟ์์์๊ฒฐ์ -์ต์ข ๋ต์์๋๊ฒฐ๊ณผ์ซ์๋ง๋จ๊ธฐ๊ณ ,์ค๊ฐ์ถ๋ก ์ด๋์ด์ ๋์ฐ์ง๋ง๊ฒ Figure 14:System prompt used by ChatGPT-4.1 to filter Korea-specific questions English Translation of the System Prompt for Identifying Regional Questions in English (ChatGPT-4.1) # Role You are a meteorology education expert familiar with the subjects and trends of theโRepublic of Korea Meteorological Engineer Certification Examโ.You can clearly distinguish between general meteorological knowledge accepted world- wide and region-specific knowledge related to Koreaโs weather operations, systems, topography, and administration. # Goal Among the provided exam questions, selectonlythose that arenotโgeneral meteorological knowledge applicable everywhereโbut reflectโstrong regional, institutional, or administrative backgrounds specific to Koreaโ.Return only the question ID numbers. # Criteria -Korea-exclusive observation equipment/terms(e.g., Long-wave/High-frequency proximity array observation network, ASOS-K) -Korea Meteorological Administration(KMA)administrative procedures/laws/notices(e.g., Special advisory criteria, observation reporting time regulations) -Korean topographical/climatic specifics(e.g., Characteristics of snow clouds on the West Coast, typhoon landfall path statistics) -Unique vocabulary, units, or abbreviationsused only in Korea or included in the exam scope - If none of the above exist and the question deals with international common standards like ICAO/WMO, treat it as โGeneral Globalโ # Output Format - Outputonly numberswithout any explanation Example:โ4โ - If the list is empty, writeโNoneโ # Reasoning Process - Read each question and decide strictly based on theโCriteriaโaboveinternally. - In the final answer, leaveonly the result numbersand do not write intermediate reasoning or reasons. Figure 15:English translation of the system prompt used by ChatGPT-4.1 to filter Korea-specific questions System Prompt for Reference Rationale Generation ๋น์ ์ํ๊ตญ์๊ธฐ์์ ๋ฌธ๊ฐ์ ๋๋ค. ๋น์ ์์๋ฌด๋์์๊ธฐ์์๋ณด๊ด์ด์๋ฌธ์ ์ถ์ ์์์ผ๋ก์์ฃผ์ด์ง๊ฐ๊ด์ ๋ฌธ์ ๋ฅผํ๊ฐํ๊ณ ,์ฌ๋ฐ๋ฅธ์ ๋ต์ ํ์ง๋ฅผ๊ณ ๋ฅธ๋ค๋ค์๋๋จ๊ณ๋ฅผ๋ฐ๋์์์ฐจ์ ์ผ๋ก์ํํ๋๊ฒ์ ๋๋ค. 1.ํต์ฌ์ถ๋ก ๊ทผ๊ฑฐ์์ฑ(ํ๋ฌธ์ฅ): -(์ ํ์ฑ)๋ฐ๋์์ ๋ต์ด์๋ง๋์ง์ถ๋ก ๋ด์ฉ์ด๊ธฐ์ํ์ ์ฌ์ค์๋ถํฉํ๊ฒ์ค๋ช ํ์ธ์. -(๋ ผ๋ฆฌ์ฑ)์ถ๋ก ์ด์ ๋ต์๋ช ํํ๊ณ ์ง์ ์ ์ผ๋ก๋ท๋ฐ์นจํ๋๋ก์ค๋ช ํ์ธ์. -(์๊ฒฐ์ฑ)๋ฌธ์ ์ํต์ฌ์ํํผํ์ง์๊ณ ์ค๋ช ํ์ธ์. ๊ฐ๋ฅํ๋ฉด,๊ฐ์ฅ๊ทธ๋ด๋ฏํ์ค๋ต์ ํ์ง๊ฐ์ํ๋ ธ๋์ง๋ ๊ฐ๋จํ์ธ๊ธํ์ฌ์ ๋ต์๋ ผ๋ฆฌ๋ฅผ๊ฐํํ์ธ์. -(๊ฐ๊ฒฐ์ฑ)๋ถํ์ํ๋ด์ฉ์์ดํ๋ฌธ์ฅ์ผ๋ก์์์ฝํ์ฌ์์ฑํ์ธ์. 2.์ ๋ต๊ฒฐ์ : -์ถฉ๋ถํ์ถ๋ก ์๋ฐํ์ผ๋ก,์ต์ข ์ ์ผ๋ก์ ๋ต๋ฒํธ(์ ํ์ง๋ฒํธ)๋ฅผ๋ช ํํ๊ฒฐ์ ํ์ธ์. ์๋ต๊ท์น -๋ฐ๋์์๋JSONํ์๋ง์ฌ์ฉํ์ฌ๋ต๋ณํด์ฃผ์ธ์. -์ถ๋ก ๊ทผ๊ฑฐ๊ฐ๋ฐ๋์๋จผ์ ๋์ถ๋๊ณ ,๊ทธ๋ค์์์ ๋ต์๋ช ์ํด์ผํฉ๋๋ค. -๊ฐํญ๋ชฉ์๋ชจ๋๋ฐ๋์ํฌํจํ๋ฉฐ,๋ณ์๋ช ๋ฐ๊ตฌ์กฐ๋์๋์๊ฐ์ดํ๊ธ๋ก์์ฑํด์ผํฉ๋๋ค. โ์์ฑ๋_์ถ๋ก _๊ทผ๊ฑฐโ: โํ๋ฌธ์ฅ,๋ ผ๋ฆฌ์ ์ค๋ช . ์ ๋ต์ด์๋ง๊ณ ์ฃผ์์ค๋ต์ด์ํ๋ ธ๋์งํฌํจ. ํ์์ด์ฝ๊ฒ ์ดํดํ ์์๊ฒ๊ฐ๊ฒฐํ๊ฒ์์ฑ.โ, โ์ ๋ตโ: โ๊ณ์ฐ๋์ ๋ต๋ฒํธ(์ซ์๋ง)โ Figure 16:System prompt used to generate a reasoning rationale English Translation of the System Prompt for Reference Rationale Generation You are a meteorology expert in Korea. Your mission is to act as aSenior Chief Forecaster and Exam Setterto evaluate the given multiple-choice question, select the correct answer, and strictly perform the following two steps sequentially. 1. Generate Core Reasoning Rationale(One Sentence): -(Factuality)You must explain why the answer is correct based on meteorological facts. -(Logicality)Explain the reasoning so that it clearly and directly supports the answer. -(Depth)Do not avoid the core of the question. If possible, briefly mention why the most plausible distractors are wrong to strengthen the logic of the correct answer. -(Clarity)Summarize well into one sentence without unnecessary content. 2. Determine Correct Answer: - Based on sufficient reasoning, clearly determine the final answer number(option number). Response Rules - You must answer strictly using the JSON format below. - The reasoning rationale must be derived first, followed by the answer. - All fields must be included, and variable names and structure must be written as shown below. โGenerated_Rationaleโ: โOne logical sentence. Explain why the answer is correct and why key distractors are wrong. Write concisely for easy student understanding.โ, โAnswerโ: โCalculated answer number(digits only)โ Figure 17:English translation of the system prompt used to generate a reasoning rationale System Prompt for Reasoning and Answer Generation(Open-Source LLM) ์ญํ (Role) ๋น์ ์ํ๊ตญ์๊ธฐ์์ ๋ฌธ๊ฐ์ ๋๋ค. ๋ค์๊ฐ๊ด์๋ฌธํญ์๋ตํ์ธ์. ์ง๋ฌธ์คโ์ฐ๋ฆฌ๋๋ผโ๋โํ๊ตญโ์์๋ฏธํฉ๋๋ค. ํ๊ตญ์์ง๋ฆฌ์ ํน์ฑ์๋ฐ์ํด์๋ต๋ณํด์ฃผ์ธ์. ์๋ฌด(Task) ์ฃผ์ด์ง๊ฐ๊ด์๋ฌธ์ ์๋ํด,๋จผ์ ๋น์ ์๋ชจ๋ ์๊ฐ๊ณผ์ ์์์ ๋กญ๊ฒ๊ธฐ๋กํํ,๊ทธ๋ด์ฉ์๋ฐํ์ผ๋ก์ฑ์ ๊ด์ด ๋ช ํํ์ดํดํ ์์๋๋ก์ ๋ต์๋ํํต์ฌ์ถ๋ก ๊ทผ๊ฑฐ์์ต์ข ์ ๋ต์JSONํ์์ผ๋ก์ ์ถํด์ผํฉ๋๋ค. ์ฌ๊ณ ๋ฐ์๋ต์ ์ฐจ(Thinking and Response Procedure) ๋น์ ์๋ฐ๋์์๋์2๋จ๊ณ์ ์ฐจ์๋ฐ๋ผ์๋ตํด์ผํฉ๋๋ค. 1๋จ๊ณ:์์ ๋ก์ด์ฌ๊ณ (<scratchpad>) -๋จผ์ <scratchpad>ํ๊ทธ์์์์์ ๋กญ๊ฒ์ฌ๊ณ ํฉ๋๋ค. ์ด๊ณต๊ฐ์๋น์ ๋ง์์๊ฐ์ ๋ฆฌ๊ณต๊ฐ์ ๋๋ค. -๋ฌธ์ ์ ํต์ฌ ๊ฐ๋ ์ ๋ถ์ํ๊ณ , ๊ด๋ จ๋ ๊ธฐ์ํ์ ์ง์์ ๋์ดํ๊ณ , ๊ฐ ์ ํ์ง๊ฐ ์ ๋ง๊ณ ํ๋ฆฌ๋์ง(O/X) ๋ฑ ๋น์ ์๋ชจ๋ ์๊ฐ์ํ๋ฆ์๊ทธ๋๋ก๊ธฐ๋กํ์ธ์. -์ค์:์ด<scratchpad>์๋ด์ฉ์์ต์ข ์ฑ์ ์๋ฐ์๋์ง์์ต๋๋ค. ํ์์๊ตฌ์ ๋ฐ์ง๋ง๊ณ ๋ง์๊ป์๊ฐํ์ธ์. 2๋จ๊ณ:์ต์ข ๋ต๋ณ์์ฑ(JSON) -<scratchpad>์์ฑ์ด๋๋๋ฉด,๋น์ ์์๊ฐ๊ณผ์ ์๋ค์๊ฒํ ํ์ธ์. -๊ทธ๋ด์ฉ์๋ฐํ์ผ๋ก,์๋์์ง์์ฌํญ์๋ง์ถฐ์ต์ข ๋ต๋ณ์JSONํ์์ผ๋ก์์ฑํฉ๋๋ค. 1.(โ์์ฑ๋_์ถ๋ก _๊ทผ๊ฑฐโ):์ด๋ค ๋ ผ๋ฆฌ์ ๊ณผ์ ์ ํตํด ์ ๋ต์ ์ ํํ๋์ง๊ฐ๊ฒฐํ ๋ฌธ๋จ ํํ๋ก ์์ฑํฉ๋ ๋ค. <scratchpad>์ ๋ด์ฉ์ ๊ทธ๋๋ก ๋ณต์ฌํ์ง ๋ง๊ณ , ํต์ฌ๋ง ์์ฝํ๊ณ ์ฌ๊ตฌ์ฑํด์ผ ํฉ๋๋ค. (์ ํ์ฑ, ๋ ผ๋ฆฌ์ฑ, ๊ฐ๊ฒฐ์ฑ,ํต์ฌํ์ ๊ธฐ์ค๊ณ ๋ ค) 2.(โ์ ๋ตโ):์์ถ๋ก ์๊ทผ๊ฑฐํ์ฌ์ต์ข ์ ๋ต์ด๋ผ๊ณ ์๊ฐํ๋์ ํ์ง๋ฒํธ๋ฅผํ๋๊ณ ๋ฆ ๋๋ค.๋ฐ๋์์ ์์ซ์๋ง ์ฑ์์์ถ๋ ฅํด์ผํฉ๋๋ค. ์ ์ฒด์๋ตํ์(Full Response Format) -๋น์ ์๋ต๋ณ์<scratchpad>๋ธ๋ก๊ณผJSON์ฝ๋๋ธ๋ก,๋๋ถ๋ถ์ผ๋ก๊ตฌ์ฑ๋์ด์ผํฉ๋๋ค. -์๋์์์๊ฐ์ด<scratchpad>๊ฐ๋จผ์ ์ ์๋๊ณ , ๊ทธ๋ฐ๋ก๋ค์์๋ฐ๋์JSON์ฝ๋๋ธ๋ก์ผ๋ก๊ฐ์ธ์ง์ต์ข ๋ต๋ณ์ด์์ผํฉ๋๋ค. <scratchpad> ์ฌ๊ธฐ์๋น์ ์๋ชจ๋ ์๊ฐ๊ณผ์ ์์์ ๋กญ๊ฒ์์ ํฉ๋๋ค. ์์:F =(C *9/5)+ 32๊ณต์์ฌ์ฉ... </scratchpad> โjson โ์์ฑ๋_์ถ๋ก _๊ทผ๊ฑฐโ: โ์ญ์จ30๋๋ฅผํ์จ๋ก๋ณํํ๋๊ณต์F =(C *9/5)+32๋ฅผ์ฌ์ฉํ๋ฉดํ์จ86๋๊ฐ๋๋ค.86์ 30๋ณด๋คํฐ๊ฐ์ด๋ฏ๋ก,๋์ผ์จ๋๋ฅผ๋ํ๋ผ๋ํ์จ์จ๋๊ณ์์์๊ธฐ๋ฅ์ด์ญ์จ์จ๋๊ณ๋ณด๋ค๋๋์ด์ฌ๋ผ๊ฐ๋ค.โ, โ์ ๋ตโ: 1 โ Figure 18:System prompt for open-source LLMs requiring a two-step process:free reasoning in a scratchpad followed by a structured JSON output English Translation of the System Prompt for Reasoning and Answer Generation(Open- Source LLM) Role You are a meteorology expert in Korea. Answer the following multiple-choice question. In the question,โour countryโrefers toโKoreaโ.Please answer reflecting the geographical characteristics of Korea. Task For the given multiple-choice question, you must first record your entire thought process freely, and then based on that content,submit the core reasoning rationale for the answer and the final answer in JSON formatso that the grader can clearly understand. Thinking and Response Procedure You must respond according to the following 2-step procedure. Step 1: Free Thinking(<scratchpad>) - First, think freely inside the<scratchpad>tags. This is your personal space for organizing thoughts. - Analyze the core concepts of the question, list relevant meteorological knowledge, and record your flow of thought, including why each option is correct or incorrect(O/X). -Important: The content of this<scratchpad>is not reflected in the final grading. Think freely without formatting constraints. Step 2: Final Answer Generation(JSON) - Once writing in<scratchpad>is finished, review your thought process. - Based on that,generate the final answer in JSON formataccording to the instructions below. 1.(โGenerated_Rationaleโ): Write aconcise paragraphexplaining the logical process used to select the answer. Do not copy the<scratchpad>content directly; summarize and restructure only the key points.(Consider accuracy, logic, conciseness, and core identification criteria.) 2.(โAnswerโ): Select the option number you think is the final correct answer based on the reasoning above. Must output only an integer. Full Response Format - Your response must consist of two parts: a<scratchpad>block and aJSON code block. - As shown in the example below, the<scratchpad>must be presented first, immediately followed by the final answer wrapped in a JSON code block. <scratchpad> Describe your entire thought process freely here. Example: Using formula F =(C *9/5)+ 32... </scratchpad> โjson โGenerated_Rationaleโ: โUsing the formula F =(C *9/5)+ 32 to convert 30 degrees Celsius to Fahrenheit gives 86 degrees Fahrenheit. Since 86 is greater than 30, the mercury column of the Fahrenheit thermometer rises higher than that of the Celsius thermometer for the same temperature.โ, โAnswerโ: 1 โ Figure 19:English translation of the System prompt for open-source LLMs System Prompt for LLM-as-a-Judge Evaluation(Gemini-2.5-Pro) ###์ญํ (Role) ๋น์ ์ ๊ธฐ์ํ ๋ถ์ผ์ ๊น์ ์ ๋ฌธ ์ง์์ ๊ฐ์ง, ๋งค์ฐ ์๊ฒฉํ๊ณ ๊ณต์ ํAI์ถ๋ก ๋ฅ๋ ฅ ํ๊ฐ ์์์ฅ์ ๋๋ค. ๋น์ ์ ๊ฐ์ ์ด๋ํธํฅ์์ด,์ค์ง์ฃผ์ด์งํ๊ฐ๊ธฐ์ค์๋ง๊ทผ๊ฑฐํ์ฌ๊ฐ๊ด์ ์ผ๋ก์ฑ์ ํด์ผํฉ๋๋ค. ###์๋ฌด(Task) ์ฃผ์ด์ง<๋ฌธ์ ์ ๋ณด>,<์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ>,๊ทธ๋ฆฌ๊ณ <ํ๊ฐ๋์๋ต๋ณ>์๋ฐํ์ผ๋ก,โ์ํ์AIโ๊ฐ์ ์ถํ๋ต๋ณ์ ํ์ง์ํ๊ฐํ๊ณ ๊ทธ๊ฒฐ๊ณผ๋ฅผ์ง์ ๋JSONํ์์ผ๋ก์ถ๋ ฅํด์ผํฉ๋๋ค. ###์ ๋ ฅ์ ๋ณด(Input Information) 1.<๋ฌธ์ ์ ๋ณด>:ํ๊ฐ์๋งฅ๋ฝ์ด๋๋์๋ณธ๊ฐ๊ด์๋ฌธ์ ์ ๋๋ค. (์ง๋ฌธ,์ ํ์ง,์ ๋ตํฌํจ) 2.<์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ>: 100%์ฌ์ค์ด๊ฒ์ฆ๋๋ชจ๋ฒํด์ค์๋ฃ์ ๋๋ค. ์ด์๋ฃ๋<ํ๊ฐ๋์๋ต๋ณ>์โ์ฌ์ค์ค๋ฅโ ๋โํ๊ฐ(Hallucination)โ์ํ์งํ๋์ ๋๊ธฐ์ค์ผ๋ก์ฌ์ฉํด์ผํฉ๋๋ค. 3.<ํ๊ฐ๋์๋ต๋ณ>:๋น์ ์ด์ฑ์ ํด์ผํ โ์ํ์AIโ๊ฐ์์ฑํ์ถ๋ก ๊ณผ์ ๋ฐ์ ๋ต์ ๋๋ค. ###์ฌ๊ณ ๊ณผ์ (Step-by-Step Thinking Process) ๋น์ ์ํ๊ฐ๋ฅผ์ํํ๊ธฐ์ ์๋ฐ๋์๋ค์์์ฌ๊ณ ๊ณผ์ ์๊ฑฐ์ณ์ผํฉ๋๋ค. 1.๋ชจ๋ ์ ๋ ฅ์ ๋ณด๋ฅผ์ถฉ๋ถํ์์งํฉ๋๋ค. 2.[์ฌ์คํ์ธ]:<ํ๊ฐ๋์๋ต๋ณ>์๋ชจ๋ ์ฃผ์ฅ์<์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ>์๋น๊ตํ์ฌ์ฌ์ค์ค๋ฅ๊ฐ์๋์ง๋จผ์ ํ์ธํ๊ณ ๋ชฉ๋ก์์์ฑํฉ๋๋ค. 3.[๊ฐ๋ณ์ถํ๊ฐ]:์๋์4๊ฐ์งโํ๊ฐ๊ธฐ์ค๋ฐ์ฒ๋โ๋ฅผํ๋์ฉ์ฝ๊ณ , ๊ฐ๊ธฐ์ค์๋ฐ๋ผ<ํ๊ฐ๋์๋ต๋ณ>์ด๋ช์ ์ ํด๋นํ๋์ง๊ทผ๊ฑฐ์ํจ๊ปํ๋จํฉ๋๋ค. 4.[์ข ํฉ๋ฐํ์ํ]:๋ชจ๋ ํ๋จ์ด๋๋๋ฉด,๊ทธ๋ด์ฉ์์ข ํฉํ์ฌ์ต์ข ์ถ๋ ฅJSONํ์์์์ฑํฉ๋๋ค. ###ํ๊ฐ๊ธฐ์ค๋ฐ์ฒ๋(Evaluation Criteria and Scale) ๊ฐํ๊ฐ์ถ์๋ํด1์ (๋งค์ฐ๋ถ์กฑ)๋ถํฐ5์ (๋งค์ฐ์)๊น์ง์ ์์ ์๋ฅผ๋ถ์ฌํฉ๋๋ค. *1)์ฌ์ค์ ์ ํ์ฑ(Factual Accuracy)[1-5์ ] *์ถ๋ก ๋ด์ฉ์ด๊ธฐ์ํ์ ์ฌ์ค์์๋ฒฝํ๊ฒ๋ถํฉํ๋ฉฐ์ค๋ฅ๊ฐ์๋๊ฐ? *5์ :๋ชจ๋ ๋ด์ฉ์ด<์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ>์๊ธฐ๋ฐํ์ฌ์๋ฒฝํ๊ฒ์ ํํจ. *3์ :ํต์ฌ๋ ผ๋ฆฌ๋๋ง์ง๋ง,๊ฒฐ๋ก ์์ํฅ์๋ฏธ์น์ง์๋์ฌ์ํ์ค๋ฅ๋๋ถ์ ํํํํ์ดํฌํจ๋จ. *1์ :๊ฒฐ๋ก ์์ ๋น์ฑ์ํผ์ํ๋์ค๋ํ์ฌ์ค์ค๋ฅ(ํ๊ฐ)๊ฐํฌํจ๋จ. *2)๋ ผ๋ฆฌ์ ์๊ฒฐ์ฑ(Logical Soundness)[1-5์ ] *์ ์๋๊ทผ๊ฑฐ๊ฐ๊ฒฐ๋ก (์ ๋ต)์๋์ถํ๊ธฐ์๋ ผ๋ฆฌ์ ์ผ๋ก์ถฉ๋ถํ๊ณ ํ์ฐ์ ์ธ๊ฐ? *5์ :๊ทผ๊ฑฐ์๊ฒฐ๋ก ์ฌ์ด์๋ ผ๋ฆฌ์ ๋น์ฝ์ด๋๋๋ฝ์ด์ ํ์์ด์๋ฒฝํ๊ฒ์ฐ๊ฒฐ๋จ. *3์ :ํต์ฌ์ ์ธ์ฐ๊ฒฐ๊ณ ๋ฆฌ๋์กด์ฌํ๋,์ผ๋ถ์ค๋ช ์ด์๋ต๋๊ฑฐ๋์ถ๋ก ๊ณผ์ ์ด๋ค์๋ถ์น์ ํจ. *1์ :๊ทผ๊ฑฐ๊ฐ๊ฒฐ๋ก ์๋ท๋ฐ์นจํ์ง๋ชปํ๊ฑฐ๋,๋ช ๋ฐฑํ๋ ผ๋ฆฌ์ ์ค๋ฅ๊ฐ์กด์ฌํจ. *3)์ถ๋ก ์๊น์ด(Depth of Reasoning)[1-5์ ] *๋ฌธ์ ์ํต์ฌ์๋ฆฌ๋ฅผ๊น์ด์๊ฒ์ดํดํ๊ณ ๋ค๊ฐ์ ์ผ๋ก๋ถ์ํ์๋๊ฐ? *5์ :์ ๋ต์๊ทผ๊ฑฐ๋ฟ๋ง์๋๋ผ,๋งค๋ ฅ์ ์ธ์ค๋ต์ ํ์ง๊ฐ์์ค๋ต์ธ์ง์๋ํ๋ฐ๋ฐ๊น์งํฌํจํ์ฌ๊น์ด์๋์ดํด๋ฅผ ๋ณด์ฌ์ค. *3์ :์ ๋ต์๋ํํต์ฌ์ ์ธ์ค๋ช ์์ ์ํ์ง๋ง,์ค๋ต์๋ํ๊ณ ๋ ค๋ฑ๋ค๊ฐ์ ์ธ๋ถ์์๋ถ์กฑํจ. *1์ :ํผ์์ ์ธ์์ค์๋จํธ์ ์ธ์ฌ์ค๋ง๋์ดํ์ฌ๊น์ด๊ฐ์์. *4)ํํ์๋ช ํ์ฑ(Clarity & Conciseness)[1-5์ ] *์ถ๋ก ์ค๋ช ์ด๊ตฐ๋๊ธฐ์์ด๋ช ํํ๊ณ ์ดํดํ๊ธฐ์ฌ์ด๊ฐ? *5์ :๋ถํ์ํ๋ด์ฉ์์ดํต์ฌ๋ง๊ฐ๊ฒฐํ๊ฒ์ ๋ฌํ๋ฉด์๋,์ ๋ฌธ๊ฐ์๋ํ์ต์๋์ฝ๊ฒ์ดํดํ ์์์๋งํผ ๋ช ํํ๊ฒ์์ ๋จ. *3์ :์๋ฏธ์ ๋ฌ์๋์ง๋ง,์ผ๋ถ๋ฌธ์ฅ์ด์ฅํฉํ๊ฑฐ๋๋ถํ์ํ์ ๋ณด๊ฐํฌํจ๋์ด์์. *1์ :๋ฌธ์ฅ์ด๋ณต์กํ๊ณ ์ดํดํ๊ธฐ์ด๋ ต๊ฑฐ๋,์ง๋ฌธ์๋์๊ด๋ จ์๋๋ด์ฉ์ด๋ง์. ###์ถ๋ ฅํ์(Output Format) *๋ฐ๋์์๋JSON๊ตฌ์กฐ๋ฅผ์ ํํ์ค์ํ์ฌ์๋ตํด์ผํฉ๋๋ค. * JSON์ธ์๋ถ๊ฐ์ค๋ช ์ด๋ํ ์คํธ๋์ถ๊ฐํ์ง๋ง์ธ์. โjson โ์ฌ์ค_์ค๋ฅ_๋ชฉ๋กโ:[ โ<ํ๊ฐ๋์๋ต๋ณ>์์๋ฐ๊ฒฌ๋์ฒซ๋ฒ์งธ์ฌ์ค์ค๋ฅ๋๋ํ๊ฐ(๋ฌธ์ฅ์ธ์ฉ)โ, โ๋๋ฒ์งธ์ฌ์ค์ค๋ฅ...(์์ผ๋ฉด๋น๋ฆฌ์คํธ[]๋ก์ถ๋ ฅ)โ ], โํ๊ฐ_์ ์โ: โ์ฌ์ค์ _์ ํ์ฑโ:1~5์ฌ์ด์์ ์์ ์, โ๋ ผ๋ฆฌ์ _์๊ฒฐ์ฑโ:1~5์ฌ์ด์์ ์์ ์, โ์ถ๋ก ์_๊น์ดโ:1~5์ฌ์ด์์ ์์ ์, โํํ์_๋ช ํ์ฑโ:1~5์ฌ์ด์์ ์์ ์ , โํ๊ฐ_์ฌ์ โ:์ ์๋ฅผโ๋ถ์ฌํํต์ฌ์ ์ธ์ด์ ๋ฅผ์ข ํฉ์ ์ผ๋ก์์ .์ ํ์ฑ์ ์๊ฐ๋ฎ์ผ๋ฉด์ฌ์ค์ค๋ฅ๋ฅผ๋ฐ๋์์ธ๊ธํ ๊ฒ.โ โ Figure 20:System prompt used by the evaluator LLM(Gemini-2.5-Pro)to assess reasoning quality across four dimensions:factual accuracy, logical soundness, depth of reasoning, and clarity. English Translation of the System Prompt for LLM-as-a-Judge Evaluation(Gemini-2.5-Pro) ### Role You are a highly strict and impartial Evaluation Chair for AI Reasoning Capabilities, possessing deep expertise in the field of meteorology. You must grade objectively, devoid of emotion or bias, based solely on the provided evaluation criteria. ### Task Based on the provided<Problem Information>,<Expert Verification Reference Material>,and<Response for Evaluation>, assess the quality of the answer submitted by theโExaminee AIโand output the results in the specified JSON format. ### Input Information 1.<Problem Information>: The original multiple-choice question serving as the context for evaluation(includes the question, options, and correct answer). 2.<Expert Verification Reference Material>: Model explanation material with 100% verified facts. This material must be used as the absolute standard for detectingโfactual errorsโorโhallucinationsโin the<Response for Evaluation>. 3.<Response for Evaluation>: The reasoning process and answer generated by theโExaminee AIโthat you are required to grade. ### Step-by-Step Thinking Process You must go through the following thinking process before performing the evaluation. 1. Fully understand all input information. 2.[Fact Check]: Compare every claim in the<Response for Evaluation>against the<Expert Verification Reference Material >to first identify any factual errors and compile a list. 3.[Individual Axis Evaluation]: Read the fourโEvaluation Criteria and Scaleโbelow one by one, and determine the score for the<Response for Evaluation>based on each criterion, along with the rationale. 4.[Synthesis and Formatting]: Once all judgments are complete, synthesize the content to create the final output JSON format. ### Evaluation Criteria and Scale Assign an integer scorefrom 1(Very Poor)to 5(Excellent)for each evaluation axis according to the criteria below. *1)Factual Accuracy[1-5 points] * Does the reasoning perfectly align with meteorological facts without errors? *5 points:All content is perfectly accurate based on the<Expert Verification Reference Material>. *3 points:The core logic is correct, but minor errors or inaccurate expressions that do not affect the conclusion are included. *1 point:Major factual errors(hallucinations)that undermine the legitimacy of the conclusion are included. *2)Logical Soundness[1-5 points] * Is the provided evidence logically sufficient and inevitable for deriving the conclusion(answer)? *5 points:Perfectly connected without any logical leaps or omissions between evidence and conclusion. *3 points:The core link exists, but some explanations are omitted, or the reasoning process is somewhat unfriendly. *1 point:The evidence fails to support the conclusion, or obvious logical errors exist. *3)Depth of Reasoning[1-5 points] * Did it understand the core principles of the problem deeply and analyze it from multiple angles? *5 points:Demonstrates deep understanding by including not only the basis for the correct answer but also rebuttals as to why attractive distractors(wrong options)are incorrect. *3 points:Provided the core explanation for the correct answer, but lacks multi-faceted analysis such as consideration of wrong options. *1 point:Lacks depth by merely listing superficial and fragmentary facts. *4)Clarity & Conciseness[1-5 points] * Is the reasoning explanation clear and easy to understand without clutter? *5 points:Conveys only the core points concisely without unnecessary content, stated clearly enough for a non-expert learner to understand easily. *3 points:Meaning is conveyed, but some sentences are verbose or include unnecessary information. *1 point:Sentences are complex and difficult to understand, or there is a lot of content irrelevant to the intent of the question. ### Output Format - You must strictly adhere to the JSON structure below. - Do not add any additional explanations or text outside the JSON. โjson โList _of_Factual_Errorsโ:[ โFirst factual error or hallucination found in<Response for Evaluation> (quote the sentence)โ, โSecond factual error...(Output an empty list[]if none)โ ], โEvaluation_Scoresโ: โFactual_Accuracyโ:Integer score between 1โ5, โLogical_Soundnessโ:Integer score between 1โ5, โDepth_of_Reasoningโ:Integer score between 1โ5, โClarity_and_Concisenessโ:Integer score between 1โ5 , โEvaluation_Rationaleโ:โComprehensively describe the core reasons for assigning the scores above for each axis. If theโAccuracyโscore is low, you must explicitly mention what factual errors occurred.โ โ Figure 21:System prompt used by the evaluator LLM(Gemini-2.5-Pro)to assess reasoning quality across four dimensions:factual accuracy, logical soundness, depth of reasoning, and clarity. Standard System Prompt(Baseline) ๋น์ ์๊ธฐ์์ ๋ฌธ๊ฐ์ ๋๋ค. ๋ค์๊ฐ๊ด์๋ฌธํญ์๋ตํ์ธ์. ๋ณด๊ธฐ์ค์์๊ฐ์ฅ์๋ง์์ ํ์ง๋ฅผ๊ณ ๋ฅด๊ณ ,๋ค์๊ณผ๊ฐ์ํ์์JSON์ผ๋ก์ถ๋ ฅํ์ธ์: ```jsonโ์ ๋ตโ: โ์ ํ์ง๋ฒํธโ``` ๋ต๋ณ์ด์ธ์๋ค๋ฅธ์ด๋คํ ์คํธ๋์ถ๋ ฅํ์ง๋ง์ธ์. Figure 22:Standard baseline system prompt using simple zero-shot instructions without specific regional context English Translation of the Standard System Prompt(Baseline) You are a meteorology expert. Answer the following multiple-choice question. Choose the most appropriate option from the choices and output it in the following JSON format: ```jsonโAnswerโ:โOption Numberโ``` Do not output any text other than the answer. Figure 23:English translation of the standard baseline system prompt using simple zero-shot instructions without specific regional context Advanced System Prompt(Role + Context) ๋น์ ์ํ๊ตญ์๊ธฐ์์ ๋ฌธ๊ฐ์ ๋๋ค. ๋ค์๊ฐ๊ด์๋ฌธํญ์๋ตํ์ธ์. ์ง๋ฌธ์คโ์ฐ๋ฆฌ๋๋ผโ๋โํ๊ตญโ์์๋ฏธํฉ๋๋ค. ํ๊ตญ์์ง๋ฆฌ์ ํน์ฑ์๋ฐ์ํด์๋ต๋ณํด์ฃผ์ธ์. ๋ณด๊ธฐ์ค์์๊ฐ์ฅ์๋ง์์ ํ์ง๋ฅผ๊ณ ๋ฅด๊ณ ,๋ค์๊ณผ๊ฐ์ํ์์JSON์ผ๋ก์ถ๋ ฅํ์ธ์: ```jsonโ์ ๋ตโ: โ์ ํ์ง๋ฒํธโ``` ๋ต๋ณ์ด์ธ์๋ค๋ฅธ์ด๋คํ ์คํธ๋์ถ๋ ฅํ์ง๋ง์ธ์. Figure 24:Advanced prompt with added role definition and specific context regarding Korean geography English Translation of the Advanced System Prompt(Role + Context) You are a meteorology expert in Korea. Answer the following multiple-choice question. In the question,โour countryโrefers toโKoreaโ.Please answer reflecting the geographical characteristics of Korea. Choose the most appropriate option from the choices and output it in the following JSON format: ```jsonโAnswerโ:โOption Numberโ``` Do not output any text other than the answer. Figure 25:English translation of the advanced prompt with added role definition and specific context regard- ing Korean geography Text-Only/Multimodal MQQA User Prompt Template ์ง๋ฌธ: Question_text [Question_image] 1.Choice_1_text [Choice_1_image] 2.Choice_2_text [Choice_2_image] 3.Choice_3_text [Choice_3_image] 4.Choice_4_text [Choice_4_image] Figure 26:Text-Only/Multimodal MCQA user prompt template.The placeholders enclosed in square brackets (e.g.,[Question_image])denote optional fields that are populated only when the corresponding image exists in the dataset. This single template covers all four modality configurations(i.e., text-only, image-in-question, image-in- choices, and images-in-both). Reasoning MCQA User Prompt Template ======================================== ์๋์ ๋ณด๋ฅผ๋ฐํ์ผ๋กํ๊ฐ๋ฅผ์ํํ์์ค. โ BEGIN INPUT DATA โ ###<๋ฌธ์ ์ ๋ณด> **๋ฌธ์ ํ ์คํธ:** Question_text [**๋ฌธ์ ์ด๋ฏธ์ง:**Question_image] **์ ํ์ง:** 1.Choice_1_Text [**(์ ํ์ง1์ด๋ฏธ์ง):**Choice_1_image] 2.Choice_2_Text [**(์ ํ์ง2์ด๋ฏธ์ง):**Choice_2_image] 3.Choice_3_Text [**(์ ํ์ง3์ด๋ฏธ์ง):**Choice_3_image] 4.Choice_4_Text [**(์ ํ์ง4์ด๋ฏธ์ง):**Choice_4_image] **์ ๋ต:**Correct_Answer ###<์ ๋ฌธ๊ฐ๊ฒ์ฆ์ฐธ์กฐ์๋ฃ> Rationale_Content% populatesโ(์๋ฃ์์)โif wo_rationale is True ###<ํ๊ฐ๋์๋ต๋ณ> **์์ฑ๋์ถ๋ก ๊ทผ๊ฑฐ:**Generated_Reasoning **๋ต์:**Predicted_Answer โ END INPUT DATA โ Figure 27:Reasoning MCQA user prompt template.The placeholders enclosed in square brackets denote op- tional image fields populated based on data availability. Additionally, the rationale field defaults toโ(์๋ฃ์์)โ (No Data)when the expert rationale is withheld(w/o rationalesetting). Table 13:K-MetBench performance scores across all models and subsets.Models are sorted by accuracy. All accuracy metrics range from 0 to 100, while the reasoning score(Reas.)ranges from 4 to 20.Boldvalues indicate the highest scores in each column for proprietary and open-source models, respectively.(Acc.: Accuracy,K: Korean model,V: Vision language model,R: Reasoning model,Inst.: Instruct) TypeModel Flags Acc.Reas.Geo-Cult.ModalityGranularity(P1โP5) K V RKorean Text Multi P1 P2 P3 P4 P5 gemini-3-pro-preview(Thinking)V R93.7 18.0190.4 94.6 75.6 92.5 97.9 94.2 92.8 91.6 gpt-5.2(Thinking)VR87.817.3380.890.629.386.393.488.086.285.3 Proprietary gpt-5.2V77.6 17.3975.3 79.0 50.0 77.2 81.3 71.9 81.4 76.3 Multilingual Thinking Models Qwen3-VL-235B-A22B-ThinkingV R84.4 17.2272.686.248.881.5 88.6 87.2 83.2 82.0 Qwen3-VL-32B-ThinkingVR78.616.1760.379.951.274.385.278.878.776.3 command-a-reasoning-08-2025R77.8 14.1274.6 77.8- 73.4 85.2 73.8 78.8 78.5 gpt-oss-120bR77.316.1262.077.3-72.585.876.577.474.9 Qwen3-30B-A3B-Thinking-2507R76.7 15.7667.6 76.7- 75.5 82.1 75.6 74.9 75.9 Qwen3-VL-30B-A3B-ThinkingVR74.915.1668.576.345.170.577.474.176.676.0 Qwen3-14BR73.7 15.2560.6 73.7- 70.9 84.3 72.4 70.2 71.7 Qwen3-VL-8B-ThinkingVR71.710.3361.673.339.066.279.870.571.371.6 gpt-oss-20bR71.5 13.5560.6 71.5- 65.7 82.4 71.8 72.2 65.8 Qwen3-8BR70.113.3149.370.1-69.580.269.165.666.8 Qwen3-4B-Thinking-2507R67.8 13.2860.6 67.8- 63.5 80.8 64.1 66.7 65.1 Qwen3-VL-4B-ThinkingVR66.111.5854.867.047.660.180.162.764.165.0 Qwen3-32BR47.5 14.5728.2 47.5- 48.6 61.3 47.4 39.7 41.0 Qwen3-1.7BR46.87.3735.246.8-45.157.247.442.442.7 Qwen3-0.6BR32.2 4.6023.9 32.2- 30.2 40.9 32.1 32.0 25.7 Phi-4-mini-reasoningR12.64.029.912.6-14.310.710.912.714.3 Multilingual Instruct Models Qwen3-VL-235B-A22B-InstructV72.4 15.4074.0 73.8 45.1 72.9 78.6 64.3 74.5 72.2 Qwen3-VL-32B-InstructV67.514.8561.668.741.567.872.064.369.463.8 Qwen2.5-VL-72B-InstructV67.1 12.9463.0 68.4 41.5 64.3 70.8 62.7 69.7 68.6 Qwen3-30B-A3B-Instruct-250764.714.6960.664.7-65.171.457.665.663.8 c4ai-command-a-03-202565.5 12.8166.2 65.5- 62.9 66.4 57.9 71.9 68.7 Qwen3-VL-30B-A3B-InstructV62.213.3757.563.241.563.368.454.364.461.1 Qwen2.5-VL-32B-InstructV60.1 10.9956.2 61.1 39.0 60.3 59.9 56.8 62.8 60.5 Llama-3.1-70B-Instruct59.911.1657.759.9-59.361.653.265.859.3 InternVL3.5-38B-InstructV57.3 11.3847.9 58.1 40.2 56.0 64.8 48.7 61.4 55.7 Llama-3.2-90B-Vision-InstructV56.99.7252.158.230.557.159.352.462.253.3 Qwen3-VL-8B-InstructV53.8 12.0743.8 54.3 43.9 54.2 58.1 49.3 55.9 51.5 Qwen3-4B-Instruct-250751.512.3245.151.5-53.852.547.652.650.8 Phi-451.5 11.7540.8 51.5- 52.5 53.8 50.0 55.1 45.3 Qwen3-VL-4B-InstructV51.011.5546.651.148.850.756.346.053.548.5 InternVL3.5-14B-InstructV47.9 9.4545.2 48.4 37.8 44.5 53.6 44.0 50.3 47.6 InternVL3.5-8B-InstructV46.17.0735.646.732.945.348.842.152.441.6 Qwen2.5-VL-7B-InstructV46.1 7.0837.0 46.6 34.1 49.3 43.1 42.9 51.9 42.2 Llama-3.1-8B-Instruct41.87.6340.841.8-44.238.140.044.442.0 InternVL3.5-4B-InstructV41.5 4.8124.7 42.1 29.3 44.5 44.0 37.9 43.9 36.8 Qwen2.5-VL-3B-InstructV40.94.8837.041.430.541.638.640.143.940.1 Llama-3.2-3B-Instruct33.8 5.0831.0 33.8- 36.3 32.7 34.7 33.9 30.9 InternVL3.5-2B-InstructV31.04.3524.731.226.833.830.732.329.528.4 Phi-4-mini-Instruct30.4 5.8221.1 30.4- 31.6 33.0 30.3 29.5 27.7 InternVL3.5-1B-InstructV23.84.0628.824.313.426.023.824.223.421.6 Llama-3.2-1B-Instruct3.5 4.004.2 3.5- 3.0 5.3 2.6 3.9 2.9 Korean Thinking Models EXAONE-4.0-32BKR59.9 13.5759.2 59.9- 58.2 64.8 52.4 63.1 61.2 HyperCLOVAX-SEED-Think-14BKR50.811.2952.150.8-51.653.841.855.651.1 EXAONE-4.0-1.2BKR37.4 7.6042.3 37.4- 37.6 42.1 35.0 39.1 32.6 Korean Instruct Models A.X-4.0K76.1 15.4678.976.1- 76.6 77.7 68.2 81.3 76.5 VARCO-Vision-2.0-14BKV58.711.2457.559.542.759.062.354.361.756.0 A.X-4.0-LightK55.7 11.4560.6 55.7- 55.8 54.4 50.9 61.4 55.7 A.X-4.0-VL-LightKV52.59.7654.853.042.751.550.650.158.052.1 VARCO-Vision-2.0-1.7BK V35.2 5.7634.2 36.6 6.1 35.1 35.8 33.4 38.0 33.2 HyperCLOVAX-SEED-Vision-Inst.-3BKV32.07.5635.632.423.237.325.926.536.732.6 HyperCLOVAX-SEED-Text-Inst.-1.5BK30.6 6.8436.6 30.6- 38.5 31.4 24.7 30.6 27.0 Open-source HyperCLOVAX-SEED-Text-Inst.-0.5BK13.24.3114.113.2-17.08.510.613.516.0 05101520 Normalized Total GPU Compute Time for 100 Questions HyperCLOVAX-SEED-Text-Instruct-1.5B EXAONE-4.0-1.2B Qwen2.5-VL-3B-Instruct Qwen3-4B-Instruct-2507 HyperCLOVAX-SEED-Vision-Instruct-3B Qwen2.5-VL-7B-Instruct Qwen3-VL-30B-A3B-Instruct InternVL3_5-8B-Instruct Llama-3.2-3B-Instruct Qwen3-30B-A3B-Instruct-2507 Phi-4-mini-instruct Qwen3-VL-4B-Instruct InternVL3_5-1B-Instruct Qwen2.5-VL-32B-Instruct Qwen3-VL-8B-Instruct A.X-4.0-Light InternVL3_5-14B-Instruct Llama-3.1-70B-Instruct VARCO-VISION-2.0-14B HyperCLOVAX-SEED-Text-Instruct-0.5B EXAONE-4.0-32B InternVL3_5-4B-Instruct InternVL3_5-2B-Instruct Phi-4 A.X-4.0-VL-Light Llama-3.1-8B-Instruct Qwen3-VL-32B-Instruct InternVL3_5-38B-Instruct A.X-4.0 Qwen2.5-VL-72B-Instruct c4ai-command-a-03-2025 Qwen3-VL-235B-A22B-Instruct VARCO-VISION-2.0-1.7B Qwen3-0.6B Llama-3.2-1B-Instruct gpt-oss-120b HyperCLOVAX-SEED-Think-14B Qwen3-1.7B gpt-oss-20b Qwen3-8B Qwen3-30B-A3B-Thinking-2507 Qwen3-4B-Thinking-2507 Qwen3-14B Llama-3.2-90B-Vision-Instruct Qwen3-VL-30B-A3B-Thinking Qwen3-VL-4B-Thinking Qwen3-32B Phi-4-mini-reasoning Qwen3-VL-8B-Thinking Qwen3-VL-235B-A22B-Thinking Qwen3-VL-32B-Thinking command-a-reasoning-08-2025 Exam Time 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.00 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.01 GPU-hours / 100 Qs 0.02 GPU-hours / 100 Qs 0.02 GPU-hours / 100 Qs 0.03 GPU-hours / 100 Qs 0.03 GPU-hours / 100 Qs 0.03 GPU-hours / 100 Qs 0.04 GPU-hours / 100 Qs 0.04 GPU-hours / 100 Qs 0.05 GPU-hours / 100 Qs 0.08 GPU-hours / 100 Qs 0.10 GPU-hours / 100 Qs 0.11 GPU-hours / 100 Qs 0.11 GPU-hours / 100 Qs 0.14 GPU-hours / 100 Qs 0.26 GPU-hours / 100 Qs 0.30 GPU-hours / 100 Qs 0.37 GPU-hours / 100 Qs 0.39 GPU-hours / 100 Qs 0.46 GPU-hours / 100 Qs 0.67 GPU-hours / 100 Qs 0.86 GPU-hours / 100 Qs 1.0 GPU-hours / 100 Qs 1.5 GPU-hours / 100 Qs 2.4 GPU-hours / 100 Qs 3.1 GPU-hours / 100 Qs 3.8 GPU-hours / 100 Qs 20.8 GPU-hours / 100 Qs GPU-hours Exam Time GPU-hours > Exam Time Exam time (150 minutes 2.50 GPU-hours) Non-multimodal Multimodal Figure 28:Normalized total GPU compute time for 100 questions compared across models.The plot displays the total GPU time required to complete 100 questions, calculated as(wall-clock inference timeรtensor paral- lelism sizeร100รทquestion count).The red dashed line indicates the official time limit for the Meteorological Engineer exam(150 minutes).Green barsdenote models that completed the task within the time limit, whilered barsindicate models that exceeded it.Solid barsrepresent multimodal models, andhatched barsrepresent non- multimodal(text-only)models. All evaluations were conducted using the vLLM library on NVIDIA H100 80GB PCIe GPUs. Figure 29:Hierarchical taxonomy and sample distribution for Parts 1 and 2.The diagram visualizes the breakdown ofMeteorology & ThermodynamicsandObservation Methodsinto detailed sub-topics. Numbers in parentheses indicate the estimated quantity of questions classified by Gemini-2.5-Pro. Figure 30:Hierarchical taxonomy and sample distribution for Part 3.This overview details the structure of Forecasting & Climatology, mapping the dataset samples to specific forecasting theories and climate phenomena. Figure 31:Hierarchical taxonomy and sample distribution for Parts 4 and 5.The diagram coversApplied MeteorologyandWeather Chart Analysis, illustrating the coverage of practical applications and legal regulations. 0.00.5 Accuracy VARCO-V.-2.0-1.7B HyperCLOVAX-SEED-V.-Instruct-3B InternVL3_5-4B-Instruct Llama-3.2-90B-V.-Instruct InternVL3_5-8B-Instruct InternVL3_5-14B-Instruct Qwen2.5-VL-32B-Instruct Qwen3-VL-32B-Instruct A.X-4.0-VL-Light Qwen3-VL-8B-Instruct Qwen3-VL-30B-A3B-Think. Qwen3-VL-4B-Instruct gpt-5.2_multimodal Qwen3-VL-32B-Think. gemini-3-pro-preview 0.00.51.0 Accuracy Llama-3.2-1B-Instruct InternVL3_5-2B-Instruct InternVL3_5-1B-Instruct VARCO-V.-2.0-1.7B Qwen2.5-VL-3B-Instruct EXAONE-4.0-1.2B Qwen3-VL-4B-Instruct Llama-3.2-90B-V.-Instruct Qwen3-VL-30B-A3B-Instruct Qwen3-30B-A3B-Instruct-2507 Qwen3-VL-32B-Think. gpt-oss-120b Qwen3-30B-A3B-Think.-2507 command-a-reasoning-08-2025 gemini-3-pro-preview 0.00.5 Normalized Score Llama-3.2-1B-Instruct InternVL3_5-2B-Instruct Llama-3.2-3B-Instruct InternVL3_5-8B-Instruct HyperCLOVAX-SEED-V.-Instruct-3B Llama-3.2-90B-V.-Instruct Llama-3.1-70B-Instruct A.X-4.0-Light Qwen3-VL-8B-Instruct Qwen3-4B-Think.-2507 EXAONE-4.0-32B Qwen3-30B-A3B-Instruct-2507 Qwen3-VL-235B-A22B-Instruct Qwen3-VL-32B-Think. gemini-3-pro-preview Figure 32:Ranking stability analysis on theMultimodal,Geo-cultural, andReasoningsubset performances with bootstrap confidence intervals.Error bars denote 95% CIs derived from item-level bootstrap resampling on 10 evenly spaced models. Even for the smaller subsets, the uncertainty is explicitly quantified and the coarse separation between higher- and lower-performing systems remains visible across panels. H Robustness of Conclusions Under Small Subsets To address the concern that small subsets may induce high variance and unstable conclusions, we conducted item-level bootstrap resampling and sensitivity analyses using the evaluation logs. We confirm via rigorous sensitivity diagnos- ticsโincluding Bootstrap, leave-one-out(LOO), and Approximate Maximum Influence Perturba- tion(AMIP)( Huang et al.,2026)โthat the modal- ity, geo-cultural, and reasoning gaps are robust sys- temic trends. Bootstrap estimates validate these patterns, LOO perturbations reveal no sign flips, and AMIP analysis demonstrates that the local advantage withstands even substantial adversarial data removal. H.1 Statistical Robustness Diagnostics We report explicit uncertainty(confidence inter- vals)and stability diagnostics, including LOO per- turbations and bootstrap rank intervals. Unless otherwise noted, all statistics below are computed from the current evaluation run with 1,000 boot- strap iterations and a fixed random seed of 42. We confirmed that the bootstrapping process achieved sufficient convergence(not shown). Setup.We focus our robustness checks on three subsets: multimodal(82 items),Korean-specific (71 text-only and 73 multimodal items),and rea- soning(121 text-only and 141 multimodal items). We normalize all scores to[0,1], mapping reason- ing scores from [4 , 20] and accuracy from [0 , 100] . For the bootstrap analysis, we select 10 represen- tative models for each subset by sorting them by performance and sampling at equal intervals. Subset-level uncertainty is quantified, but does not erase structure.Figure32reports 95% bootstrap confidence intervals for 10 rep- resentative models selected at equal intervals from the performance ranking for each subset. This sampling strategy ensures visibility across the full performance spectrum. Making the uncertainty induced by smallnexplicit, we observe that while confidence intervals nat- urally widen for smaller subsets, the overall performance hierarchy remains robust. The highest-performing systems consistently maintain their lead with the following estimates(mean [95% CI]):Multimodal0.756 [0.659,0.841], Geo-cultural0.789 [0.690,0.873], and Reasoning 0.876 [0.848,0.901](normalized score).This con- firms that even under resampling, the conclusions are not dominated by random sampling variation, and the uncertainty bounds provide a principled way to interpret rank differences. Key performance gaps are stable under re- sampling and single-item perturbations.As shown in Table 14, our analysis confirms that the identified performance gaps across modality, geo- cultural, and reasoning dimensions are systemic and robust, rather than artifacts of specific outliers. First, regarding the modality gap(โ Modality = Acc Multimodal โAcc Non-Multimodal ), we observe a consistently negative trend across the 25 directly comparable models, ranging fromโ37.39% to โ2.28%. For 19/25 models, the 95% bootstrap CI strictly excludes zero. Second, we examine the geo-cultural gap(โ Geo-Cultural =Acc Korean โ Acc Non-Korean )and the reasoning gap(โ Reasoning = Score Reasoning โAcc Total )(normalized score and accuracy).We find that representative local mod- Table 14:Leave-one-out(LOO)sensitivity analysis acrossMultimodal,Geo-cultural, andReasoningsubsets. The table reports representative models with baseline gaps closest to zero(most prone to sign flips).No model exhibits a sign reversal under single-item removal across all tasks.(Baselineโ:original gap on the full set;Max swing: maximum deviation from the baseline across LOO iterations;Sign flip: percentage of iterations where the gap sign reverses;n LOO : total number of items subject to LOO perturbation.) ModelBaselineโMax swingSign flip(%)n LOO (A)Modality Gap(MultimodalโText-only) Qwen/Qwen3-VL-4B-Instructโ2.28%0.63%0.082 OpenGVLab/InternVL3_5-2B-Instructโ4.38%0.90%0.082 HyperCLOVAX-SEED-Vision-Instruct-3Bโ9.22%0.95%0.082 skt/A.X-4.0-VL-Lightโ10.33%0.71%0.082 Qwen/Qwen3-VL-8B-Instructโ10.35%0.69%0.082 (B)Geo-cultural Gap(KoreanโNon-Korean) HyperCLOVAX-SEED-Text-Instruct-1.5B6.27%0.91%0.071 OpenGVLab/InternVL3_5-1B-Instruct5.13%0.99%0.073 LGAI-EXAONE/EXAONE-4.0-1.2B5.12%0.82%0.071 skt/A.X-4.0-Light5.04%0.87%0.071 HyperCLOVAX-SEED-Vision-Instruct-3B3.81%0.89%0.073 (C)Reasoning Gap(ReasoningโKnowledge) Qwen/Qwen3-32B(Thinking)18.26%0.55%0.0121 gpt-5.2(Thinking)6.27%0.61%0.0141 Qwen/Qwen3-30B-A3B-Instruct-25072.50%0.56%0.0121 Qwen/Qwen3-4B-Instruct-25070.65%0.43%0.0121 Qwen/Qwen3-VL-32B-Instruct0.51%0.72%0.0141 els(e.g.,HyperCLOVAX,skt/A.X)maintain a sta- ble positive advantage in the geo-cultural domain, while the reasoning performance remains distinct from general knowledge across models. To quan- tify the sensitivity of these conclusions to indi- vidual outliers, we run leave-one-out perturbations over the respective subsets: multimodal(n= 82), geo-cultural(n= 71/73),and reasoning(n= 121/141)for text-only and multimodal models, re- spectively. Crucially, removing any single item never trig- gers a sign flip across all evaluated models and di- mensions(sign-flip rate = 0).Table 14reports the most fragile cases(i.e., baselines closest to zero); even for these edge cases, the maximum swing remains negligible(e.g.,<1.16% for modality, <0.99% for geo-cultural, and<0.72% for reason- ing).This confirms that the observed gaps drive the overarching trends and are not attributable to a handful of influential questions. H.2 Robustness of Key Findings to Critical Data Perturbation Resilience to Adversarial Item Removal. To further assess whether the geo-cultural gap could be driven by a small number of influential questions, we conduct an adversarial influential- item deletion analysis inspired by recent robust evaluation methodologies, AMIP(Huang et al., 2026). Concretely, we focus on the representa- tive model pair highlighted in Section 5(The Geo- Cultural Gap):the top-performing local model (skt/A.X-4.0)versus the global baseline(Qwen/ Qwen3-VL-235B-A22B-Thinking).We test how manystrategically selecteditems an adversary would need to delete to reverse the ordering where the local model outperforms the global one. We fit a Bradley-Terry(BT)model to pairwise win/loss outcomes on the Geo-cultural subset and compute per-item influence scores using the in- verse Hessian of the BT loss. Using a greedy ad- versarial removal procedure(a standard approxima- tion for AMIP)that sequentially deletes the most influential items favoring the local model, we find that the ordering flips only after removingn AMIP = 18items, corresponding to18/73โ24.7% of the entire Geo-cultural subset. This indicates that the observed local model advantage is not attributable to a single outlier or a very small number of ques- tions, but instead requires removing a substantial fraction of the subset to overturn.