Paper deep dive
Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination
Isak Hwang, Yoon Pyo Lee
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/2/2026, 12:57:01 PM
Summary
This study benchmarks a 31-billion-parameter multimodal language model (Gemma 4 31B-IT) on the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator Generic Fundamentals Examination to evaluate its capability in nuclear engineering. The research compares eight configurations involving supervised fine-tuning (SFT), retrieval-augmented generation (RAG), and retrieval-augmented fine-tuning (RAFT), alongside two chunking strategies (fixed-size and structure-aware). Results indicate that SFT combined with fixed-size chunking RAG is the most effective strategy, passing 8 of 14 examinations against an 80% human passing criterion, while RAFT underperformed compared to standard SFT.
Entities (11)
Relation Signals (6)
Gemma-4-31B-it → isevaluatedon → NRC Reactor Operator Licensing Examination
confidence 98% · This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on its capacity to apply nuclear knowledge by benchmarking... against the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator licensing examination.
SFT with Fixed-size Chunking RAG → outperforms → all alternatives
confidence 95% · The SFT configuration with fixed-size chunking RAG met the criterion on 8 of the 14 examinations, outperforming all alternatives
RAFT → underperforms → SFT
confidence 92% · RAFT underperforms compared to standard SFT in matching search environments.
Gemini 3 Flash → generates → Chain-of-Thought Rationales
confidence 90% · The stem, the four options and the ground-truth answer are submitted to Gemini 3 Flash... The teacher is instructed to explain step by step... The rationale is stored...
DOE Fundamentals Handbook → issourcefor → Retrieval-Augmented Generation
confidence 90% · retrieval-augmented generation (RAG) with BM25 sparse retrieval over the U.S. Department of Energy Fundamentals Handbook
BM25 → isusedfor → Retrieval-Augmented Generation
confidence 90% · retrieval-augmented generation (RAG) with BM25 sparse retrieval
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on its capacity to apply nuclear knowledge by benchmarking eight model-retrieval configurations against the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator licensing examination. We evaluate 14 Generic Fundamentals Examinations (GFE) from the 2015-2021 March sittings (seven pressurized and seven boiling water reactor exams) using the standard 80% human passing criterion. The base model is compared against configurations utilizing supervised fine-tuning (SFT) on Gemini-distilled chain-of-thought (CoT) rationales, retrieval-augmented generation (RAG) with BM25 sparse retrieval over the U.S. Department of Energy Fundamentals Handbook, and retrieval-augmented fine-tuning (RAFT). Within the retrieval pipeline, we compare fixed-size sliding-window chunking against structure-aware chunking. The SFT configuration with fixed-size chunking RAG met the criterion on 8 of the 14 examinations, outperforming all alternatives, whereas no configuration without fine-tuning passed any. Aggregate accuracy reached 79.7%, with a confidence interval spanning the threshold, and 80.2% on PWR items specifically. Furthermore, two regularities emerged: the preferred chunking strategy reverses depending on the model's training state, and RAFT underperforms compared to standard SFT in matching search environments. These results demonstrate which combination of fine-tuning and search approaches achieves operator-level capabilities.
Tags
Links
- Source: https://arxiv.org/abs/2607.22067v1
- Canonical: https://arxiv.org/abs/2607.22067v1
Trouble viewing inline? Open PDF directly →
Full Text
76,333 characters extracted from source content.
Expand or collapse full text
Highlights Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination Isak Hwang, Yoon Pyo Lee • A 31B open model passes 8 of 14 NRC reactor operator licensing examinations. • Eight model-retrieval configurations benchmarked on 700 NRC PWR and BWR items. • Two chunking methods are compared, structure-aware and fixed-size. • RAFT underperforms plain SFT under matched chunking and retrieval conditions. arXiv:2607.22067v1 [cs.CL] 24 Jul 2026 Benchmarking Fine-tuning and Retrieval Strategies for a Multimodal Language Model on the NRC Reactor Operator Licensing Examination Isak Hwang a , Yoon Pyo Lee b,∗ a Department of Nuclear Engineering, Hanyang University, 222 Wangsimni-ro, 04763, Seongdong-gu, Seoul, South Korea b The Grainger College of Engineering, Nuclear, Plasma & Radiological Engineering, University of Illinois Urbana-Champaign, Urbana, IL, USA A R T I C L E I N F O Keywords: Nuclear Engineering Reactor Operator Licensing Large Language Models Retrieval-Augmented Generation Supervised Fine-Tuning Multimodal Models A B S T R A C T The integration of large language models (LLMs) into the nuclear power industry requires outputs grounded in domain-specific knowledge. This study evaluates a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) on its capacity to apply nuclear knowledge by benchmarking eight model-retrieval configurations against the U.S. Nuclear Regulatory Commission (NRC) Reactor Operator licensing examination. We evaluate 14 Generic Fundamentals Examinations (GFE) from the 2015–2021 March sittings (seven pressurized and seven boiling water reactor exams) using the stan- dard 80% human passing criterion. The base model is compared against configurations utilizing super- vised fine-tuning (SFT) on Gemini-distilled chain-of-thought (CoT) rationales, retrieval-augmented generation (RAG) with BM25 sparse retrieval over the U.S. Department of Energy Fundamentals Handbook, and retrieval-augmented fine-tuning (RAFT). Within the retrieval pipeline, we compare fixed-size sliding-window chunking against structure-aware chunking. The SFT configuration with fixed-size chunking RAG met the criterion on 8 of the 14 examinations, outperforming all alternatives, whereas no configuration without fine-tuning passed any. Aggregate accuracy reached 79.7%, with a confidence interval spanning the threshold, and 80.2% on PWR items specifically. Furthermore, two regularities emerged: the preferred chunking strategy reverses depending on the model’s training state, and RAFT underperforms compared to standard SFT in matching search environments. These results demonstrate which combination of fine-tuning and search approaches achieves operator-level capabilities. 1. Introduction The safe operation of nuclear power plants depends on highly trained human operators who must apply complex nuclear engineering knowledge and procedural reasoning under safety-critical conditions and under time pressure [1, 2]. In the United States (U.S.), the Nuclear Regulatory Commission (NRC) certifies these operators through the Reactor Operator (RO) license examination [3], a rigorous assessment spanning reactor physics, plant systems, ther- modynamics, instrumentation and control, and emergency operating procedures. The qualification criteria are defined in the NRC operator licensing exam standards. To obtain a license, a candidate must receive a total score of at least 80% on the written exam, and rounding the score up to meet this passing mark is not permitted [4]. The generic engineering component of that written examination is the Generic Fundamentals Examination (GFE), which is the instrument used throughout this study. Because this 80% threshold is the same standard applied to every human RO candidate, the examination lets us assess a model against a human on identical terms. In recent years, large language models (LLMs) have advanced rapidly in language understanding, multi-step rea- soning, and information retrieval. A growing body of work uses standardized professional licensing examinations, most prominently in medicine and law. Singhal et al. [5] showed ∗ Corresponding author yoonpyo2@illinois.edu (Y.P. Lee) ORCID(s): 0009-0004-4044-4447 (Y.P. Lee) that Med-PaLM 2 could approach expert clinician perfor- mance on questions styled after the United States Medi- cal Licensing Examination (USMLE), and Nori et al. [6] demonstrated that GPT-4 attained passing scores across all three USMLE steps without any medical-specific fine- tuning. In the legal domain, Katz et al. [7] reported that GPT- 4 passed the Uniform Bar Examination at approximately the 90th percentile among human test-takers. Because these examinations are verified, professionally administered, and graded against fixed passing criteria, strong performance on them is meaningful evidence of domain expertise rather than superficial pattern matching, which motivates our use of the NRC RO examination as a benchmark. Despite this progress, adoption of LLMs within the nuclear industry has remained limited. Existing applications of natural language processing in nuclear engineering have concentrated on document classification, anomaly detection, predictive maintenance, and natural-language interfaces to plant data [8–13], rather than on demonstrating engineering knowledge competence on the standardized licensure assess- ments that define operator qualification. The main obstacle is the well-documented tendency of LLMs to produce fluent but factually incorrect outputs, commonly called halluci- nations, a failure mode difficult to reconcile with reactor operations, where a single incorrect procedural recommen- dation can carry severe safety consequences. Closing this gap requires capable base models together with fine-tuning strategies that ground generation in authoritative, domain- specific sources. I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 1 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams We therefore ask whether an LLM can reach a reliable standard of competence on the NRC RO license exami- nation. To address the question we evaluate a candidate model on the exam, the same instrument applied to human candidates, and hold it to the same regulatory 80% passing standard. Performance at or above this threshold would be concrete evidence that a model possesses operationally rele- vant knowledge of reactor physics and regulatory procedure, supporting its potential deployment as a digital co-pilot that augments, rather than replaces, the judgment of a licensed human operator. Two complementary techniques are commonly used to ground general-purpose LLMs in specialized domains. The first, supervised fine-tuning (SFT) [14], updates model pa- rameters on domain-specific knowledge. Its effectiveness increases substantially when the training targets are step- by-step chain-of-thought (CoT) rationales distilled from a stronger teacher model, a strategy that transfers reasoning ability to smaller students far more efficiently than label-only supervision [15]. The second, retrieval-augmented genera- tion (RAG) [16], leaves parameters untouched and instead injects authoritative external context at inference time, aug- menting the model’s parametric knowledge with retrieved evidence to suppress hallucination on knowledge-intensive queries. Retrieval-augmented fine-tuning (RAFT) combines the two [17], training the model on retrieved context so that it learns to attend to relevant passages while remaining robust to imperfect or distracting retrieval. A critical design parameter in all RAG-based methodologies is chunking, which dictates how source documents are segmented into discrete, retrievable units. Chunk granularity is known to af- fect retrieval quality [18], and its interaction with the model’s training state has received little attention, particularly when retrieval is paired with fine-tuning. A further complication specific to technical and reg- ulatory domains is that the source material is intrinsi- cally multimodal. NRC examinations and the U.S. De- partment of Energy (DOE) Fundamentals Handbooks that support them contain many equations [19–21], piping-and- instrumentation diagrams (P&IDs), system schematics, and annotated graphs that are integral to the questions they ac- company. Answering such items therefore requires a model that reasons over both text and images, not text alone. This motivates our use of an open-weight multimodal model. The instruction-tuned Gemma 4 family [22] provides a publicly available 31-billion-parameter variant with native image understanding, well suited to reproducible domain- adaptation studies on image-bearing regulatory content. Building on these observations, this study presents a sys- tematic benchmark of a 31-billion-parameter open-weight multimodal model (Gemma 4 31B-IT) across eight model– retrieval configurations on the NRC Reactor Operator Generic Fundamentals Examination, covering both reactor types. The contributions of this work are as follows. We introduce a reproducible evaluation pipeline that adapts a 31B mul- timodal model to nuclear regulatory question answering, including image-bearing items such as P&IDs and system schematics, through Gemini-distilled CoT training data, merged low-rank adaptation (LoRA) adapters, and an MLX- based inference engine. We empirically determine which combination of adaptation and retrieval strategies is suffi- cient to meet the regulatory 80% NRC passing threshold, finding that SFT combined with fixed-size chunking RAG satisfies the criterion on 8 of the 14 administered exam- inations, more than any of the other seven configurations evaluated, and that no configuration without fine-tuning sat- isfies it on any examination. We report a chunking-strategy reversal phenomenon in which structure-aware chunking is preferred for the base model whereas fixed-size sliding- window chunking is preferred for the fine-tuned model, which indicates that effective retrieval design is contingent on the underlying model’s training state. Finally, we show that the more elaborate RAFT paradigm underperforms plain SFT under matched chunking, which indicates that, for this task, the gains attributable to CoT distillation outweigh those attributable to retrieval-conditioned training. 2. Methodology This section describes the evaluation framework, the dataset and the computational pipeline. Figure 1 summarizes the workflow. 2.1. Problem Formulation and Evaluation Standard The task is a closed-ended, four-choice question-answering problem drawn from the GFE. Let 푋 denote an input query, consisting of a technical scenario, the options 퐴,퐵,퐶,퐷, and any accompanying images, and let푌 ∈ 퐴,퐵,퐶,퐷 de- note the ground-truth answer. The objective is to maximize 푃 (푌 ∣ 푋) for unaugmented configurations and 푃 (푌 ∣ 푋,퐶) for retrieval-augmented configurations, where 퐶 denotes external technical context retrieved from a reference corpus. The passing criterion is fixed by regulation rather than chosen by the authors. A candidate must score at least 80%, and rounding up to reach that mark is not permitted. This criterion applies to a single administered examination and not to an average across sittings. We therefore adopt the individual administered examination, rather than a pooled item set, as the primary unit of evaluation. Each examination consists of 50 items, except for the two 2017 papers, which contain 49 items owing to the exclusion of an unrecoverable figure. Scoring is strictly binary. The model receives 1 point for a correct answer and 0 for an incorrect one. We mirror the actual examination convention here, under which a candidate who leaves an item unan- swered receives no credit for it. For a given examination 푒, the pass indicator is Pass(푒) = ⎧ ⎪ ⎨ ⎪ ⎩ 1 if Correct answers in 푒 Total items in 푒 ≥ 0.80 0 otherwise (1) I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 2 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams NRC question banks PWR 2,140 BWR 2,149 Administered papers March 2015 to 2021 14 papers DOE handbooks 7 volumes Training Evaluation Deduplication by normalized question text CoT distillation Gemini 3 Flash teacher LoRA fine-tuning SFT and RAFT Chunking fixed-size and structure-aware BM25 index top-4, 7,000 chars Model inference 8 configurations greedy decoding Answer extraction cascaded matcher Pass count and pooled accuracy 80% criterion text match 700 items 3,577 items fine-tuned model retrieved context Figure 1: Pipeline of the proposed framework. The question banks and the administered papers are obtained independently from the NRC and the papers overlap the banks. Deduplication removes from the pooled banks every item whose normalized question text matches an evaluation item, leaving 3,577 training questions that then receive distilled CoT rationales and are used for low-rank fine-tuning. The DOE handbook corpus is segmented under two chunking policies and indexed for BM25 retrieval. Eight configurations are evaluated on the fourteen administered papers, each scored independently against the 80% criterion. so that Equation (1) evaluates to 1 only when at least 40 of 50 items are answered correctly. The requirement is 40 of 49 for the two 2017 papers. Our primary metric is the number of examinations passed out of the 14 administered papers, reported in Table 4. As a secondary metric we report pooled accuracy, the percentage of correct answers across all items from all examinations combined. Pooled accuracy offers a broader statistical summary but has no direct regulatory interpreta- tion, since candidates are never assessed on a pooled item set drawn from multiple papers. The scope of the instrument is limited in one further respect. The GFE section does not include questions on site- specific regulations, technical specifications, or emergency operating procedures. Because full operator licensing also requires a separate site-specific written examination and an operating test, the claims in this paper are strictly limited to the generic knowledge component of licensure. 2.2. Evaluation Dataset and Training Corpus 2.2.1. Data Sources Three publicly available NRC and DOE sources form the experimental corpus. The first is the pair of NRC GFE question banks for pres- surized water reactors [23] and boiling water reactors [24], the authoritative pools from which examination questions are drawn. Both are used to construct the training corpus. The PWR question bank contains 2,140 questions and the corresponding BWR question bank contains 2,149 ques- tions. Each bank is organized across three subject areas. These are Components under Topic 191x, Reactor The- ory under Topic 192x and Thermodynamics under Topic 193x. Each item carries a unique question identifier and a knowledge tag encoding the competency tested. A subset of I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 3 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams items includes engineering diagrams, schematics or piping and instrumentation figures as integral parts of the question stem. These are extracted as native images and processed through the multimodal interface of the model. The second is the set of administered examinations that forms the evaluation set, comprising every GFE adminis- tered at the March sitting from 2015 through 2021 for both reactor types [25–38]. Each examination consists of 50 four- choice items. This yields 350 PWR items and 350 BWR items across fourteen administrations. The March sitting is selected because it is the only sitting available in every year of the study period. The study period terminates in 2021 because the last GFE was administered in September of that year, the NRC having discontinued the program on 16 March 2022 [39]. The evaluation set is therefore a census of the March administrations of a closed program rather than a sample of an ongoing series. The third is the external retrieval corpus, built from seven volumes of the DOE Fundamentals Handbook series, which the NRC recommends as preparatory reading. They comprise Thermodynamics, Heat Transfer, and Fluid Flow in volumes 1 to 3, Instrumentation and Control in volumes 1 and 2 and Nuclear Physics and Reactor Theory in volumes 1 and 2. The handbooks are agnostic to reactor type. They cover the subject areas tested for both PWR and BWR examinations. We exploit this property in Section 3.3. 2.2.2. Train and Evaluation Split A deduplication procedure prevents overlap between the training corpus and the evaluation set. The two question banks are pooled into a single set of 4,289 items. Each pooled item and each evaluation item is reduced to a normalized signature by deleting all non-alphanumeric characters and lower-casing the remainder. Every pooled item whose signa- ture matches that of an evaluation item is removed. Matching is therefore on question text rather than on question identi- fier. The procedure removed 711 items, corresponding to 472 distinct evaluation questions, and 3,578 items were retained. Table 1 summarizes the composition. The number of items removed exceeds the number of evaluation questions matched because many questions ap- pear in both banks. Only 3,227 of the 4,289 pooled signa- tures are distinct, and 2,069 items share a signature with at least one other item. A single evaluation item therefore deletes every pooled item carrying the same text. The proce- dure does not remove duplication internal to the training cor- pus, so a question common to both reactor types contributes more than one training example. Text matching removes an evaluation question only where its wording is reproduced up to punctuation and case. The NRC constructs each administered examination from three sources [39]. Forty items are drawn directly from the bank. Five are derived by modifying bank items through one or more significant modifications. Five are newly authored. Verbatim reuse is therefore eliminated. A derived item and its bank ancestor differ in wording and both survive. We treat the residual similarity between such pairs as a bounded but non-zero source of optimism and return to it in Section 4.7. Table 1 Dataset composition. Deduplication removed 711 of the 4,289 pooled bank items, and one further item was discarded during rationale generation, leaving the training corpus shown. The training rows are stated after both removals, so differencing a bank row against its training row accounts for 712 items in total, of which 711 are deduplication removals and one is the discarded item. SplitSourceItems Question banksPWR2,140 BWR2,149 Bank total4,289 Training (SFT / RAFT) PWR question bank1,797 BWR question bank1,780 Training total3,577 EvaluationPWR, March 2015 to 2021 350 BWR, March 2015 to 2021 350 Evaluation total700 2.2.3. Chain-of-Thought Distillation The 3,578 retained items are augmented with step-by- step rationales produced by a stronger teacher model. One item produced no output because the teacher API call ter- minated with an error, and no retry was attempted, so that item was dropped. The corpus used for fine-tuning therefore comprises 3,577 examples, of which 1,797 originate in the PWR bank and 1,780 in the BWR bank. No other item was excluded at this stage. For each question the stem, the four options and the ground-truth answer are submitted to Gemini 3 Flash together with reference passages retrieved from the handbook corpus. The teacher is instructed to explain step by step why the indicated option is correct using engineering principles present in the supplied references. It concludes with an explicit "Therefore, the correct answer is (X)" statement. The rationale is stored as a two-turn dialog. The user turn is the question with options and the assistant turn is the distilled rationale. For items containing diagrams the images are passed to the teacher alongside the text so that the rationale can reference visual content. The same images are retained in the training sample and routed through the vision encoder of the student. The teacher-side retriever differs from the retriever used at evaluation. The passages supplied to the teacher are re- trieved by a dense index built once over the same handbook texts. It uses FAISS over all-MiniLM-L6-v2 embeddings and returns the top three passages. The evaluation-time retriever is sparse and returns the top four chunks, as described in Section 2.5. For the SFT configurations this difference is inconsequential. SFT training examples contain no retrieved context. The retriever of the teacher therefore affects only the content of the rationale that the student learns to reproduce from parametric knowledge. For RAFT the situation differs. The training example pairs a sparse-retrieved context block I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 4 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams with a rationale written against a differently retrieved pas- sage set. The supervision signal may therefore assert facts that are absent from the context on which the student is simultaneously conditioned. We record this asymmetry here because it bears on the interpretation of the RAFT results in Section 4.4. It was not a designed manipulation but a property of the pipeline that we identify retrospectively. 2.2.4. Evaluation Protocol Each test item is presented as a structured prompt con- taining the question stem, the four labeled options and the corresponding image where a figure is present. Images are passed through the multimodal interface as native pixel inputs. No answer choices are masked or modified. For retrieval-augmented configurations the full item text, comprising the question stem and all four options, is submit- ted as a query against the chunked handbook corpus before generation. The top four chunks are jointly capped at 7,000 characters and prepended as a [Retrieved Reference Context] block. The prompt template is shared across all retrieval- augmented conditions. All generation is deterministic at temperature zero with greedy decoding and up to 4,096 new tokens. The free- form response is post-processed by a cascaded answer ex- tractor applying three strategies in order of priority. The first matches explicit answer phrases such as “the correct answer is X” or “Final answer: X”. The second takes the last paren- thesized option token within the final 300 characters. The third takes a terminal standalone option token. Responses yielding no extractable option are scored as incorrect. Two items are excluded from all conditions because their stems reference figures that were unrecoverable from the source documents. These are P-17_Q14 and B-17_Q15. Both fall in the 2017 papers. Those two papers are therefore scored over 49 items and all others over 50. The criterion is a proportion, so the exclusion raises the number of correct answers required to pass the 2017 papers from 40 of 50 to 40 of 49. Section 3.7 reports a sensitivity analysis treating the excluded items as incorrect. For the statistical analysis, pooled accuracy is computed over 698 scored items with 349 per reactor type. For each configuration we report the point estimate with a Wilson score 95% confidence interval for a binomial proportion. Individual examinations comprise only 50 items. The cor- responding interval is approximately 11 percentage points in each direction. We therefore do not interpret the pass or failure of any single examination as evidence about that examination. Per-examination results are confined to the aggregated count ∑ 푒 Pass(푒). That count is reported over the complete set of fourteen administrations without selection. No examination or subgroup was chosen after inspection of results. 2.3. Corpus Segmentation The external knowledge base is constructed from the seven handbook volumes. Two segmentation strategies are compared. Both operate directly on the source documents. 2.3.1. Fixed-Size Sliding-Window Chunking A sliding-window function partitions each handbook into overlapping chunks 푐 푗 ∈. Each chunk is a fixed- length span of characters, and each chunk shares its trailing 200 characters with the next chunk’s leading 200 characters, as illustrated below. |푐 푗 | = 푁 size ,|푐 푗 ∩ 푐 푗+1 | = 푁 overlap (2) with 푁 size = 1000 characters and 푁 overlap = 200 characters, so consecutive chunks overlap by 20% of their length. The overlap causes any given assertion in the source text to appear in more than one retrievable unit. This is the property Section 4.3 appeals to in explaining why the scheme suits the fine-tuned model. Section 3.4 examines the consequences of the redundancy. Pages dominated by exercise or problem content are discarded before chunking so that the corpus contains only expository theory text. A page is discarded if it contains four or more leading tokens of the form Eq., Ex., Q., Prob. or numbered problem listings. A page is also discarded if the keywords problems, exercises or homework appear within its first 500 characters. 2.3.2. Structure-Aware Chunking The second strategy partitions documents along their in- herent structural boundaries rather than at arbitrary character offsets. The procedure proceeds in three stages. 1. Base-style estimation. The font-size histogram of the first fifteen pages of each handbook is computed. The modal font size is taken as the body-text size 푓 base . 2. Heading detection. A text span is a candidate head- ing if two conditions hold. Its rendered font size must satisfy 푓 span ≥ 푓 base + 0.8 pt. It must also either match a programmatically extracted table-of- contents entry or begin with a numbered prefix of the form 푛 1 .푛 2 ... followed by a capitalized word. White-on-white spans are discarded to avoid spurious matches against running headers. 3. Section accumulation. Body text between consecu- tive headings is accumulated into a single chunk. Fig- ures whose captions begin with Figure, Fig., FIGURE or FIG. are extracted as 200 DPI rasters from the region above the caption and attached to the enclosing section chunk. The same problem-page filter is applied before accumula- tion. The resulting chunks are of variable length with a mean of approximately 1,435 characters. They respect the original organization of the source handbooks. We use the term structure-aware chunking deliberately to distinguish this scheme from embedding-based semantic chunking, which we do not employ. Segmentation is driven entirely by typo- graphic and table-of-contents cues in the source documents. No embedding model is involved. The procedure is fully deterministic given the sources. I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 5 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams 2.4. Fine-Tuning Paradigms 2.4.1. Supervised Fine-Tuning Parameter-efficient fine-tuning uses LoRA [40] on the 3,577 distilled training items of Section 2.2.3. Only LoRA parameters are updated. Base weights remain frozen. LoRA modules are restricted to the language-model side of the architecture. The vision tower and the multi-modal projec- tor remain frozen. This preserves the pretrained image en- coder while permitting language-side specialization. Target modules are all n.Linear leaves matching q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj outside the vision and projector subgraphs. Training minimizes the standard next-token prediction loss, but computed only over the assistant’s rationale, not over the question itself. Concretely, the model is shown the question 푋 and the teacher’s distilled rationale, one token at a time, and at each position 푡 it is penalized in proportion to how much probability it assigned to the wrong next token 푦 푡 given everything before it. The user’s question tokens are excluded from this penalty by masking them at label −100, a standard convention that tells the loss function to skip them, so the model is trained to produce the rationale but not to reproduce the question. SFT (휃) = − 푇 ∑ 푡=1 log푃 휃 ( 푦 푡 ∣ 푦 <푡 , 푋 ) (3) Here 휃 denotes the trainable LoRA parameters, 푇 is the number of rationale tokens, 푦 푡 is the 푡-th rationale token, and 푦 <푡 denotes every rationale token generated before it. Minimizing SFT is equivalent to maximizing the proba- bility the model assigns to the teacher’s rationale, token by token, conditioned on the question. Hyperparameters appear in Table 2. Two safeguards protect against silent training failure. The trainable parameter count is asserted non-zero immediately after LoRA injection. A callback aborts train- ing if the gradient norm is exactly zero across the first five logged steps. The training corpus pools PWR and BWR items. A single fine-tuned model is therefore produced and evaluated on both reactor types. 2.4.2. Retrieval-Augmented Fine-Tuning RAFT trains the model to operate in the presence of imperfect retrieval. Each training example is augmented with context produced by the retrieval pipeline used at inference time. Training-time and test-time context distributions there- fore match. For every training question the top four chunks are jointly capped at 7,000 characters and prepended under the block format used at evaluation. The assistant target remains the distilled rationale of Section 2.2.3. The RAFT objective is identical to the SFT objective of Equation (3), with one addition. The retrieved context block 퐶 retr is also visible to the model when it predicts each rationale token. The model is therefore trained to use that context rather than Table 2 LoRA hyperparameters. Identical settings are used for SFT and the two RAFT variants. Only the training corpus differs. HyperparameterValue Base model google/gemma-4-31b-it LoRA rank (푟)16 LoRA scaling (훼)32 LoRA dropout0.05 Target modulesq/k/v/o_proj, gate/up/down_proj (language model only) Frozen subgraphsvision tower, multi-modal projector Epochs3 Per-device batch size1 Gradient accumulation steps8 Effective batch size8 Learning rate2 × 10 −4 LR schedulercosine Warmup ratio0.05 Weight decay0.01 Maximum sequence length 2,048 Precisionbfloat16 OptimizerAdamW Gradient checkpointingenabled to ignore it. RAFT (휃) = − 푇 ∑ 푡=1 log푃 휃 ( 푦 푡 ∣ 푦 <푡 , 푋, 퐶 retr ) (4) where 퐶 retr is the block of retrieved reference passages described below, concatenated with the question exactly as it appears at inference time. Two variants are trained, one per chunking strategy. Each is evaluated with the matching strategy so that chunking statistics are consistent between training and evaluation. All other hyperparameters are iden- tical to SFT. RAFT presumes a retrieval component. The No Retrieval condition is therefore not applicable to it. Section 2.2.3 noted that the rationale targets were gener- ated against dense-retrieved passages while 퐶 retr is sparse- retrieved. The context distribution is matched between train- ing and inference but the supervision is not matched to the context. We treat this as an explanatory hypothesis rather than a controlled variable. 2.5. Retrieval at Inference All retrieval-augmented configurations use sparse lexi- cal retrieval with BM25 [41]. BM25 scores a chunk against a query by rewarding query terms that occur in the chunk, weighted by how rare each term is across the corpus, so that terms occurring almost everywhere contribute little. For a query 푄 tokenized into distinct terms and a candidate chunk 퐷, BM25(푄,퐷) = ∑ 푞∈푄 IDF(푞) (푘 1 + 1)푓(푞,퐷) 푓(푞,퐷) + 퐾(퐷) (5) 퐾(퐷) = 푘 1 ( 1 − 푏 + 푏 |퐷|∕ 푑 ) (6) I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 6 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams Here 푓(푞,퐷) is the frequency of term 푞 in 퐷, |퐷| is the length of 퐷 in tokens, and푑 is the mean chunk length across the corpus, so that Equation (6) normalizes for chunk length. Terms repeated within a query contribute only once. The inverse document frequency is IDF(푞) = ln ( 1 + (푁 − 푛 푞 + 0.5)∕(푛 푞 + 0.5) ) , where 푁 is the number of indexed chunks and 푛 푞 the number of chunks containing 푞. It is larger for terms confined to few chunks. Such terms carry more discriminating power. The fraction saturates in 푓(푞,퐷). Each additional occurrence therefore raises the score by a diminishing amount, so that a chunk mentioning a term ten times is not scored ten times as relevant as one mentioning it once. The term 퐾(퐷) penalizes long chunks that accumulate matches by containing more text. The constants 푘 1 = 1.5 and 푏 = 0.75 are the standard defaults governing, respectively, the rate of saturation and the strength of the length penalty. The retrieval query is the full item text, comprising the question stem together with all four answer options. Queries and chunks are lower-cased and tokenized with an alphanumeric regular expression. No stemming, stopword removal or query expansion is applied. Chunks shorter than 80 characters are excluded from the index, as they are typically headers or page artefacts rather than substantive content. The four highest-scoring chunks are concatenated in rank order under a cumulative budget of 7,000 characters. Internal whitespace is normalized. The chunk that exceeds the budget is truncated at the character level. BM25 is used in place of a dense retriever because lexical retrieval is fully reproducible from the corpus alone, without any dependence on a particular embedding model or its version. 2.6. Computational Framework All experiments run on Apple Silicon hardware with an M5 Max chip and 128 GB of unified memory. The entire pipeline is therefore executable within an air-gapped environment on commodity workstation hardware. Fine-tuning executes under PyTorch with Hugging Face transformers and peft on the Metal Performance Shaders device. Training uses bfloat16 with gradient checkpointing. For inference, each LoRA module is merged into the base parameters on CPU. This avoids additional memory pressure on the Metal device. The resulting standalone safetensors model is converted to the Apple MLX format [42]. Evalu- ation uses mlx-vlm, which exhibits substantially more stable memory reclamation. All weights and activations remain in bfloat16 through- out fine-tuning, merging, conversion and inference. Decod- ing settings follow Section 2.2.4 and are identical across all eight conditions. Accuracy differences are therefore at- tributable to the model and the retrieval pipeline rather than to decoding stochasticity. 2.7. Experimental Configurations Table 3 summarizes the eight configurations. They cross three model states of Base, SFT and RAFT with three retrieval conditions of No Retrieval, RAG-Fixed and RAG- Structure-Aware. Every configuration is evaluated on all fourteen administered examinations. 3. Results This section reports the performance of all eight config- urations on the fourteen administered papers. The primary quantity is the number of examinations satisfying the crite- rion of Equation (1). Pooled accuracy is reported as a sec- ondary summary. All accuracies follow the strict convention of Section 2.1, under which an unextractable response is scored as incorrect and the denominator is fixed. Section 3.7 reports the alternative convention. Every evaluation item was removed from the training corpus by the deduplication of Section 2.2. Reported scores therefore measure general- ization to unseen questions rather than recall of the question bank. 3.1. Examinations Passed Table 4 reports the number of administered examina- tions on which each configuration met the criterion. Two outcomes hold across all configurations. No configuration without fine-tuning met the criterion on any of the fourteen examinations, whether or not retrieval was supplied. The best such configuration peaked at 76.0% on its strongest paper. The strongest fine-tuned configuration was SFT with fixed-size chunking retrieval. It met the criterion on 8 of 14 examinations and on more examinations than any alterna- tive. The ordering among fine-tuned configurations follows the pooled ordering of Section 3.2 but separates them more sharply. 3.2. Pooled Accuracy Table 5 reports pooled accuracy by reactor type under the strict convention. The accuracy reaches the criterion on the PWR items alone. SFT with fixed-size chunking retrieval scores 80.23% on PWR, which corresponds to 280 of 349 items and clears the mark by a single item, since 279 correct answers would give 79.94%. The same configuration reaches 79.08% on BWR and 79.66% pooled, both below the criterion. The Wilson interval on the pooled figure spans the threshold. The pooled comparison is therefore not decisive in either direction. Section 3.7 shows that the sign of the pooled comparison also depends on the scoring convention. The per-examination count of Section 3.1 is the quantity on which we rest the substantive claim. Distilled CoT fine-tuning is the largest single contribu- tor to accuracy. Holding the retrieval condition fixed, SFT raises combined accuracy from 51.86% to 75.50% without retrieval for a gain of 23.6 percentage points. Under fixed- size chunking it raises accuracy from 62.18% to 79.66% for a gain of 17.5 points. Under structure-aware chunking it raises accuracy from 64.47% to 77.51% for a gain of 13.0 points. The fine-tuned model without retrieval reaches 75.50%. This exceeds the best base configuration at 64.47% by 11.0 points. Retrieval provides a consistent secondary benefit that diminishes as fine-tuning increases. On the base model I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 7 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams Table 3 Experimental matrix. Checkmarks indicate configurations evaluated in this study. RAFT requires retrieval by design. Each RAFT variant is paired at evaluation time with the chunking strategy used during its training. Model StateNo RetrievalRAG (Fixed)RAG (Structure-Aware) Base✓ SFT✓ RAFTN/A✓ Table 4 Number of administered examinations meeting the 80% crite- rion out of seven per reactor type and fourteen in total. Per- examination scores are given in Table 7, in which the dagger count of each row reproduces the corresponding row of this able. ConfigurationPWR BWR Total Base, no retrieval0/70/70/14 Base + RAG (fixed)0/70/70/14 Base + RAG (structure-aware)0/70/70/14 SFT, no retrieval3/71/74/14 SFT + RAG (fixed)3/75/78/14 SFT + RAG (structure-aware)3/72/75/14 RAFT (fixed)2/73/75/14 RAFT (structure-aware)3/71/74/14 retrieval adds 10.3 points under fixed-size chunking and 12.6 points under structure-aware chunking. On the fine- tuned model the increment falls to 4.2 points and 2.0 points respectively. The two interventions are individually large but jointly sub-additive. Their separate gains of 23.6 and 12.6 points sum to 36.3 points. The observed gain from the base model to the best configuration is 27.8 points. Fine-tuning on bank questions therefore supplies implicitly a substantial portion of the handbook knowledge that retrieval would otherwise contribute. 3.3. Retrieval Gain by Reactor Type Retrieval benefits the BWR items more than the PWR items, and the direction holds in all four available com- parisons. On the base model, fixed-size chunking adds 7.2 points on PWR against 13.5 points on BWR, and structure- aware chunking adds 8.9 points against 16.3 points. On the fine-tuned model the two gaps narrow to 0.9 and 0.6 points. The retrieval corpus offers no mechanism for the pattern. The seven handbook volumes are agnostic to reactor type and cover the subject areas tested on both examinations, so neither reactor type is better served by the retrieved text. Nor is the pattern an artefact of the base rates, which differ by 0.6 points between the two reactor types before retrieval is supplied. We report the asymmetry without an explanation. Any single one of the four differences is within sampling error at 349 items per reactor type, so we do not claim a reactor-type effect. Locating the pattern would require a decomposition of accuracy by topic rather than by reactor type, which we leave to future work. 3.4. Chunking-Strategy Reversal The preferred chunking strategy reverses with the train- ing state of the model. The reversal holds without exception across the comparisons in Table 6. Structure-aware chunking is preferred for the base model at 64.47% against 62.18%. Fixed-size chunking is preferred for both fine-tuned paradigms. SFT reaches 79.66% against 77.51% and RAFT reaches 77.36% against 75.36%. The reversal is reproduced separately within the BWR items. On the PWR items it holds for the SFT family, at 80.23% against 78.22%, but the two RAFT variants are exactly tied at 77.08%, corresponding to 269 of 349 items in both cases. Section 4.3 interprets the pattern. 3.5. RAFT Against SFT RAFT underperforms the corresponding SFT configu- ration under matched chunking and matched inference-time retrieval. It reaches 77.36% against 79.66% under fixed-size chunking, a deficit of 2.3 points, and 75.36% against 77.51% under structure-aware chunking, a deficit of 2.2 points. The deficit is close to constant across chunking strategies and appears in both reactor types, at 3.2 and 1.1 points on PWR and 1.4 and 3.2 points on BWR. Because the two paradigms share identical rationale targets and identical retrieval at inference time, and differ only in whether retrieved context is present during training, the deficit is attributable to retrieval conditioning during fine-tuning. 3.6. Difficulty Structure Across Examinations Aggregation across examinations conceals substantial heterogeneity. Table 7 reports every configuration on every administration, so that the dispersion underlying the pooled figures of Table 5 is visible in full. The spread is wide. The highest single-paper score in the study is 90.0% on the March 2021 PWR paper, and the lowest is 42.0% on the March 2020 BWR paper, a range of 48 percentage points across the matrix. The highest score is not recorded by the configuration that passes the most examinations. It belongs to SFT with structure-aware chunking, whereas SFT with fixed-size chunking peaks at 88.0% on the 2015 and 2021 PWR papers. At 50 items per paper the difference between these cells is one item and lies well inside the sampling interval of Section 2.2.4. We record the observation to illustrate why the best single-paper score is not an admissible basis for selecting a configuration, and we rest no claim on it. I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 8 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams Table 5 Pooled accuracy in percent across all administered examinations by reactor type, under the strict convention. Denominators are fixed at 349 items per reactor type and 698 pooled. The interval is the Wilson score 95% interval on the combined figure. ConfigurationPWRBWRCombined95% CI Base, no retrieval52.1551.5851.86[48.2, 55.5] Base + RAG (fixed)59.3165.0462.18[58.5, 65.7] Base + RAG (structure-aware)61.0367.9164.47[60.8, 67.9] SFT, no retrieval76.5074.5075.50[72.2, 78.5] SFT + RAG (fixed)80.2379.0879.66[76.5, 82.5] SFT + RAG (structure-aware)78.2276.7977.51[74.3, 80.4] RAFT (fixed)77.0877.6577.36[74.1, 80.3] RAFT (structure-aware)77.0873.6475.36[72.0, 78.4] Table 6 Combined accuracy in percent by chunking strategy and model state, under the strict convention. The preferred strategy inverts after fine-tuning. Model state Fixed Structure-aware Preferred Base + RAG 62.18 64.47Structure-aware SFT + RAG 79.66 77.51Fixed RAFT77.36 75.36Fixed Counting passes by sitting rather than by configuration exposes a pronounced year effect. The matrix contains six- teen configuration-paper cells per sitting, and the fourteen papers together yield twenty-six passing cells. The March 2019 sitting accounts for nine of them, more than any other year. Every fine-tuned configuration meets the criterion on the 2019 PWR paper, at between 80.0% and 84.0%, and four of the five also meet it on the 2019 BWR paper, while no base configuration exceeds 74.0% on either. The March 2021 sitting follows with seven. At the other extreme the March 2017 sitting is the only one on which no configuration meets the criterion on either paper. The best result there is 79.6%, which corresponds to 39 of the 49 scored items and falls one item short of the 40 required. The 2015, 2018 and 2020 sittings yield two passing cells each. The difficulty ordering is not uniform across configu- rations. Taking each configuration’s weakest single paper, no sitting accounts for more than three of the eight. Three configurations are weakest in 2016, two in 2015, two in 2020 and one in 2017. Year difficulty is therefore partly dependent on model state rather than an intrinsic property of the paper, and a sitting that is hard for an off-the-shelf model need not be hard for its fine-tuned descendant. The March 2015 BWR paper behaves differently again. The base model scores 48.0% on it. That is only 3.6 points below its BWR average, so the paper is not intrinsically hard. Fine-tuned configurations fall far below their own averages on it. SFT falls by 10.5 points to 64.0% and SFT with fixed- size retrieval falls by 11.1 points to 68.0%. The deficit grows with fine-tuning rather than shrinking. This is the signature of a training-coverage gap rather than of item difficulty. The paper is also consequential. It accounts for the entire gap between BWR and PWR pooled accuracy. With it excluded, BWR accuracy under the best configuration would be 80.9% against 80.23% on PWR. 3.7. Scoring Convention and Sensitivity The strict convention scores an unextractable response as incorrect and holds the denominator fixed. The alternative convention conditions accuracy on successful extraction and removes such items from the denominator. Under the alter- native the effective denominator varies between conditions, which both raises accuracy and makes configurations non- comparable. Table 8 reports both. The choice of convention does not affect the order- ing of configurations at the reported precision. It does af- fect the pooled comparison against the threshold. Under the extraction-conditioned convention the best configuration reaches 80.1% pooled and clears the mark. Under the strict convention it reaches 79.66% and falls below it. The PWR figure of 80.23% clears the mark under both. The swing on the pooled figure is driven by four items out of 700. We therefore rest no claim on the position of the pooled figure relative to the threshold. The per-examination pass count is unaffected by the choice. The two excluded items require separate treatment. They are removed from all conditions for the reason given in Section 2.2.4, so the 2017 papers are scored over 49 items. Treating both as incorrect and restoring the denominator to 350 per reactor type changes the best configuration to 80.00% on PWR, 78.86% on BWR and 79.43% pooled. The PWR figure remains at the criterion, which the NRC standard treats as passing since no rounding is required. No per-examination pass status changes under this treatment. The width of the confidence intervals bounds what can be claimed. At 698 items the Wilson score 95% interval for the best configuration is 79.66% with bounds of 76.5% and 82.5%. This interval spans the threshold. The corresponding interval for the base model without retrieval is 51.86% with bounds of 48.2% and 55.5%. The separation between fine- tuned and base configurations is far larger than the interval width. The position of the best configuration relative to the 80% mark is not resolved at this sample size. Intervals among the four fine-tuned configurations overlap substantially. We therefore treat the chunking reversal and the RAFT deficit I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 9 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams Table 7 Accuracy in percent of all eight configurations on each administered examination, under the strict convention. The 2017 papers are scored over 49 items and all others over 50. Daggers mark examinations meeting the 80% criterion. Configuration2015201620172018201920202021Total (a) PWR Base, no retrieval56.050.053.146.056.048.056.052.15 Base + RAG (fixed)68.052.053.160.064.062.056.059.31 Base + RAG (structure-aware)64.046.067.364.064.062.060.061.03 SFT, no retrieval76.082.0 † 73.566.084.0 † 72.082.0 † 76.50 SFT + RAG (fixed)88.0 † 74.079.678.080.0 † 74.088.0 † 80.23 SFT + RAG (structure-aware)84.0 † 78.071.472.082.0 † 70.090.0 † 78.22 RAFT (fixed)78.074.077.670.082.0 † 74.084.0 † 77.08 RAFT (structure-aware)74.070.079.672.082.0 † 80.0 † 82.0 † 77.08 (b) BWR Base, no retrieval48.048.046.970.056.042.050.051.58 Base + RAG (fixed)60.064.073.566.072.060.060.065.04 Base + RAG (structure-aware)60.070.063.376.074.058.074.067.91 SFT, no retrieval64.076.079.676.080.0 † 74.072.074.50 SFT + RAG (fixed)68.084.0 † 77.682.0 † 82.0 † 80.0 † 80.0 † 79.08 SFT + RAG (structure-aware)70.086.0 † 69.482.0 † 78.076.076.076.79 RAFT (fixed)70.080.0 † 79.678.088.0 † 66.082.0 † 77.65 RAFT (structure-aware)72.070.075.572.080.0 † 72.074.073.64 Table 8 Combined accuracy in percent under the two scoring conventions, with the per-examination pass count. The extraction-conditioned figures use a denominator that varies between 689 and 698 across configurations. ConfigurationExtraction-cond.StrictDifferencePasses Base, no retrieval52.251.86−0.30/14 Base + RAG (fixed)62.462.18−0.20/14 Base + RAG (structure-aware)65.064.47−0.50/14 SFT, no retrieval75.575.50±0.04/14 SFT + RAG (fixed)80.179.66−0.48/14 SFT + RAG (structure-aware)77.677.51−0.15/14 RAFT (fixed)77.677.36−0.25/14 RAFT (structure-aware)75.375.36+0.14/14 as robust in direction, on the strength of their consistency across reactor types and chunking strategies. We treat them as modest in magnitude. We claim no statistically significant separation among fine-tuned configurations from the pooled figures alone. 4. Discussion This section interprets the pass count, the choice of evaluation unit, the two negative findings on retrieval design and the implications for deployment. 4.1. What the Result Establishes An openly available 31-billion-parameter multimodal model fine-tuned with distilled CoT supervision and paired with fixed-size chunking retrieval satisfied the 80% criterion on 8 of the 14 papers administered at the March sitting from 2015 to 2021, while no configuration without fine- tuning satisfied it on any. The appropriate reading is that the pipeline reaches the vicinity of the criterion rather than clearing it with margin. Pooled accuracy is 80.1% under one defensible scoring convention and 79.66% under another. The confidence interval spans the threshold under both. A human candidate who passed 8 of 14 sittings would not be described as reliably qualified. The result establishes that the gap between an off-the- shelf open model at 51.86% and a criterion of 80% is closable by techniques that run entirely on a single workstation. Data governance, auditability and on-premise deployment are fre- quently prerequisites in this sector, and the deployed artifact satisfies them: the merged adapter and the sparse index require no proprietary model, no external inference endpoint and no network access at run time. One external dependency exists at construction time, since the CoT rationales were distilled from a proprietary teacher, but it is confined to a single offline step and leaves no residual requirement in the deployed system. I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 10 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams 4.2. The Unit of Evaluation We report per-examination outcomes because that is the unit at which the criterion operates. A candidate passes or fails a specific administration and never a weighted average across years. Pooled accuracy ranks PWR above BWR at 80.23% against 79.08%. It would support the statement that the model fails on BWR content. The per-examination count reverses this. The model met the criterion on 5 of 7 BWR papers against 3 of 7 PWR papers. The pooled ordering is produced almost entirely by one anomalous BWR paper. Where a benchmark inherits its criterion from an external standard we suggest that it should also inherit the unit of assessment of that standard. Reporting only pooled accuracy can invert the qualitative conclusion. 4.3. Retrieval Design Depends on the Training State Structure-aware chunking is preferable before fine-tuning and fixed-size chunking after. This holds without excep- tion in our pooled comparisons. A plausible mechanism follows from what the two schemes supply. Structure-aware chunks are section-level and self-contained. They carry a complete expository treatment of a topic. Fixed-size chunks are shorter and more numerous and overlap by 20%. Any given assertion therefore recurs across several retrievable units. The base model lacks an internal organization of the domain. It must reconstruct an explanation from the re- trieved text and therefore benefits from narrative coherence. The fine-tuned model already possesses that organization. It needs only the specific quantity or relation on which the item turns. For that purpose the redundancy of overlapping windows raises the probability that the decisive fact appears somewhere in the context. The loss of narrative continuity costs little. The practical implication is that chunking cannot be tuned against an off-the-shelf model and then inherited by its fine-tuned descendant. That practice is nonetheless common. 4.4. Why Retrieval-Conditioned Training Did Not Help RAFT trails plain SFT by 2.3 and 2.2 points under the two chunking strategies and in both reactor types. The near-constancy of the deficit is inconsistent with sampling noise. It is also inconsistent with explanations that predict interaction with chunk granularity. We advance three candi- date mechanisms in decreasing order of fit to the observed pattern. The first concerns a mismatch internal to the RAFT training example identified in Section 2.2.3. The context block shown to the student is sparse-retrieved. The rationale it is trained to reproduce was authored by the teacher against dense-retrieved passages, under an instruction to ground the explanation in those passages. The student is therefore trained on examples in which the target text cites evidence that may be absent from the accompanying context. Such supervision does not teach robustness to distracting retrieval. It instead associates confident assertion with contexts that do not license it. The predicted consequence is a uniform degradation, which is what we observe. The hypothesis is testable by measuring passage overlap between the teacher context and the student context across the training set. We identify that measurement as the primary follow-up to this study. We state it as a hypothesis because the pipeline was not designed to vary this factor and no experiment reported here isolates it. The second concerns the retrieval corpus. Sparse lexi- cal retrieval over an authoritative and tightly scoped hand- book corpus returns relatively on-topic passages. RAFT is designed to confer robustness against distracting retrieval. Where retrieval is already clean, that robustness yields little benefit. Its cost in model capacity is unchanged. The third concerns capacity itself. At 3,577 items and rank 16 the capacity allocated to retrieval robustness may trade against capacity for the target reasoning. Under noisier or broader corpora the balance between the two paradigms may differ. We do not claim the ordering generalizes beyond this setting. 4.5. Distinguishing Hard Examinations from Coverage Gaps The March 2020 PWR paper and the March 2015 BWR paper both depart from the aggregate pattern, and they depart from it in different ways. The comparison that separates them is the deficit of a paper relative to a configuration’s own average, evaluated before and after fine-tuning. On the 2020 PWR paper both the base model and the strongest configuration score below their own PWR aver- ages, by 4.2 and 6.2 points respectively. The paper is harder for the adapted model in roughly the measure that it is harder for the unadapted one. That is the signature of item difficulty. On the 2015 BWR paper the base model sits 3.6 points below its BWR average, which is unremarkable, while SFT falls 10.5 points and SFT with fixed-size retrieval falls 11.1 points below theirs. The deficit widens by a factor of three once fine-tuning is applied. Difficulty intrinsic to the items cannot produce that asymmetry, since it would depress the base model as well. What the pattern points to instead is content underrepresented in the bank-derived training corpus. The diagnostic generalizes. Comparing the deficit of a paper under the base model against its deficit under the fine- tuned model separates item difficulty from training-coverage gaps using only quantities a benchmark already produces, with no additional experiments. It also carries an operational reading. A deployed assistant whose training corpus under- represents a topic will not signal that fact through degraded confidence. Here it manifested as an accuracy drop of more than 11 points on one paper with no corresponding drop for the base model. Coverage auditing of the fine-tuning corpus is therefore required in a safety-critical deployment. 4.6. Implications for Deployment These results indicate that fine-tuning and retrieval to- gether can bring an open model near operator-level com- mand of engineering fundamentals. They do not indicate I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 11 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams readiness for autonomous action in a control room. The instrument used here measures generic technical knowledge under idealized conditions. It contains no questions on regu- lations, technical specifications or emergency operating pro- cedures. It does not exercise real-time judgment, situational awareness or accountability. The appropriate role at this competence level is assistive. Such a system can surface relevant reference material, cross- check reasoning and retrieve authoritative sources while a licensed operator retains decision authority. Roughly one item in five is answered incorrectly even in the best configu- ration. This indicates that human oversight is required. Any deployment must ground assertions in retrievable sources so that they can be traced and verified. 4.7. Limitations Several limitations bound these findings. The instrument covers generic fundamentals only, so the results say nothing about competence on plant-specific systems, regulations or procedures, which the remaining stages of licensing as- sess. Deduplication by normalized question text removes verbatim overlap but cannot remove the residual similarity between a derived item and its bank ancestor. By NRC con- struction 5 of 50 items per paper are such derivations [39]. A bounded optimism therefore remains. The evaluation covers the March sitting only. The pass count would differ under a different selection of administrations. The deduplication procedure matches on normalized question text, and exact text matching is defeated by format- ting differences between the two sources. The 711 removals correspond to 472 of the 700 evaluation questions, leaving 228 without a textual match in the banks. By NRC construc- tion only 140 evaluation items, five derived and five newly authored per paper, should lack a bank counterpart [39]. The excess of 88 bounds the number of evaluation questions whose bank counterparts may remain in the training corpus, or roughly 2.5% of it. A further 121 training items carry fewer than four parsed answer options and 108 carry none, the option text having remained inside the question stem. The teacher received an empty option block for those items and the same items were used for fine-tuning. We study one model family and one sparse retriever. The efficacy of dense or hybrid retrieval and of other model scales is therefore unknown. The chunking reversal in particular may not survive a change of retriever. Individual examinations carry 50 items. The 95% interval at that size is roughly 11 points in each direction. Per-examination outcomes are therefore informative only in aggregate. The closed four- choice format does not capture open-ended reasoning or confidence calibration. We do not quantify hallucination rates or the faithfulness of generated rationales. An item an- swered correctly with faulty reasoning is scored identically to one answered correctly with sound reasoning. The CoT targets derive from a single teacher. Systematic reasoning biases in that teacher may therefore propagate to the student undetected. 5. Conclusion This study benchmarked a 31-billion-parameter open- weight multimodal model across eight configurations on the NRC Generic Fundamentals Examination. Each of the fourteen administered March papers from 2015 to 2021 was scored against the same 80% criterion required of human candidates. The pipeline combines distilled CoT fine-tuning, BM25 retrieval over the DOE Fundamentals Handbook cor- pus and a hybrid PyTorch and MLX back-end executable on a single workstation. The strongest configuration was SFT with fixed-size chunking retrieval. It satisfied the criterion on 8 of 14 examinations. No configuration without fine-tuning satisfied it on any. Pooled accuracy of 79.66% sits below the threshold with a confidence interval that spans it. The PWR figure of 80.23% sits above it. The result is therefore properly read as approaching operator-level command of engineering fundamentals rather than reliably achieving it. The central contribution is a transformation rather than a measurement. The same open-weight model answers 51.86% of the items correctly out of the box, a score at which no human candidate would be licensed, and satisfies the criterion on none of the fourteen papers. After domain adaptation it reaches 80.23% on the PWR item set and meets the criterion on 8 of the 14 papers. The distance between an off-the-shelf model that fails the assessment outright and one that meets a regulator’s own criterion on a majority of administrations is therefore closable by fine-tuning and retrieval alone, without recourse to a larger model or to a different architecture. The criterion is not one we chose. It is the threshold the NRC applies to every human candidate, and the papers are the instruments actually administered, so the comparison is made on terms the regulator already set rather than on a benchmark constructed for the purpose. The second contribution concerns what the result costs to obtain. Fine-tuning, adapter merging, format conversion and inference all execute on a single commodity workstation with unified memory, at a parameter count small enough that the weights, the training corpus and the retrieval index can be held and audited in one place. Retrieval is served by a sparse index over seven publicly available DOE handbooks, so the deployed system depends on no proprietary model, no external inference endpoint and no network access at run time. One external dependency does exist and we state it plainly. The CoT rationales were distilled from a proprietary teacher model. That dependency is confined to a single offline step in training-data construction, and the resulting adapter and index carry no residual requirement for it, so the artifact that would be deployed is self-contained even though the pipeline that produced it was not. For a sector in which data governance, auditability and on-premise op- eration are frequently prerequisites rather than preferences, this property is a substantive part of the result rather than an implementation detail. Three further findings refine how such systems should be built. First, the optimal chunking strategy reverses with the training state of the model. Structure-aware chunking favors I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 12 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams the base model and fixed-size sliding windows favor the fine- tuned model. Retrieval components must therefore be co- designed with the fine-tuning strategy rather than selected in advance. Second, RAFT underperforms plain SFT under matched retrieval by a margin that is small but consistent across chunking strategies and reactor types. We attribute this provisionally to a mismatch between the retriever that supplied the evidence of the teacher and the retriever that supplies the context of the student. A passage-overlap mea- surement would settle the question. Third, comparing the deficit of a paper under the base model against its deficit after fine-tuning separates intrinsic item difficulty from gaps in training coverage. The diagnostic requires no additional experiments. It surfaced a coverage gap invisible to pooled accuracy. These results position a fine-tuned and retrieval-grounded language model as an assistive tool that augments rather than replaces the judgment of a licensed operator. Roughly one item in five is still answered incorrectly in the best configuration, and the instrument covers generic fundamen- tals only. Future work will measure teacher and student context overlap to test the RAFT hypothesis directly. It will also decompose accuracy across text-only and image- bearing items to characterize the multimodal contribution and compare dense and hybrid retrieval against the sparse baseline established here. It will also seek an explanation for the reactor-type asymmetry in retrieval gain reported in Section 3.3, which a decomposition by topic rather than by reactor type would localize. Establishing trust in safety-critical settings requires evaluating not only whether answers are correct but whether the reasoning supporting them is faithful. We identify calibration measurement and faithfulness metrics such as entailment checking and citation accuracy as the necessary next step. Data availability The evaluation items, the question banks and the re- trieval corpus are drawn entirely from publicly available NRC and DOE publications cited in Section 2.2. The prompt templates, chunking parameters and evaluation scripts needed to reproduce the reported results are available from the corresponding author on reasonable request. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Acknowledgements The authors received no specific grant for this research from any funding agency in the public, commercial or not- for-profit sectors. CRediT authorship contribution statement Isak Hwang: Methodology, Software, Investigation, Writing – original draft. Yoon Pyo Lee: Conceptualization, Supervision, Writing – review & editing. References [1] S. Pakarinen, M. Sallinen, Workload management measures for sup- porting nuclear industry main control room operators and emergency response organization personnel during crises—a scoping review, Industrial health 63 (2025) 214–241. [2] U.S. Nuclear Regulatory Commission, A Human Error Handbook and Verification of IDHEAS-DATA, Research Information Letter 2025- 02, U.S. Nuclear Regulatory Commission, Washington, DC, 2025. [3] U.S. Nuclear Regulatory Commission, Title 10, code of federal regulations, part 55—Operators’ Licenses, Electronic Code of Fed- eral Regulations (eCFR), 2024. URL: https://w.ecfr.gov/current/ title-10/chapter-I/part-55, sections 55.40 and 55.41. [4] U.S. Nuclear Regulatory Commission, Operator Licensing Exami- nation Standards for Power Reactors, Final Report NUREG-1021, Rev. 12, U.S. Nuclear Regulatory Commission, Office of Nuclear Reactor Regulation, Washington, DC, USA, 2021. [5] K. Singhal, S. Azizi, T. Tu, S. S. Mahdavi, J. Wei, H. W. Chung, N. Scales, A. Tanwani, H. Cole-Lewis, S. Pfohl, et al., Large language models encode clinical knowledge, Nature 620 (2023) 172–180. [6] H. Nori, N. King, S. M. McKinney, D. Carignan, E. Horvitz, Ca- pabilities of GPT-4 on medical challenge problems, arXiv preprint arXiv:2303.13375 (2023). [7] D. M. Katz, M. J. Bommarito, S. Gao, P. Arredondo, GPT-4 passes the bar exam, Philosophical Transactions of the Royal Society A 382 (2024) 20230254. [8] C. Wang, D. Mandelli, J. Cogliati, Technical language processing of nuclear power plants equipment reliability data, Energies 17 (2024) 1785. [9] L. Wang, J. Hwang, K. Han, A. Gupta, Large language model- driven workflow for regulatory reporting in nuclear construction, in: Transactions of the International Conference on Structural Mechanics in Reactor Technology (SMiRT), IASMiRT, 2025. [10] B. Qi, X. Xiao, J. Liang, H. Zhou, W. Wu, X. Zhang, Synergistic ai for nuclear safety: Pinns, xai, dts, and llms in fault diagnosis, Nuclear Technology (2026) 1–18. [11] Y. Fayyaz, W. Elouataoui, Y. Gahi, K. El-Khatib, G. Harvel, K. Sankaranarayanan, Natural language processing in the nuclear industry: Opportunities and challenges, Nuclear Technology (2025) 1–21. [12] Y. P. Lee, J. Cha, Y. Yu, S. G. Kim, Large language model agent for nuclear reactor operation assistance, Nuclear Engineering and Technology 57 (2025) 103842. [13] L. Lin, T. Reeves, V. Agarwal, Using a Large Language Model for Ac- curate Technical Language Generation in the Predictive Maintenance of Circulating Water Systems in Nuclear Power Plants, Technical Report, Idaho National Laboratory (INL), Idaho Falls, ID (United States), 2025. [14] C.-Y. Hsieh, C.-L. Li, C.-K. Yeh, H. Nakhost, Y. Fujii, A. Ratner, R. Krishna, C.-Y. Lee, T. Pfister, Distilling step-by-step! outperform- ing larger language models with less training data and smaller model sizes, in: Findings of the Association for Computational Linguistics: ACL 2023, Association for Computational Linguistics, 2023, p. 8003–8017. doi:10.18653/v1/2023.findings-acl.507. [15] L. C. Magister, J. Mallinson, J. Adamek, E. Malmi, A. Severyn, Teaching small language models to reason, in: Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 2023, p. 1773–1781. [16] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, S. Riedel, D. Kiela, Retrieval-augmented generation for knowledge-intensive NLP tasks, I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 13 of 14 Benchmarking LLMs on NRC Reactor Operator Licensing Exams in: Advances in Neural Information Processing Systems, volume 33, Curran Associates, Inc., 2020, p. 9459–9474. [17] T. Zhang, S. G. Patil, N. Jain, S. Shen, M. Zaharia, I. Stoica, J. E. Gonzalez, RAFT: Adapting language model to domain specific RAG, arXiv preprint arXiv:2403.10131 (2024). [18] LlamaIndex Team, Evaluating the ideal chunk size for a RAG system using LlamaIndex, https://w.llamaindex.ai/blog/ evaluating-the-ideal-chunk-size-for-a-rag-system-using-llamaindex-6207e5d3fec5, 2024. Accessed: June 2026. [19] U.S. Department of Energy, DOE Fundamentals Handbook: Thermodynamics, Heat Transfer, and Fluid Flow, Volumes 1–3, Handbook DOE-HDBK-1012/1-92, DOE-HDBK-1012/2-92, DOE-HDBK-1012/3-92, U.S. Department of Energy, Washington, DC, 1992. [20] U.S. Department of Energy, DOE Fundamentals Handbook: In- strumentation and Control, Volumes 1–2, Handbook DOE-HDBK- 1013/1-92, DOE-HDBK-1013/2-92, U.S. Department of Energy, Washington, DC, 1992. [21] U.S. Department of Energy, DOE Fundamentals Handbook: Nu- clear Physics and Reactor Theory, Volumes 1–2, Handbook DOE- HDBK-1019/1-93, DOE-HDBK-1019/2-93, U.S. Department of En- ergy, Washington, DC, 1993. [22] Gemma Team, Google DeepMind, Gemma 4 model card, https:// ai.google.dev/gemma/docs/core/model_card_4, 2026. Accessed: June 2026. [23] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination question bank — pressurized water reactor, https://w.nrc.gov/reactors/operator-licensing/ history-rulemaking-activities/generic-fundamentals-examinations/ pwr, 2020. 2,140 questions organized by topic. Accessed: June 2026. [24] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination question bank — boiling water reactor, https://w.nrc. gov/reactors/operator-licensing/history-rulemaking-activities/ generic-fundamentals-examinations/bwr, 2020. 2,149 questions organized by topic. Accessed: June 2026. [25] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2015, https://w. nrc.gov/docs/ML1617/ML16173A029.pdf, 2015. ADAMS Accession No. ML16173A029. [26] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2016, https://w. nrc.gov/docs/ML1630/ML16300A005.pdf, 2016. ADAMS Accession No. ML16300A005. [27] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2017, https://w. nrc.gov/docs/ML1807/ML18074A026.pdf, 2017. ADAMS Accession No. ML18074A026. [28] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2018, https://w. nrc.gov/docs/ML1826/ML18269A212.pdf, 2018. ADAMS Accession No. ML18269A212. [29] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2019, https://w. nrc.gov/docs/ML1929/ML19290H601.pdf, 2019. ADAMS Accession No. ML19290H601. [30] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2020, https://w. nrc.gov/docs/ML2110/ML21103A276.pdf, 2020. ADAMS Accession No. ML21103A276. [31] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — pressurized water reactor, march 2021, https://w. nrc.gov/docs/ML2200/ML22007A311.pdf, 2021. ADAMS Accession No. ML22007A311. [32] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2015, https://w. nrc.gov/docs/ML1617/ML16173A025.pdf, 2015. ADAMS Accession No. ML16173A025. [33] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2016, https://w. nrc.gov/docs/ML1630/ML16300A003.pdf, 2016. ADAMS Accession No. ML16300A003. [34] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2017, https://w. nrc.gov/docs/ML1807/ML18074A024.pdf, 2017. ADAMS Accession No. ML18074A024. [35] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2018, https://w. nrc.gov/docs/ML1826/ML18269A207.pdf, 2018. ADAMS Accession No. ML18269A207. [36] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2019, https://w. nrc.gov/docs/ML1929/ML19290H604.pdf, 2019. ADAMS Accession No. ML19290H604. [37] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2020, https://w. nrc.gov/docs/ML2110/ML21103A277.pdf, 2020. ADAMS Accession No. ML21103A277. [38] U.S. Nuclear Regulatory Commission, NRC generic fundamentals examination — boiling water reactor, march 2021, https://w. nrc.gov/docs/ML2200/ML22007A312.pdf, 2021. ADAMS Accession No. ML22007A312. [39] U.S. Nuclear Regulatory Commission, Generic fundamentals examinations for reactor operators, https://w.nrc.gov/ reactors/operator-licensing/history-rulemaking-activities/ generic-fundamentals-examinations, 2022. Page last reviewed 31 March 2022. [40] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, LoRA: Low-rank adaptation of large language models, in: International Conference on Learning Representations, OpenRe- view.net, 2022. URL: https://openreview.net/forum?id=nZeVKeeFYf9. [41] S. E. Robertson, H. Zaragoza, The probabilistic relevance framework: BM25 and beyond, Foundations and Trends in Information Retrieval 3 (2009) 333–389. [42] A. Hannun, J. Digani, A. Katharopoulos, R. Collobert, MLX: Efficient and flexible deep learning on Apple Silicon, https://github.com/ ml-explore/mlx, 2023. Accessed: June 2026. I. Hwang and Y.P. Lee: Preprint submitted to ElsevierPage 14 of 14