Paper deep dive
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study
Yosuke Yamagishi, Atsushi Takamatsu, Yasunori Hamaguchi, Tomohiro Kikuchi, Shouhei Hanaoka, Takeharu Yoshikawa, Osamu Abe
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/3/2026, 12:18:53 AM
Summary
This study evaluates the educational suitability of LLM-generated Japanese translations of chest CT reports by comparing them with human-edited translations. Using 150 reports from the CT-RATE-JPN dataset, the study found that while LLM-generated translations were fluent, there was negligible agreement between human radiologists and LLM-as-a-judge evaluations. LLM judges consistently favored LLM-generated output, whereas human radiologists showed significant inter-rater disagreement and often preferred human-edited versions for clinical accuracy and authenticity, suggesting that automated LLM evaluation is currently insufficient for clinical or educational validation.
Entities (5)
Relation Signals (3)
DeepSeek-V3.2 → generatedtranslationsfor → CT-RATE-JPN
confidence 100% · LLM-generated translation produced by DeepSeek-V3.2
Radiologist → disagreedwith → LLM-as-a-judge
confidence 95% · Agreement between radiologists and LLM judges was near zero
LLM-as-a-judge → evaluated → CT-RATE-JPN
confidence 95% · three LLM judges... evaluated the same pairs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background: Accurate translation of radiology reports is important for multilingual research, clinical communication, and radiology education, but the validity of LLM-based evaluation remains unclear. Objective: To evaluate the educational suitability of LLM-generated Japanese translations of chest CT reports and compare radiologist assessments with LLM-as-a-judge evaluations. Methods: We analyzed 150 chest CT reports from the CT-RATE-JPN validation set. For each English report, a human-edited Japanese translation was compared with an LLM-generated translation by DeepSeek-V3.2. A board-certified radiologist and a radiology resident independently performed blinded pairwise evaluations across 4 criteria: terminology accuracy, readability, overall quality, and radiologist-style authenticity. In parallel, 3 LLM judges (DeepSeek-V3.2, Mistral Large 3, and GPT-5) evaluated the same pairs. Agreement was assessed using QWK and percentage agreement. Results: Agreement between radiologists and LLM judges was near zero (QWK=-0.04 to 0.15). Agreement between the 2 radiologists was also poor (QWK=0.01 to 0.06). Radiologist 1 rated terminology as equivalent in 59% of cases and favored the LLM translation for readability (51%) and overall quality (51%). Radiologist 2 rated readability as equivalent in 75% of cases and favored the human-edited translation for overall quality (40% vs 21%). All 3 LLM judges strongly favored the LLM translation across all criteria (70%-99%) and rated it as more radiologist-like in >93% of cases. Conclusions: LLM-generated translations were often judged natural and fluent, but the 2 radiologists differed substantially. LLM-as-a-judge showed strong preference for LLM output and negligible agreement with radiologists. For educational use of translated radiology reports, automated LLM-based evaluation alone is insufficient; expert radiologist review remains important.
Tags
Links
- Source: https://arxiv.org/abs/2604.02207v1
- Canonical: https://arxiv.org/abs/2604.02207v1
Trouble viewing inline? Open PDF directly →
Full Text
59,474 characters extracted from source content.
Expand or collapse full text
Blinded Radiologist and LLM-Based Evaluation of LLM-Generated Japanese Translations of Chest CT Reports: Comparative Study Yosuke Yamagishi, MD, MSc 1 ; Atsushi Takamatsu, MD, PhD 2, 3 ; Yasunori Hamaguchi, MD 1 ; Tomohiro Kikuchi, MD, PhD, MPH 4, 5 ; Shouhei Hanaoka, MD, PhD 1, 2 ; Takeharu Yoshikawa, MD, PhD 4 ; Osamu Abe, MD, PhD 1, 2 1 Division of Radiology and Biomedical Engineering, Graduate School of Medicine, The University of Tokyo, Tokyo, Japan 2 Department of Radiology, The University of Tokyo Hospital, Tokyo, Japan 3 Department of Radiology, Kanazawa University Graduate School of Medical Sciences, Ishikawa, Japan 4 Department of Computational Diagnostic Radiology and Preventive Medicine, The University of Tokyo Hospital, Tokyo, Japan 5 Department of Radiology, School of Medicine, Jichi Medical University, Tochigi, Japan Abstract Background: Accurate translation of radiology reports is important not only for multilingual research and clinical communication but also for radiology education and trainee learning. Recent large language models (LLMs) have demonstrated strong performance in medical translation, but the validity of LLM-based evaluation for educationally relevant translation quality remains unclear. Objective: To evaluate the educational suitability of LLM-generated Japanese translations of chest CT reports using blinded radiologist assessment and to compare expert ratings with LLM-as-a-judge evaluations. Methods: We used 150 chest CT reports from the validation set of CT-RATE-JPN. For each original English report, we compared a human-edited Japanese translation with an LLM-generated translation produced by DeepSeek-V3.2. A board-certified radiologist (radiologist 1) and a radiology resident (radiologist 2) independently performed blinded pairwise evaluations across four criteria: terminology accuracy, readability and fluency, overall report quality, and radiologist-style authenticity (which translation appeared more like a report written by a radiologist). In parallel, three LLM judges, DeepSeek-V3.2, Mistral Large 3, and GPT-5, evaluated the same translation pairs using identical criteria. Agreement between radiologists and between radiologists and LLM judges was assessed using quadratic weighted kappa (QWK) and percentage agreement. Results: Inter-rater agreement between radiologists and LLM judges was near zero (QWK = -0.04 to 0.15), whereas some LLM-LLM pairs showed higher agreement (up to QWK = 0.29). Agreement between the two radiologists remained poor (QWK = 0.01 to 0.06). The radiologist 1 found the two translations nearly equivalent in terminology (59% tie) but favored the LLM translation for readability (51%) and overall quality (51%). The radiologist 2 found translations equivalent in readability (75% tie) but favored the human-edited translation for overall quality (40% vs 21%). All three LLM judges overwhelmingly and systematically favored the LLM translation across all criteria, with LLM preference rates ranging from 70% to 99% depending on the criterion, yielding near-zero agreement with either radiologist. For radiologist-style authenticity, all three LLM judges overwhelmingly rated the LLM translation as more radiologist-like (>93% of cases); the radiologist 2 more often rated the human-edited translation as more radiologist-like (66% of definitive responses; p = .002), whereas the radiologist 1 more often rated the LLM translation as more radiologist-like (64% of definitive responses; p = .002). Conclusions: LLM-generated translations of chest CT reports were rated highly natural and fluent by both LLM judges and one human evaluator, but the two radiologists differed substantially in their assessments. LLM-as-a-judge evaluations showed a strong directional preference for the LLM translation and negligible agreement with radiologists. For educational deployment of translated radiology reports, automated LLM-based evaluation alone is insufficient; expert radiologist review remains important. Keywords: artificial intelligence; radiology reports; machine translation; computed tomography; multilingual datasets; large language models; DeepSeek; GPT Introduction Accurate translation of medical documentation is a cornerstone of modern globalized healthcare, facilitating not only cross-border clinical communication but also medical education and patient empowerment [1–3]. In the era of data-driven medicine, the availability of high-quality translated datasets has become even more critical. Large-scale English-language medical datasets are increasingly being leveraged to develop AI models in other languages through machine translation, creating a pressing need for translation methods that maintain clinical nuance and technical precision [4–7]. Large Language Models (LLMs) have recently emerged as transformative tools in the field of medical natural language processing [8–10]. Recent studies suggest that state-of-the-art LLMs can achieve translation quality comparable to professional translators in general medical contexts, including patient-facing instructional materials and the simplification of discharge documentation [11,12]. However, the application of LLMs to highly specialized fields like radiology remains relatively limited [13]. Radiology reports contain dense technical terminology, substantial stylistic variation despite ongoing efforts toward standardization, and highly consequential clinical information for which even subtle wording differences may affect interpretation [14,15]. To our knowledge, research specifically evaluating the English-to-Japanese translation of chest CT reports remains scarce, warranting a more rigorous assessment of the latest models. A significant bottleneck in advancing medical translation is the challenge of quality evaluation. Traditional automated metrics, such as Bilingual Evaluation Understudy (BLEU) [16] and Recall-Oriented Understudy for Gisting Evaluation (ROUGE) [17], rely primarily on lexical or n-gram overlap and often fail to capture semantic equivalence, factual consistency, or clinical accuracy [18]. In medical text evaluation, quality assessment is inherently challenging because semantically equivalent and clinically acceptable outputs may differ considerably in wording, making lexical- overlap-based evaluation inherently limited [19], and evaluating whether a translation preserves radiologist-aligned phrasing and reporting style may be even more difficult. Consequently, the “LLM-as-a-judge” framework, in which an LLM evaluates the output of another model, has gained traction as a scalable alternative to labor-intensive human review [20,21]. However, its validity in high-stakes radiology translation, particularly its alignment with the nuanced judgments of human specialists, remains insufficiently validated. In this study, we evaluate the performance of DeepSeek-V3.2, a state-of-the-art open-weight LLM [22], in translating chest CT reports from English to Japanese. By comparing model-generated translations with human-edited versions from the CT- RATE-JPN dataset [23], we aim to assess the clinical utility of the latest LLMs for medical education and research. Furthermore, we critically examine whether the LLM-as-a-judge framework provides a reliable surrogate for expert evaluation, or whether it introduces systematic biases that diverge from the professional standards of a board-certified radiologist and resident. Methods Ethical Considerations Given that this study utilized a publicly available dataset with deidentified patient information, and that our research focused on the translation and linguistic analysis of the existing dataset without accessing any additional patient data, institutional review board approval was not required for this research. Dataset We used 150 chest CT reports from the validation set of CT-RATE-JPN [23], which is derived from CT-RATE [24], a publicly available dataset comprising chest CT imaging studies with corresponding free-text radiology reports. For each report, two Japanese translations were available: • Human-edited translation: a translation produced through a multi-stage quality control pipeline. An initial machine translation was generated by GPT-4o mini, then reviewed and corrected by a radiology resident, and finally refined by a board-certified radiologist. This multi-stage process serves as the ground-truth Japanese translation in CT-RATE-JPN. • LLM-generated translation: produced by DeepSeek-V3.2, a large language model with strong multilingual and medical domain performance. Each evaluation instance consisted of the original English report, the human-edited Japanese translation, and the LLM-generated Japanese translation. The presentation order of the two Japanese translations (left/right or A/B) was randomized and counterbalanced. Translation Generation LLM-generated translations were produced using DeepSeek-V3.2 [22]. DeepSeek- V3.2 was selected for its strong performance on multilingual medical tasks and its permissive redistribution license (Apache-2.0), which enables future dataset sharing. Translations were generated at temperature 0 using a standardized prompt instructing the model to produce an idiomatic Japanese radiology report from the English source. The full prompt is provided in Supplementary Material 1. Linguistic Analysis of Translation Pairs To quantify text-level differences between the two translation sets, we performed a series of paired analyses on all 150 report pairs. Character count was computed after removing whitespace and newlines. Sentence segmentation was performed by splitting on Japanese sentence-ending punctuation and newlines, retaining segments of more than two characters. For morphological analysis, we used the Janome tokenizer [25] to tokenize each report and computed type-token ratio (TTR), the ratio of unique token types to total tokens, as a measure of lexical diversity, using both all parts of speech and content words only (nouns, verbs, adjectives, and adverbs). Lemma-level base forms were used where available. Statistical comparisons between paired human-edited and DeepSeek translations were performed using the Wilcoxon signed-rank test and paired t-test; two-tailed p- values are reported. Blinded Radiologist Evaluation The translation pairs were independently evaluated by two radiologists: one board- certified radiologist (radiologist 1) and one radiology resident (postgraduate year 5, radiologist 2). Both evaluators were blinded to whether each translation was human-edited or LLM-generated. Evaluation Procedure For each of the 150 report pairs, evaluators were shown: 1. The original English chest CT report 2. Japanese Translation A 3. Japanese Translation B The assignment of human-edited and LLM-generated translations to positions A and B was randomized. Evaluators selected their preferred translation or indicated equivalence for each criterion independently. Evaluation Criteria Evaluators assessed four criteria: 1. Terminology accuracy: accuracy of medical and anatomical terminology 2. Readability and fluency: naturalness and coherence of the translated text 3. Overall clinical suitability: overall appropriateness for use as a clinical radiology report 4. Radiologist-style authenticity: whether the translation more closely resembled the style and expression of a report written by a radiologist For criteria 1–3, evaluators chose one of three options: Translation A is better, Translation B is better, or Equivalent. For criterion 4, evaluators chose Translation A, Translation B, or Equivalent, as this criterion was intended to capture stylistic impression rather than objective superiority. Responses were decoded post hoc to determine whether the human-edited or LLM-generated translation was preferred. LLM-as-a-Judge Evaluation Three LLM judges evaluated the same 150 translation pairs using an identical prompt structure: • DeepSeek-V3.2 • Mistral Large 3 • GPT-5 These models were selected to represent different evaluation perspectives: DeepSeek-V3.2 was included as a self-judge because it was also used as the translation model, GPT-5 was included as a high-end commercial model [26], and Mistral Large 3 was included as a comparatively neutral open-weight model [27]. Each judge received the original English report and the two Japanese translations in randomized order, with the translation source hidden. The judge was instructed to evaluate the same four criteria and provide a structured JSON response indicating which translation was preferred for each criterion (A / B / TIE), along with a brief justification. Temperature was set to 0 for all judge evaluations. The full judge prompt is provided in the Supplementary Material. Statistical Analysis For each criterion and each evaluator, we report: • Count and percentage favoring the human-edited translation • Count and percentage favoring the LLM translation • Count and percentage of equivalent (TIE) responses Inter-rater agreement between the two radiologists, and between each radiologist and each LLM judge, was quantified using quadratic weighted kappa (QWK). QWK was chosen because the three response categories (human-edited preferred / equivalent / LLM preferred) form a natural ordinal scale; QWK appropriately penalizes larger disagreements more than adjacent-category disagreements. The ordinal order used was: human-edited preferred < equivalent < LLM preferred. Raw percentage agreement is also reported. For criterion 4 (radiologist-style authenticity), we treated definitive responses (excluding Uncertain) as binary outcomes and tested whether preference for the human-edited translation exceeded chance (50%) using a one-sided binomial test. This tests whether the human-edited translation was more consistently judged as radiologist-like, not whether evaluators could correctly identify the translation method. Results Linguistic Comparison of Translation Pairs Human-edited translations were significantly longer than DeepSeek-generated translations in total character count (median 517.5 vs. 481.5 characters; Wilcoxon signed-rank test, W = 729.5, p < .001), with DeepSeek translations averaging approximately 94% of the length of their human-edited counterparts (mean ratio 0.940). Despite this overall brevity, DeepSeek translations contained significantly more sentences per report (mean 20.3 vs. 19.6; p < .001), while individual sentences were significantly shorter (mean sentence length 24.7 vs. 27.3 characters; p < .001). This pattern indicates that DeepSeek-generated translations segmented content into a greater number of shorter sentences compared with the human-edited versions. Lexical diversity, measured by TTR, differed significantly between the two translation sets when computed across all parts of speech: DeepSeek translations showed higher TTR (mean 0.402 vs. 0.381; p < .001), indicating greater lexical variety relative to total token count. This seemingly counterintuitive finding is partly attributable to the shorter overall length of DeepSeek translations, as TTR tends to increase with shorter texts. When TTR was restricted to content words only (nouns, verbs, adjectives, and adverbs), no significant difference was observed (mean 0.550 vs. 0.549; p = .665), suggesting that the two translation methods employed a comparable breadth of clinically meaningful vocabulary. Radiologist Evaluation Results Figure 1 summarizes the blinded radiologist evaluation results. Terminology Accuracy The radiologist 1 rated 59% (n=89) of report pairs as equivalent in terminology, favoring the LLM translation in 23% (n=35) and the human-edited translation in 17% (n=26) of cases. The radiologist 2 rated 51% (n=76) as equivalent, favoring the human-edited translation in 35% (n=53) and the LLM translation in 14% (n=21) of cases. Readability and Fluency The radiologist 1 favored the LLM translation for readability in 51% (n=77) of cases, with human-edited preferred in 25% (n=38) and equivalent in 23% (n=35). The radiologist 2 found translations equivalent in readability in 75% (n=112) of cases, favoring LLM in 15% (n=22) and human-edited in 11% (n=16). Overall Report Quality The radiologist 1 favored the LLM translation for overall quality in 51% (n=77) of cases, with human-edited preferred in 29% (n=43) and equivalent in 20% (n=30). The radiologist 2 favored the human-edited translation in 40% (n=60) of cases, found translations equivalent in 39% (n=58), and favored the LLM translation in 21% (n=32). Radiologist-Style Authenticity When asked which translation more closely resembled the style of a report written by a radiologist, the radiologist 2 more often rated the human-edited translation as more radiologist-like (66% of definitive responses; binomial test p = .002). In contrast, the radiologist 1 more frequently rated the LLM translation as more radiologist-like (64% of definitive responses; p = .002). Figure 1. Blinded radiologist pairwise evaluation results. Stacked bar charts for radiologist 1 and radiologist 2 across four criteria. Inter-Rater Agreement Inter-rater agreement between the two radiologists was poor across all four criteria (confusion matrices in Figure 2). QWK ranged from 0.012 (radiologist-style authenticity) to 0.059 (readability/fluency), and raw agreement ranged from 28% (readability) to 37% (terminology accuracy). These values are consistent with slight-to-no agreement, well below the threshold conventionally considered as fair agreement (QWK ≥ 0.20). The direction of disagreement was notably systematic: the radiologist 1 consistently favored the LLM translation more than the radiologist 2 did, particularly for readability and overall quality. For overall quality, the radiologist 1 favored the LLM in 51% of cases versus 21% for the radiologist 2; the radiologist 2 favored the human-edited in 40% versus 29% for the radiologist 1. Figure 2. Inter-rater confusion matrices between radiologist 1 and radiologist 2. LLM-as-a-judge Results All three LLM judges strongly and systematically favored the LLM-generated translation across all criteria (Figure 3). Terminology Accuracy Judges favored the LLM translation in 79–91% of cases (DeepSeek: 91%, n=137; Mistral: 79%, n=119; GPT-5: 81%, n=122). TIE rates were low (7–13%), and the human-edited translation was preferred in only 2–18% of cases. Readability and Fluency Preference for the LLM translation was nearly universal: 70% (DeepSeek, n=105), 88% (Mistral, n=132), and 95% (GPT-5, n=142). Equivalent ratings were near zero. Overall Report Quality Judges favored the LLM translation in 83–95% of cases (DeepSeek: 95%, n=142; Mistral: 91%, n=137; GPT-5: 83%, n=125). Human-edited was preferred in 5–17%. Radiologist-Style Authenticity All three judges almost universally rated the LLM translation as more radiologist- like: DeepSeek favored LLM in 99% (n=148), Mistral in 93% (n=140), and GPT-5 in 95% (n=143) of cases. The human-edited translation was rated as more radiologist- like in <7% of responses for all three judges. Figure 3. LLM-as-a-judge evaluation results for all three judge models. Radiologist vs LLM Judge Agreement Agreement between radiologists and LLM judges was near zero across all criteria and criterion–evaluator combinations (Table 1, Figure 4). QWK between either radiologist and any LLM judge ranged from −0.038 to 0.148. Agreement among the LLM judges themselves ranged from poor to moderate depending on the criterion pair, with QWK values ranging from −0.010 to 0.286, but these judge–judge agreements did not translate into meaningful agreement with human radiologists. Figure 4. Pairwise agreement heatmap (QWK and percent agreement) for all evaluator pairs. Table 1. Pairwise inter-rater agreement across radiologists and LLM judges for each evaluation criterion, expressed as QWK and percent agreement. Criterion Rater 1 Rater 2 QWK % Agreement Terminology Accuracy Radiologist 1 Radiologist 2 0.056 36.7 Radiologist 1 DeepSeek V3.2 -0.022 26 Radiologist 1 Mistral Large 3 -0.038 26 Radiologist 1 GPT-5 0.084 24.7 Radiologist 2 DeepSeek V3.2 0.026 20 Radiologist 2 Mistral Large 3 0.015 25.3 Radiologist 2 GPT-5 0.025 18.7 DeepSeek V3.2 Mistral Large 3 0.286 82 DeepSeek V3.2 GPT-5 0.136 77.3 Mistral Large 3 GPT-5 0.169 70 Readability / Fluency Radiologist 1 Radiologist 2 0.059 28 Radiologist 1 DeepSeek V3.2 0.148 50 Radiologist 1 Mistral Large 3 0.128 53.3 Radiologist 1 GPT-5 0.063 52 Radiologist 2 DeepSeek V3.2 0.105 17.3 Radiologist 2 Mistral Large 3 0.016 15.3 Radiologist 2 GPT-5 0.007 14.7 DeepSeek V3.2 Mistral Large 3 0.071 67.3 DeepSeek V3.2 GPT-5 0.026 68.7 Mistral Large 3 GPT-5 0.177 86.7 Overall Quality Radiologist 1 Radiologist 2 0.021 31.3 Radiologist 1 DeepSeek V3.2 0.016 50.7 Radiologist 1 Mistral Large 3 0.055 51.3 Radiologist 1 GPT-5 0.119 52 Radiologist 2 DeepSeek V3.2 0.021 23.3 Radiologist 2 Mistral Large 3 0.05 25.3 Radiologist 2 GPT-5 0.076 28 DeepSeek V3.2 Mistral Large 3 0.235 90 DeepSeek V3.2 GPT-5 0.044 80.7 Mistral Large 3 GPT-5 -0.01 77.3 Radiologist-Style Authenticity Radiologist 1 Radiologist 2 0.012 32.7 Radiologist 1 DeepSeek V3.2 0.016 51.3 Radiologist 1 Mistral Large 3 0.079 52.7 Radiologist 1 GPT-5 0.105 54 Radiologist 2 DeepSeek V3.2 0.016 22 Radiologist 2 Mistral Large 3 0.069 26 Radiologist 2 GPT-5 0.063 25.3 DeepSeek V3.2 Mistral Large 3 0.072 92.7 DeepSeek V3.2 GPT-5 0.229 95.3 Mistral Large 3 GPT-5 0.191 91.3 Qualitative Analysis of Radiologist and LLM Judge Disagreements Representative Disagreement Cases To illustrate the nature of disagreements between radiologist evaluators and LLM judges, we identified cases in which both radiologists preferred the human-edited translation while all three LLM judges unanimously preferred the LLM-generated translation for the same criterion. The number of such cases was 5 (3.3%) for terminology accuracy, 4 (2.7%) for readability and fluency, 10 (6.7%) for overall quality, and 12 (8.0%) for radiologist-style authenticity; no case met this criterion across all four domains simultaneously. Representative examples are instructive, and are drawn from cases flagged for terminology accuracy. In one case, the LLM-generated translation rendered "sequela fibrotic changes" as a term approximating "prognostic symptom," a clinically meaningful mistranslation. Despite this error, GPT-5 did not acknowledge it in its evaluation rationale, while DeepSeek-V3.2 and Mistral Large 3 explicitly characterized the terminology as accurate and praised it positively. In a second example, the Japanese term for the right lobe of the thyroid gland was rendered with reversed word order, deviating from standard Japanese anatomical nomenclature, yet all three LLM judges overlooked this error entirely. A third case involved the translation of "tree-in-bud appearance," a standard radiological term, as a literal Japanese rendering approximating "budding tree-like appearance." This phrasing is not used in clinical radiology reports in Japan, where the term is universally retained in English. Despite this, one LLM judge rated this translation favorably, while the remaining two did not explicitly penalize it. Both radiologists rated the human-edited translation as superior for this case. Patterns in LLM Judge Rationales A review of the judges’ free-text rationales in these disagreement cases suggested that stylistic preferences played a substantial role in LLM-based evaluations (Table 2, 3). Across cases where radiologists preferred the human-edited translation but LLM judges preferred the LLM-generated version, the terms concise and natural appeared repeatedly in the judge explanations. This pattern was especially prominent for readability/fluency and radiologist-style authenticity, whereas these expressions were rarely invoked for terminology accuracy and overall quality. As shown in Table 2, 3, concise was frequently cited when judges favored LLM- generated translations for readability/fluency and radiologist-style authenticity, particularly by GPT-5 and Mistral Large 3. A similar tendency was observed for natural, which was again concentrated in readability/fluency judgments, especially for GPT-5 and Mistral Large 3. Table 2. Frequency of the term “concise” in LLM judges’ rationales for disagreement cases. Judge DeepSeek-V3.2 Mistral Large 3 GPT-5 Winner Human- edited TIE LLM Human -edited TIE LLM Human -edited TIE LLM Terminology Accuracy 0 0 0 0 0 0 0 0 0 Readability / Fluency 1 0 20 0 0 84 2 0 87 Overall Quality 0 0 3 0 0 16 0 0 0 Radiologist-Style Authenticity 0 0 70 4 0 119 1 0 23 Table 3. Frequency of the term “natural” in LLM judges’ rationales for disagreement cases. Judge DeepSeek-V3.2 Mistral Large 3 GPT-5 Winner Huma n- edited TIE LLM Human- edited TIE LLM Human- edited TIE LLM Terminology Accuracy 0 0 0 0 0 21 0 1 10 Readability / Fluency 17 0 42 11 0 55 8 0 105 Overall Quality 0 0 10 0 0 4 0 0 2 Radiologist-Style Authenticity 0 0 4 1 0 25 1 0 5 Discussion Principal Results This study found that LLM-generated Japanese translations of chest CT reports received divergent assessments depending on the evaluator. The radiologist 1 rated the LLM translation favorably for readability and overall quality, more so than the radiologist 2. The radiologist 2, by contrast, found translations frequently equivalent and showed a slight preference for the human-edited version overall. Crucially, all three LLM judges showed strong and systematic preference for the LLM translation, far exceeding even the more favorable of the two human assessments, and showed negligible agreement with either human rater. The inter-rater agreement between the two radiologists was strikingly low (QWK = 0.01–0.06 for all criteria), underscoring that even expert radiologists do not apply identical standards when evaluating translation quality. Importantly, this low agreement should be interpreted in context rather than viewed simply as annotation failure. In post hoc debriefing, both radiologists reported that, in many cases, the two translations were nearly indistinguishable in overall quality, and that final preferences were often determined by subtle wording differences or minor nuances. Under such conditions, some degree of variability in pairwise preference judgments is expected. Nevertheless, this level of variability has important methodological implications: studies relying on a single human rater to validate LLM translation quality may produce conclusions that are specific to that rater's preferences rather than reflecting a shared clinical standard. LLM Translation Quality and Its Implications for Education and AI Development The linguistic analysis revealed that LLM-generated translations were systematically shorter, contained more sentences, and had shorter individual sentences compared with human-edited counterparts. Despite this, content-word lexical diversity was comparable between the two methods, suggesting that the core clinical vocabulary was equivalently represented. These structural differences likely underlie the divergent stylistic impressions reported by the two radiologists. From an educational standpoint, this study has direct relevance to the broader use of translated medical materials in training and cross-lingual learning contexts [1–3]. For educational contexts requiring precise clinical terminology and faithful representation of diagnostic uncertainty, an essential feature of radiology reporting [28,29], conciseness and natural phrasing, although desirable, should not be assumed to be the most important indicators of translation quality in medical texts. Beyond medical education, the quality of machine-translated radiology reports has important implications for AI model development. Large-scale multilingual datasets are increasingly used to train and evaluate AI models for non-English clinical environments [4–7]. The fidelity with which LLM translation preserves clinical nuance, terminology conventions, and uncertainty expressions directly determines the quality of the training signal these datasets provide. Our findings suggest that LLM-generated translations are broadly adequate for high-volume dataset enrichment, such as pre-training or language model fine-tuning, where breadth and fluency are prioritized. However, for datasets intended to support high-stakes downstream tasks such as automated radiology report generation [30–32], expert review of at least a representative subset remains necessary to ensure that subtle clinical conventions are faithfully preserved. Why LLM Judges Diverge from Radiologists By contrast, the LLM-based judges showed a much stronger and more directional preference pattern, favoring the LLM translation in 79–95% of cases across all criteria. This discrepancy suggests that the LLM judges may not have been detecting clinically meaningful superiority alone, but may also have reflected model-specific stylistic preferences or alignment with particular lexical and syntactic patterns, consistent with prior reports of self-preference bias in LLM-as-a-judge settings [33]. In other words, when candidate translations are broadly similar in quality, LLM-as- a-judge may amplify small textual differences into disproportionately confident preferences. Several mechanisms may underlie this systematic bias. Prior work has shown that LLM-as-a-judge systems are susceptible to multiple forms of bias [34], raising the possibility that our three LLM judges were influenced by similar evaluation tendencies. As a result, they may have rewarded text that appeared polished, smooth, and lexically precise, even when those characteristics did not necessarily correspond to greater clinical appropriateness. A related explanation is that fluency may be implicitly conflated with overall translation quality [35]. Because contemporary LLMs are optimized in part through human preference signals that often reward naturalness and readability [36], they may systematically assign higher scores to outputs that sound more fluent on the surface, creating a structural advantage for LLM-generated translations [37]. Another likely factor is limited sensitivity to clinical register. Japanese radiology reports follow highly conventionalized patterns in wording, uncertainty marking, anatomic description, and overall sentence style. Even when a translation is grammatically correct and semantically plausible, small deviations from these conventions may render it less acceptable to radiologists. LLM judges may not be sufficiently grounded in radiology-specific reporting norms to recognize and appropriately weight such distinctions. This interpretation is further supported by the marked discrepancy in judgments of radiologist-like style. All three LLM judges almost uniformly rated the LLM translation as more characteristic of radiologist writing (>93% of cases). The two radiologists, however, not only disagreed with the LLM judges but also with each other: radiologist 2 showed an above-chance preference for the human-edited translation (p = .002), while radiologist 1 showed the opposite preference (p = .002). This pattern suggests that the criteria used to assess radiologist-style authenticity vary substantially across evaluators—whether human or LLM—and that surface fluency alone may not capture what expert radiologists consider professionally authentic writing. Practical Recommendation These findings point toward a tiered approach for deploying LLM-translated radiology reports. For applications where breadth and readability are prioritized, such as parallel corpora for language learning, multilingual teaching collections, or large-scale pre-training datasets for AI development, LLM-generated translations may be used with confidence. For higher-stakes applications, including model reports, teaching files, standardized assessment materials, or training data for clinical AI systems in which faithful representation of uncertainty is critical, expert radiologist review remains an important quality gate. Critically, these findings indicate that LLM-as-a-judge should not serve as the sole quality gate for either educational content or AI training data. Given the near-zero agreement with radiologist assessments and the systematic tendency to favor LLM- generated output, automated judging appears to substantially overestimate educational and clinical suitability. A more practical approach is therefore a hybrid workflow: LLM translation and automated evaluation for initial scaling and screening, followed by targeted expert review of cases intended for high-stakes applications. This workflow balances scalability with the clinical rigor that radiology-specific use cases demand. Limitations Several limitations should be noted. First, our dataset comprised only chest CT reports, and findings may not generalize to other modalities or report styles. Second, only Japanese translations were evaluated; results may differ for other target languages. Third, we had only two human raters, limiting statistical power for inter-rater analyses and precluding adjudication of disagreements. Fourth, the LLM- as-a-judge analysis was based on a single prompt framework and did not include systematic exploration of alternative judge prompts, rubric structures, or prompt optimization strategies. Fifth, we did not assess downstream educational outcomes—whether trainees who learn from LLM-translated versus human-edited reports show differences in knowledge acquisition or terminology proficiency. Such outcomes studies would be necessary to establish the ultimate educational utility of LLM-translated reports. Future Directions Future work should investigate: (1) the impact of translation quality on trainee learning outcomes in controlled educational settings; (2) systematic exploration and optimization of LLM-as-a-judge prompts and rubric designs grounded in radiology terminology standards; (3) extension to other languages, modalities, and report types; and (4) the development of radiologist-aligned automated evaluation metrics for medical translation quality. Conclusions LLM-generated Japanese translations of chest CT reports received divergent assessments from the two radiologists, with poor inter-rater agreement across all criteria. However, both human evaluators differed markedly from the three LLM judges, who showed systematic and extreme preference for the LLM translation far beyond either radiologist's assessment. LLM-as-a-judge evaluations showed near- zero QWK agreement with radiologist judgments and rated the LLM-generated translation as more radiologist-like in >93% of cases. These findings indicate that LLM-generated radiology report translations may hold promise for educational support, but that LLM-as-a-judge evaluation is insufficient for quality assurance in this domain. Expert radiologist review remains an important component of quality control when translating radiology reports for educational use. Data Availability The dataset used in the present analyses was CT-RATE-JPN [38], a derived dataset based on CT-RATE [39]. CT-RATE-JPN is publicly available on Hugging Face. The original CT-RATE dataset is also publicly available on Hugging Face. Acknowledgements The authors used GPT (OpenAI) and Claude (Anthropic) to assist with English- language editing and manuscript refinement. The authors reviewed and revised all AI-assisted outputs and take full responsibility for the final content of the manuscript. Conflicts of Interest The Department of Computational Diagnostic Radiology and Preventive Medicine of The University of Tokyo Hospital is sponsored by HIMEDIC Inc and Siemens Healthcare K. Author Contributions Y contributed to the conceptualization, study design, implementation, formal analysis, and manuscript writing. AT and YH served as radiologist evaluators. TK, SH, TY, and OA provided supervision and critically reviewed the manuscript for important intellectual content. All authors reviewed, approved the final manuscript, and agree to be accountable for all aspects of the work. Abbreviations BLEU: Bilingual Evaluation Understudy CT: computed tomography LLM: large language model QWK: quadratic weighted kappa ROUGE: Recall-Oriented Understudy for Gisting Evaluation TTR: type-token ratio References 1. Al Shamsi H, Almutairi AG, Al Mashrafi S, Al Kalbani T. Implications of Language Barriers for Healthcare: A Systematic Review. Oman Med J 2020 Mar;35(2):e122. PMID:32411417 2. Merx R, Phillips C, Suominen H. Machine Translation Technology in Health: A Scoping Review. Stud Health Technol Inform 2024 Sept 24;318:78–83. PMID:39320185 3. Albrecht U-V, Behrends M, Matthies HK, Jan U von. Usage of Multilingual Mobile Translation Applications in Clinical Settings. JMIR mHealth and uHealth JMIR Publications Inc., Toronto, Canada; 2013 Apr 23;1(1):e2268. doi: 10.2196/mhealth.2268 4. Phan L, Dang T, Tran H, Trinh TH, Phan V, Chau LD, Luong M-T. Enriching Biomedical Knowledge for Low-resource Language Through Large-scale Translation. In: Vlachos A, Augenstein I, editors. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics Dubrovnik, Croatia: Association for Computational Linguistics; 2023. p. 3131–3142. doi: 10.18653/v1/2023.eacl-main.228 5. Kobayashi K, Wan Z, Cheng F, Tsuta Y, Zhao X, Jiang J, Huang J, Huang Z, Oda Y, Yokota R, Arase Y, Kawahara D, Aizawa A, Kurohashi S. Leveraging High- Resource English Corpora for Cross-lingual Domain Adaptation in Low-Resource Japanese Medicine via Continued Pre-training. In: Christodoulopoulos C, Chakraborty T, Rose C, Peng V, editors. Findings of the Association for Computational Linguistics: EMNLP 2025 Suzhou, China: Association for Computational Linguistics; 2025. p. 11469–11488. doi: 10.18653/v1/2025.findings-emnlp.615 6. Jiang J, Huang J, Aizawa A. JMedBench: A Benchmark for Evaluating Japanese Biomedical Large Language Models. In: Rambow O, Wanner L, Apidianaki M, Al- Khalifa H, Eugenio BD, Schockaert S, editors. Proceedings of the 31st International Conference on Computational Linguistics Abu Dhabi, UAE: Association for Computational Linguistics; 2025. p. 5918–5935. Available from: https://aclanthology.org/2025.coling-main.395/ [accessed Mar 23, 2026] 7. Gaschi F, Fontaine X, Rastin P, Toussaint Y. Multilingual Clinical NER: Translation or Cross-lingual Transfer? In: Naumann T, Ben Abacha A, Bethard S, Roberts K, Rumshisky A, editors. Proceedings of the 5th Clinical Natural Language Processing Workshop Toronto, Canada: Association for Computational Linguistics; 2023. p. 289–311. doi: 10.18653/v1/2023.clinicalnlp-1.34 8. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med Nature Publishing Group; 2023 Aug;29(8):1930–1940. doi: 10.1038/s41591-023-02448-8 9. Zhou J, Li H, Chen S, Chen Z, Han Z, Gao X. Large language models in biomedicine and healthcare. npj Artif Intell Nature Publishing Group; 2025 Dec 1;1(1):44. doi: 10.1038/s44387-025-00047-1 10. Wang D, Zhang S. Large language models in medical and healthcare fields: applications, advances, and challenges. Artif Intell Rev 2024 Sept 20;57(11):299. doi: 10.1007/s10462-024-10921-0 11. Ray M, Kats DJ, Moorkens J, Rai D, Shaar N, Quinones D, Vermeulen A, Mateo CM, Brewster RCL, Khan A, Rader B, Brownstein JS, Hron JD. Evaluating a Large Language Model in Translating Patient Instructions to Spanish Using a Standardized Framework. JAMA Pediatr 2025 Sept 1;179(9):1026–1033. doi: 10.1001/jamapediatrics.2025.1729 12. Zaretsky J, Kim JM, Baskharoun S, Zhao Y, Austrian J, Aphinyanaphongs Y, Gupta R, Blecker SB, Feldman J. Generative Artificial Intelligence to Transform Inpatient Discharge Summaries to Patient-Friendly Language and Format. JAMA Netw Open 2024 Mar 11;7(3):e240357. PMID:38466307 13. Meddeb A, Lüken S, Busch F, Adams L, Ugga L, Koltsakis E, Tzortzakakis A, Jelassi S, Dkhil I, Klontzas ME, Triantafyllou M, Kocak B, Yüzkan S, Zhang L, Hu B, Andreychenko A, Yurievich EA, Logunova T, Morakote W, Angkurawaranon S, Makowski MR, Wattjes MP, Cuocolo R, Bressem K. Large Language Model Ability to Translate CT and MRI Free-Text Radiology Reports Into Multiple Languages. Radiology Radiological Society of North America; 2024 Dec;313(3):e241736. doi: 10.1148/radiol.241736 14. Gunn AJ, Gilcrease-Garcia B, Mangano MD, Sahani DV, Boland GW, Choy G. JOURNAL CLUB: Structured Feedback From Patients on Actual Radiology Reports: A Novel Approach to Improve Reporting Practices. AJR Am J Roentgenol 2017 June;208(6):1262–1270. PMID:28402133 15. Perlis N, Finelli A, Lovas M, Berlin A, Papadakos J, Ghai S, Bakas V, Alibhai S, Lee O, Badzynski A, Wiljer D, Lund A, Di Meo A, Cafazzo J, Haider M. Creating patient- centered radiology reports to empower patients undergoing prostate magnetic resonance imaging. Can Urol Assoc J 2021 Apr;15(4):108–113. PMID:33007175 16. Papineni K, Roukos S, Ward T, Zhu W-J. Bleu: a Method for Automatic Evaluation of Machine Translation. In: Isabelle P, Charniak E, Lin D, editors. Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics Philadelphia, Pennsylvania, USA: Association for Computational Linguistics; 2002. p. 311–318. doi: 10.3115/1073083.1073135 17. Lin C-Y. ROUGE: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out Barcelona, Spain: Association for Computational Linguistics; 2004. p. 74–81. Available from: https://aclanthology.org/W04- 1013/ [accessed Mar 23, 2026] 18. Kocak B, Klontzas ME, Stanzione A, Meddeb A, Demircioğlu A, Bluethgen C, Bressem K, Ugga L, Mercaldo N, Díaz O, Cuocolo R. Evaluation metrics in medical imaging AI: fundamentals, pitfalls, misapplications, and recommendations. European Journal of Radiology Artificial Intelligence 2025 Sept 1;3:100030. doi: 10.1016/j.ejrai.2025.100030 19. Gupta M, Aizawa A, Shah R. Med-CoDE: Medical Critique based Disagreement Evaluation Framework. In: Ebrahimi A, Haider S, Liu E, Haider S, Leonor Pacheco M, Wein S, editors. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop) Albuquerque, USA: Association for Computational Linguistics; 2025. p. 112–119. doi: 10.18653/v1/2025.naacl-srw.11 20. Gu J, Jiang X, Shi Z, Tan H, Zhai X, Xu C, Li W, Shen Y, Ma S, Liu H, Wang S, Zhang K, Lin Z, Zhang B, Ni L, Gao W, Wang Y, Guo J. A survey on LLM-as-a-Judge. The Innovation 2026 Jan 9;101253. doi: 10.1016/j.xinn.2025.101253 21. Xu J, Zhang X, Abderezaei J, Bauml J, Boodoo R, Haghighi F, Ganjizadeh A, Brattain E, Van Veen D, Meng Z, Eyre DW, Delbrouck J-B. RadEval: A framework for radiology text evaluation. In: Habernal I, Schulam P, Tiedemann J, editors. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations Suzhou, China: Association for Computational Linguistics; 2025. p. 546–557. doi: 10.18653/v1/2025.emnlp- demos.40 22. DeepSeek-AI, Liu A, Mei A, Lin B, Xue B, Wang B, Xu B, Wu B, Zhang B, Lin C, Dong C, Lu C, Zhao C, Deng C, Xu C, Ruan C, Dai D, Guo D, Yang D, Chen D, Li E, Zhou F, Lin F, Dai F, Hao G, Chen G, Li G, Zhang H, Xu H, Li H, Liang H, Wei H, Zhang H, Luo H, Ji H, Ding H, Tang H, Cao H, Gao H, Qu H, Zeng H, Huang J, Li J, Xu J, Hu J, Chen J, Xiang J, Yuan J, Cheng J, Zhu J, Ran J, Jiang J, Qiu J, Li J, Song J, Dong K, Gao K, Guan K, Huang K, Zhou K, Huang K, Yu K, Wang L, Zhang L, Wang L, Zhao L, Yin L, Guo L, Luo L, Ma L, Wang L, Zhang L, Di MS, Xu MY, Zhang M, Zhang M, Tang M, Zhou M, Huang P, Cong P, Wang P, Wang Q, Zhu Q, Li Q, Chen Q, Du Q, Xu R, Ge R, Zhang R, Pan R, Wang R, Yin R, Xu R, Shen R, Zhang R, Liu SH, Lu S, Zhou S, Chen S, Cai S, Chen S, Hu S, Liu S, Hu S, Ma S, Wang S, Yu S, Zhou S, Pan S, Zhou S, Ni T, Yun T, Pei T, Ye T, Yue T, Zeng W, Liu W, Liang W, Pang W, Luo W, Gao W, Zhang W, Gao X, Wang X, Bi X, Liu X, Wang X, Chen X, Zhang X, Nie X, Cheng X, Liu X, Xie X, Liu X, Yu X, Li X, Yang X, Li X, Chen X, Su X, Pan X, Lin X, Fu X, Wang YQ, Zhang Y, Xu Y, Ma Y, Li Y, Zhao Y, Sun Y, Wang Y, Qian Y, Yu Y, Zhang Y, Ding Y, Shi Y, Xiong Y, He Y, Zhou Y, Zhong Y, Piao Y, Wang Y, Chen Y, Tan Y, Wei Y, Ma Y, Liu Y, Yang Y, Guo Y, Wu Y, Wu Y, Cheng Y, Ou Y, Xu Y, Wang Y, Gong Y, Wu Y, Zou Y, Li Y, Xiong Y, Luo Y, You Y, Liu Y, Zhou Y, Wu ZF, Ren Z, Zhao Z, Ren Z, Sha Z, Fu Z, Xu Z, Xie Z, Zhang Z, Hao Z, Gou Z, Ma Z, Yan Z, Shao Z, Huang Z, Wu Z, Li Z, Zhang Z, Xu Z, Wang Z, Gu Z, Zhu Z, Li Z, Zhang Z, Xie Z, Gao Z, Pan Z, Yao Z, Feng B, Li H, Cai JL, Ni J, Xu L, Li M, Tian N, Chen RJ, Jin RL, Li S, Zhou S, Sun T, Li XQ, Jin X, Shen X, Chen X, Song X, Zhou X, Zhu YX, Huang Y, Li Y, Zheng Y, Zhu Y, Ma Y, Huang Z, Xu Z, Zhang Z, Ji D, Liang J, Guo J, Chen J, Xia L, Wang M, Li M, Zhang P, Chen R, Sun S, Wu S, Ye S, Wang T, Xiao WL, An W, Wang X, Sun X, Wang X, Tang Y, Zha Y, Zhang Z, Ju Z, Zhang Z, Qu Z. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. arXiv.org. 2025. Available from: https://arxiv.org/abs/2512.02556v1 [accessed Mar 23, 2026] 23. Yamagishi Y, Nakamura Y, Kikuchi T, Sonoda Y, Hirakawa H, Kano S, Nakamura S, Hanaoka S, Yoshikawa T, Abe O. Development of a Large-Scale Dataset of Chest Computed Tomography Reports in Japanese and a High-Performance Finding Classification Model: Dataset Development and Validation Study. JMIR Medical Informatics JMIR Publications Inc., Toronto, Canada; 2025 Aug 28;13(1):e71137. doi: 10.2196/71137 24. Hamamci IE, Er S, Wang C, Almas F, Simsek AG, Esirgun SN, Dogan I, Durugol OF, Hou B, Shit S, Dai W, Xu M, Reynaud H, Dasdelen MF, Wittmann B, Amiranashvili T, Simsar E, Simsar M, Erdemir EB, Alanbay A, Sekuboyina A, Lafci B, Kaplan A, Lu Z, Polacin M, Kainz B, Bluethgen C, Batmanghelich K, Ozdemir MK, Menze B. Generalist foundation models from a multimodal dataset for 3D computed tomography. Nat Biomed Eng Nature Publishing Group; 2026 Feb 12;1–19. doi: 10.1038/s41551-025-01599-y 25. Uchida T. mocobeta/janome. 2026. Available from: https://github.com/mocobeta/janome [accessed Mar 23, 2026] 26. Singh A, Fry A, Perelman A, Tart A, Ganesh A, El-Kishky A, McLaughlin A, Low A, Ostrow AJ, Ananthram A, Nathan A, Luo A, Helyar A, Madry A, Efremov A, Spyra A, Baker-Whitcomb A, Beutel A, Karpenko A, Makelov A, Neitz A, Wei A, Barr A, Kirchmeyer A, Ivanov A, Christakis A, Gillespie A, Tam A, Bennett A, Wan A, Huang A, Sandjideh AM, Yang A, Kumar A, Saraiva A, Vallone A, Gheorghe A, Garcia AG, Braunstein A, Liu A, Schmidt A, Mereskin A, Mishchenko A, Applebaum A, Rogerson A, Rajan A, Wei A, Kotha A, Srivastava A, Agrawal A, Vijayvergiya A, Tyra A, Nair A, Nayak A, Eggers B, Ji B, Hoover B, Chen B, Chen B, Barak B, Minaiev B, Hao B, Baker B, Lightcap B, McKinzie B, Wang B, Quinn B, Fioca B, Hsu B, Yang B, Yu B, Zhang B, Brenner B, Zetino CR, Raymond C, Lugaresi C, Paz C, Hudson C, Whitney C, Li C, Chen C, Cole C, Voss C, Ding C, Shen C, Huang C, Colby C, Hallacy C, Koch C, Lu C, Kaplan C, Kim C, Minott-Henriques CJ, Frey C, Yu C, Czarnecki C, Reid C, Wei C, Decareaux C, Scheau C, Zhang C, Forbes C, Tang D, Goldberg D, Roberts D, Palmie D, Kappler D, Levine D, Wright D, Leo D, Lin D, Robinson D, Grabb D, Chen D, Lim D, Salama D, Bhattacharjee D, Tsipras D, Li D, Yu D, Strouse DJ, Williams D, Hunn D, Bayes E, Arbus E, Akyurek E, Le EY, Widmann E, Yani E, Proehl E, Sert E, Cheung E, Schwartz E, Han E, Jiang E, Mitchell E, Sigler E, Wallace E, Ritter E, Kavanaugh E, Mays E, Nikishin E, Li F, Such FP, Peres F de AB, Raso F, Bekerman F, Tsimpourlas F, Chantzis F, Song F, Zhang F, Raila G, McGrath G, Briggs G, Yang G, Parascandolo G, Chabot G, Kim G, Zhao G, Valiant G, Leclerc G, Salman H, Wang H, Sheng H, Jiang H, Wang H, Jin H, Sikchi H, Schmidt H, Aspegren H, Chen H, Qiu H, Lightman H, Covert I, Kivlichan I, Silber I, Sohl I, Hammoud I, Clavera I, Lan I, Akkaya I, Kostrikov I, Kofman I, Etinger I, Singal I, Hehir J, Huh J, Pan J, Wilczynski J, Pachocki J, Lee J, Quinn J, Kiros J, Kalra J, Samaroo J, Wang J, Wolfe J, Chen J, Wang J, Harb J, Han J, Wang J, Zhao J, Chen J, Yang J, Tworek J, Chand J, Landon J, Liang J, Lin J, Liu J, Wang J, Tang J, Yin J, Jang J, Morris J, Flynn J, Ferstad J, Heidecke J, Fishbein J, Hallman J, Grant J, Chien J, Gordon J, Park J, Liss J, Kraaijeveld J, Guay J, Mo J, Lawson J, McGrath J, Vendrow J, Jiao J, Lee J, Steele J, Wang J, Mao J, Chen K, Hayashi K, Xiao K, Salahi K, Wu K, Sekhri K, Sharma K, Singhal K, Li K, Nguyen K, Gu-Lemberg K, King K, Liu K, Stone K, Yu K, Ying K, Georgiev K, Lim K, Tirumala K, Miller K, Ahmad L, Lv L, Clare L, Fauconnet L, Itow L, Yang L, Romaniuk L, Anise L, Byron L, Pathak L, Maksin L, Lo L, Ho L, Jing L, Wu L, Xiong L, Mamitsuka L, Yang L, McCallum L, Held L, Bourgeois L, Engstrom L, Kuhn L, Feuvrier L, Zhang L, Switzer L, Kondraciuk L, Kaiser L, Joglekar M, Singh M, Shah M, Stratta M, Williams M, Chen M, Sun M, Cayton M, Li M, Zhang M, Aljubeh M, Nichols M, Haines M, Schwarzer M, Gupta M, Shah M, Huang M, Dong M, Wang M, Glaese M, Carroll M, Lampe M, Malek M, Sharman M, Zhang M, Wang M, Pokrass M, Florian M, Pavlov M, Wang M, Chen M, Wang M, Feng M, Bavarian M, Lin M, Abdool M, Rohaninejad M, Soto N, Staudacher N, LaFontaine N, Marwell N, Liu N, Preston N, Turley N, Ansman N, Blades N, Pancha N, Mikhaylin N, Felix N, Handa N, Rai N, Keskar N, Brown N, Nachum O, Boiko O, Murk O, Watkins O, Gleeson O, Mishkin P, Lesiewicz P, Baltescu P, Belov P, Zhokhov P, Pronin P, Guo P, Thacker P, Liu Q, Yuan Q, Liu Q, Dias R, Puckett R, Arora R, Mullapudi RT, Gaon R, Miyara R, Song R, Aggarwal R, Marsan RJ, Yemiru R, Xiong R, Kshirsagar R, Nuttall R, Tsiupa R, Eldan R, Wang R, James R, Ziv R, Shu R, Nigmatullin R, Jain S, Talaie S, Altman S, Arnesen S, Toizer S, Toyer S, Miserendino S, Agarwal S, Yoo S, Heon S, Ethersmith S, Grove S, Taylor S, Bubeck S, Banesiu S, Amdo S, Zhao S, Wu S, Santurkar S, Zhao S, Chaudhuri SR, Krishnaswamy S, Shuaiqi, Xia, Cheng S, Anadkat S, Fishman SP, Tobin S, Fu S, Jain S, Mei S, Egoian S, Kim S, Golden S, Mah SQ, Lin S, Imm S, Sharpe S, Yadlowsky S, Choudhry S, Eum S, Sanjeev S, Khan T, Stramer T, Wang T, Xin T, Gogineni T, Christianson T, Sanders T, Patwardhan T, Degry T, Shadwell T, Fu T, Gao T, Garipov T, Sriskandarajah T, Sherbakov T, Kaftan T, Hiratsuka T, Wang T, Song T, Zhao T, Peterson T, Kharitonov V, Chernova V, Kosaraju V, Kuo V, Pong V, Verma V, Petrov V, Jiang W, Zhang W, Zhou W, Xie W, Zhan W, McCabe W, DePue W, Ellsworth W, Bain W, Thompson W, Chen X, Qi X, Xiang X, Shi X, Dubois Y, Yu Y, Khakbaz Y, Wu Y, Qian Y, Lee YT, Chen Y, Zhang Y, Xiong Y, Tian Y, Cha Y, Bai Y, Yang Y, Yuan Y, Li Y, Zhang Y, Yang Y, Jin Y, Jiang Y, Wang Y, Wang Y, Liu Y, Stubenvoll Z, Dou Z, Wu Z, Wang Z. OpenAI GPT-5 System Card. arXiv; 2025. doi: 10.48550/arXiv.2601.03267 27. Liu AH, Khandelwal K, Subramanian S, Jouault V, Rastogi A, Sadé A, Jeffares A, Jiang A, Cahill A, Gavaudan A, Sablayrolles A, Héliou A, You A, Ehrenberg A, Lo A, Eliseev A, Calvi A, Sooriyarachchi A, Bout B, Rozière B, Monicault BD, Lanfranchi C, Barreau C, Courtot C, Grattarola D, Dabert D, Casas D de las, Chane-Sane E, Ahmed F, Berrada G, Ecrepont G, Guinet G, Novikov G, Kunsch G, Lample G, Martin G, Gupta G, Ludziejewski J, Rute J, Studnia J, Amar J, Delas J, Roberts JS, Yadav K, Chandu K, Jain K, Aitchison L, Fainsin L, Blier L, Zhao L, Martin L, Saulnier L, Gao L, Buyl M, Jennings M, Pellat M, Prins M, Poirée M, Guillaumin M, Dinot M, Futeral M, Darrin M, Augustin M, Chiquier M, Schimpf M, Grinsztajn N, Gupta N, Raghuraman N, Bousquet O, Duchenne O, Wang P, Platen P von, Jacob P, Wambergue P, Kurylowicz P, Muddireddy PR, Chagniot P, Stock P, Agrawal P, Torroba Q, Sauvestre R, Soletskyi R, Menneer R, Vaze S, Barry S, Gandhi S, Waghjale S, Gandhi S, Ghosh S, Mishra S, Aithal S, Antoniak S, Scao TL, Cachet T, Sorg TS, Lavril T, Saada TN, Chabal T, Foubert T, Robert T, Wang T, Lawson T, Bewley T, Bewley T, Edwards T, Jamil U, Tomasini U, Nemychnikova V, Phung V, Maladière V, Richard V, Bouaziz W, Li W-D, Marshall W, Li X, Yang X, Ouahidi YE, Wang Y, Tang Y, Ramzi Z. Ministral 3. arXiv; 2026. doi: 10.48550/arXiv.2601.08584 28. Audi S, Pencharz D, Wagner T. Behind the hedges: how to convey uncertainty in imaging reports. Clin Radiol 2021 Feb;76(2):84–87. PMID:32883516 29. Bruno MA, Petscavage-Thomas J, Abujudeh H. Communicating Uncertainty in the Radiology Report. AJR Am J Roentgenol 2017 Nov;209(5):1006–1008. PMID:28705061 30. Blankemeier L, Kumar A, Cohen JP, Liu J, Liu L, Van Veen D, Gardezi SJS, Yu H, Paschali M, Chen Z, Delbrouck J-B, Reis E, Holland R, Truyts C, Bluethgen C, Wu Y, Lian L, Jensen MEK, Ostmeier S, Varma M, Valanarasu JMJ, Fang Z, Huo Z, Nabulsi Z, Ardila D, Weng W-H, Junior EA, Ahuja N, Fries J, Shah NH, Zaharchuk G, Willis M, Yala A, Johnston A, Boutin RD, Wentland A, Langlotz CP, Hom J, Gatidis S, Chaudhari AS. Merlin: a computed tomography vision–language foundation model and dataset. Nature Nature Publishing Group; 2026 Mar 4;1–11. doi: 10.1038/s41586-026-10181-8 31. Lee S, Youn J, Kim H, Kim M, Yoon SH. CXR-LLaVA: a multimodal large language model for interpreting chest X-ray images. Eur Radiol 2025 July 1;35(7):4374– 4386. doi: 10.1007/s00330-024-11339-6 32. Agrawal K, Liu L, Lian L, Nercessian M, Harguindeguy N, Wu Y, Mikhael P, Lin G, Sequist LV, Fintelmann F, Darrell T, Bai Y, Chung M, Yala A. Pillar-0: A New Frontier for Radiology Foundation Models. arXiv; 2025. doi: 10.48550/arXiv.2511.17803 33. Wataoka K, Takahashi T, Ri R. Self-Preference Bias in LLM-as-a-Judge. arXiv; 2025. doi: 10.48550/arXiv.2410.21819 34. Chen GH, Chen S, Liu Z, Jiang F, Wang B. Humans or LLMs as the Judge? A Study on Judgement Bias. In: Al-Onaizan Y, Bansal M, Chen Y-N, editors. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Miami, Florida, USA: Association for Computational Linguistics; 2024. p. 8301– 8327. doi: 10.18653/v1/2024.emnlp-main.474 35. Zhou H, Huang H, Long Y, Xu B, Zhu C, Cao H, Yang M, Zhao T. Mitigating the Bias of Large Language Model Evaluation. In: Maosong S, Jiye L, Xianpei H, Zhiyuan L, Yulan H, editors. Proceedings of the 23rd Chinese National Conference on Computational Linguistics (Volume 1: Main Conference) Taiyuan, China: Chinese Information Processing Society of China; 2024. p. 1310–1319. Available from: https://aclanthology.org/2024.ccl-1.101/ [accessed Mar 24, 2026] 36. Ouyang L, Wu J, Jiang X, Almeida D, Wainwright CL, Mishkin P, Zhang C, Agarwal S, Slama K, Ray A, Schulman J, Hilton J, Kelton F, Miller L, Simens M, Askell A, Welinder P, Christiano P, Leike J, Lowe R. Training language models to follow instructions with human feedback. 37. Martindale M, Carpuat M. Fluency Over Adequacy: A Pilot Study in Measuring User Trust in Imperfect MT. In: Cherry C, Neubig G, editors. Proceedings of the 13th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track) Boston, MA: Association for Machine Translation in the Americas; 2018. p. 13–25. Available from: https://aclanthology.org/W18- 1803/ [accessed Mar 24, 2026] 38. YYama0/CT-RATE-JPN · Datasets at Hugging Face. 2025. Available from: https://huggingface.co/datasets/YYama0/CT-RATE-JPN [accessed Mar 23, 2026] 39. ibrahimhamamci/CT-RATE · Datasets at Hugging Face. 2025. Available from: https://huggingface.co/datasets/ibrahimhamamci/CT-RATE [accessed Mar 23, 2026] Supplementary Material Translation Prompt for DeepSeek-V3.2 Translations were generated using the following system prompt at temperature 0. The model received the English report text as the user message, with no additional instructions. System prompt (Japanese original) あなたは放射線科の読影レポートを翻訳する専門家です。 以下の英語の読影レポートを、日本語として自然で臨床的に正確な表現に翻訳してくださ い。 制約・注意: - 医学的意味を変えない - 放射線科で一般的な日本語表現を用いる - コロンやセミコロンは使用せずに、「。」や「、」を使うようにする - 「断層内」という表現を用いず、「撮像範囲内」や「検査範囲内」とする - 造影剤を使用していない場合、「非造影」と訳す - 「CTO」は「Cardiothoracic outline」のことで心臓のサイズや構造に言及している - 「maximal physiological limit」は「生理的上限」と訳す - 「natural」は「正常」と訳す - 「lung parenchyma window」は「肺野条件」と訳す - 「consolidation」はそのま「コンソリデーション」と訳す System prompt (English translation) You are an expert translator specializing in radiology reports. Translate the following English radiology report into natural, clinically accurate Japanese. Constraints and notes: - Do not alter the medical meaning - Use standard Japanese radiology terminology - Do not use colons or semicolons; use "。" and "、" instead - Do not use the expression "断層内" (within the tomographic slice); use "撮像範囲内" or "検査 範囲内" (within the imaging range) instead - If no contrast agent was used, translate as "非造影" (non-contrast) - "CTO" refers to "Cardiothoracic outline" and pertains to cardiac size and structure - Translate "maximal physiological limit" as "生理的上限" - Translate "natural" as "正常" (normal) - Translate "lung parenchyma window" as "肺野条件" - Translate "consolidation" as "コンソリデーション" (untranslated loan word) LLM Judge Prompt All three LLM judges (DeepSeek-V3.2, Mistral Large 3, and GPT-5) used the following system prompt and user prompt template at temperature 0, with JSON response format enforced. For each report pair, the two Japanese translations were assigned to positions A and B in randomized order; the assignment was recorded to allow post hoc normalization. System prompt (Japanese original) あなたは放射線科読影レポート翻訳の評価者です。 与えられた英語原文と2つの日本語訳(A/B)を比較し、どちらが優れているか、または引 き分けかを判定してください。 重要: - 英語原文への忠実性と医療安全性を最重視する - Findings と Impressions は結合された文章として扱う 出力は厳密なJSONのみ。 System prompt (English translation) You are an evaluator of radiology report translations. Compare the given English source text and two Japanese translations (A/B), and judge which is superior or whether they are equivalent. Important: - Prioritize fidelity to the English source and medical safety above all - Treat Findings and Impressions as a combined text Output strictly in JSON format only. User prompt template (Japanese original) 以下を評価してください。 # English (source) en_text # Japanese A ja_a # Japanese B ja_b 評価項目: 1) 医療用語の正確さ 2) 日本語としての読みやすさ 3) 臨床でそのま使用するとしたら、どちらがより適切か(総合) 4) 文体や表現の点で、どちらがより放射線科医が書いたレポートらしいか 出力要件: - 1)〜4) の winner をすべて "A" / "B" / "TIE" から選ぶ - 各項目に短い理由(日本語、1文程度)を付ける - 出力は以下のJSON形式(キー名を厳守) JSON形式: "term_accuracy": "winner": "A|B|TIE", "reason": "短い理由", "readability": "winner": "A|B|TIE", "reason": "短い理由", "overall": "winner": "A|B|TIE", "reason": "短い理由", "radiologist_like": "winner": "A|B|TIE", "reason": "短い理由" User prompt template (English translation) Please evaluate the following. # English (source) en_text # Japanese A ja_a # Japanese B ja_b Evaluation criteria: 1) Accuracy of medical terminology 2) Readability and fluency in Japanese 3) Overall clinical suitability (which would be more appropriate for direct clinical use) 4) Radiologist-style authenticity (which more closely resembles a report written by a radiologist in terms of style and expression) Output requirements: - Select the winner for criteria 1)–4) from "A", "B", or "TIE" - Provide a brief reason for each criterion (approximately one sentence in Japanese) - Output strictly in the following JSON format (key names must match exactly) JSON format: "term_accuracy": "winner": "A|B|TIE", "reason": "brief reason", "readability": "winner": "A|B|TIE", "reason": "brief reason", "overall": "winner": "A|B|TIE", "reason": "brief reason", "radiologist_like": "winner": "A|B|TIE", "reason": "brief reason"