Paper deep dive
Can "AI" Be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs
Mariano Barone, Francesco Di Serio, Roberto Moio, Marco Postiglione, Giuseppe Riccio, Antonio Romano, Vincenzo Moscato
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/26/2026, 6:46:01 PM
Summary
This study evaluates the communicative alignment of Large Language Models (LLMs) in healthcare, focusing on semantic fidelity, readability, and affective resonance. By comparing general-purpose models (GPT-5, Claude, Mixtral) and domain-specialized models (Med-PaLM) against physician-authored responses from the MedQuAD and iCliniqQAs datasets, the researchers found that baseline LLMs often exhibit higher linguistic complexity and extreme affective polarity compared to human physicians. While empathy-oriented prompting improves emotional tone and reduces complexity, 'Collaborative Rewriting' (Rephrase Prompting) yielded the strongest alignment with physician-authored content. The study concludes that LLMs are most effective as collaborative communication enhancers rather than direct replacements for clinical expertise.
Entities (10)
Relation Signals (4)
MedQuAD → usedforevaluationof → Large Language Models
confidence 100% · To conduct our experiments, we used the MedQuAD dataset
Rephrase Prompting → improves → Semantic Similarity
confidence 95% · Rephrase configurations achieve the highest semantic similarity to physician answers (up to mean = 0.93)
GPT-5 → exhibitshighcomplexity → Flesch-Kincaid Grade Level
confidence 90% · in larger architectures such as GPT-5 and Claude, produce substantially higher linguistic complexity (FKGL up to 16.91-17.60)
Empathy Prompting → reduces → Linguistic Complexity
confidence 90% · Empathy-oriented prompting reduces extreme negativity and lowers grade-level complexity
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) are increasingly deployed in healthcare, yet their communicative alignment with clinical standards remains insufficiently quantified. We conduct a multidimensional evaluation of general-purpose and domain-specialized LLMs across structured medical explanations and real-world physician-patient interactions, analyzing semantic fidelity, readability, and affective resonance. Baseline models amplify affective polarity relative to physicians (Very Negative: 43.14-45.10% vs. 37.25%) and, in larger architectures such as GPT-5 and Claude, produce substantially higher linguistic complexity (FKGL up to 16.91-17.60 vs. 11.47-12.50 in physician-authored responses). Empathy-oriented prompting reduces extreme negativity and lowers grade-level complexity (up to -6.87 FKGL points for GPT-5) but does not significantly increase semantic fidelity. Collaborative rewriting yields the strongest overall alignment. Rephrase configurations achieve the highest semantic similarity to physician answers (up to mean = 0.93) while consistently improving readability and reducing affective extremity. Dual stakeholder evaluation shows that no model surpasses physicians on epistemic criteria, whereas patients consistently prefer rewritten variants for clarity and emotional tone. These findings suggest that LLMs function most effectively as collaborative communication enhancers rather than replacements for clinical expertise.
Tags
Links
- Source: https://arxiv.org/abs/2604.20791v1
- Canonical: https://arxiv.org/abs/2604.20791v1
Trouble viewing inline? Open PDF directly →
Full Text
108,448 characters extracted from source content.
Expand or collapse full text
Can “AI” be a Doctor? A Study of Empathy, Readability, and Alignment in Clinical LLMs Mariano Barone 1* , Francesco Di Serio 1† , Roberto Moio 2† , Marco Postiglione 3† , Giuseppe Riccio 1† , Antonio Romano 1† , Vincenzo Moscato 1† 1* Department of Electrical Engineering and Information Technology, University of Naples Federico I, Via Claudio 21, Naples, 80125, Italy. 2 Department of Translational Medical Sciences, University of Campania ”Luigi Vanvitelli”, Via Leonardo Bianchi, Naples, 80131, Italy. 3 Department of Computer Science, McCormick School of Engineering and Applied Science, Northwestern University, 2309 Sheridan Rd, Evanston, 60201, IL, United States. *Corresponding author(s). E-mail(s): mariano.barone@unina.it; Contributing authors: francesco.diserio@unina.it; roberto.moio@unicampania.it; marco.postiglione@northwestern.edu; giuseppe.riccio3@unina.it; antonio.romano5@unina.it; vmoscato@unina.it; † These authors contributed equally to this work. Abstract Large Language Models (LLMs) are increasingly deployed in healthcare, yet their communicative alignment with clinical standards remains insufficiently quantified. We conduct a multidimensional evaluation of general-purpose and domain-specialized LLMs across structured medical explanations and real-world physician–patient interactions, analyzing semantic fidelity, readability, and affec- tive resonance. Baseline models amplify affective polarity relative to physicians (Very Negative: 43.14–45.10% vs. 37.25%) and, in larger architectures such as GPT-5 and Claude, produce substantially higher linguistic complexity (FKGL up to 16.91–17.60 vs. 11.47–12.50 in physician-authored responses). Empathy- oriented prompting reduces extreme negativity and lowers grade-level complexity (up to -6.87 FKGL points for GPT-5) but does not significantly increase semantic fidelity. Collaborative rewriting yields the strongest overall alignment. Rephrase configurations achieve the highest semantic similarity to physician answers (up 1 arXiv:2604.20791v1 [cs.CL] 22 Apr 2026 to μ = 0.93) while consistently improving readability and reducing affective extremity. Dual stakeholder evaluation shows that no model surpasses physicians on epistemic criteria, whereas patients consistently prefer rewritten variants for clarity and emotional tone. These findings suggest that LLMs function most effectively as collaborative communication enhancers rather than replacements for clinical expertise. Keywords: Large Language Models, Healthcare AI, Empathy in AI, Readability, Medical Communication, MedQuAD, Medical Question-Answering 1 Introduction As AI systems become increasingly capable and autonomous, ensuring their alignment with human values has become a practical concern rather than a purely theoretical one [1, 2]. In patient-facing medical applications, misalignment may result in unclear communication, inappropriate reassurance, or unsafe recommendations, with direct consequences for patient trust and clinical decision-making. This issue is particularly critical in healthcare, where Large Language Models (LLMs) are increasingly deployed in patient-facing settings, often in contexts characterized by vulnerability and emo- tional distress[3]. Recent evidence shows that over 150,000 clinicians across more than 150 institutions already rely on AI-powered systems to assist with patient messaging 1 , thus reshaping everyday clinical communication practices [4–6]. While existing research has largely focused on the factual accuracy of medical AI systems, particularly addressing hallucinations and reliability [7, 8], other determi- nants of effective clinical communication remain comparatively underexplored. In this work, we focus on three communicative dimensions that are central in clinical interac- tions: semantic correctness (preservation of medical meaning), readability (linguistic accessibility to non-expert users), and affective appropriateness (alignment of emo- tional tone with patient needs). In the absence of these values, patients may experience confusion, anxiety, and erosion of trust in their healthcare providers [9–11]. Despite the existence of well-established communication frameworks such as SPIKES [12] and the Calgary-Cambridge Guide [13], little empirical work has investigated whether modern LLMs reproduce or deviate from these communicative principles when interacting with patients[14]. We address this gap through a systematic evaluation of three communica- tive dimensions that are particularly relevant in clinical contexts: (1) semantic fidelity, how faithfully AI responses match expert clinical judgment; (2) readability, whether AI answers remain comprehensible across diverse literacy levels and cultural back- grounds; and (3) affective resonance, the extent to which AI responses acknowledge patients’ emotional needs. To this end, we conduct our analysis on a large-scale med- ical question-answering corpus comprising 47,457 entries derived from authoritative healthcare sources [15], and investigate the following research questions: 1 https://w.nytimes.com/2024/09/24/health/ai-patient-messages-mychart.html 2 1. RQ1 (Empathy): How do LLMs compare to human physicians in expressing empathy and emotional awareness in clinical communication? 2. RQ2 (Readability): Do LLM-generated responses differ from physician- authored answers in terms of linguistic readability? 3. RQ3 (Prompt-based Alignment): Does the use of empathy-oriented prompt- ing improve the emotional tone and readability of LLM outputs while preserving semantic fidelity? 4. RQ4 (Human-AI Collaboration): Can LLMs improve the clarity and emo- tional appropriateness of physician-authored responses through collaborative rewriting? 5. RQ5 (Expert-Patient Value Alignment): To what extent do different LLM configurations satisfy the distinct preferences expressed by medical experts and patients? Our study addresses these questions through one of the first large-scale empirical comparisons between LLM-generated and physician-authored medical communication. We identify systematic differences between models and human clinicians in emo- tional tone, readability, and semantic consistency, with no single approach consistently outperforming the others across all dimensions. A central contribution of this work lies in the nature of the data examined. In contrast to prior studies that predominantly rely on social media content, patient self- reports, or synthetic dialogues, we analyze expert-authored medical communication produced by practicing clinicians in institutional settings. This enables an evaluation of alignment against high-standard clinical language as used in real healthcare workflows, rather than informal user-generated text. To the best of our knowledge, such data have been rarely leveraged in large-scale evaluations of LLMs for healthcare communication. 2 Related Work Empathy is widely recognized as a central component of effective clinical communi- cation. In medical contexts, it is not merely a matter of emotional warmth but of calibrated emotional engagement, often described as detached concern [16]. Recent studies suggest that LLMs can generate emotionally resonant medical responses, although the empirical findings remain mixed. Ayers et al. [17] reported that 79% of Reddit AskDocs users preferred ChatGPT responses over those written by physi- cians, largely due to perceived empathy and tone. However, the informal context of online forums limits the generalizability of these findings to real clinical practice. More recently, analogous findings have been reported in oncological settings, where patients consistently rated chatbot responses as more empathetic than physician responses, further highlighting the divergence between lay and clinical perceptions of affective tone[18]. Similarly, Luo et al. [19] proposed EMRank as a metric for quantifying empa- thy in LLM responses, reporting higher empathy scores for ChatGPT compared to physicians, though their analysis focused primarily on emotional expression rather than overall clinical adequacy. 3 Beyond emotional tone, readability and linguistic simplification represent further challenges in patient–provider communication. Roy et al. [20] showed that GPT- 5 can improve comprehension of medical information; similar benefits have been demonstrated when AI is used to simplify surgical consent forms through human-AI collaborative approaches [21], yet excessive simplification may reduce interpretability in complex clinical scenarios. A primary mechanism through which these dimensions are implicitly adjusted is prompt engineering. Prompt engineering has emerged as a key modulator of LLM behavior in medical settings. Prior work shows that techniques such as chain-of- thought prompting improve reasoning transparency and task performance [22], while role conditioning can increase perceived empathy in generated responses. However, few studies have operationalized established clinical communication frameworks such as SPIKES [12] or the Calgary–Cambridge Guide [13] within prompt templates, limiting the transferability of these findings to structured clinical environments. These variations in prompt design not only affect how models generate responses, but also complicate direct comparisons with physician-authored communication. Sev- eral meta-analyses report that LLMs offer broader informational coverage, particularly in identifying symptoms, potential diagnoses, and treatment options, though concerns remain regarding precision and contextual appropriateness [23]. Model performance varies substantially across tasks, prompting strategies, and domain settings [24]. In addition, commonly used benchmarks often rely on public forums or synthetic datasets, which lack domain realism and rarely include parallel expert-authored con- tent. These limitations have motivated a shift toward evaluation frameworks that better reflect clinical stakeholders and real-world deployment conditions. [25–27]. Nev- ertheless, existing implementations remain limited in scale and scope, with relatively few studies integrating emotional, linguistic, and semantic dimensions within a unified evaluation framework[28]. Overall, the literature reveals persistent methodological limitations in assessing LLMs for clinical communication. Most evaluations adopt isolated or unidimensional perspectives, examining empathy, readability, or correctness separately rather than in combination. Many studies rely on small-scale human evaluations without rigorous inter-rater validation, which reduces statistical robustness and reproducibility. More- over, limited control over semantic preservation complicates interpretation of whether improvements reflect genuine clinical adequacy or superficial stylistic variation. Our work seeks to address these gaps through a multidimensional evaluation framework that jointly examines emotional tone, readability, and semantic fidelity across mul- tiple LLM configurations. In addition, we assess both AI-generated responses and LLM-assisted rewriting of physician-authored content, combining simulated expert assessment with human patient evaluation to capture distinct stakeholder perspectives. 3 Methodology This section introduces a multidimensional evaluation framework designed to assess the case study across three core communicative dimensions: semantic fidelity, read- ability, and affective resonance (sentiment and empathy). The framework supports 4 "How can botulism be prevented?" Question Answer Generation PHASE Semantic Fidelity Analysis Readability Analysis Empathy and Sentiment Analysis Final comparison Human vs LLM Base Answer vs Empathy Answer LLM Rephrased Answer vs Physician base answer Base Answer Empathy Answer Base Prompt AI Empathy Prompt Rephrase Answer Rephrase Prompt Doctors Authoritative Response Evaluation PHASE Sentiment Distribution Emotion Distribution Fig. 1: The framework compares LLM-generated and physician-authored answers across semantic similarity, readability, sentiment, and emotion. It includes both direct generation and LLM-based revision of expert responses, enabling evaluation of AI models as autonomous communicators and collaborative assistants in clinical settings. both the autonomous generation of responses by LLMs and the collaborative revi- sion of expert-authored responses, simulating hybrid human–AI interaction scenarios. Figure 1 illustrates the overall pipeline. 3.1 Background Each communicative dimension is conceptually defined and formally operationalized using established computational metrics, as detailed below. Definition 1 (Semantic Fidelity) Let r h be the human-authored response and r m the model- generated response. Let φ : T → R d be a sentence embedding function mapping a text sequence into a d-dimensional semantic space. We define semantic fidelity as SF(r h ,r m ) = cos φ(r h ),φ(r m ) , where cosine similarity quantifies the conceptual proximity between the embeddings. Higher values indicate stronger alignment in global meaning. Semantic fidelity evaluates the degree to which model-generated responses preserve the conceptual content expressed by clinicians. We operationalize this dimension using cosine similarity computed over embeddings produced by the BioBERT-mnli-snli-scinli-scitail-mednli-stsb encoder. Because this metric does not capture fine-grained factual inaccuracies (e.g., incorrect dosages or omitted clinical entities), we interpret it as a measure of conceptual fidelity, supplemented by domain-specific analyses reported in Section 5. 5 Definition 2 (Readability) Let r be a text consisting of W words, S sentences, Sy syllables, and C complex words. We define readability as the vector Read(r) = FKGL(r), GFI(r) , where the Flesch–Kincaid Grade Level (FKGL) [29] is FKGL(r) = 0.39× W S + 11.8× Sy W − 15.59, and the Gunning Fog Index (GFI) [30] is GFI(r) = 0.4× W S + 100× C W . Lower values correspond to more linguistically accessible text. In the definition above, FKGL estimates the U.S. school grade level required for comprehension, while GFI estimates the years of formal education needed to under- stand a text on first reading. These metrics are widely used in health communication research because they capture syntactic complexity, lexical difficulty, and overall patient-facing accessibility. Definition 3 (Affective Resonance) Let r be a textual response. Let σ(r) be a sentiment clas- sification function mapping r to Very Negative, Negative, Neutral, Positive, Very Positive, and let ε(r) be an emotion classifier that outputs a probability distribution over a set of emotions E . We define affective resonance as AR(r) = σ(r), ε(r) , representing the affective polarity and fine-grained emotional profile of the response. Affective resonance quantifies the emotional characteristics of a text, provid- ing a structured approximation of empathetic tone. We operationalize it using two complementary affective signals: • SentimentClassification:The tabularisai/robust-sentiment-analysis model [31] assigns each response to one of five sentiment classes, offering a coarse-grained measure of affective polarity. • Emotion Classification: The SamLowe/roberta-base-goemotions model [32] provides probabilities over 28 fine-grained emotions, including caring, a signal associated with supportive or empathetic tone. We compare emotion distributions across human and model responses via contingency analysis. 3.2 Dataset To conduct our experiments, we used the MedQuAD dataset [15, 33], publicly available through the Hugging Face Hub 2 . MedQuAD contains 47,457 question–answer pairs extracted from 12 authoritative NIH websites, including MedlinePlus, cancer.gov, and 2 https://hf.co/datasets/keivalya/MedQuad-MedicalQnADataset 6 niddk.nih.gov 3 . Each entry includes metadata such as UMLS Concept Unique Identi- fiers (CUIs), semantic types, question focus (e.g., disease, drug, test), and topic type (e.g., treatment, side effects, diagnosis). The dataset is distributed in XML format with structured tags for questions, answers, and metadata. For this study, we selected a subset of 16,400 QA pairs covering 37 question types. We excluded three sections of the original corpus: A.D.A.M. Medical Encyclopedia, MedlinePlus Drugs, and MedlinePlus Herbal Supplements. These sections account for approximately 31,000 entries. They follow editorial standards that differ from the core NIH sources and show substantial variation in writing style, clinical depth, and structure. Their exclusion ensures consistency in tone, source reliability, and clinical framing. This controlled subset enables fair comparisons across models. The resulting MedQuAD subset is suitable for evaluating patient-facing language models. It primarily includes symptom-, treatment-, and diagnosis-oriented ques- tions authored by NIH experts. Each entry contains a question that reflects common medical concerns and a corresponding answer derived from expert-curated institu- tional content. The dataset provides standardized and clinically grounded medical explanations. In addition to MedQuAD, we used the iCliniqQAs subset from the medical- question-answer-data repository 4 . This subset contains 465 real-world physician– patient question–answer pairs collected from an online medical consultation platform. The questions are written by patients and reflect spontaneous descriptions of symp- toms, concerns, and contextual information. The answers are authored by licensed physicians and follow a conversational clinical style. The inclusion of iCliniqQAs introduces naturally occurring medical dialogue into our evaluation setting. Unlike MedQuAD, which provides institutionally curated explanations, iCliniqQAs captures authentic patient concerns and real consultation dynamics. The combination of these datasets allows us to evaluate model behavior across both standardized medical communication and real-world clinical interaction scenarios. 3.3 Models Evaluated We selected multiple large language models (LLMs) to capture different design philosophies and degrees of domain specialization. Mixtral [34] represents a general- purpose model trained on diverse conversational and web-scale corpora without explicit biomedical fine-tuning. Conversely, Med-PaLM [35] is a domain-adapted model optimized for clinical reasoning and medical question answering through supervised instruction on biomedical literature and expert-annotated data. This contrast enables a controlled investigation of how domain specialization influences the communicative quality of medical responses. To further test whether observed trends generalize across distinct architectures and training paradigms, we evaluated GPT-5 5 , Gemini 2.5 Pro [36], and Claude Sonnet 4.5 6 . 3 https://medlineplus.gov/ — https://cancer.gov/ — https://niddk.nih.gov/ 4 https://github.com/LasseRegin/medical-question-answer-data 5 https://openai.com/index/introducing-gpt-5/ 6 https://w.anthropic.com/claude/sonnet 7 4681012141618 FleschKincaid 6 8 10 12 14 16 18 20 Gunning Fog Clustering-based representative selection (50 clusters) All physician answers Cluster representatives (50) Fig. 2: Readability-based selection of 50 representative MedQuAD questions after outlier removal. The corpus is visualized in the Gunning Fog vs. Flesch–Kincaid space. Extreme values were excluded using the interquartile range (IQR) criterion prior to clustering. The selected representatives correspond to the centroids of each k-means cluster (k = 50). The resulting subset spans the full readability distribution of the cleaned corpus while avoiding anomalous texts that could distort clustering geometry. MedQuAD Subset Construction Due to the high inference cost of commercial models, evaluation was conducted on a reduced subset of 50 MedQuAD questions. This subset reflects explicit budget constraints and follows a structured selection protocol designed to preserve linguis- tic diversity. The experiment is framed as a controlled cross-architecture comparison rather than a population-scale benchmark. The subset was constructed through a readability-driven clustering procedure: 1. Extract linguistic features: For each question in dataset D, compute the Flesch–Kincaid Grade Level (FKGL), the Gunning Fog Index (GFI), lexical repre- sentativeness (cosine similarity between TF–IDF vectors and the corpus centroid), and answer length. 2. Remove extreme outliers: Apply interquartile range (IQR) filtering indepen- dently to FKGL and GFI. 8 46810121416 FleschKincaid 6 8 10 12 14 16 18 20 Gunning Fog Severity 1 (White) Severity 2 (Green) Severity 3 (Yellow) Severity 4 (Orange) Severity 5 (Red) Fig. 3: Severity-aware selection of 50 representative iCliniqQAs samples. The corpus is visualized in the Gunning Fog vs. Flesch–Kincaid space. Samples are stratified into five clinical severity levels (White, Green, Yellow, Orange, Red). Ten representatives are selected per severity class after clustering in the normalized linguistic feature space. The resulting subset preserves both urgency distribution and readability variability. 3. Normalize features: Apply z-score normalization: z = x− μ σ 4. Cluster by linguistic properties: Perform k-means clustering with k = 50 in the normalized feature space (FKGL, GFI, lexicalrepr,|answer|). 5. Select representatives: Select the question closest to each cluster centroid. 6. Aggregate subset: Collect selected samples into D MedQuAD 50 . Figure 2 confirms that the retained subset spans the full readability range of the corpus. 9 iCliniqQAs Subset Construction For iCliniqQAs, subset construction combined linguistic stratification with clinical severity balancing. The goal was to ensure representation across urgency levels while maintaining variability in linguistic complexity. 1. Extract linguistic features: Compute FKGL, GFI, lexical representativeness, and answer length for each sample. 2. Normalize features: Apply z-score normalization to all linguistic features. 3. Severity labeling: Assign each question to one of five triage levels (White, Green, Yellow, Orange, Red). Labels were generated using PalMed-2 to ensure medically coherent classification. 4. Stratify by severity: Partition the dataset into five severity groups. 5. Cluster within each group: Perform clustering in the normalized linguistic feature space within each severity class. 6. Select balanced representatives: Select 10 samples per severity class based on centroid proximity. 7. Aggregate subset: Combine selected samples into D iCliniq 50 . Figure 3 shows that the resulting subset preserves the full range of clinical urgency while maintaining diversity in readability and lexical density. Both subsets support controlled cross-model comparison under resource con- straints. The reduced sample size limits statistical generalization but maximizes coverage across linguistic and clinical dimensions. 3.3.1 Prompting strategies To evaluate how expertise and communication style affect outputs, we tested each model under three distinct response-generation settings: • Base Prompt - Clinical Baseline Mode (Appendix A.1): this label empha- sizes that the prompt represents the model’s default, unconditioned clinical behavior, serving as a neutral reference point for all comparisons. • Empathy Prompt - Empathy-Driven Generation (Appendix A.2): this name highlights that the model is explicitly instructed to generate responses with enhanced emotional awareness and patient-centered tone, framing the prompt as an affective alignment strategy rather than a mere style change. • Rephrase Prompt - AI-Assisted Clinical Editing (Appendix A.3): this formulation clarifies that the model operates as a collaborative editor, reframing its role from content generator to clinical communication enhancer[37]. Together, these three configurations isolate complementary aspects of communica- tive alignment: the Base Prompt captures factual generation, the Empathy Prompt evaluates stylistic modulation during autonomous generation, and the Rephrase Prompt measures the model’s capacity to enhance existing human-authored content through collaborative refinement. 10 4 Experiments and Results In this section, we present the experimental setup, describe the evaluation procedures, and report the results for each research question (RQ). Before addressing the RQs individually, we first evaluate semantic fidelity across systems to establish a baseline understanding of how closely LLM-generated responses align with physician-written content. 4.1 Experimental Setup Our experimental setup was structured in two main phases. 4.1.1 Phase 1 – Model Comparison In the first phase of the study, we generated four responses for each question in the MedQuAD and iCliniqQAs subsets by combining two language models - Mixtral and Med-PaLM 2 - with two prompting strategies: the Base Prompt and the Empathy Prompt. This setup yielded four distinct outputs, representing general-purpose and medical-domain generations under both standard and empathy-enhanced conditions. Each output was then systematically compared with the physician-authored reference answer, resulting in a structured five-way evaluation for every question. To further examine the generalizability of the observed trends, we extended the evaluation to additional architectures-GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4.5-using a representative subset of 50 questions selected through a readability-based clustering procedure. For Gemini 2.5 Pro, however, a complete evaluation was not feasible: the model frequently produced limited or incomplete answers when contextual information was insufficient, underscoring its reliance on external context to generate medically grounded responses. This behavior also reflected an ethical safeguard, as the model tended to refrain from producing potentially unreliable clinical information in the absence of adequate medical context. 4.1.2 Phase 2 – Physician Answer Rephrase To investigate whether large language models can assist or refine physician- authored responses, we employed the same set of models as in the previous experiments: Mixtral, Med-PaLM 2, GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4.5, to rewrite each original medical answer. The resulting generations are denoted as ModelRephrase, following a consistent naming convention across models (e.g., MixtralRephrase, Med-PaLMRephrase, etc.). Using a dedicated rewriting prompt, each model was instructed to enhance the emotional tone, clarity, and accessibility of the physician’s message while preserving its medical accuracy and factual consistency. This phase simulated a human–AI co-authoring process distinct from the Base Prompt and Empathy Prompt (Empathy Prompt) configurations used in Phase 1, emphasizing collaborative refinement rather than autonomous response generation. All outputs were generated with low-temperature sampling (temperature = 0.1) to limit stochastic variation and promote consistency across runs. 11 4.2 Preliminary Evaluation – Semantic Fidelity Semantic fidelity was evaluated as a prerequisite validation step to verify that LLM- generated responses are semantically aligned with physician-authored answers before conducting downstream analyses. Cosine similarity was computed between sentence embeddings obtained with the BioBERT-mnli-snli-scinli-scitail-mednli-stsb 7 model [38]. Descriptive statistics for each configuration are reported in Figure 4 and Figure 6. All evaluated systems exhibit strong conceptual alignment with physician responses across both datasets. Average cosine similarity values are consistently above 0.78 in the first dataset and above 0.75 in the iCliniqQAs dataset. In the first dataset as we can see in figure 5, the highest semantic fidelity is achieved by GPT5Rephrase (μ = 0.92), followed by MixtralRephrase (μ = 0.91) and MedPaLMRephrase (μ = 0.89), while GeminiRephrase and ClaudeRephrase reach μ = 0.87 and μ = 0.85, respectively. In contrast, on the iCliniqQAs dataset (Figure 7), the highest semantic fidelity is achieved by MedPaLMRephrase (μ = 0.93), followed by GPTRephrase (μ = 0.91) and MixtralRephrase (μ = 0.89), with GeminiRephrase (μ = 0.86) and ClaudeRephrase (μ = 0.84) showing comparatively lower performance. Notably, the separation between domain-specialized and general-purpose models is more pronounced in iCliniqQAs, where MedPaLMRephrase consistently outperforms all other configurations. Among baseline architectures, performance remains tightly clustered in both datasets. In the first dataset, MedPaLMBase (μ = 0.82), MixtralBase (μ = 0.80), and GPT5Base (μ = 0.79) show closely matched alignment. In iCliniqQAs, Med- PaLMBase (μ = 0.80) and MixtralBase (μ = 0.78) remain comparable, while GPTBase (μ = 0.85) exhibits slightly higher raw similarity but does not consistently match the gains observed in domain-adapted rephrasing configurations. Prompt-based variants yield comparable distributions across both datasets. In the first dataset, GPT5Empathy (μ = 0.81), MedPaLMEmpathy (μ = 0.80), ClaudeEmpathy (μ = 0.80), and MixtralEmpathy (μ = 0.78) remain aligned with baseline levels. Similarly, in iCliniqQAs, GPTEmpathy (μ = 0.83), Med- PaLMEmpathy (μ = 0.80), ClaudeEmpathy (μ = 0.78), and MixtralEmpathy (μ = 0.78) show limited deviation from their respective base configurations, confirming that empathy prompting alone does not substantially increase semantic fidelity. Statistical significance between model configurations was assessed via two-sided paired t-tests with False Discovery Rate (FDR) correction using the Benjamini– Hochberg procedure [39]. Each statistical population corresponds to the distribution of cosine similarity scores produced by a model across all evaluated questions. Let m 1 and m 2 denote two distinct model configurations and μ m the associated mean similarity. For each pairwise comparison, the null hypothesis is defined as H 0 : μ m 1 = μ m 2 . No statistically significant difference is observed between MixtralBase and Med- PaLMBase in either dataset, confirming comparable semantic fidelity at baseline. Rephrasing configurations introduce systematic improvements across both datasets; however, the effect is particularly pronounced in the iCliniqQAs dataset, where Med- PaLM Rephrase achieves statistically significant improvements over both its baseline 7 https://huggingface.co/pritamdeka/BioBERT-mnli-snli-scinli-scitail-mednli-stsb 12 Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase -0.02 * +0.11 *** +0.14 *** +0.01 +0.03 ** -0.11 *** -0.01 +0.01 -0.13 *** -0.02 +0.08 *** +0.11 *** -0.03 *** +0.08 *** +0.10 *** -0.01 +0.01 -0.13 *** -0.02 * +0.00 -0.10 *** +0.00 +0.02 *** -0.11 *** -0.01 +0.01 -0.08 *** +0.01 +0.12 *** +0.14 *** +0.01 +0.11 *** +0.13 *** +0.04 *** +0.13 *** +0.12 *** +0.07 *** +0.10 *** -0.04 ** +0.07 *** +0.08 *** -0.01 +0.08 *** +0.07 *** -0.05 *** -0.01 +0.01 -0.13 *** -0.02 -0.00 -0.10 *** -0.00 -0.02 -0.13 *** -0.08 *** -0.01 +0.01 -0.13 *** -0.02 -0.00 -0.10 *** -0.00 -0.01 -0.13 *** -0.08 *** +0.00 +0.05 *** +0.07 *** -0.06 *** +0.04 *** +0.06 *** -0.03 ** +0.06 *** +0.05 *** -0.07 *** -0.02 +0.06 *** +0.06 *** 0.10 0.05 0.00 0.05 0.10 Mean Difference (Semantic Fidelity) Fig. 4: Pairwise comparison of language models in terms of semantic fidelity on the MedQuAD dataset. Each cell reports the mean difference in semantic fidelity between model pairs (Model i−j), where positive values indicate higher similarity to the medi- cal reference for the model reported on the row. Color intensity encodes the magnitude of the difference, while statistical significance after FDR correction is indicated by asterisks ( ∗ p < 0.05, ∗ p < 0.01, ∗ p < 0.001). and empathy variants as well as over multiple general-purpose counterparts (p < 0.01, FDR-corrected). The full matrices of mean differences and FDR-adjusted p-values are depicted in Figure 4 and Figure 6, highlighting statistically significant contrasts across multiple model pairs, especially those involving domain-specialized rephrasing configurations. 13 Mixtral Base Med-PaLM Base Mixtral Empathy Med-PaLM Empathy Mixtral Rephrase Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Model 0.2 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Cosine Similarity Cosine Similarity Across Models Mean ± Std Fig. 5: Cosine similarity between physician-written and model-generated answers on the MedQuAD dataset. Higher values reflect closer semantic alignment. Gemini 2.5 Pro appears only in the Rephrase configuration because, in our experiments, the model frequently refused to generate Base or Empathy responses without sufficient clinical context, exhibiting strong safety guardrails similar to those observed in ClaudeBase (left in the comparison as an explicit example of this behaviour). Takeaway 0☞ Rewriting consistently yields the highest semantic fidelity across both datasets. In MedQuAD, GPT5 Rephrase achieves the strongest alignment with physician-authored answers (μ = 0.92), while in iCliniqQAs the best performance is obtained by MedPaLMRephrase (μ = 0.93). Across architectures, rephrase config- urations systematically outperform both baseline and empathy-prompted variants, confirming collaborative rewriting as the most effective strategy for maximizing conceptual overlap with clinical experts. 4.3 RQ1 – Empathy and Sentiment Analyses This research question evaluates whether LLMs can produce responses that match physician-authored texts in terms of emotional attunement. To this end, we con- duct a two-step evaluation using both general sentiment classification and fine-grained emotion detection, using the models described in the Background section. Hypothesis 1 Let R LLM be a response generated by an LLM, and let R Phys be a physician- authored response. Let E(·) denote the affective resonance function introduced in Section 3.1, which captures both sentiment polarity and fine-grained emotional expression. We hypothesize that LLM-generated responses exhibit comparable affective resonance to physician-authored ones, i.e., E[E(R LLM )] = E[E(R Phys )]. 14 Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase +0.00 +0.15 *** +0.14 *** +0.02 +0.02 -0.13 *** +0.02 +0.02 -0.13 *** +0.00 +0.23 *** +0.22 *** +0.08 *** +0.21 *** +0.21 *** +0.07 *** +0.07 *** -0.08 *** +0.05 *** +0.05 *** -0.16 *** +0.05 ** +0.05 ** -0.10 *** +0.03 ** +0.03 * -0.18 *** -0.02 * +0.17 *** +0.17 *** +0.02 +0.15 *** +0.15 *** -0.06 *** +0.10 *** +0.12 *** +0.11 *** +0.11 *** -0.04 * +0.09 *** +0.09 *** -0.12 *** +0.04 +0.06 ** -0.06 ** -0.02 -0.02 -0.17 *** -0.04 ** -0.04 * -0.25 *** -0.09 *** -0.07 *** -0.19 *** -0.13 *** +0.00 +0.00 -0.14 *** -0.02 -0.02 -0.22 *** -0.06 *** -0.05 *** -0.16 *** -0.11 *** +0.02 +0.13 *** +0.13 *** -0.01 +0.12 *** +0.12 *** -0.09 *** +0.07 *** +0.09 *** -0.03 +0.03 +0.16 *** +0.13 *** 0.2 0.1 0.0 0.1 0.2 Mean Difference (Semantic Fidelity) Fig. 6: Pairwise comparison of language models in terms of semantic fidelity on the iCliniqQAs dataset. Each cell reports the mean difference in semantic fidelity between model pairs (Model i−j), where positive values indicate higher similarity to the medi- cal reference for the model reported on the row. Color intensity encodes the magnitude of the difference, while statistical significance after FDR correction is indicated by asterisks ( ∗ p < 0.05, ∗ p < 0.01, ∗ p < 0.001). 4.3.1 Sentiment Distribution We categorized each response into one of 5 sentiment classes: Very Negative, Nega- tive, Neutral, Positive, and Very Positive. As shown in Figures 8 and 9, physicians’ responses predominantly fall into the Neutral category. From Table 1, physician answers in the MedQuAD dataset concentrate predom- inantly in the Neutral category (49.02%), with a substantial proportion of Very Negative responses (37.25%) and virtually no positive affect. As illustrated in Figure 8, this distribution reflects a clinically restrained tone typical of institutional medi- cal communication. In contrast, baseline LLM configurations on MedQuAD tend to amplify polarity. Both Mixtral and Med-PaLM increase the proportion of Very 15 Mixtral Base Mixtral Empathy Mixtral Rephrase Med-PaLM Base Med-PaLM Empathy Med-PaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Model 0.2 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Cosine Similarity (normalized) Cosine Similarity Across Models Mean ± Std Fig. 7: Cosine similarity between physician-written and model-generated answers on the iCliniqQAs dataset. Higher values reflect closer semantic alignment. Gemini 2.5 Pro appears only in the Rephrase configuration because, in our experiments, the model frequently refused to generate Base or Empathy responses without sufficient clinical context, exhibiting strong safety guardrails similar to those observed in ClaudeBase (left in the comparison as an explicit example of this behaviour). Negative responses (43.14% and 45.10%, respectively), indicating a sharper affec- tive framing than physicians. Prompt-based and rephrased configurations mitigate this effect, systematically shifting outputs toward higher Neutral rates and reducing extreme negativity. Notably, Gemini Rephrase is the only configuration exhibiting a non-negligible proportion of Positive sentiment (8.0%), suggesting a mild but distinct tendency toward affective reinforcement absent from physician-authored texts. A different pattern emerges in the second dataset (iCliniqQAs), where physician responses are even more strongly dominated by Neutral sentiment (84.0%) and contain markedly lower levels of Very Negative content (6.0%). This reflects the conversational and patient-facing nature of the dataset, in which clinicians adopt a less confronta- tional and more stabilizing tone. In this setting, baseline models do not systematically amplify extreme negativity as observed in MedQuAD; instead, they display greater variability in the distribution of Negative and Neutral responses. Rephrasing strategies generally increase Neutral proportions (e.g., up to 90–92% in several configurations), further aligning outputs with physician affective restraint. However, GeminiRephrase again stands out, exhibiting a substantially higher proportion of Positive sentiment (18.0%), a level not observed in physician responses in either dataset. Pairwise chi-square analyses with Benjamini–Hochberg correction confirm that these deviations are not uniform across systems (Figures 10 and 11). In the MedQuAD setting, Claude (Base) exhibits the strongest divergence from physicians (V = 0.45, p < 0.001). This result, however, does not reflect polarity amplification of the same kind observed in Mixtral and Med-PaLM : as noted in Section 4, Claude Base fre- quently produced cautious, hedged responses in the absence of sufficient clinical context-analogous to the safety-driven refusals observed in Gemini. These outputs 16 were nonetheless classified by the sentiment model, yielding a disproportionately high Very Negative rate (82.0%) that reflects classifier sensitivity to evasive or uncertainty-laden language rather than affectively charged clinical content. By con- trast, Med-PaLMBase remains closest to physician distributions (V = 0.08, p > 0.05), representing the only baseline configuration whose sentiment profile is not statisti- cally distinguishable from that of physician-authored responses. In the iCliniqQAs dataset, effect sizes are generally more moderate, indicating closer overall alignment with human-authored sentiment patterns, though statistically significant differences persist for selected configurations. Very Negative Negative Neutral Positive Very Positive 0 10 20 30 40 50 60 70 80 Percentage (%) Sentiment Distribution Across Models Models Physicians Mixtral Base Mixtral Empathy Mixtral Rephrase Med-PaLM Base Med-PaLM Empathy Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Claude Base Claude Empathy Claude Rephrase Gemini Rephrase Fig. 8: Distribution of sentiment expressed by models on the MedQuAD dataset. Very Negative Negative Neutral Positive Very Positive 0 20 40 60 80 Percentage (%) Sentiment Distribution Across Models Models Physicians Mixtral Base Mixtral Empathy Mixtral Rephrase Med-PaLM Base Med-PaLM Empathy Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Claude Base Claude Empathy Claude Rephrase Gemini Rephrase Fig. 9: Distribution of sentiment expressed by models on the iCliniqQAs dataset. 17 Table 1: Percentage distribution of sentiment labels per system on MedQuAD dataset. Arrows indicate comparison to Doctor: ↑ = higher, ↓ = lower, - = similar. SystemVery Negative (%) Negative (%) Neutral (%) Positive (%) Very Positive (%) Physician Answer37.2513.7349.020.000.00 Mixtral43.14 (↑)0.00 (↓)56.86 (↑)0.00 (—)0.00 (—) Med-PaLM45.10 (↑)13.73 (—)41.18 (↓)0.00 (—)0.00 (—) MixtralEmpathy Prompt23.53 (↓)1.96 (↓)74.51 (↑)0.00 (—)0.00 (—) Med-PaLMEmpathy Prompt25.49 (↓)5.88 (↓)68.63 (↑)0.00 (—)0.00 (—) MixtralRephrase21.57 (↓)7.84 (↓)70.59 (↑)0.00 (—)0.00 (—) Med-PaLMRephrase19.61 (↓)7.84 (↓)72.55 (↑)0.00 (—)0.00 (—) GPT-5 (BASE)40.00 (↑)2.00 (↓)58.00 (↑)0.00 (—)0.00 (—) GPT-5Empathy Prompt22.00 (↓)4.00 (↓)74.00 (↑)0.00 (—)0.00 (—) GPT-5Rephrase22.00 (↓)10.00 (↓)68.00 (↑)0.00 (—)0.00 (—) GeminiRephrase16.00 (↓)8.00 (↓)68.00 (↑)8.00 (↑)0.00 (—) Claude (BASE)82.00 (↑)4.00 (↓)14.00 (↓)0.00 (—)0.00 (—) ClaudeEmpathy Prompt48.00 (↑)20.00 (↑)32.00 (↓)0.00 (—)0.00 (—) ClaudeRephrase56.00 (↑)10.00 (↓)34.00 (↓)0.00 (—)0.00 (—) Table 2: Percentage distribution of sentiment labels per system on the iCliniqQAs dataset. Arrows indicate comparison to Physician: ↑ = higher, ↓ = lower, - = similar. SystemVery Negative (%) Negative (%) Neutral (%) Positive (%) Very Positive (%) Physician Answer6.010.084.00.00.0 Mixtral14.0 (↑)2.0 (↓)84.0 (—)0.0 (—)0.0 (—) Med-PaLM10.0 (↑)8.0 (↓)82.0 (↓)0.0 (—)0.0 (—) MixtralEmpathy Prompt2.0 (↓)10.0 (—)84.0 (—)4.0 (↑)0.0 (—) Med-PaLMEmpathy Prompt4.0 (↓)12.0 (↑)84.0 (—)0.0 (—)0.0 (—) MixtralRephrase0.0 (↓)10.0 (—)90.0 (↑)0.0 (—)0.0 (—) Med-PaLMRephrase2.0 (↓)8.0 (↓)90.0 (↑)0.0 (—)0.0 (—) GPT-5 (BASE)28.0 (↑)8.0 (↓)62.0 (↓)2.0 (↑)0.0 (—) GPT-5Empathy Prompt16.0 (↑)16.0 (↑)62.0 (↓)4.0 (↑)2.0 (↑) GPT-5Rephrase0.0 (↓)12.0 (↑)88.0 (↑)0.0 (—)0.0 (—) GeminiRephrase4.0 (↓)26.0 (↑)52.0 (↓)18.0 (↑)0.0 (—) Claude (BASE)6.0 (—)2.0 (↓)92.0 (↑)0.0 (—)0.0 (—) Claude Empathy Prompt10.0 (↑)10.0 (—)80.0 (↓)0.0 (—)0.0 (—) ClaudeRephrase0.0 (↓)8.0 (↓)92.0 (↑)0.0 (—)0.0 (—) Taken together, the results suggest that affective misalignment is more pronounced in institutionally curated medical explanations (MedQuAD) than in conversational clinical exchanges (iCliniqQAs). While empathy prompting and rephrasing con- sistently reduce extreme negativity and increase neutrality across both datasets, certain architectures introduce an independent tendency toward positive reinforce- ment, revealing a systematic stylistic shift rather than strict replication of physician affective norms. 4.3.2 Emotion Distribution Beyond general sentiment, we analyzed the presence of 28 fine-grained emotional cat- egories. Figures 12 and 13 report the five most frequent dominant emotions across systems in the MedQuAD and iCliniqQAs datasets, respectively. Across both datasets, two emotions consistently dominate model-generated out- puts: approval and caring. However, their relative balance differs substantially between datasets, reflecting the distinct communicative setting. In MedQuAD, physician- authored responses are strongly approval-oriented, with approval as the dominant emotion in 78.4% of cases, while caring and disapproval each account for 7.8%, and realization for 5.9%. This pattern is consistent with institutional medical explanations, 18 Physicians (Humans) Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Physicians (Humans) Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase 0.27 0.29 * 0.22 0.220.29 * 0.14 0.080.29 * 0.36 ** 0.30 * 0.210.240.110.060.28 * 0.240.30 * 0.140.020.32 * 0.08 0.200.100.170.240.240.170.26 0.250.250.060.090.33 * 0.070.090.20 0.190.29 * 0.170.050.270.080.070.240.12 0.26 * 0.31 * 0.210.170.31 * 0.190.170.27 * 0.190.17 0.28 * 0.090.160.240.32 * 0.190.250.120.190.250.27 * 0.190.37 ** 0.44 ** 0.40 ** 0.110.37 ** 0.42 ** 0.34 * 0.43 ** 0.36 ** 0.37 ** 0.41 ** 0.180.30 * 0.41 ** 0.39 ** 0.120.34 * 0.42 ** 0.270.40 ** 0.36 ** 0.37 ** 0.36 ** 0.14 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 Cramér's V (Effect Size) Fig. 10: Pairwise comparison of sentiment distributions across systems on the MedQuAD dataset. Cells report Cram ́er’s V effect size for each model pair; darker color indicates larger divergence. Asterisks denote FDR-corrected significance (* p < 0.05, ** p < 0.01, *** p < 0.001). where clinicians primarily convey validation and guidance, with occasional corrective or reflective cues. In contrast, iCliniqQAs exhibits a marked shift toward affective support. Here, physician answers are predominantly caring -oriented (33.3%), while approval becomes secondary (2.0%). The remaining dominant emotions appear at much lower rates, including gratitude (2.0%), curiosity (3.9%), and optimism (2.0%). This difference indicates that conversational consultations elicit a substantially more supportive and relational emotional style than standardized institutional explanations, even in physician-written content. Several systematic model behaviors emerge across both datasets. First, LLMs display high variability in the expression of approval on MedQuAD. Some base config- urations are more approval-heavy than physicians, such as GPT5 BASE (92.23%) and ClaudeBASE (86.31%), whereas others substantially reduce approval when prompted for caring or rewriting: MixtralRephrase (23.5%) and MedPaLMRephrase (19.60%) 19 Physicians (Humans) Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Physicians (Humans) Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase 0.21 0.140.25 0.180.31 * 0.14 0.080.150.180.23 0.050.250.130.150.13 0.110.250.120.110.170.09 0.260.210.30 * 0.34 ** 0.210.28 * 0.31 * 0.200.230.210.27 * 0.180.210.250.15 0.180.33 * 0.140.000.230.140.120.34 ** 0.26 * 0.33 ** 0.41 *** 0.28 * 0.36 ** 0.35 ** 0.32 * 0.36 ** 0.37 ** 0.230.35 ** 0.170.130.200.240.160.200.170.29 * 0.26 * 0.260.41 *** 0.070.170.180.230.040.120.180.210.170.230.33 ** 0.19 0.180.30 * 0.150.000.230.160.100.35 ** 0.28 * 0.050.38 ** 0.220.24 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 Cramér's V (Effect Size) Fig. 11: Pairwise comparison of sentiment distributions across systems on the iClin- iqQAs dataset. Cells report Cram ́er’s V effect size for each model pair; darker color indicates larger divergence. Asterisks denote FDR-corrected significance (* p < 0.05, ** p < 0.01, *** p < 0.001). illustrate a strong reallocation away from validation toward more explicitly support- ive framing. In iCliniqQAs, approval is instead generally attenuated across systems and rarely becomes dominant; when it appears among the top emotions, it does so at modest levels (e.g., MixtralEmpathy 16.00%, GPT5Rephrase 18.00%), consistent with the dataset’s baseline emphasis on reassurance rather than endorsement. Second, caring is systematically amplified in LLM outputs relative to physicians in both datasets, but the magnitude of amplification depends on the conversational context. In MedQuAD, caring is dominant in only 7.8% of physician responses, yet it becomes one of the primary emotions in most model settings, especially under empa- thy prompting and rewriting: MixtralEmpathy (52.90%) and MedPaLMEmpathy (41.20%) strongly exceed physicians, while rewriting further accentuates caring, with Mixtral Rephrase and MedPaLMRephrase reaching 76.5%. In iCliniqQAs, the same tendency persists but starts from a substantially higher human baseline (33.30%). Many models push caring to near-saturation levels, particularly in base configurations 20 Approval Caring Sadness Realization Disapproval Emotion 0 20 40 60 80 Percentage (%) Top 5 Dominant Emotions Across Models Model Physicians Mixtral Base Med-PaLM Base Mixtral Rephrase Med-PaLM Rephrase Mixtral Empathy Med-PaLM Empathy GPT-5 Base GPT-5 Rephrase GPT-5 Empathy Gemini Rephrase Claude Base Claude Rephrase Claude Empathy Fig. 12: Emotions most frequently expressed by models on the MedQuAD dataset. (e.g., MixtralRephrase = 92.00%, ClaudeBase = 82.00%, GPT5Base = 92.0%), indi- cating that in naturally emotional patient narratives, models converge toward a highly supportive stance regardless of whether they are explicitly prompted for empathy. Third, negative or corrective emotions are attenuated in model outputs, especially in MedQuAD. In the first dataset, disapproval is consistently present in physician texts (7.8%) yet appears marginally or disappears in most LLM configurations, rarely exceeding 5.9% and often remaining absent in the top emotions. This aligns with an avoidance of negatively directive stances in machine-generated clinical communication. In iCliniqQAs, disapproval is not among the dominant emotions for physicians or mod- els; instead, low-frequency positive-affiliative emotions such as gratitude, optimism, and curiosity emerge among the top categories, but remain limited in prevalence (gen- erally below∼5%), suggesting that the overall affective profile is still largely governed by caring. Taken together, fine-grained emotion analysis shows that LLMs do not reproduce physician affective behavior verbatim. Rather, they exhibit a systematic reweighting of emotional cues that amplifies affiliative signals such as caring and, depending on the dataset, either preserves or reduces approval. Importantly, the direction of this shift is dataset-dependent: institutional explanations (MedQuAD) highlight a transition from approval-dominant physician discourse toward caring-heavy model outputs, whereas real-world consultations (iCliniqQAs) already start from a caring-oriented physician baseline and are further pushed by LLMs toward near-uniform supportive affect. This difference reflects variation in emotional style and emphasis across contexts, rather than an absolute improvement in communication quality. 21 caring approval gratitude curiosity optimism Emotion 0 20 40 60 80 100 Percentage (%) Top 5 Dominant Emotions Across Models Model Physicians Mixtral Base Med-PaLM Base Mixtral Rephrase Med-PaLM Rephrase Mixtral Empathy Med-PaLM Empathy GPT-5 Base GPT-5 Rephrase GPT-5 Empathy Gemini Rephrase Claude Base Claude Rephrase Claude Empathy Fig. 13: Emotions most frequently expressed by models on the iCliniqQAs dataset. Takeaway 1☞ R LLM can approximate R Phys in overall emotional restraint across both datasets. However, fine-grained emotion analysis reveals a systematic reweight- ing rather than faithful replication: LLMs consistently amplify affiliative signals such as caring, while attenuating corrective or discordant cues. This shift is dataset- dependent: models move from approval-dominant discourse in institutional texts to near-saturated caring in conversational settings, indicating stylistic modulation rather than improved clinical alignment. 4.4 RQ2 – Readability Analysis This research question explores whether LLMs can produce more readable responses than those authored by physicians, who may rely on complex phrasing and technical jargon. To evaluate the readability of each response type, we applied the FKGL and GFI metrics previously introduced in the Background section. Figures 14 and 15 report average scores for physician-written content and base (i.e., non–prompt-engineered, non–rephrased) LLM generations. Hypothesis 2 Let R base LLM be a zero-shot (baseline) LLM response without domain prompting or rewriting, and let Read(·) be a function that measures the readability of a text where lower scores indicate greater accessibility. We hypothesize that baseline LLM generations will exhibit equal or higher readability compared to physician-authored content R Phys , i.e., E[Read(R base LLM )]≤ E[Read(R Phys )]. Overall, physician-authored responses exhibit moderate complexity in both datasets. On MedQuAD, physicians show FKGL = 11.47 and GFI = 12.82. On 22 iCliniqQAs, physicians exhibit FKGL = 12.50 and GFI = 12.60, indicating that con- versational physician responses are not substantially simpler than institutional ones in terms of formal grade-level metrics. In MedQuAD, GPT5Base displays substantially higher complexity (FKGL = 16.91, GFI = 20.39), and ClaudeBase also exceeds physician readability (FKGL = 14.26, GFI = 16.54). MixtralBase and MedPaLMBase remain closer to physician levels. For the iCliniqQAs dataset, GPT5Base produces highly complex text (FKGL = 17.60, GFI = 17.60), while ClaudeBase reaches the highest GFI values overall (GFI = 20.30). As shown in Figures 18 and 19, on iCliniqQAs GPT5Base is signif- icantly less readable than physicians across both metrics (∆FKGL = +5.44, ∆GFI = +7.57, all p < 0.001, FDR-corrected). ClaudeBase also produces significantly more complex text than physicians (∆FKGL = +2.78, ∆GFI = +3.71, p < 0.01). In contrast, MixtralBase and MedPaLMBase do not show statistically significant deviations from physician readability on either dataset, confirming that their baseline lexical complexity is broadly aligned with expert-authored responses. Across both datasets, empathy prompting and rephrasing systematically reduce linguistic complexity relative to baseline models. On MedQuAD (Figures 16 and 17), MixtralEmpathy and MedPaLMEmpathy reduce FKGL by −2.41 and −2.41 points respectively compared to their base variants, with analogous improvements in GFI (−1.79 and −2.31, all p < 0.001). For GPT5, the reduction is even larger: GPT5Empathy and GPT5Rephrase lower FKGL by−6.87 and−6.61 points and GFI by −8.69 and −7.98 points relative to GPT5Base (all p < 0.001). ClaudeEmpathy and ClaudeRephrase also improve readability relative to ClaudeBase (∆FKGL =−2.95 and −1.02; ∆GFI =−3.11 and −0.86, p < 0.01). A comparable pattern is observed in iCliniqQAs (Figures 18 and 19). GPT5Empathy reduces FKGL by −4.54 and GFI by −5.39 relative to GPT5Base (both p < 0.001), and GPT5Rephrase yields even larger improvements (∆FKGL = −6.95, ∆GFI = −7.81, p < 0.001). The effect is especially pronounced for Claude: ClaudeEmpathy lowers FKGL by −7.87 and GFI by −9.56 relative to ClaudeBase, while ClaudeRephrase further reduces complexity (∆FKGL = −9.39, ∆GFI =−11.08, all p < 0.001). GeminiRephrase also shows statistically significant improvements relative to more complex baselines (e.g., ∆FKGL = −3.16, ∆GFI = −3.50 vs. GPT5Base on iCliniqQAs, p < 0.001). Taken together, these findings confirm that improved readability is not an intrinsic property of LLM output. Baseline generations from Mixtral and MedPaLM approxi- mate physician readability across both datasets, whereas GPT5Base and ClaudeBase produce significantly more complex prose. Consistent readability gains emerge primar- ily when models are explicitly instructed or used as rewriting assistants, indicating that accessibility depends strongly on prompting strategy rather than architecture alone. 23 Physicians Mixtral Base Med-PaLM Base Mixtral Empathy Med-PaLM Empathy Mixtral Rephrase Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 22.5 Readability Score FleschKincaid Grade Level (a) FKGL. Lower values = better readability. Physicians Mixtral Base Med-PaLM Base Mixtral Empathy Med-PaLM Empathy Mixtral Rephrase Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 22.5 Readability Score Gunning Fog Index (b) GFI. Higher values = harder text. Fig. 14: Readability analysis: Flesch–Kincaid Grade Level and Gunning Fog Index scores across physician and LLM outputs on the MedQuAD dataset. Takeaway 2☞ Baseline LLM generations do not systematically improve accessibil- ity across datasets. While Mixtral and Med-PaLM approximate physician readability in both MedQuAD and iCliniqQAs, GPT5Base and ClaudeBase consistently pro- duce significantly more complex text than clinicians. Substantial and statistically significant readability gains emerge only under empathy prompting or collaborative rewriting. Accessibility therefore appears to be a controllable property induced by alignment strategies rather than an intrinsic characteristic of large language models. 4.5 RQ3 – Effect of Prompt Engineering on AI Alignment This research question assesses whether empathy-enhancing prompt design can steer LLMs toward more emotionally appropriate and readable outputs. 24 Physicians Mixtral Base Mixtral Empathy Mixtral Rephrase Med-PaLM Base Med-PaLM Empathy Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 22.5 Readability Score FleschKincaid Grade Level (a) FKGL. Lower values = better readability. Physicians Mixtral Base Mixtral Empathy Mixtral Rephrase Med-PaLM Base Med-PaLM Empathy Med-PaLM Rephrase GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 22.5 Readability Score Gunning Fog Index (b) GFI. Higher values = harder text. Fig. 15: Readability analysis: Flesch–Kincaid Grade Level and Gunning Fog Index scores across physician and LLM outputs on the iCliniqQAs dataset. Hypothesis 3 Let R LLM Empathy Prompt denote responses generated under the empathy- enhanced Empathy Prompt condition. We hypothesize that Empathy Prompt primarily affects affective and communicative style rather than technical content, leading to (i) increased affec- tive support and (i) improved readability due to indirect stylistic simplification rather than explicit textual optimization, relative to zero-shot outputs, i.e., (i) E[E(R LLMEmpathy Prompt )] > E[E(R LLMBase )] and (i) E[Read(R LLM Empathy Prompt )] < E[Read(R LLM Base )]. Readability outcomes are reported in Figures 14 and 15, which present aver- age Flesch–Kincaid Grade Level (FKGL) and Gunning Fog Index (GFI) scores for physician-authored responses and for each model configuration across both datasets. 25 Physicians (Humans) Mixtral Base Mixtral Empathy MedPaLM Base MedPaLM Empathy GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Physicians (Humans) Mixtral Base Mixtral Empathy MedPaLM Base MedPaLM Empathy GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase +1.44 -0.97-2.41 *** +1.65+0.21+2.63 *** -0.75-2.19 *** +0.22-2.41 *** +5.44 *** +4.00 *** +6.41 *** +3.79 *** +6.19 *** -1.64-2.94 *** -0.46-3.07 *** -0.65 * -6.87 *** -1.22-2.54 *** -0.10-2.83 *** -0.40-6.61 *** +0.41 -0.42-1.84 *** +0.54-2.04 *** +0.32-5.83 *** +1.05 ** +0.80 ** +2.78 * +1.34 ** +3.76 *** +1.13 * +3.54 *** -2.66 *** +4.35 *** +3.92 *** +3.11 *** -0.17-1.60 *** +0.81 * -1.82 *** +0.59-5.60 *** +1.27 *** +0.98 * +0.25-2.95 *** +1.76+0.32+2.73 *** +0.11+2.51 *** -3.68 *** +3.23 *** +2.96 *** +2.17 *** -1.02 ** +1.93 *** Pairwise Model Comparison (FleschKincaid) 6 4 2 0 2 4 6 Mean Difference Fig. 16: Pairwise differences in Flesch–Kincaid Grade Level (FKGL) across systems on the MedQuAD dataset. Each cell reports the mean difference between row and column models (row minus column); negative values indicate better readability for the row model. Statistical significance is assessed via paired t-tests with Benjamini– Hochberg FDR correction (* p < 0.05, ** p < 0.01, *** p < 0.001). Across models, Empathy Prompt consistently lowers FKGL and GFI scores rela- tive to their corresponding base variants in MedQuAD. For instance, MixtralEmpathy Prompt reduces FKGL from 12.91 to 11.88 and GFI from 13.66 to 12.50, while Med- PaLMEmpathy Prompt shows similar reductions (FKGL 13.13→ 10.72; GFI 14.47→ 12.16). For larger architectures, the effect is even more pronounced: GPT5Empathy lowers FKGL from 16.91 to 10.04 and GFI from 20.39 to 11.70, representing reductions of −6.87 and −8.69 points respectively (all p < 0.001). These results indicate that empathy-oriented prompting encourages simpler sentence construction and reduced lexical density in institutionally curated medical explanations. A comparable but context-sensitive pattern emerges in iCliniqQAs. Here, physi- cians exhibit FKGL = 12.50 and GFI = 12.60. GPT5Base produces substantially higher complexity (FKGL = 17.60, GFI = 17.60), whereas GPT5Empathy reduces these scores to FKGL = 13.06 and GFI = 12.21, yielding reductions of −4.54 and 26 Physicians (Humans) Mixtral Base Mixtral Empathy MedPaLM Base MedPaLM Empathy GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Physicians (Humans) Mixtral Base Mixtral Empathy MedPaLM Base MedPaLM Empathy GPT5 Base GPT5 Empathy GPT5 Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase +0.84 -0.95-1.79 *** +1.64+0.80 ** +2.59 *** -0.66-1.50 *** +0.28-2.31 *** +7.57 *** +6.72 *** +8.51 *** +5.92 *** +8.23 *** -1.33-2.00 *** -0.13-2.77 *** -0.41-8.69 *** -0.49-1.20 * +0.63-2.06 *** +0.25-7.98 *** +0.81 +0.18-0.61+1.15 ** -1.41 *** +0.87 * -7.36 *** +1.32 *** +0.65 +3.71 ** +2.87 *** +4.66 *** +2.07 *** +4.38 *** -3.85 *** +4.99 *** +4.10 *** +3.42 *** +0.60-0.24+1.55 *** -1.04 ** +1.26 ** -6.96 *** +1.73 *** +1.01 * +0.40-3.11 *** +2.85 ** +2.01 *** +3.80 *** +1.21 *** +3.51 *** -4.71 *** +4.03 *** +3.30 *** +2.65 *** -0.86 * +2.25 *** Pairwise Model Comparison (GunningFog) 8 6 4 2 0 2 4 6 8 Mean Difference Fig. 17: Pairwise differences in Gunning Fog Index (GFI) across systems on the MedQuAD dataset. Values represent mean score differences (row minus column); lower values correspond to easier-to-read text. Significance is evaluated using paired t-tests with FDR correction (* p < 0.05, ** p < 0.01, *** p < 0.001). −5.39 points respectively (both p < 0.001). ClaudeEmpathy similarly improves read- ability compared to ClaudeBase (∆FKGL = −7.87, ∆GFI = −9.56, p < 0.001). Thus, across both institutional (MedQuAD) and conversational (iCliniqQAs) settings, empathy prompting systematically reduces linguistic complexity. In terms of sentiment alignment (Tables 1 and 2), Empathy Prompt shifts model responses toward more neutral and less confrontational phrasing in MedQuAD. Mix- tralEmpathy Prompt increases Neutral responses from 56.86% to 74.51% while reduc- ing Very Negative sentiment from 43.14% to 23.53%. MedPaLMEmpathy Prompt shows a comparable shift (Very Negative 45.10% → 25.49%; Neutral 41.18% → 68.63%). GPT5Empathy Prompt reduces Very Negative sentiment from 40.00% to 22.00% and increases Neutral responses from 58.00% to 74.00%. In iCliniqQAs, baseline physician sentiment is already strongly Neutral (84.00%) with low Very Negative content (6.00%). In this setting, empathy prompting reduces extreme negativity but does not universally increase neutrality. For example, Mix- tralEmpathy Prompt reduces Very Negative responses from 14.00% to 2.00%, while 27 Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase -4.46 *** -6.33 *** -1.87 *** -1.33 *** +3.13 *** +5.00 *** -3.62 *** +0.83 ** +2.71 *** -2.30 *** -4.66 *** -0.21+1.66 *** -3.34 *** -1.04 *** -0.56+3.90 *** +5.77 *** +0.77 * +3.06 *** +4.11 *** -5.10 *** -0.64 * +1.23 *** -3.77 *** -1.47 *** -0.43-4.54 *** -7.51 *** -3.05 *** -1.18 *** -6.18 *** -3.88 *** -2.84 *** -6.95 *** -2.41 *** -8.26 *** -3.80 *** -1.93 *** -6.93 *** -4.64 *** -3.60 *** -7.70 *** -3.16 *** -0.75 ** +1.71 *** +6.16 *** +8.04 *** +3.03 *** +5.33 *** +6.37 *** +2.27 *** +6.80 *** +9.21 *** +9.97 *** -6.16 *** -1.70 *** +0.17-4.83 *** -2.54 *** -1.49 *** -5.60 *** -1.06 *** +1.35 *** +2.10 *** -7.87 *** -7.69 *** -3.23 *** -1.36 *** -6.36 *** -4.06 *** -3.02 *** -7.13 *** -2.59 *** -0.18+0.57 ** -9.39 *** -1.53 *** Pairwise Model Comparison (FleschKincaid) 7.5 5.0 2.5 0.0 2.5 5.0 7.5 Mean Difference Fig. 18: Pairwise differences in Flesch–Kincaid Grade Level (FKGL) across systems on the iCliniqQAs dataset. Each cell reports the mean difference between row and column models (row minus column); negative values indicate better readability for the row model. Statistical significance is assessed via paired t-tests with Benjamini– Hochberg FDR correction (* p < 0.05, ** p < 0.01, *** p < 0.001). maintaining Neutral at 84.00%. GPT5Empathy Prompt, however, reduces Very Negative from 28.00% to 16.00% while shifting distribution toward both Negative (16.00%) and Neutral (62.00%), indicating that affective modulation in conversational data is more architecture-dependent than in MedQuAD. ClaudeEmpathy Prompt slightly increases Very Negative responses from 6.00% to 10.00%, demonstrating that empathy prompting does not uniformly guarantee improved sentiment alignment in patient-facing dialogue. Empathy Prompt increases supportive emotional cues without artificially inflating Positive sentiment in MedQuAD, where Positive remains at 0.00% for most systems. In contrast, in iCliniqQAs, certain configurations introduce modest Positive proportions (e.g., MixtralEmpathy Prompt = 4.00%, GPT5Empathy Prompt = 4.00%), reflecting the conversational tone of the dataset. Fine-grained emotion analysis (Figures 12 and 13) further clarifies this divergence. In MedQuAD, empathy prompting substantially amplifies caring relative to physicians (from 7.80% to over 40.00% in several configurations), whereas in iCliniqQAs, where 28 Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase Mixtral Base Mixtral Empathy Mixtral Rephrase MedPaLM Base MedPaLM Empathy MedPaLM Rephrase GPT Base GPT Empathy GPT Rephrase Gemini Rephrase Claude Base Claude Empathy Claude Rephrase -5.48 *** -7.46 *** -1.98 *** -1.67 *** +3.81 *** +5.79 *** -4.78 *** +0.71 * +2.69 *** -3.10 *** -5.62 *** -0.14+1.84 *** -3.95 *** -0.84 * -0.77 * +4.72 *** +6.70 *** +0.91 * +4.01 *** +4.85 *** -6.16 *** -0.68 * +1.30 *** -4.49 *** -1.38 *** -0.54-5.39 *** -8.58 *** -3.09 *** -1.11 *** -6.90 *** -3.80 *** -2.96 *** -7.81 *** -2.42 *** -9.66 *** -4.17 *** -2.20 *** -7.99 *** -4.88 *** -4.04 *** -8.89 *** -3.50 *** -1.08 *** +1.93 *** +7.42 *** +9.40 *** +3.61 *** +6.71 *** +7.55 *** +2.70 *** +8.09 *** +10.51 *** +11.59 *** -7.62 *** -2.14 *** -0.16-5.95 *** -2.85 *** -2.00 *** -6.86 *** -1.46 *** +0.95 ** +2.03 *** -9.56 *** -9.15 *** -3.66 *** -1.68 *** -7.47 *** -4.37 *** -3.53 *** -8.38 *** -2.99 *** -0.57 * +0.51 * -11.08 *** -1.52 *** Pairwise Model Comparison (GunningFog) 10 5 0 5 10 Mean Difference Fig. 19: Pairwise differences in Gunning Fog Index (GFI) across systems on the iCliniqQAs dataset. Values represent mean score differences (row minus column); lower values correspond to easier-to-read text. Significance is evaluated using paired t-tests with FDR correction (* p < 0.05, ** p < 0.01, *** p < 0.001). physician responses are already caring-dominant (33.30%), models often push caring toward near-saturation levels (e.g., MixtralRephrase = 92.00%). Thus, the emotional effect of empathy prompting is additive in institutional discourse but saturating in conversational settings. Statistical testing via paired t-tests with Benjamini–Hochberg FDR correction confirms that Empathy Prompt produces significant improvements over baseline gen- erations in readability across both datasets (all p < 0.01 for major architectures). However, sentiment improvements are dataset-dependent: reductions in extreme neg- ativity are systematic in MedQuAD but more variable in iCliniqQAs, where baseline physician affect is already strongly neutral and supportive. Taken together, these results indicate that empathy prompting robustly enhances readability in both institutional and conversational medical communication. Its effect on affective alignment, however, is moderated by the underlying discourse context: it corrects polarity amplification in formal explanatory texts but produces more architecture-specific shifts in already supportive patient–physician dialogue. 29 Takeaway 3☞ Empathy-enhanced prompting (Empathy Prompt) systematically improves readability and modulates affective tone across both datasets. It reduces extreme negativity and increases affiliative cues — particularly caring — while low- ering linguistic complexity relative to baseline generations. Readability gains are especially pronounced for larger architectures such as GPT5, whereas sentiment shifts are more dataset-dependent: polarity correction is consistent in institutional texts (MedQuAD) but more variable in conversational dialogue (iCliniqQAs). Overall, Empathy Prompt acts as a controllable alignment mechanism, steering emotional tone toward clinical norms while improving accessibility without compromising semantic fidelity. 4.6 RQ4 – Human–AI Collaboration This research question evaluates LLMs not only as content generators, but also as edi- tors capable of revising expert-authored medical texts to improve clarity and emotional resonance. Hypothesis 4 Let R LLM Rephrase denote the physician-authored response rewritten by an LLM using the Rephrase prompt, and R Phys the original physician response. We hypothesize that collaborative rewriting will produce responses that are (i) more readable and (i) more affectively supportive than the original physician-authored text, i.e., (i) E[Read(R LLM Rephrase )] < E[Read(R Phys )] and (i) E[E(R LLMRephrase )] > E[E(R Phys )]. Rewriting systematically shifts emotional polarity toward more supportive and less confrontational phrasing across both datasets. In MedQuAD (Table 1), rephrase variants markedly reduce Very Negative senti- ment while increasing Neutral responses. MedPaLMRephrase achieves the strongest moderation effect (Very Negative = 19.61%, Neutral = 72.55%), followed closely by MixtralRephrase (Very Negative = 21.57%, Neutral = 70.59%). GPT5Rephrase reduces Very Negative sentiment from 40.00% to 22.00% while increasing neutrality to 68.00%. Although ClaudeRephrase remains more polarity-heavy than other mod- els (Very Negative = 56.00%), it still substantially moderates its baseline behavior relative to ClaudeBase (82.00%). A different but structurally consistent pattern emerges in iCliniqQAs (Table 2). Here, physician responses already exhibit strong neutrality (Neutral = 84.00%, Very Negative = 6.00%), reflecting the conversational nature of the dataset. In this context, rewriting primarily compresses extreme negativity and increases neutral dominance rather than correcting polarity amplification. MixtralRephrase achieves Very Negative = 0.00% and Neutral = 90.00%, while MedPaLMRephrase yields Very Negative = 2.00% and Neutral = 90.00%. GPT5Rephrase reduces Very Negative from 28.00% to 0.00% and increases Neutral to 88.00%. ClaudeRephrase similarly eliminates extreme negativity (Very Negative = 0.00%) and raises Neutral to 92.00%. Emotional tone analysis (Figures 12 and 13) further supports these trends. In MedQuAD, rewriting substantially increases caring relative to physicians (7.80%), 30 with MixtralRephrase and MedPaLMRephrase reaching 76.50%. In iCliniqQAs, where physician responses are already caring-dominant (33.30%), rewriting pushes affective support toward near-saturation levels (e.g., MixtralRephrase = 92.00%). Thus, rewriting acts as polarity correction in institutional discourse and as affective amplification in conversational dialogue. Rewriting also enhances linguistic accessibility in both datasets. On MedQuAD, MixtralRephrase and MedPaLMRephrase reduce FKGL from 12.91 to 11.18 and from 13.13 to 10.56, respectively, and GFI from 13.66 to 12.45 and from 14.47 to 12.16. GPT5Rephrase lowers FKGL from 16.91 to 10.30 and GFI from 20.39 to 12.41, yielding statistically significant improvements (all p < 0.001). In iCliniqQAs, similar reductions are observed: GPT5Rephrase decreases FKGL by−6.95 and GFI by−7.81 relative to GPT5Base, while ClaudeRephrase reduces FKGL by −9.39 and GFI by −11.08 relative to ClaudeBase (all statistically significant under FDR correction). Notably, the magnitude of readability improvement is comparable across datasets, but the emotional effect differs in function: in MedQuAD, rewriting primarily miti- gates excessive polarity, whereas in iCliniqQAs it consolidates an already supportive conversational tone. Paired statistical tests with Benjamini–Hochberg False Discovery Rate correction confirm that rewriting introduces statistically significant improvements over baseline generations in sentiment distribution, fine-grained emotion profiles, and readability across multiple model families (all p < 0.01). Takeaway 4☞ Collaborative rewriting consistently improves emotional align- ment and linguistic accessibility across both datasets. It reduces extreme negativity, increases neutral and supportive phrasing, and lowers readability scores relative to baseline generations. Open-source models such as MedPaLMRephrase and Mix- tralRephrase exhibit the most stable cross-dataset gains, while larger architectures (e.g., GPT5 Rephrase, ClaudeRephrase) demonstrate substantial reductions in polar- ity and complexity compared to their base variants. Rewriting therefore emerges as a robust post-hoc alignment mechanism that enhances clarity and affective appropriateness without degrading semantic fidelity. 4.7 RQ5 – Value Alignment between Experts and Patients To address RQ5, we evaluate whether different LLM configurations align with expert and patient communication values in clinical settings. Hypothesis 5 Let R LLM be a response generated under a given configuration and R Phys the physician-authored version. Let V exp (·) and V pat (·) denote alignment with expert and patient preferences. We hypothesize that collaboratively rewritten responses (R LLMRephrase ) achieve higher alignment than physician-authored content: E[V exp, pat (R LLMRephrase )] > E[V exp, pat (R Phys )] 31 Table 3: Mean ( ̄ x) and standard deviation (σ) of physician and patient (human) ratings on the MedQuAD. Arrows indicate comparison to Doctor: ↑ = higher, ↓ = lower, - = similar. Bold values denote best-performing variants per metric (excluding Physician baseline). Response Variant Expert Evaluation (Human)Patient Evaluation (Human) ̄ x Accuracy ± σ ̄ x Style ± σ ̄ x Precision ± σ ̄ x Trust ± σ ̄ x Compr. ± σ ̄ x Emot. Tone ± σ Physician Answer5.00± 0.005.00± 0.005.00± 0.002.50± 0.502.10± 0.602.30± 0.50 Mixtral5.00± 0.00 (−)4.50± 0.00 (↓) 5.00± 0.00 (−)3.80± 0.30 (↑)4.10± 0.20 (↑)3.70± 0.30 (↑) Med-PaLM5.00± 0.00 (−)4.50± 0.00 (↓) 5.00± 0.00 (−)3.50± 0.30 (↑)4.00± 0.30 (↑)3.60± 0.20 (↑) Claude5.00± 0.00 (−)4.50± 0.00 (↓) 5.00± 0.00 (−)3.40± 0.20 (↑)3.20± 0.10 (↑)3.10± 0.10 (↑) GPT54.00± 1.41 (↓)4.00± 0.00 (↓)4.50± 0.71 (↓)3.60± 0.20 (↑)3.20± 0.20 (↑)2.30± 0.50 (−) MixtralEmpathyPrompt4.00± 1.41 (↓)4.50± 0.71 (↓)4.00± 1.41 (↓)4.60± 0.20 (↑) 4.70± 0.10 (↑)4.80± 0.10 (↑) Med-PaLMEmpathyPrompt4.50± 0.71 (↓)4.50± 0.71 (↓)4.50± 0.71 (↓)4.40± 0.20 (↑)4.50± 0.20 (↑)4.50± 0.20 (↑) ClaudeEmpathyPrompt4.50± 0.71 (↓)4.50± 0.71 (↓)4.50± 0.71 (↓)3.80± 0.20 (↑)3.60± 0.20 (↑)3.80± 0.50 (↑) GPT5EmpathyPrompt4.00± 1.41 (↓)4.50± 0.71 (↓)4.50± 0.71 (↓)4.20± 0.10 (↑)3.50± 0.30 (↑)3.80± 0.50 (↑) MixtralRephrase5.00± 0.00 (−) 5.00± 0.00 (−)4.50± 0.71 (↓)4.50± 0.20 (↑)4.60± 0.20 (↑)4.60± 0.10 (↑) Med-PaLMRephrase4.50± 0.71 (↓)4.00± 1.41 (↓)3.50± 0.71 (↓)4.70± 0.10 (↑)4.60± 0.20 (↑)4.70± 0.10 (↑) GeminiRephrase4.00± 1.41 (↓)3.00± 1.41 (↓)4.00± 1.41 (↓)4.70± 0.20 (↑)4.60± 0.30 (↑) 4.90± 0.10 (↑) ClaudeRephrase4.00± 1.41 (↓)4.00± 0.00 (↓)4.00± 1.41 (↓)4.60± 0.10 (↑)4.60± 0.20 (↑)4.50± 0.10 (↑) GPT5Rephrase4.00± 1.41 (↓)4.00± 0.00 (↓)4.50± 0.71 (↓)4.50± 0.30 (↑)4.40± 0.10 (↑)4.20± 0.20 (↑) Table 4: Mean ( ̄ x) and standard deviation (σ) of physician and patient (human) ratings on the iCliniqQAs. Arrows indicate comparison to Doctor: ↑ = higher, ↓ = lower, - = similar. Bold values denote best-performing variants per metric (excluding Physician baseline). Response Variant Expert Evaluation (Human)Patient Evaluation (Human) ̄ x Accuracy ± σ ̄ x Style ± σ ̄ x Precision ± σ ̄ x Trust ± σ ̄ x Compr. ± σ ̄ x Emot. Tone ± σ Physician Answer5.00± 0.005.00± 0.005.00± 0.004.60± 0.504.70± 0.454.65± 0.48 Mixtral3.00± 1.41 (↓)3.50± 2.12 (↓)4.00± 1.41 (↓)4.30± 0.70 (↓)4.40± 0.65 (↓)4.25± 0.72 (↓) Med-PaLM2.50± 2.12 (↓)3.00± 0.00 (↓)2.00± 1.41 (↓)4.20± 0.75 (↓)4.30± 0.70 (↓)4.15± 0.80 (↓) Claude5.00± 0.00 (−)1.00± 0.00 (↓) 5.00± 0.00 (−)4.65± 0.48 (↑)4.75± 0.44 (↑)4.70± 0.46 (↑) GPT54.50± 0.71 (↓)3.00± 1.41 (↓)4.00± 0.00 (↓)4.70± 0.46 (↑)4.80± 0.40 (↑)4.75± 0.43 (↑) MixtralEmpathyPrompt2.00± 1.41 (↓)3.50± 0.71 (↓)2.00± 1.41 (↓)4.40± 0.65 (↓)4.50± 0.60 (↓)4.35± 0.68 (↓) Med-PaLMEmpathyPrompt1.50± 0.71 (↓)4.00± 0.00 (↓)1.50± 0.71 (↓)4.55± 0.55 (↓)4.60± 0.52 (↓)4.50± 0.58 (↓) ClaudeEmpathyPrompt1.50± 0.71 (↓)3.00± 0.00 (↓)1.50± 0.71 (↓)4.75± 0.44 (↑)4.85± 0.36 (↑)4.80± 0.40 (↑) GPT5EmpathyPrompt3.50± 0.71 (↓)4.00± 1.41 (↓)3.00± 0.00 (↓)4.85± 0.35 (↑)4.90± 0.30 (↑)4.88± 0.32 (↑) MixtralRephrase3.00± 2.83 (↓)4.50± 0.71 (↓)3.00± 2.83 (↓)4.75± 0.43 (↑)4.85± 0.36 (↑)4.80± 0.40 (↑) Med-PaLMRephrase3.00± 2.83 (↓) 5.00± 0.00 (−)4.00± 2.83 (↓)4.88± 0.32 (↑)4.92± 0.28 (↑)4.90± 0.30 (↑) ClaudeRephrase3.00± 2.83 (↓)4.50± 0.71 (↓)3.00± 2.83 (↓)4.90± 0.30 (↑)4.95± 0.22 (↑)4.93± 0.25 (↑) GPT5Rephrase3.00± 2.83 (↓)4.50± 0.71 (↓)3.50± 2.12 (↓)4.95± 0.22 (↑) 4.98± 0.15 (↑) 4.96± 0.20 (↑) We evaluate two axes: epistemic values (accuracy, stylistic appropriateness, preci- sion) and relational values (trust, comprehensibility, emotional tone). All scores use 5-point Likert scales. In MedQuAD, as we can see in Table 3, physician answers receive maximum expert scores (5.00) across accuracy, style, and precision. No LLM configuration surpasses the physician baseline on epistemic criteria. Rewriting configurations improve rela- tional metrics without exceeding physician epistemic performance. MixtralRephrase achieves the highest stylistic score (5.00) while maintaining strong patient trust (4.50) and emotional tone (4.60). MedPaLMRephrase preserves high expert precision (4.00) and patient trust (4.70). GPT5 Rephrase shows balanced performance (expert accuracy = 4.00; patient trust = 4.50; emotional tone = 4.20). Empathy prompt- ing increases patient-oriented metrics but reduces expert scores relative to physician answers. 32 In iCliniqQAs, as we can se in Table 4, physician answers obtain expert scores of 5.00 across epistemic criteria and strong patient ratings (trust = 4.60; emotional tone = 4.65). Baseline LLM configurations diverge more strongly from physicians than in MedQuAD. GPT5Base reaches high patient trust (4.70) but lower stylistic alignment (3.00). ClaudeBase achieves strong expert alignment (accuracy = 5.00; precision = 5.00) and high patient tone (4.70). Rewriting configurations yield the largest relational gains. GPT5Rephrase reaches near-ceiling patient scores (trust = 4.95; comprehensibility = 4.98; emotional tone = 4.96). ClaudeRephrase shows similar relational alignment (trust = 4.90; tone = 4.93). No configuration exceeds physicians on expert accuracy. Across datasets, rewriting improves relational alignment more strongly in conversa- tional data than in institutional explanations. MedQuAD remains expert-dominated, with physicians as the epistemic reference. iCliniqQAs emphasizes relational values, where rewriting yields larger measurable gains. Empathy prompting produces mod- erate improvements in both datasets. Epistemic superiority over physician-authored answers is not observed. Takeaway 5☞ Collaborative rewriting consistently improves relational alignment across datasets but does not surpass physician-authored responses on epistemic criteria. The largest gains emerge in conversational clinical data. Rewriting acts as a communication enhancement mechanism rather than a substitute for clinical expertise. 5 Discussion Our findings provide a structured perspective on the role of large language models (LLMs) in patient-directed clinical communication. The results must be interpreted across two distinct settings: institutionally curated medical explanations (MedQuAD) and real-world physician–patient consultations (iCliniqQAs). Regarding RQ1, LLMs do not reproduce physician affective distributions. In MedQuAD, physician answers concentrate in the Neutral category with substan- tial Very Negative content. Baseline LLM configurations increase affective polarity, particularly Very Negative sentiment. Empathy prompting and collaborative rewrit- ing reduce extreme negativity and increase Neutral proportions. Gemini Rephrase introduces non-negligible Positive sentiment, which is absent in physician-authored content. In iCliniqQAs, physicians exhibit strong Neutral dominance and minimal Very Neg- ative content. LLM outputs show lower polarity amplification than in MedQuAD but still exhibit systematic emotional shifts. Rephrase configurations further increase Neu- tral proportions, often exceeding physician baselines. Fine-grained emotion analysis confirms a consistent amplification of caring signals across models, while disapproval and corrective cues are attenuated. This behavior aligns with prior observations that LLMs tend to generate warmer and more supportive language in clinical contexts [40]. The pattern reflects a shift from detached concern [41] toward regulated empathy [42], but it does not imply faithful reproduction of physician affective norms. 33 For RQ2, baseline LLM generations do not systematically improve readability. In both datasets, GPT5 Base and Claude Base produce significantly higher FKGL and GFI scores than physician-authored responses. Mixtral Base and MedPaLM Base remain closer to physician readability levels. Empathy prompting and collaborative rewriting reduce linguistic complexity across architectures. These reductions are sta- tistically significant and consistent in both datasets. The results confirm that stylistic accessibility depends on alignment strategies rather than intrinsic model properties. This pattern is consistent with findings that domain specialization increases termino- logical density and lexical complexity [43]. Readability improvements therefore emerge primarily through explicit control mechanisms[44]. For RQ3, empathy-oriented prompting modifies communicative style but does not substantially alter semantic fidelity. Across both datasets, cosine similarity remains stable between base and empathy configurations. The main effect of prompting con- cerns affective distribution and moderate readability reduction. This supports prior work showing that stylistic control through prompting influences surface-level struc- ture and tone [43]. Prompt design acts as a lightweight alignment intervention but does not fundamentally reshape epistemic alignment. RQ4 shows that collaborative rewriting produces the most robust improvements across dimensions. Rephrase variants consistently achieve the highest semantic fidelity. In MedQuAD, GPT5 Rephrase reaches the strongest conceptual alignment. In iClin- iqQAs, MedPaLM Rephrase achieves the highest similarity scores and significantly outperforms baseline variants. Rewriting improves readability and reduces affective extremity without degrading semantic overlap. These results align with studies report- ing that guided rewriting outperforms prompt-only stylistic control [45]. Similar findings in AI-assisted documentation show that editing assistance enhances coher- ence and clarity without replacing clinician expertise [46]. The evidence supports a human–AI collaborative model rather than autonomous substitution. RQ5 highlights systematic divergence between expert and patient preferences. In MedQuAD, expert ratings remain anchored to physician-level epistemic standards. No LLM configuration surpasses physicians on accuracy, style, or precision. Relational gains appear primarily in patient evaluations. In iCliniqQAs, relational metrics dom- inate stakeholder differentiation. Rephrase configurations achieve near-ceiling patient trust and emotional tone scores, while expert ratings remain bounded by physician baselines. These findings confirm that stakeholder alignment is multidimensional and strongly dependent on communicative context and dataset characteristics. Across both datasets, rewriting consistently improves relational alignment with- out demonstrating epistemic superiority over physician-authored content. Gains are larger in conversational data than in institutional explanations. The difference suggests that communicative context mediates the magnitude of alignment effects. Institutional explanations impose stronger epistemic constraints. Conversational exchanges allow greater stylistic modulation. Limitations. This study relies on controlled subsets rather than full-corpus eval- uation. The MedQuAD subset is readability-stratified and the iCliniqQAs subset is severity-balanced. The reduced sample size limits statistical generalization. Sentiment 34 and emotion classifiers are general-domain models and may not capture all medical dis- course nuances [47]. Human expert evaluations were conducted by a panel of medical professionals using a structured questionnaire, though the limited number of evalua- tors constrains the statistical power of the human assessment. The study is limited to English-language data and selected architectures. Ethical Considerations. Emotional amplification may increase perceived sup- port while masking epistemic limitations. Stylistic alignment must not compromise factual rigor or induce overconfidence. Human oversight remains necessary. LLMs func- tion most effectively as communication enhancers rather than independent clinical authorities [48]. 6 Conclusion This work provides a multidimensional evaluation of large language models in clinical communication across two distinct settings: institutionally curated medical explana- tions (MedQuAD) and real-world physician–patient consultations (iCliniqQAs). We analyze semantic fidelity, readability, affective resonance, and stakeholder alignment under baseline generation, empathy prompting, and collaborative rewriting. Results show that baseline LLMs do not systematically improve accessibility or affective alignment relative to physician-authored content. Linguistic complex- ity often exceeds clinician levels, particularly in larger general-purpose architectures. Readability gains emerge primarily under explicit alignment strategies. Empathy-oriented prompting reduces affective extremity and moderately improves readability without significantly altering semantic fidelity. However, collaborative rewriting consistently yields the strongest overall improvements. Rephrase configura- tions achieve the highest semantic similarity to physician answers across both datasets. In MedQuAD, GPT5 Rephrase reaches the strongest conceptual alignment, while in iCliniqQAs MedPaLMRephrase achieves the highest semantic fidelity. Rewriting also produces the largest reductions in linguistic complexity and the most consistent gains in patient-rated trust and emotional tone. Expert evaluations confirm that no configuration surpasses physicians on epistemic criteria such as accuracy and precision. Relational improvements do not translate into epistemic superiority. Patient evaluations reveal stronger preference for rewritten variants, particularly in conversational clinical contexts, where clarity and emotional support are central. Taken together, the findings indicate that LLMs function most effectively as collaborative editing tools rather than autonomous communicators. Human–AI co- authorship improves clarity and relational alignment while preserving clinical meaning, but it does not replace physician expertise. The code and data supporting this work are publicly available at https://github. com/PRAISELab-PicusLab/CanAIBeADoctor. Future work should extend this framework to multi-turn interactions, integrate domain-adapted affective models, involve certified clinicians in structured evaluation, and expand analysis to multilingual and low-resource healthcare contexts. In addition, future research should investigate the impact of clinical question criticality on LLM 35 behavior by stratifying responses across severity levels. This would enable a fine- grained analysis of semantic fidelity, readability, and affective alignment as a function of clinical urgency. A key hypothesis is that LLMs may exhibit stronger alignment with physician responses in low-criticality scenarios, while showing degradation in high-criticality contexts that require precise reasoning, risk calibration, and cautious communication. Such an analysis would clarify whether current models are robust across the full spectrum of clinical demands or disproportionately reliable in lower- stakes settings. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Grant number: Not applicable. Appendix A Prompt Templates This appendix details the prompt formulations used across experimental conditions. Prompts differ in objective: (i) producing a direct medical answer, (i) improving clarity without stylistic modification, and (i) collaboratively rewriting physician- authored content while preserving meaning. A.1 Base Prompt (Formal Clinical Answer) Used for generating direct medical responses from general-purpose models such as Mixtral. Emphasizes accuracy, formal tone, and discursive structure without enumeration. Base Prompt [INST] <<SYS>> You are a helpful, respectful, and accurate medical doctor. Always answer using the provided context. Answer in a formal, scientific tone, in the third person. Do not enumerate or provide bullet points. Write in a continuous, discursive manner. <<SYS>> Question: query [/INST] A.2 Empathy Prompt (Clarity-Focused Prompt) Applied to Mixtral and other general-purpose models to enhance readability and accessibility. This version prioritizes simplicity and comprehension without stylistic emotional bias. 36 Empathy Prompt [INST] <<SYS>> You are a medical professional. Reformulate answers using sim- ple and accessible language while maintaining scientific accuracy. Avoid dense jargon. Use short sentences and common vocabulary. Avoid lists or numbered items. Respond in a fluent, natural way. Optimize readability (e.g., lower Flesch-Kincaid / Fog scores). <<SYS>> Question: query [/INST] A.3 Rephrase Prompt (Collaborative Human–LLM Editing) This prompt is used to rewrite physician-authored responses, ensuring clarity, warmth, and accessibility while preserving meaning. It corresponds to models such as MixtralRephrase, Med-PaLMRephrase, and GPT5Rephrase. Rephrase Prompt [INST] <<SYS>> You are a medical expert. Reformulate the provided answer to make it easier to understand for a non-expert audience while keeping the mean- ing identical. Use clear language, short sentences, and accessible vocabulary. Maintain a professional, supportive tone. Do not add new information. Avoid enumeration or bullet points. Write in continuous prose. <<SYS>> Original question: question Original answer: original-answer [/INST] Appendix B Evaluation Questionnaire Structure B.1 Human Evaluation (Patients) Patients were shown approximately 30 clinical questions, each followed by responses generated under the model configurations described in Section 3.3 and Section 4. For each response, participants rated: • Comprehensibility: The response was easy to understand. • Perceived Trustworthiness: I would trust this response in a real medical context. • Emotional Tone: The tone felt supportive and reassuring. Ratings were collected on a 5-point Likert scale: 37 Score Interpretation 1Strongly disagree 2Disagree 3Neutral 4Agree 5Strongly agree Responses appeared in randomized order to reduce position bias, and participants were blinded to whether a response originated from physicians or LLMs. B.2 Expert Evaluation (Human Panel) Expert evaluations were collected through a structured Google Form administered to medical professionals. Each evaluator independently rated model-generated responses using the same 5-point Likert scale employed in the main study. Experts assessed each response along three criteria: • Clinical Accuracy • Stylistic Appropriateness • Linguistic Precision Scores ranged from 1 (strongly disagree) to 5 (strongly agree). Each response was evaluated independently without exposure to model identity to reduce bias. The evaluation form presented the clinical question followed by the generated response. Experts were instructed to provide numerical ratings only. Expert Evaluation Form (Google Form Structure) Question: question Answer: model-generated-answer Please assign a score from 1 (strongly disagree) to 5 (strongly agree): Clinical Accuracy; Stylistic Appropriateness; Linguistic Precision Evaluations were aggregated by computing the mean and standard deviation for each model configuration. Appendix C Qualitative Example Table C1 reports representative Base configuration responses to a sample question. Gemini and Claude exhibit the safety-driven, hedged behavior discussed in Section 4, in which the absence of sufficient clinical context leads to evasive or non-specific outputs rather than direct medical answers. 38 Table C1: Representative Base configuration responses to the question: “What are the side effects of using ibuprofen?” SystemResponse Physician AnswerIbuprofen can cause gastrointestinal issues such as stomach pain, nausea, or ulcers, especially when taken in high doses or over long periods. Mixtral (Base)Ibuprofen may lead to nausea, headache, dizziness, or stomach pain. Prolonged use increases the risk of ulcers. Med-PaLM (Base)Common adverse effects include gastrointestinal discomfort, nausea, headache, and dizziness, with rare instances of ulceration or renal complications. Gemini (Base)I’m not able to provide medical advice or information about specific drug side effects. It’s best to speak with a healthcare professional or consult official medical sources for guidance. Claude (Base)Ibuprofen is commonly used for pain relief, but how someone reacts can vary. Some individuals may notice discomfort or sensitivity after taking it. If you’re unsure how it affects you or have concerns, discussing it with a healthcare provider can help clarify what to expect. GPT-5 (Base)Some people may feel mild stomach discomfort, nausea, or dizziness when taking ibupro- fen. These effects are usually temporary, but check with a doctor if symptoms persist, especially with long-term use. References [1] Oliver Klingefjord, Ryan Lowe, and Joe Edelman. What are human values, and how do we align ai to them? 2024. [2] Paul R ̈ottger, Fabio Pernisi, Bertie Vidgen, and Dirk Hovy. Safetyprompts: a systematic review of open datasets for evaluating and improving large language model safety. 2025. [3] Antonis A. Armoundas and Joseph Loscalzo. Patient agency and large language models in worldwide encoding of equity. npj Digit. Medicine, 8(1), 2025. [4] Jiyeon Han, Jimin Park, Jinyoung Huh, Uran Oh, Jaeyoung Do, and Daehee Kim. Ascleai: A llm-based clinical note management system for enhancing clinician productivity. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pages 1–7, 2024. [5] Sina Shool, Sara Adimi, Reza Saboori Amleshi, Ehsan Bitaraf, Reza Golpira, and Mahmood Tara. A systematic review of large language model (llm) evaluations in clinical medicine. BMC Medical Informatics and Decision Making, 25(1):117, 2025. [6] Marium M. Raza, Kaushik P. Venkatesh, and Joseph C. Kvedar. Generative AI and large language models in health care: pathways to implementation. npj Digit. Medicine, 7(1), 2024. [7] Carlos Garcia-Fernandez, Luis Felipe, Monique Shotande, Muntasir Zitu, Aakash Tripathi, Ghulam Rasool, Issam El Naqa, Vivek Rudrapatna, and Gilmer Valdes. Trustworthy ai for medicine: Continuous hallucination detection and elimination with check. 2025. 39 [8] Elham Asgari, Nina Monta ̃na Brown, Magda Dubois, Saleh Khalil, Jasmine Bal- loch, Joshua Au Yeung, and Dominic Pimenta. A framework to assess clinical safety and hallucination rates of llms for medical text summarisation. npj Digit. Medicine, 8(1), 2025. [9] Matthew K Wynia and Chandra Y Osborn. Health literacy and communi- cation quality in health care organizations. Journal of health communication, 15(S2):102–115, 2010. [10] Martina Horvat, Ivan Erˇzen, and Dominika Vrbnjak. Barriers and facilitators to medication adherence among the vulnerable elderly: a focus group study. In Healthcare, volume 12, page 1723. MDPI, 2024. [11] Lucille ML Ong, Johanna CJM De Haes, Alaysia M Hoos, and Frits B Lammes. Doctor-patient communication: a review of the literature.Social science & medicine, 40(7):903–918, 1995. [12] Walter F Baile, Robert Buckman, Renato Lenzi, Gary Glober, Estela A Beale, and Andrzej P Kudelka. Spikes—a six-step protocol for delivering bad news: application to the patient with cancer. The oncologist, 5(4):302–311, 2000. [13] Suzanne Kurtz, Jonathan Silverman, John Benson, and Juliet Draper. Marrying content and process in clinical method teaching: enhancing the calgary–cambridge guides. Academic Medicine, 78(8):802–809, 2003. [14] Monica Agrawal, Irene Y. Chen, Freya Gulamali, and Shalmali Joshi. The eval- uation illusion of large language models in medicine. npj Digit. Medicine, 8(1), 2025. [15] Asma Ben Abacha, Eugene Agichtein, Yuval Pinter, and Dina Demner-Fushman. Overview of the medical question answering task at TREC 2017 liveqa. In Ellen M. Voorhees and Angela Ellis, editors, Proceedings of The Twenty-Sixth Text REtrieval Conference, TREC 2017, Gaithersburg, Maryland, USA, Novem- ber 15-17, 2017, volume 500-324 of NIST Special Publication. National Institute of Standards and Technology (NIST), 2017. [16] Ingrid M Nembhard, Guy David, Iman Ezzeddine, David Betts, and Jennifer Radin. A systematic review of research on empathy in health care. Health services research, 58(2):250–263, 2023. [17] John W. Ayers, Adam Poliak, Mark Dredze, Eric C. Leas, and et al. Comparing physician and artificial intelligence chatbot responses to patient questions posted to a public social media forum. JAMA Internal Medicine, 183:589–596, 2023. [18] David Chen, Kabir Chauhan, Rod Parsa, Zhihui Amy Liu, Fei-Fei Liu, Ernie Mak, Lawson Eng, Breffni Louise Hannon, Jennifer Croke, Andrew Hope, Nazanin Fallah-Rad, Phillip Wong, and Srinivas Raman. Patient perceptions of empathy in 40 physician and artificial intelligence chatbot responses to patient questions about cancer. npj Digit. Medicine, 8(1), 2025. [19] Man Luo, Christopher J Warren, Lu Cheng, Haidar M Abdul-Muhsin, and Imon Banerjee. Assessing empathy in large language models with real-world physician-patient interactions. In 2024 IEEE International Conference on Big Data (BigData), pages 6510–6519. IEEE, 2024. [20] Joanna M. Roy, Elias Atallah, Keenan Piper, Shyam Majmundar, Nikolaos Mouchtouris, D. Mitchell Self, Anand Kaul, Saman Sizdahkhani, Basel Musmar, Stavropoula I. Tjoumakaris, Michael R. Gooch, Robert H. Rosenwasser, and Pascal M. Jabbour. Comparison of quality, empathy and readability of physi- cian responses versus chatbot responses to common cerebrovascular neurosurgical questions on a social media platform. Clinical Neurology and Neurosurgery, 255:108986, 2025. [21] Rohaid Ali, Ian D. Connolly, Oliver Y. Tang, Fatima N. Mirza, Benjamin John- ston, Hael F. Abdulrazeq, Paul F. Galamaga, Tiffany J. Libby, Neel R. Sodha, Michael W. Groff, Ziya L. Gokaslan, Albert E. Telfeian, John H. Shin, Wael F. Asaad, James Zou, and Curtis E. Doberstein. Bridging the literacy gap for sur- gical consents: an ai-human expert collaborative approach. npj Digit. Medicine, 7(1), 2024. [22] Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. [23] Ling Wang, Jinglin Li, Boyang Zhuang, Shasha Huang, Meilin Fang, Cunze Wang, Wen Li, Mohan Zhang, and Shurong Gong. Accuracy of large language models when answering clinical research questions: Systematic review and network meta- analysis. J Med Internet Res, 27:e64486, Apr 2025. [24] Moaiz Abrar, Yusuf Sermet, and Ibrahim Demir. An empirical evaluation of large language models on consumer health questions. 2024. [25] Kexin Ding, Mu Zhou, Akshay Chaudhari, Shaoting Zhang, and Dimitris N. Metaxas.Aligning large language models with healthcare stakeholders: A pathway to trustworthy ai integration. 2025. [26] Charumathi Raghu Subramanian, Daniel A Yang, and Raman Khanna. Enhanc- ing health care communication with large language models—the role, challenges, and future directions. JAMA Network Open, 7(3):e240347–e240347, 2024. [27] Thomas Yu Chow Tam, Sonish Sivarajkumar, Sumit Kapoor, Alisa V. Stolyar, Katelyn Polanska, Karleigh R. McCarthy, Hunter Osterhoudt, Xizhi Wu, Shyam Visweswaran, Sunyang Fu, Piyush Mathur, Giovanni E. Cacciamani, Cong Sun, 41 Yifan Peng, and Yanshan Wang. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digit. Medicine, 7(1), 2024. [28] Shirui Wang, Zhihui Tang, Huaxia Yang, Qiuhong Gong, Tiantian Gu, Hongyang Ma, Yongxin Wang, Wubin Sun, Zeliang Lian, Kehang Mao, Yinan Jiang, Zhicheng Huang, Lingyun Ma, Wenjie Shen, Yajie Ji, Yunhui Tan, Chunbo Wang, Yunlu Gao, Qianling Ye, Rui Lin, Mingyu Chen, Lijuan Niu, Zhihao Wang, Peng Yu, Mengran Lang, Yue Liu, Huimin Zhang, Haitao Shen, Long Chen, Qiguang Zhao, Si-Xuan Liu, Lina Zhou, Hua Gao, Dongqiang Ye, Lingmin Meng, Youtao Yu, Naixin Liang, and Jianxiong Wu. A novel evaluation benchmark for medical llms illuminating safety and effectiveness in clinical domains. npj Digit. Medicine, 9(1), 2026. [29] J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. Derivation of new readability formulas (automated readability index, fog count and flesch reading ease formula) for navy enlisted personnel. –, 1975. [30] Philip P. Gross and Karen Sadowski. Fogindex: A readability formula program for microcomputers. Journal of Reading, 28(7):614–618, 1985. [31] Vadim Borisov and Richard H. Schreiber. robust-sentiment-analysis (revision c542a28). 2025. [32] Sam Lowe. roberta-base-go emotions (revision 58b6c5b). 2024. [33] Asma Ben Abacha and Dina Demner-Fushman. A question-entailment approach to question answering. BMC Bioinform., 20(1):511:1–511:23, 2019. [34] Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lam- ple, L ́elio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Th ́eophile Gervet, Thibaut Lavril, Thomas Wang, Timoth ́e Lacroix, and William El Sayed. Mixtral of experts. CoRR, abs/2401.04088, 2024. [35] Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Nathaneal Scharli, Aakanksha Chowdhery, Philip Mansfield, Blaise Aguera y Arcas, Dale Webster, Greg S. Corrado, Yossi Matias, Katherine Chou, Juraj Gottweis, Nenad Tomasev, Yun Liu, Alvin Rajkomar, Joelle Barral, Christopher Semturs, Alan Karthike- salingam, and Vivek Natarajan. Large language models encode clinical knowledge, 2022. 42 [36] Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, Marcel Blistein, Ori Ram, Dan Zhang, Evan Rosen, Luke Marris, Sam Petulla, Colin Gaffney, Asaf Aharoni, Nathan Lintz, Tiago Cardal Pais, Henrik Jacobsson, Idan Szpektor...., and Wesley Helmholz. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities, 2025. [37] Soumik Mandal, Batia Mishan Wiesenfeld, Adam Szerencsy, William R. Small, Vincent J. Major, Safiya Richardson, Antoinette M. Schoenthaler, Devin M. Mann, and Oded Nov. Utilization of generative ai-drafted responses for managing patient-provider communication. npj Digit. Medicine, 8(1), 2025. [38] Pritam Deka, Anna Jurek-Loughrey, et al. Evidence extraction to validate medical claims in fake news detection. In International Conference on Health Information Science, pages 3–15. Springer, 2022. [39] Yoav Benjamini and Yosef Hochberg. Controlling the false discovery rate: a prac- tical and powerful approach to multiple testing. Journal of the Royal statistical society: series B (Methodological), 57(1):289–300, 1995. [40] Xiangbin Meng, Xiangyu Yan, Kuo Zhang, Da Liu, Xiaojuan Cui, Yaodong Yang, Muhan Zhang, Chunxia Cao, Jingjia Wang, Xuliang Wang, et al. The application of large language models in medicine: A scoping review. Iscience, 27(5), 2024. [41] Clarissa Guidi and Chiara Traversa. Empathy in patient care: from ‘clinical empa- thy’to ‘empathic concern’. Medicine, Health Care and Philosophy, 24(4):573–585, 2021. [42] Yoon Kyung Lee, Jina Suh, Hongli Zhan, Junyi Jessy Li, and Desmond C. Ong. Large language models produce responses perceived to be empathic, 2024. [43] Zihao Li, Samuel Belkadi, Nicolo Micheletti, Lifeng Han, Matthew Shardlow, and Goran Nenadic. Investigating large language models and control mechanisms to improve text readability of biomedical abstracts. In 2024 IEEE 12th International Conference on Healthcare Informatics (ICHI), pages 265–274. IEEE, 2024. [44] Mengting Wang, Haoming Ma, and Meihua Piao. Effectiveness of large language models in preoperative and discharge education: a systematic review based on an evaluation framework. npj Digit. Medicine, 9(1), 2026. [45] Avanti Bhandarkar, Ronald Wilson, Anushka Swarup, and Damon Woodard. Emulating author style: a feasibility study of prompt-enabled text stylization with off-the-shelf llms. In Proceedings of the 1st Workshop on Personalization of Generative AI Systems (PERSONALIZE 2024), pages 76–82, 2024. 43 [46] Archana Reddy Bongurala, Dhaval Save, Ankit Virmani, and Rahul Kashyap. Transforming health care with artificial intelligence: redefining medical documen- tation. Mayo Clinic Proceedings: Digital Health, 2(3):342–347, 2024. [47] Zixiao Zhu and Kezhi Mao. Knowledge-based bert word embedding fine-tuning for emotion recognition. Neurocomputing, 552:126488, 2023. [48] Lars Riedemann, Maxime Labonne, and Stephen Gilbert. The path forward for large language models in medicine is open. npj Digit. Medicine, 7(1), 2024. 44