Paper deep dive
Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology
Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/23/2026, 2:04:26 AM
Summary
This paper introduces an evaluation-first, retrieval-augmented generation (RAG) framework for auditing Large Language Model (LLM) hypotheses generated from longitudinal Cell Painting morphology data. The framework processes a 9-week time course of RPE-1 cells exposed to low-dose-rate ionizing radiation, combining quantitative morphology deltas with retrieved perturbation neighbors, pathway context, and literature evidence. It employs two quantitative auditing tests: V1 (citation validity) to ensure evidence grounding, and V2 (proxy-based morphology compatibility) to verify biological consistency. The study demonstrates that the framework produces auditable, falsifiable hypotheses, identifying adaptive phenotypes involving metabolic reprogramming and proteostatic stress at lower dose rates.
Entities (12)
Relation Signals (8)
Cell Painting â usedfor â Longitudinal Morphology Profiling
confidence 95% · High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state
LLM â augmentedby â RAG
confidence 92% · retrieval-augmented generation (RAG) improves provenance by coupling generation to external evidence stores
V2 â evaluates â Morphology Compatibility
confidence 90% · V2 proxy-based morphology compatibility, which evaluates consistency between predicted biological processes and the most altered morphology features
V1 â verifies â Citation Validity
confidence 90% · V1 citation validity, which verifies that cited evidence identifiers exist in the prompt
Reactome â provides â Pathway Context
confidence 88% · pathway context... Reactome provide standardized pathway context
JUMP â provides â Perturbation Neighbors
confidence 88% · retrieved perturbation neighbors... JUMP Cell Painting Consortium
Low-dose-rate ionizing radiation â causes â Metabolic reprogramming
confidence 85% · adaptive phenotype involving metabolic reprogramming and proteostatic stress at lower dose rates
Low-dose-rate ionizing radiation â causes â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology remains difficult, especially for weak, chronic perturbations such as low-dose-rate ionizing radiation. Large language models (LLMs) can synthesize heterogeneous evidence into biological narratives, yet their scientific use requires quantitative auditing. We present an evaluation-first, retrieval-augmented interpretation framework for longitudinal Cell Painting morphology, applied to a 9-week RPE-1 time course across five dose rates (0.003--6.0 mGy/hr). Week-matched treated-control morphology deltas are combined with retrieved perturbation neighbors, pathway context, and literature evidence through stable evidence identifiers, enabling an LLM to generate structured, evidence-linked hypotheses that are hierarchically summarized while preserving provenance. We introduce two quantitative auditing tests: V1 citation validity, which verifies that cited evidence identifiers exist in the prompt, and V2 proxy-based morphology compatibility, which evaluates consistency between predicted biological processes and the most altered morphology features. In our experiments, V1 detected no invalid evidence references, while V2 showed meaningful morphology compatibility that increased with perturbation strength and was positively associated with an independent morphology drift summary. The framework produces auditable, falsifiable biological hypotheses, including an adaptive phenotype involving metabolic reprogramming and proteostatic stress at lower dose rates (0.003--0.3 mGy/hr). Current limitations include proxy-based evaluation and the lack of ground-truth mechanism labels.
Tags
Links
- Source: https://arxiv.org/abs/2607.19415v1
- Canonical: https://arxiv.org/abs/2607.19415v1
Trouble viewing inline? Open PDF directly â
Full Text
66,499 characters extracted from source content.
Expand or collapse full text
Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo Brookhaven National Laboratory Upton, New York, USA gpark,gzhao,byoon,sjyoo@bnl.gov Abstract High-content morphological profiling (Cell Painting) yields sensi- tive, high-dimensional signatures of cellular state, but translating longitudinal morphology trajectories into interpretable biology re- mains difficultâespecially for weak, chronic perturbations such as low-dose-rate ionizing radiation. Large language models (LLMs) can synthesize heterogeneous evidence into narratives, yet sci- entific use requires quantitative auditing: fluent outputs may be ungrounded or inconsistent with measured morphology. We present an evaluation-first, retrieval-augmented interpre- tation framework for longitudinal Cell Painting morphology. To demonstrate the frameworkâs utility in a highly challenging real- world scenario characterized by weak signals and sparse refer- ence data, we apply it to a 9-week RPE-1 time course spanning five dose rates (0.003â6.0 mGy/hr). We compute week-matched treatedâcontrol morphology deltas and construct evidence-rich prompt payloads that bind (i) quantitative deltas to (i) retrieved per- turbation neighbors, pathway context, and literature snippets via stable evidence identifiers. An LLM generates structured controlled- vocabulary hypotheses at the weekĂdose level, which are then inte- grated hierarchically into dose-time phases and global summaries while preserving evidence provenance. We operationalize auditing with two quantitative tests: V1 ci- tation validity, which verifies that every cited evidence identifier is present in the prompt payload, and V2 proxy-based morphol- ogy compatibility, which measures whether predicted process la- bels are consistent with the top changed morphology features un- der an explicit proxy table. In the audited run, V1 detected no invalid evidence references, and V2 showed non-trivial morphol- ogy compatibility that increased with perturbation strength. At weekĂdose resolution, V2 also showed a modest positive associa- tion with an LLM-external morphology drift summary computed from the same Cell Painting data distribution. Crucially, this au- ditable pipeline distilled complex trajectories into falsifiable hy- potheses, including a distinct adaptive phenotype associated with metabolic reprogramming and proteostatic stress at lower dose rates (0.003-0.3 mGy/hr). Limitations include proxy-based evalua- tion granularity and the absence of ground-truth mechanism labels. The source code and datasets for this study are publicly available at https://github.com/cellpainting-llm-auditor. CCS Concepts âą Applied computingâ Bioinformatics. Keywords Cell Painting, morphological profiling, longitudinal phenotyping, low-dose ionizing radiation (LDIR), LDR, Large Language Models (LLMs), retrieval-augmented generation, LLM auditing, hypothesis generation 1 Introduction High-content imaging assays provide a sensitive readout of how perturbations reshape cellular state. Cell Paintingâstyle morpholog- ical profiling converts multiplexed fluorescence images into quan- titative single-cell feature vectors spanning size, shape, intensity, and texture [3]. Repeated measurement enables longitudinal tra- jectories, but mechanistic interpretation is particularly challenging for weak, chronic perturbations such as low-dose and low-dose- rate ionizing radiation, where responses can be multifactorial and context dependent [22, 31]. The interpretability bottleneck is largely downstream of feature extraction. Morphological profiles are high-dimensional and sensi- tive to technical variation, often requiring robust normalization and aggregation before comparisons are meaningful across plates and timepoints [4,5]. Even after preprocessing, translating a trajectory in morphology space into biological hypotheses typically remains manual. Public reference resources partially mitigate this gap via âprofile matching,â where new signatures are compared to large perturbation libraries (e.g., JUMP) to suggest candidate gene pro- grams or pathways [5]. However, similarity alone rarely yields an end-to-end explanationâespecially in longitudinal settings where dominant processes may shift by dose or phase. Large language models (LLMs) offer an appealing interface for synthesizing heterogeneous evidence, but scientific deployment requires auditing: fluent narratives can be unfaithful to evidence and may include fabricated or incorrect citations [11,28]. Retrieval- augmented generation (RAG) improves provenance by coupling generation to external evidence stores [15], and curated resources such as Reactome provide standardized pathway context [20]. In this work, we introduce an LLM-centered, retrieval-augmented interpretation framework for longitudinal Cell Painting morphol- ogy that is evaluation-first by design. To demonstrate the practical utility of this framework in a highly challenging real-world scenario, we apply it to the longitudinal profiling of chronic low-dose-rate (LDR) ionizing radiationâa domain where morphological signals are notoriously weak and established reference mechanisms are scarce. Rather than treating LLM narratives as conclusions, we treat them as auditable scientific artifacts: structured hypotheses grounded in (i) measured morphology deltas, and (i) retrieved exter- nal evidence (perturbation neighbors, pathway context, literature snippets). Specifically, we operationalize auditing through two rigorous quantitative tests. First, we enforce citation integrity (V1) by veri- fying that every piece of evidence cited by the LLM strictly maps arXiv:2607.19415v1 [q-bio.QM] 17 Jul 2026 Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo Figure 1: Architectural overview of the evaluation-first LLM-driven morphological reasoning and validation pipeline. The framework consists of four integrated modules: (A) Data Acquisition & Processing, where raw microscopy images undergo feature extraction and delta Z-score normalization to yield quantitative morphology profiles; (B) Static RAG Backend, which constructs an integrated semantic knowledge payload by synthesizing morphological reference profiles (JUMP-CP), mechanistic biological pathways (Reactome), and scientific literature contexts (EuropePMC); (C) Hierarchical LLM Reasoning, employing the Gemini large language model to iteratively consolidate local condition-specific inferences (weekĂ dose) into temporal trajectories and unified global hypotheses; and (D) Quantitative Auditing, the primary evaluation module where hypotheses are rigorously anchored against grounding integrity (V1: citation validity) and morphological alignment (V2: quantitative predictive hit-rate) to produce audited, falsifiable biological discoveries. Furthermore, the V2 auditing metric is meta-validated against an independent external reference of trajectory-based PCA morphology drift. (Note: The visual representation of this diagram was generated with the assistance of the Gemini.) back to the provided prompt payload, effectively eliminating ref- erence hallucinations. Second, we assess biological fidelity via a proxy-based alignment metric (V2), which evaluates the consis- tency between the LLMâs predicted biological processes and the most prominent morphological feature shifts. To further substanti- ate these findings, we establish an independent, morphology-only PCA drift summary. This drift metric serves as a crucial secondary quantitative anchor, ensuring that the semantic trajectories pro- posed by the LLM are fundamentally grounded in the principal axes of phenotypic variation observed in the raw data (Figure 1). Contributions: âąEvaluation-first auditing for LLM interpretation of morphology: Formal validation of hypothesis grounding (V1) and quantitative proxy alignment (V2), designed specif- ically for settings lacking ground-truth mechanism labels. âą Traceable retrieval-augmented prompting: A robust prompt payload schema that binds quantitative morphology deltas to retrieved perturbation and literature evidence via stable IDs, enabling automatic citation verification. âą Hierarchical reasoning with provenance: A multi-step synthesis approach (dose-time and global summaries) that preserves longitudinal structure and reliably propagates evidence references, preventing the loss of context typical of flattened weekĂ dose outputs. 2 Related Work 2.1 Cell Painting and morphological profiling Cell Painting is a widely adopted foundation for image-based mor- phological profiling, yielding multi-channel fluorescence images that can be summarized into quantitative single-cell feature vectors [3]. A recurring theme in the literature is that interpretability is limited less by feature extraction than by downstream processing and contextualization: profiles often require robust normalization, Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology aggregation, and careful similarity computation before they can be compared across batches or experimental conditions [5]. Recently, large public reference resources have emphasized both the promise and complexity of morphological profiling. The JUMP Cell Painting Consortium released a benchmark-scale dataset span- ning matched chemical and genetic perturbations to support simi- larity evaluation and representation learning [5]. While early ap- proaches relied on classical image analysis pipelines, modern ef- forts increasingly leverage deep representation learningâsuch as weakly supervised learning and vision transformersâto extract more robust phenotypic embeddings [4, 5]. Such resources enable âprofile matchingâ as a practical mechanism-discovery heuristic, but nearest-neighbor similarity typically requires additional path- way and literature context. This remains especially challenging for longitudinal settings where phenotypes evolve dynamically over time. 2.2 LLMs, RAG systems, and biological interpretation Biomedical LLMs have become central tools for text mining and syn- thesis, building on domain-adapted pretraining (e.g., BioBERT, Pub- MedBERT) and subsequent generative models tailored for biomedi- cal knowledge tasks [8,14,17]. More recently, state-of-the-art mod- els like Med-PaLM have demonstrated expert-level performance on medical and biological reasoning benchmarks, signaling a shift toward highly capable, domain-specific reasoning engines [29]. In parallel, retrieval-augmented generation (RAG) couples a gen- erator with an evidence store to improve provenance and factu- ality for knowledge-intensive generation [15]. In biology, curated resources such as Reactome provide standardized pathway defini- tions and geneâpathway mappings that are well suited for evidence- aware interpretation pipelines [20]. The integration of RAG into sci- entific workflows is now moving beyond simple question-answering toward autonomous hypothesis generation and complex data inter- pretation [2]. 2.3 Grounding, evaluation, and interpretation auditing Despite their capabilities, neural generation in scientific settings can be fluent yet unfaithful to underlying evidence, motivating explicit grounding constraints and automated checks [1,19]. Prior work proposes faithfulness evaluations and truthfulness benchmarks (e.g., TruthfulQA) [16] and scientific claim verification benchmarks (e.g., SciFact) [32]. More broadly, documentation and auditing frame- works (e.g., model cards, datasheets) motivate treating AI artifacts as objects requiring structured reporting and evaluation [7, 21]. To the best of our knowledge, no existing computational frame- work natively integrates longitudinal Cell Painting trajectories with RAG-driven LLM synthesis and quantitative auditing. Traditional morphological profiling largely relies on single-time-point prox- imity matching (e.g., using JUMP-CP benchmarks), which limits its ability to capture weak, continuous perturbations unfolding over multiple weeks. Conversely, standard RAG applications in biology primarily operate on textual data and lack direct grounding in quantitative, image-derived morphological changes. Our work bridges this gap by introducing an evaluation-first auditing framework for LLM-generated hypotheses in longitudi- nal morphology. Rather than assuming that retrieval augmenta- tion alone guarantees correctness, we formalize measurable con- straintsâsuch as grounding integrityâand introduce quantitative alignment tests, including proxy-based morphology consistency. Crucially, these LLM-centric evaluation metrics are anchored to an independent morphology-only reference used strictly for valida- tion, enabling a new paradigm of auditable morphological reasoning where direct empirical comparison to existing approaches remains challenging. 3 Methods 3.1 Experimental design and dataset We studied long-term cellular responses to chronic low-dose ioniz- ing radiation in RPE-1 (human retinal pigment epithelial) cells using the Cell Painting assay [3]. Cells were exposed at five dose rates (mGy/hr): 0.003, 0.03, 0.3, 3.0, and 6.0, and profiled longitudinally from week 1 through week 9 (9 weeks total; Table A1). Paired treatedâcontrol design. For each week, the treated plate (designated as Plate 3) was paired with a same-week control plate (Plate C). Within each weekí€and doseí, we analyzed four treated replicate wells (180 treated wells total). Control wells were defined on Plate C (2 wells in week 1; 4 wells in weeks 2â9; 34 control wells total), yielding 214 wells analyzed overall (Table A1). Cell-level file structure. Cell segmentation and feature extrac- tion produced per-image tables of single-cell features. For each well, images were acquired iní â 1, . . .,9fields of view (FOVs) andí â 1, . . .,15focal planes (135 images per well nominal). Across the study, 28,886 cell-level tables were analyzed (expected 28,890; four files missing), comprising 6,375,699 single-cell obser- vations (Table A1). Single-cell segmentation was performed using the CellSAM model [18] to extract instance-level masks prior to morphological feature computation. 3.2 Morphology preprocessing pipeline (Steps 1â3) The preprocessing pipeline converts cell-level observations into well-level treatedâcontrol deltas suitable for retrieval and LLM prompting. 3.2.1Metadata construction (Step 1). We parse the directory struc- ture and file naming convention to map each cell-level file to a well identifier and experimental context(í€,í, condition)under the paired Plate 3/Plate C design. The result is stored as a metadata index table. 3.2.2 Cell-level to well-level aggregation (Step 2). For each well, we pool all single-cell observations across its available FOVĂplane files and compute robust summary statistics for each base morphol- ogy feature: median, median absolute deviation (MAD), 10th per- centile, and 90th percentile. We exclude geometry/QC-like columns orientation,centroid_x, andcentroid_yduring aggregation. For the nine base features (area,perimeter,eccentricity, solidity,mean_intensity,glcm_contrast,glcm_correlation, glcm_energy,glcm_homogeneity), this yields 36 well-level sum- mary features. Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo 3.2.3Robust within-week normalization and treatedâcontrol deltas (Step 3). Letí„ í,í denote the aggregated value for wellíand feature í, and letí€(í)denote the week of wellí. We perform robust within- week normalization using the per-week feature median and MAD computed across all wells (treated and control) in the same week. Lettingí(í)denote the set of all wells in weekí€(í), and with í= 10 â6 , the normalized value is í§ í,í = í„ í,í â median íâí(í) í„ í,í MAD íâí(í) í„ í,í +í .(1) We then define the week-specific control baseline vector as the feature-wise median of normalized control wellsí¶ í€ in weekí€. The treatedâcontrol delta for well í is Î í,í = í§ í,í â median íâí¶ í€(í) í§ í,í .(2) To ensure stable baselines, we require at least two control wells per week and record control coverage diagnostics. Dataset scale and preprocessing exclusions are summarized in Table A1. 3.3 LLM-external morphology drift summary (Step 4) We compute an LLM-external summary of morphology drift over time for validation-oriented sanity checks only. This module is inde- pendent of the retrieval and generation stack, but it is still derived from the same Cell Painting data distribution and is therefore not treated as independent biological validation. 3.3.1PCA fitting and drift definition. We fit PCA on the normalized well profiles (í§) restricted to control wells (Plate C) across all weeks. We useí PC =12 components and apply whitening. We then project each delta profileÎ í into this PCA space and define drift magnitude as the Euclidean norm of the whitened PCA score vector: í· í = â„ PCA whiten (Î í ) â„ 2 .(3) For each(í€,í), we aggregate drift by averagingí· í across the four treated replicate wells. 3.3.2 WeekĂdose and dose-level summaries. We summarize drift at two resolutions. First, for each(í€,í)we retain the weekĂdose mean drift as a higher-resolution morphology-only summary. Sec- ond, for each doseí, we summarize the week-by-week drift trajec- tory using (i) a linear slope computed by least-squares regression of mean drift against week index, and (i) an area-under-the-curve (AUC) computed by trapezoidal integration over weeks. These sum- maries are used only as LLM-external quantitative anchors for sanity checks on score behavior. 3.4 Retrieval-augmented prompt construction (Steps 5â6) We construct a structured evidence payload for each(í€,í)to sup- port grounded, traceable LLM hypothesis generation. 3.4.1 Reduced morphology neighbor index. We build a reduced- feature index over a public Cell Painting perturbation reference (JUMP). We compute a 9-dimensional embedding corresponding to the base morphology feature set above by mapping each feature to matching Cell Painting columns and averaging across available channels/scales. Reference embeddings are z-scored using mean and standard deviation withí=10 â8 . The index uses exact L2 neighbor search (FAISS when available), otherwise a scikit-learn nearest-neighbor backend with Euclidean distance [12, 25]. 3.4.2 Neighbor search and selection. For each(í€,í)observation, we query the reference index withí=100 candidate neighbors and retain up to 10 for prompting, enforcing that at least three neighbors correspond to gene perturbations when possible. Neighbor records are assigned stable evidence IDs of the formjump:<jump_doc_id>. 3.4.3 Reactome pathway retrieval. We construct a retrieval index over Reactome pathway descriptions using TFâIDF (1â2 grams; cosine similarity;min_df=2;max_features=250000) [20]. For each (í€,í)payload, we aggregate enriched pathways for Entrez genes associated with the candidate neighbors and retain the top 12 path- ways by frequency. For dose-time and global aggregation (Step 8), pathway references can be cited usingpath:<pathway_id_or_name> and are validated against the payload; for weekĂdose generation (Step 7), enforceable retrieval IDs are limited tojump:andlit: (see below). 3.4.4 Literature snippets and quantitative observations. When en- abled, we retrieve brief literature snippets per prompt via the Eu- rope PMC REST API [6], using radiation-focused queries augmented with the cell line and additional terms (dose rate,protracted, long-term). We retrieve up to 10 hits per query and optionally pull and truncate selected full-text sections when available (sections methods,results,discussion,supplementary; section character cap 5000). Each snippet is assigned a stable IDlit:q<query_idx>: <lit_id>. Each payload also includes a quantitative observation block with a unique IDobs_w<week>_d<dose>and the top changed morphology features by absolute mean delta. 3.4.5Prompt payload schema and traceability. Each JSONL prompt record stores: (i) the observation block (obs ID and top-changed features), (i) selected neighbors (withjump:IDs), (i) literature snippets (withlit:IDs), and (iv) pathway evidence (pathway IDs). These evidence identifiers form the reference set used for grounding validation (V1). 3.5 Structured LLM hypothesis generation (Step 7) We run an LLM reasoner per(í€,í)prompt with strict structured- output constraints. 3.5.1Model and decoding. We use Googleâs Gemini model via Ver- tex AI [30] with temperature 0.2, thinking levelHIGH, maximum output length 8192 tokens, and a prompt length cap of 300,000 char- acters. The runner retries up to three times on schema or grounding failures. The specific model revision served by the endpoint was gemini-3-pro-preview. 3.5.2 JSON-onlyschemaandcontrolledvocabulary. The model is instructed to output JSON only, conforming to a schema with one or more hypotheses per prompt. Each hypothesis includes:hypothesis_id,biological_process, summary,quant_evidence_refs,retrieval_refs, testable_predictions,confidence,andacategorical Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology process_labelconstrained to an 11-label controlled vocab- ulary (Table A2). Schema validation is enforced programmatically. 3.5.3Grounding enforcement. For weekĂdose hypotheses, ground- ing validation requires that allquant_evidence_refsare of the formobs:<obs_id>and match the promptâs observation ID, and that allretrieval_refsare of the formjump:<jump_id>or lit:<lit_id>and appear in the payload. We also require that each hypothesis cites at least one observation and one retrieved item. 3.6 Hierarchical integration across scales (Step 8) Longitudinal interpretation benefits from hierarchical reasoning that preserves temporal structure. 3.6.1Dose-time aggregation. We construct dose-time prompts by grouping weekĂdose outputs across weeks for each dose and adding a timepoint catalog that retains per-week observation IDs, selected neighbors, literature snippets, and pathway evidence. A second LLM pass produces phase-level summaries (e.g., early/mid/late) with evidence references that must cite at least one observation and one retrieved item per phase. 3.6.2 Why weekĂdose first (and why not flat aggregation). WeekĂdose reasoning is performed first because it is the small- est unit in which (i) quantitative deltasÎ í,í are directly defined against the matched weekly control baseline, and (i) retrieval evi- dence (nearest neighbors, pathway context, literature snippets) is conditioned on the same quantitative signature. Aggregating all timepoints into a single prompt would mix heterogeneous tem- poral regimes and can obscure non-monotonic dynamics: feature directions can cancel when averaged across weeks, phase-specific patterns can be lost, and trends can reverse when marginalizing over time (a Simpsonâs paradox). Flat prompts also dilute prove- nance by coupling a large evidence pool to a small number of claims, increasing ambiguity about which observations support which statements. To preserve traceability, we keep evidence identifiers stable across scales. Each weekĂdose prompt assigns a unique obser- vation ID (e.g.,obs:obs_wweek_ddose) and retains retrieval IDs (jump:,lit:, and Reactome IDs). Dose_time prompts embed a per-week catalog of these IDs; global prompts embed the per-dose catalogs. Downstream, V1 grounding checks operate on these iden- tifier sets to verify that every phase- or global-level statement cites concrete week-level measurements and retrieved evidence. 3.6.3 Global summarization. We further summarize across doses by constructing a global prompt that preserves the dose-time evi- dence catalog and produces a small set of global hypotheses with evidence references. Hierarchical reasoning avoids flattening het- erogeneous timepoints and enables phase-specific claims while maintaining provenance. 3.7 Validation framework We distinguish (i) measured morphology quantities (Î, drift), (i) retrieval evidence (neighbors, snippets, pathways), and (i) LLM- generated hypotheses. We quantify two complementary validation layers. 3.7.1V1: citation validity (grounding integrity). V1 checks citation integrity: do evidence references in an LLM hypothesis correspond to evidence IDs present in the associated prompt payload? Let O(í)be the observation ID set for payloadíandE(í)the set of retrieval IDs (jump:andlit:for weekĂdose). For a hypothesis âwith reference setsQ(â)(quantitative) andR(â)(retrieval), we define V1(â,í)= I h Q(â) â O(í) â§ R(â) â E(í)â§ |Q(â)|> 0 â§ |R(â)|> 0 i . (4) V1 therefore evaluates identifier-level provenance validity within a closed evidence set. It does not test whether a cited item semanti- cally entails or fully supports a claim. 3.7.2V2: proxy-based processâmorphology compatibility. V2 tests whether a predictedprocess_labelis compatible with the promptâs quantitative morphology signature. To operationalize this comparison, we use an explicit proxy mapping between the con- trolled vocabulary of biological processes and expected directional shifts in core morphological features (Table A6). Because morpho- logical profiles are highly multiplexed, we restricted our proxy definitions to well-established, macroscopic cellular hallmarks. The mapping is intentionally simplified, but it is fully specified in the manuscript and programmatically applied, making the evaluation transparent and reproducible. For instance, cellular senescence is classically characterized by cellular hypertrophy (significantly increased area and perimeter) alongside increased cytoplasmic granularity and vacuolization (re- flected as increasedglcm_contrastand decreasedglcm_energy) [10]. Conversely, apoptosis and cell death typically manifest as severe cell shrinkage (decreased area) and chromatin condensation or membrane blebbing, which drive high textural contrast [13]. Fur- thermore, cell-cycle arrest frequently leads to continued macro- molecular synthesis without division, resulting in increased overall cell size [23]. Processes such as oxidative stress and DNA damage response, while lacking a uniform macroscopic size shift, frequently induce cytoskeletal remodeling and nuclear texture alterations iden- tifiable via GLCM metrics [26]. We emphasize that this predefined proxy mapping is not in- tended as a definitive single-cell mechanism classifier or a ground- truth benchmark. Rather, it serves as a conservative, literature- grounded consistency check designed to penalize biologically con- tradictory LLM outputs (e.g., an LLM predicting âapoptosisâ when the observation payload clearly indicates massive cell enlargement). Formally, for each labelâ, letP â be its defined proxy set over the reduced feature vocabulary. LetTdenote the set of top features provided in the prompt (ranked by|Î|). We define Hit@10 as the fraction of top features belonging to the proxy set: Hit@10(â)= |T â©P â | |T| .(5) Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo We also compute a magnitude-aware weighted proxy score as the fraction of total absolute delta mass inT explained by proxy hits: WP(â)= Ă íâTâ©P â |Î í | Ă íâT |Î í | .(6) 3.7.3 Consistency with the drift summary. To relate LLM-centric validation to an LLM-external morphology-only signal, we perform two correlation analyses. First, at weekĂdose resolution we aggre- gate per-hypothesis V2 scores to the observation level by mean (with median and max summaries reported in supplementary di- agnostics) and correlate these summaries with the corresponding weekĂdose drift mean from Step 4. Second, we compute a coarser dose-level comparison using per-dose V2 summaries and dose-level drift summaries (mean drift, slope, and AUC). Pearson and Spear- man correlations are reported with nonparametric bootstrap un- certainty (10,000 resamples). These analyses are intended as sanity checks on score behavior rather than as independent biological validation. 3.7.4 Supplementary diagnostics. Additional quantitative and re- trieval diagnostics referenced throughout the paper are provided in Supplementary Materials in Appendix. 4 Results 4.1 Grounding integrity (V1) Across the full longitudinal design (9 weeksĂ5 dose rates), the weekĂdose LLM reasoning stage produced 45/45 successful prompt records (0 failures), yielding 136 total hypotheses. Under the im- plemented citation-integrity check, no grounding failures were detected: 0 invalid evidence-reference cases and 0 records missing grounding statistics. V1 grounding validity was perfect under strict grounding criteria (Methods, Eq. 4). Both the mean observation-reference validity rate and the mean retrieval-reference validity rate were 1.0. Furthermore, all hypotheses contained at least one observation reference and one retrieval reference (coverage rate = 1.0). Aggregate V1 metrics are summarized in Table 3; additional reference-count diagnostics are provided in Table A5 and Figure A2. By construction, these V1 results establish provenance validity within the closed prompt evidence set; they do not by themselves show that a cited item semantically supports a claim. 4.2 Representative auditable hypotheses To make the audit-first outputs concrete, we highlight one represen- tative weekĂdose hypothesis alongside its supporting quantitative signature. Table 1 displays the input prompt payload for week 4 at 3.0 mGy/hr, which reports the top-|Î|morphology deltas (group- meanÎí ; values rounded to three decimals). Based on this signature and the retrieved context, the LLM generated hypothesishyp_w4_d3.0_1, classifying the state as senescence_likewith high confidence (Table 2). We present this example as an auditable biological hypothesis: its cited evidence is traceable, and its declared label can be checked for morphology- compatibility under V2. Table 1: Quantitative prompt payload (top morphology deltas) for week 4 at 3.0 mGy/hr. Featureí«í mean_intensity_median +2.145 glcm_energy_median â1.786 solidity_median+1.364 glcm_homogeneity_median â1.204 area_median+1.007 glcm_correlation_median â0.756 eccentricity_median +0.621 glcm_contrast_median â0.268 Table 2: Representative weekĂdose hypothesis generated by the LLM. FieldContent Hypothesis ID hyp_w4_d3.0_1 Process Label senescence_like (Confidence: high) SummaryâThe observed increase in cell area and texture het- erogeneity, characteristic of cellular hypertrophy, aligns with a senescence-like state. This is supported by morphological similarity to the perturbation of RRAGD, a regulator of mTOR signaling, which is a key driver of senescent cell enlargement and meta- bolic activity.â Quant Evidence obs:obs_w4_d3.0 Retrieval Refs jump:row_916481, lit:q1:4, lit:q1:6 This example demonstrates identifier-level citation validity and strong proxy compatibility under our evaluation framework. It passes V1 because all cited evidence identifiers are strictly present in the associated payload (record-level validity:obs_valid_rate=1.0, retrieval_valid_rate=1.0; 0 invalid references). Furthermore, it passes V2 proxy compatibility: under thesenescence_likemap- ping (Table A6), 7 top features are proxy hits (Hit@10=0.875), and these hits account for 86.8% of the absolute delta mass (WP=0.868; Methods, Eq. 5â6). This does not by itself prove that the cited evidence semantically entails the claim, but it shows that the hy- pothesis is auditable under both V1 and V2. A complete structured example of a global-level claim is provided in Table A7. 4.3 Proxy alignment performance (V2) We next evaluated whether the LLM-generatedprocess_label hypotheses were quantitatively compatible with observed morphol- ogy deltas using the predefined proxy mapping (Methods, Eq. 5â6; proxy definitions in Table A6). At the global (doseĂlabel) level, the mean Hit@10 over rows was 0.4335 (í dose_label_rows =14), with 0 missing-delta records. Dose-stratified results showed a clear dose dependence in best- case compatibility. The per-dose best mean Hit@10 values were 0.2361 (0.003 mGy/hr), 0.4306 (0.03 mGy/hr), 0.4861 (0.3 mGy/hr), 0.8750 (3.0 mGy/hr), and 0.8889 (6.0 mGy/hr). Figure 2 visualizes how V2 alignment varies with dose rate for both hierarchical scales (dose-time and global), reporting both the Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology 10 2 10 1 10 0 Dose rate (mGy/hr) 0.0 0.2 0.4 0.6 0.8 1.0 Hit@10 V2 proxy alignment across dose rates Dose-time (mean Hit@10) Dose-time (max Hit@10) Global (mean Hit@10) Global (max Hit@10) Figure 2: V2 proxy alignment dose-response (Hit@10). mean Hit@10 across rows and the best-case (max) Hit@10 at each dose. Quantitatively, the global best-case Hit@10 increases from 0.2361 at 0.003 mGy/hr to 0.8889 at 6.0 mGy/hr, while the global mean Hit@10 increases from 0.2292 to 0.5787 over the same range. Dose_time summaries show a similar upward trend (mean Hit@10 from 0.2662 to 0.4464; max Hit@10 from 0.5000 to 0.9167). In the evaluation-first framing, this dose-response matters because it in- dicates that the compatibility metric is sensitive to perturbation strength: when quantitative deltas are weak, proxy alignment is limited, whereas stronger perturbations permit higher compatibil- ity between declaredprocess_labels and observed morphology signatures. Table 3: Validation summary (V1 grounding, V2 proxy align- ment). MetricValue Notes V1 grounding integrity WeekĂdose records (successful)45Total prompts processed Hypotheses generated136Across all records Invalid evidence-reference cases0V1 citation failures Mean obs-reference validity1.000Valid obs:<id> / total Mean retrieval-reference validity1.000Valid jump:/lit: / total Coverage (both obs + retrieval refs)1.000Fraction of hypotheses V2 proxy alignment Global mean Hit@10 (over rows)0.4335DoseĂlabel rows (í= 14) Missing-delta records0Should be 0 Per-dose best mean Hit@10 (best row per dose) Dose 0.003 mGy/hr0.2361 Dose 0.03 mGy/hr0.4306 Dose 0.3 mGy/hr0.4861 Dose 3.0 mGy/hr0.8750 Dose 6.0 mGy/hr0.8889 4.4 Hierarchical longitudinal summaries Beyond weekĂdose hypotheses, the system produces hierarchical summaries at the dose_time and global levels while preserving evidence IDs for traceability (Methods). Dose-time summaries par- tition each dose trajectory into a small set of temporal phases with dominant process labels and supporting evidence. Figure 3 shows the dose-time integration output as a phase time- line for each dose rate, where contiguous week ranges are grouped into 3â4 phases and assigned dominantprocess_labels. This vi- sualization is not used as a biological claim; rather, it demonstrates the intended behavior of hierarchical reasoning: the model pro- duces compact, temporally structured summaries that preserve within-dose phase structure instead of collapsing all weeks into a single undifferentiated narrative. Because phases cite week-level observation IDs and retrieval IDs, they remain auditable under V1 and quantifiable under V2. Figure 4 summarizes the within-week distribution of process_labelassignments across all weekĂdose hypothe- ses. The non-uniform label usage over time provides a descriptive check that the controlled vocabulary is used in a temporally structured manner (rather than collapsing to a single default label). In the evaluation-first framing, this matters because it motivates why phase-aware aggregation (Figure 3) is preferable to flat averaging: the label distribution can shift across weeks, and hierarchical summaries provide a traceable representation of these shifts. 4.5 Higher-resolution consistency with an LLM-external morphology drift summary To test whether V2 behaves like a non-arbitrary score, we com- pared it with the Step 4 morphology drift summary computed without any LLM or retrieval components. We first performed a weekĂdose analysis by aggregating per-hypothesis V2 scores to the corresponding observation level and comparing them with the matched weekĂdose drift mean (í=45 observations). For the mean weighted-proxy summary, the association was positive but modest (Pearsoní=0.307 with 95% bootstrap CI[â0.001,0.566]; Spearman í=0.293 with 95% CI[â0.021,0.558]; bootstrapí boot =10000, seed 42). Mean Hit@10 showed a similar but slightly weaker pattern (Pearson í= 0.282; Spearman í= 0.265). Figure 5 shows this higher-resolution comparison for the mean weighted-proxy summary. The effect size is intentionally inter- preted cautiously: both quantities derive from the same Cell Paint- ing distribution, and the purpose of this analysis is only to test whether V2 co-varies with an LLM-external morphology summary rather than behaving as an arbitrary label-matching score (See Ap- pendix A.1 for a complementary coarse dose-level sanity check.). 4.6Biological interpretation under quantitative constraints. To provide a more concrete biological interpretation, we explicitly examine how LLM-generated hypotheses vary across dose regimes and how these patterns relate to observed morphological signatures. At higher dose rates (3.0â6.0 mGy/hr), the dominant hypothe- ses consistently correspond to senescence-like and DNA damage response (DDR)-related processes. This regime is characterized by strong and coherent morphological shifts, including increases in cell area and perimeter along with pronounced changes in texture Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo 123456789 Week 0.003 mGy/hr 0.03 mGy/hr 0.3 mGy/hr 3 mGy/hr 6 mGy/hr Dose rate early_metabolic_adjustment metabolic_reprogramming transient_recovery adaptation_recovery oxidative_adaptation oxidative_stress late_membrane_remodeling metabolic_reprogramming early_stress_metabolic oxidative_stress structural_remodeling_arrest mitochondrial_dysfunction adaptation_homeostasis adaptation_recovery early_stress_response oxidative_stress metabolic_proteostatic_adaptation metabolic_reprogramming cumulative_arrest_quiescence cell_cycle_arrest phase_1 metabolic_reprogramming phase_2 senescence_like phase_3 apoptosis_cell_death phase_4 apoptosis_cell_death hypertrophic_senescence senescence_like lysosomal_stress_transition proteostasis_upr atrophic_inflammatory_collapse apoptosis_cell_death Dose-time phase summaries from hierarchical, grounded LLM integration adaptation_recovery apoptosis_cell_death cell_cycle_arrest dna_damage_response inflammation_innate_immune metabolic_reprogramming mitochondrial_dysfunction other_uncertain oxidative_stress proteostasis_upr senescence_like Figure 3: Dose-time integrated phase timeline (hierarchical reasoning output). 123456789 Week dna_damage_response oxidative_stress inflammation_innate_immune cell_cycle_arrest senescence_like apoptosis_cell_death mitochondrial_dysfunction proteostasis_upr metabolic_reprogramming adaptation_recovery other_uncertain process_label 0.190.270.130.130.330.200.200.140.00 0.190.000.130.130.070.000.000.070.00 0.000.130.000.000.200.000.070.070.06 0.000.000.130.000.000.130.070.070.12 0.120.070.200.130.000.000.000.070.06 0.000.000.000.000.000.130.000.000.25 0.000.070.000.070.000.000.000.000.06 0.060.000.070.200.270.200.070.210.06 0.250.270.130.070.070.130.270.210.25 0.120.200.200.200.070.200.270.140.12 0.060.000.000.070.000.000.070.000.00 LLM hypothesis process_label distribution over time 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Proportion of hypotheses (within week) Figure 4: Temporal heatmap of LLMprocess_labelassign- ments. features. Such changes are consistent with stress-associated mor- phological remodeling and hypertrophic phenotypes commonly ob- served during cellular senescence and growth dysregulation [9,24]. Importantly, prior work in image-based profiling has demon- strated that morphological features extracted from Cell Painting assays can reliably capture biologically meaningful variation and map to underlying cellular processes [4,27]. In line with this, we observe substantially higher V2 proxy alignment scores in this high-dose regime, indicating strong quantitative consistency be- tween predicted biological processes and observed morphological changes. In contrast, at lower dose rates (0.003â0.3 mGy/hr), the model produces more heterogeneous hypotheses, including processes such as metabolic reprogramming, proteostatic stress, and adaptive re- sponse programs. In this regime, morphological changes are com- paratively subtle and less dominated by canonical stress signatures, Figure 5: Higher-resolution consistency check between weekĂdose V2 and the LLM-external morphology drift sum- mary. leading to lower proxy alignment scores and reduced interpretabil- ity of processâmorphology relationships. This pattern is consistent with prior radiobiology literature, which suggests that chronic low-dose radiation often induces non- linear and adaptive cellular responses rather than acute DNA damage- driven phenotypes [22, 31]. 5 Discussion This work is motivated by a practical gap in high-content biol- ogy: morphology is information-rich yet difficult to interpret, while Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology LLMs can generate fluent interpretations that require rigorous au- diting before scientific use. We therefore frame LLM outputs as auditable artifactsâstructured hypotheses with explicit evidence pointersâand emphasize evaluation-first validation via citation validity (V1) and proxy-based morphology compatibility (V2). Scope and limits of quantitative auditing. V1 enforces identifier- level grounding: hypotheses may only cite evidence items that are actually present in the prompt payload. This directly targets out-of- payload references, a known failure mode of neural generation in scientific settings [16,19,32]. However, V1 does not guarantee that a cited item semantically supports the claim; it is a necessary but not sufficient condition for scientific faithfulness. V2 complements this by evaluating whether predictedprocess_labels are compati- ble with the measured morphology deltas under a predefined proxy mapping. This is intentionally conservative: it does not require that the label is true in a mechanistic sense, only that it is quantitatively compatible with the observed signature. Because the proxy table is hand-specified, V2 should be interpreted as a transparent and reproducible diagnostic rather than as an objective mechanistic benchmark. Hierarchical reasoning as an auditing strategy. A core de- sign choice is to generate hypotheses at the most atomic unit of measurement (weekĂdose) before aggregation. In longitudinal settings, flat aggregation can obscure phase structure and dilute evidence, and it can yield misleading global summaries when trends differ across subgroups (a form of Simpsonâs paradox). By preserv- ing evidence IDs from week-level hypotheses into dose_time phases and global summaries, the system maintains traceability while pro- ducing compact narratives suitable for downstream review. Limitations and extensibility. First, V1 evaluates citation va- lidity but not semantic support; a stronger external assessment would require claimâevidence judgments or expert annotation, which we do not yet have. Second, V2 relies on a predefined proxy mapping fromprocess_labelto expected morphology feature buckets; this mapping is necessarily simplified and may not capture context-dependent morphologyâprocess relationships. Third, the controlled vocabulary constrains expressivity and may collapse dis- tinct biological mechanisms into coarse categories. Fourth, retrieval coverage and bias (e.g., limitations of the reference library, neighbor annotation quality, and pathway text granularity) can shape the evidence available to the LLM and thus the hypotheses produced. We further emphasize that the use of morphology-only quanti- tative anchors reflects a practical constraint of the current dataset rather than a limitation of the framework itself. In many longitudi- nal biological studiesâincluding low-dose radiation experimentsâ ground-truth mechanism labels and orthogonal modalities are lim- ited or unavailable. Our framework is therefore designed to operate under this constraint, using morphology as a primary quantitative signal while enforcing strict grounding and consistency checks. The auditing structure is inherently extensible: additional modalities such as transcriptomics, proteomics, or targeted biomarkers can be incorporated into the same evidence and validation pipeline to provide stronger mechanistic validation. Finally, the Step 4 drift summary is derived from the same mor- phology distribution and should be interpreted only as an LLM- external sanity check, not as independent biological validation. Implications for low-dose radiation studies. Chronic low- dose and low-dose-rate radiation remains an area where effects can be subtle and context dependent [22,31]. Our results do not estab- lish mechanisms; rather, they demonstrate that an LLM-centered interpretation workflow can be instrumented with quantitative au- diting so that any proposed biological narrative remains tethered to measured morphology and retrievable evidence. Real-world applicability. Our study is conducted on a real- world longitudinal Cell Painting dataset involving chronic low- dose-rate radiation exposure in RPE-1 cells. This setting reflects a particularly challenging regime in radiobiology, where effects are subtle, context-dependent, and lack well-defined ground-truth labels. In such scenarios, traditional supervised or benchmarking ap- proaches are not feasible. Our framework is specifically designed for this setting: it leverages limited experimental data and external knowledge to generate structured, auditable hypotheses that can guide downstream experimental validation. This makes the approach directly applicable to other real-world biological contexts characterized by limited labels and complex temporal dynamics, such as drug response profiling, aging studies, and environmental exposure analysis. 6 Conclusion We presented an LLM-centered, retrieval-augmented interpreta- tion framework for longitudinal Cell Painting morphology data that is evaluation-first by design. By enforcing structured JSON outputs with controlled labels and explicit evidence references, and by validating outputs using grounding integrity (V1) and proxy- based processâmorphology alignment (V2), the framework enables traceable hypothesis generation that can be quantitatively audited against independent morphology signals. More broadly, this work contributes a practical template for integrating LLMs into computational biology workflows where nar- rative interpretation must remain accountable to measured data: hypotheses are generated with explicit provenance, assessed quan- titatively, and summarized hierarchically without sacrificing trace- ability. 7 Acknowledgment This material is based upon works for the LUCID: Low-dose Un- derstanding, Cellular Insights, and Molecular Discoveries program under Award Number DE-AC02-06CH11357 supported by the U.S. Department of Energy, Office of Science, Office of Biological and Environmental Research. References [1] Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610â623. [2]Daniil A Boiko, Robert MacKnight, and Gabe Gomes. 2023. Emergent au- tonomous scientific research capabilities of large language models. arXiv preprint arXiv:2304.05332 (2023). [3]Mark-Anthony Bray, Shantanu Singh, Han Han, Chad T. Davis, Blake Borgeson, Carol Hartland, Maria Kost-Alimova, Sigrun M. Gustafsdottir, Christopher C. Gibson, and Anne E. Carpenter. 2016. Cell Painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes. Nature Protocols 11, 9 (2016), 1757â1774. doi:10.1038/nprot.2016.105 Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo [4]Juan C Caicedo, Sam Cooper, Florian Heigwer, Scott Warchal, Peng Qiu, Csaba Molnar, Aliaksei S Vasilevich, Joseph D Barry, Harmanjit Singh Bansal, Oren Kraus, et al.2017. Data-analysis strategies for image-based cell profiling. Nature methods 14, 9 (2017), 849â863. [5]Srinivas Niranj Chandrasekaran, Beth A. Cimini, Amy Goodale, Lisa Miller, Maria Kost-Alimova, Nasim Jamali, John G. Doench, Briana Fritchman, Adam Skepner, Michelle Melanson, Alexandr A. Kalinin, John Arevalo, Marzieh Haghighi, Juan C. Caicedo, Daniel Kuhn, Desiree Hernandez, James Berstler, Hamdah Shafqat- Abbasi, David E. Root, Susanne E. Swalley, Sakshi Garg, Shantanu Singh, and Anne E. Carpenter. 2024. Three million images and morphological profiles of cells treated with matched chemical and genetic perturbations. Nature Methods 21 (2024), 1114â1121. doi:10.1038/s41592-024-02241-6 Published online 2024-04-09. [6]Europe PMC Consortium. 2015. Europe PMC: a full-text literature database for the life sciences and platform for innovation. Nucleic acids research 43, D1 (2015), D1042âD1048. [7]Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal DaumĂ© I, and Kate Crawford. 2021. Datasheets for Datasets. Commun. ACM 64, 12 (2021), 86â92. doi:10.1145/3458723 Originally released as arXiv:1803.09010. [8]Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. 2020. PubMedBERT: Domain-Specific Language Model Pretraining for Biomedical Text Mining. arXiv:2007.15779 [cs.CL] doi:10.48550/arXiv.2007.15779 TO BE VERIFIED: if citing a published version, replace arXiv with venue details. [9]Alejandra Hernandez-Segura, Jamil Nehme, and Marco Demaria. 2018. Hallmarks of cellular senescence. Trends in cell biology 28, 6 (2018), 436â453. [10]Arturo Hernandez-Segura, Justine Nehme, and Marco Demaria. 2018. New hallmarks of cellular senescence. Trends in Cell Biology 28, 6 (2018), 436â453. [11]Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Hao- tian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2023. A Survey on Hallucination in Large Language Models: Prin- ciples, Taxonomy, Challenges, and Open Questions. arXiv:2311.05232 [cs.CL] doi:10.48550/arXiv.2311.05232 arXiv:2311.05232 is a hallucination survey (not specifically a fabricated-citations-only study).. [12] Jeff Johnson, Matthijs Douze, and HervĂ© JĂ©gou. 2019. Billion-scale similarity search with GPUs. IEEE transactions on big data 7, 3 (2019), 535â547. [13] Guido Kroemer, Lorenzo Galluzzi, Peter Vandenabeele, John Abrams, Emad S Alnemri, Eric H Baehrecke, Mikhail V Blagosklonny, Wafik S El-Deiry, Pierre Gol- stein, Douglas R Green, et al.2009. Classification of cell death: recommendations of the Nomenclature Committee on Cell Death 2009. Cell Death & Differentiation 16, 1 (2009), 3â11. [14]Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chanho So, and Jaewoo Kang. 2020. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics 36, 4 (2020), 1234â1240. doi:10.1093/bioinformatics/btz682 [15] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich KĂŒttler, Mike Lewis, Wen-tau Yih, Tim RocktĂ€schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Informa- tion Processing Systems 33 (NeurIPS 2020). https://proceedings.neurips.c/paper/ 2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html Proceedings ver- sion (NeurIPS 2020). arXiv:2005.11401 is the preprint.. [16]Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers). 3214â3252. [17] Renqian Luo, Litao Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu. 2022. BioGPT: generative pre-trained transformer for biomedical text generation and mining. arXiv:2210.10341 [cs.CL] doi:10.48550/arXiv.2210.10341 [18]Markus Marks, Uriah Israel, Rohit Dilip, Qilin Li, Changhua Yu, Emily Laubscher, Ahamed Iqbal, Elora Pradhan, Ada Ates, Martin Abt, et al.2025. CellSAM: a foundation model for cell segmentation. Nature Methods (2025), 1â9. [19]Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020. On faithfulness and factuality in abstractive summarization. In Proceedings of the 58th annual meeting of the association for computational linguistics. 1906â1919. [20]Marija Milacic, Deidre Beavers, Patrick Conley, Chuqiao Gong, Marc Gille- spie, Johannes Griss, Robin Haw, Bijay Jassal, Lisa Matthews, Bruce May, Robert Petryszak, Eliot Ragueneau, Karen Rothfels, Cristoffer Sevilla, Veron- ica Shamovsky, Ralf Stephan, Krishna Tiwari, Thawfeek Varusai, Joel Weiser, Adam Wright, Guanming Wu, Lincoln Stein, Henning Hermjakob, and Peter DâEustachio. 2024. The Reactome Pathway Knowledgebase 2024. Nucleic Acids Research 52, D1 (2024), D672âD678. doi:10.1093/nar/gkad1025 [21]Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. In Proceedings of the Conference on Fairness, Ac- countability, and Transparency (FAT* â19). 220â229. doi:10.1145/3287560.3287596 [22]National Research Council. 2006. Health Risks from Exposure to Low Levels of Ionizing Radiation: BEIR VII Phase 2. The National Academies Press, Washington, DC. doi:10.17226/11340 [23]Gabriel E Neurohr, Rachel L Terry, Jette Lengefeld, Megan Bonney, Gregory P Brittingham, Fabrice Moretto, Teemu P Miettinen, Laura P Vaites, Lidia M Soares, Joao A Paulo, et al.2019. Excessive cell growth causes cytoplasm dilution and contributes to senescence. Cell 176, 5 (2019), 1083â1097. [24]Gabriel E Neurohr, Rachel L Terry, Jette Lengefeld, Megan Bonney, Gregory P Brittingham, Fabien Moretto, Teemu P Miettinen, Laura Pontano Vaites, Luis M Soares, Joao A Paulo, et al.2019. Excessive cell growth causes cytoplasm dilution and contributes to senescence. cell 176, 5 (2019), 1083â1097. [25] Fabian Pedregosa, GaĂ«l Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, et al.2011. Scikit-learn: Machine learning in Python. the Journal of machine Learning research 12 (2011), 2825â2830. [26] Mohammad Hasan Rohban, Shantanu Singh, Xiaoyun Wu, Julia B Berthet, Mark- Anthony Bray, Matthew D Shair, Lee L Rubin, and Anne E Carpenter. 2017. Systematic morphological profiling of human gene and allele function via Cell Painting. eLife 6 (2017), e24060. [27]Mohammad Hossein Rohban, Shantanu Singh, Xiaoyun Wu, Julia B Berthet, Mark-Anthony Bray, Yashaswi Shrestha, Xaralabos Varelas, Jesse S Boehm, and Anne E Carpenter. 2017. Systematic morphological profiling of human gene and allele function via Cell Painting. Elife 6 (2017), e24060. [28]Maria Janina Sarol, Shufan Ming, Shruthan Radhakrishna, Jodi Schneider, and Halil Kilicoglu. 2024. Assessing citation integrity in biomedical publications: corpus annotation and NLP models. Bioinformatics 40, 7 (2024), btae420. doi:10. 1093/bioinformatics/btae420 Published 2024-06-26. [29]Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, et al.2023. Large language models encode clinical knowledge. Nature 620, 7972 (2023), 172â180. [30]Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al.2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805 (2023). [31]United Nations Scientific Committee on the Effects of Atomic Radiation. 2022. Sources, Effects and Risks of Ionizing Radiation: UNSCEAR 2020/2021 Report to the General Assembly, with Scientific Annexes, Volume I â Scientific Annex C: Biological mechanisms relevant for the inference of cancer risks from low-dose and low-dose-rate radiation. Technical Report. United Nations. Scientific Annex C (low-dose/low-dose-rate). A corrigendum was issued in Sep 2022.. [32] David Wadden, Shanchuan Lin, Kyle Lo, Lucy Lu Wang, Madeleine van Zuylen, Arman Cohan, and Hannaneh Hajishirzi. 2020. Fact or Fiction: Verifying Scientific Claims. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 7534â7550. doi:10.18653/v1/2020.emnlp-main.609 Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology A Supplementary Materials Table A1: Dataset scale and preprocessing summary. ItemValueNotes Biological Setup & Experimental Design Cell lineRPE-1Human retinal pigment ep- ithelial AssayCell PaintingHigh-content morphology Plate format96-well Time courseWeeks 1â99 longitudinal time points Dose rates (mGy/hr) 0.003, 0.03, 0.3, 3.0, 6.0 5 treated dose conditions Paired controlsPlate C (weekly)Paired control plate per week Treated replicates 4 wells / (week, dose) Fixed plate map Total treated wells180 9 weeksĂ5 dosesĂ4 wells Total control wells34Week 1: 2; Weeks 2â9: 4 Total wells analyzed214Treated + control well pro- files Data Dimensions Fields of view (FOV) / well 9í â 1,..., 9 Z-planes / FOV15í â 1,..., 15 Nominal CSV files / well1359Ă 15 (when complete) Total cell-level CSV files28,886(4 files excluded via QC) Total cell-level observa- tions 6,375,699Sum of per-well cell counts Features & Preprocessing Excluded columns orientation, centroid_x, centroid_y Excluded at well-level ag- gregation Base morphology fea- ture set 9 features area,perimeter, eccentricity, solidity, mean_intensity, and 4glcmtexture features Well-level summary fea- tures 369 baseĂmedian, MAD, q10, q90 Within-week normaliza- tion Robust í -score (í„ â median)/(MAD + 10 â6 ) Delta designTreated â Control (weekly paired) Îí=í â median(í control , week) Table A2: System configuration summary (retrieval and LLM settings). ComponentParameterValue Retrieval (JUMP) Reduced embedding 9 morphology features; z-score standardization (í= 10 â8 ) Neighbor search back- end FAISSIndexFlatL2(if available);elsesklearn NearestNeighbors (Euclidean) Candidate neighbors (í ) 100 Final neighbors re- tained 10 Min. gene-like neigh- bors 3 Retrieval (Re- actome) Geneâpathway top í 12 Pathwaydocument search TFâIDF (1â2 grams); cosine sim- ilarity TFâIDF max_features 250,000 TFâIDF min_df2 Document retrieval í6 Reactome releaseDownloaded from/current/ endpoint LLM ReasonerModel gemini-3-pro-preview Temperature0.2 Max output tokens8192 Thinking level HIGH Output formatJSON-only; schema-enforced Controlled vocabulary11 canonical process_labels GroundingStrict groundingEnabled (invalid/missing refer- encesâ reject & retry) Citationtypes (weekĂdose) obs:<id>,jump:<id>, lit:<id> Citation types (hierar- chical) obs: ,jump:,lit:, path:<Reactome_id> HierarchyweekĂdoseâ dose_time Evidence preserved with per- week catalogs dose_timeâ globalCross-dose summarization with evidence-ID propagation Table A3: Delta and baseline diagnostics. ControlTreated Feature (Î z-score)Median |Î| q95 |Î| Median |Î| q95 |Î| area_median0.5202.7382.0806.484 perimeter_median0.3742.5502.3226.008 eccentricity_median0.5451.5171.3794.745 solidity_median1.0412.3831.3994.758 mean_intensity_median0.5331.4731.2293.892 glcm_contrast_median0.7282.5031.2604.500 glcm_correlation_median0.8993.1241.4895.199 glcm_energy_median0.4252.7362.1645.933 glcm_homogeneity_median0.4331.5701.8714.544 Note: Computed on well-level delta profiles. (Î=í§ â median(í§ control , week) ). í control = 34,í treated = 180. Gilchan Park, Guang Zhao, Byung-Jun Yoon, Shinjae Yoo Table A4: Process label usage summary. Process labelweekĂdose dose_time Global Global dose hypotheses phasesclaims summaries dna_damage_response24 (17.6%)1 (5.9%)1 (25.0%)2 (40.0%) oxidative_stress9 (6.6%)4 (23.5%)1 (25.0%)0 (0.0%) inflammation_innate_immune8 (5.9%)3 (17.6%)1 (25.0%)1 (20.0%) cell_cycle_arrest8 (5.9%)6 (35.3%)0 (0.0%)2 (40.0%) senescence_like10 (7.4%)2 (11.8%)1 (25.0%)2 (40.0%) apoptosis_cell_death6 (4.4%)4 (23.5%)1 (25.0%)2 (40.0%) mitochondrial_dysfunction3 (2.2%)3 (17.6%)1 (25.0%)0 (0.0%) proteostasis_upr17 (12.5%)5 (29.4%)1 (25.0%)2 (40.0%) metabolic_reprogramming25 (18.4%)8 (47.1%)1 (25.0%)2 (40.0%) adaptation_recovery23 (16.9%)4 (23.5%)1 (25.0%)1 (20.0%) other_uncertain3 (2.2%)2 (11.8%)0 (0.0%)0 (0.0%) Note: Denominators: weekĂdose hypothesesí= 136; dose_time phasesí= 17; global claimsí= 4; global dose summariesí= 5. Table A5: Grounding reference statistics (evidence-reference coverage and counts). MetricValueNotes Run scale WeekĂdose records45Prompt instances Hypotheses (total)136Across all records Hypotheses per record (mean)3.022 Hypotheses per record (minâmax)2â4Distribution: 2:1, 3:42, 4:2 Evidence references per hypothesis Quant refs per hypothesis (mean)1.000 obs:<id> citations Quant refs per hypothesis (minâmax)1â1 Retrieval refs per hypothesis (mean)2.838 jump:/lit: citations Retrieval refs per hypothesis (minâmax)1â5 Retrieval evidence usage Hypotheses citing JUMP neighbors124 (91.2%)At least one jump: ref Hypotheses citing literature102 (75.0%)At least one lit: ref Hypotheses citing both jump + lit90 (66.2%) Total retrieval refs (jump)219Count across all hypotheses Total retrieval refs (lit)167Count across all hypotheses Grounding caps (per record) Allowed obs IDs1 Allowed JUMP refs10 Allowed literature refs10 Table A6: V2 proxy mapping (process labels to proxy feature sets). Process labelProxy expectations (reduced morphology feature set) dna_damage_response glcm_contrast(+), glcm_correlation(0), mean_intensity(0), area(0) oxidative_stress glcm_contrast(+), mean_intensity(0), area(0) inflammation_innate_immune glcm_contrast(+), area(0) cell_cycle_arrest area(+) ,perimeter(+), eccentricity(0), solidity(0) senescence_like area(+),perimeter(+),solidity(0), eccentricity(0),glcm_contrast(+), glcm_correlation(0), glcm_energy(-), mean_intensity(0) apoptosis_cell_death area(-) ,perimeter(-), mean_intensity(0), glcm_contrast(+), glcm_energy(-) mitochondrial_dysfunction mean_intensity(0), glcm_contrast(+) proteostasis_upr glcm_contrast(+), glcm_energy(-) metabolic_reprogramming mean_intensity(0), area(0) adaptation_recovery area(0), glcm_contrast(0) other_uncertain(none) Note: Direction codes: (+) expected increase; (-) expected decrease; (0) either direc- tion / weak proxy. Feature mappings are derived from canonical morphological hallmarks of cell stress [10,13,26]. Used strictly as a quantitative consistency check (V2), not as a biological ground truth. Used as a consistency check, not biological ground truth. 0.0 0.2 0.4 0.6 0.8 1.0 Hit@10 (V2) Supplementary: Distribution of V2 alignment metrics by dose 0.0030.030.336 Dose rate (mGy/hr) 0.0 0.2 0.4 0.6 0.8 1.0 Weighted proxy (V2) n=26n=28n=27n=27n=28 Figure A1: V2 metric distributions. A.1 Coarse Dose-Level Sanity Check As a complementary coarse summary to the higher-resolution weekĂdose analysis (í=45) presented in the main text (Section 4.5), we also evaluated a dose-level comparison (í=5 doses). As shown Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology Obs refs (obs:...) JUMP neighbors (jump:...) Literature (lit:...) Reactome (path:...) 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 Count per hypothesis mean=1.00 mean=1.61 mean=1.23 mean=0.00 Hypotheses analyzed: 136 Supplementary: Evidence reference usage in grounded hypotheses Figure A2: Evidence reference composition. 123456789 Week 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 Drift mean (Step4, PCA-whitened norm) Supplementary: Independent drift anchor trajectories 0.003 mGy/hr 0.03 mGy/hr 0.3 mGy/hr 3 mGy/hr 6 mGy/hr Figure A3: Step 4 drift trajectories. 0.0030.030.33.06.0 Dose rate (mGy/hr) 1 2 3 4 5 6 L2 distance in reduced morphology space A. JUMP neighbor distance (top-10) 0.0030.030.33.06.0 Dose rate (mGy/hr) 0.0 0.2 0.4 0.6 0.8 1.0 Mean pairwise Jaccard (top pathways) B. Reactome pathway overlap across weeks Retrieval diagnostics (Supplementary) Figure A4: Retrieval diagnostics. dna_damage_response oxidative_stress inflammation_innate_immune cell_cycle_arrest senescence_like apoptosis_cell_death mitochondrial_dysfunction proteostasis_upr metabolic_reprogramming adaptation_recovery other_uncertain 0 5 10 15 20 25 Count A. weekĂdose hypotheses (Step 7) dna_damage_response oxidative_stress inflammation_innate_immune cell_cycle_arrest senescence_like apoptosis_cell_death mitochondrial_dysfunction proteostasis_upr metabolic_reprogramming adaptation_recovery other_uncertain 0 2 4 6 8 Count B. dose_time phases (Step 8) dna_damage_response oxidative_stress inflammation_innate_immune cell_cycle_arrest senescence_like apoptosis_cell_death mitochondrial_dysfunction proteostasis_upr metabolic_reprogramming adaptation_recovery other_uncertain 0.0 0.2 0.4 0.6 0.8 1.0 Count C. global key claims (Step 8) Controlled process_label usage across reasoning scales (Supplementary) low medium high Figure A5: Controlled vocabulary usage across reasoning scales. Table A7: Structured Example of a Global-Level Integrated Claim FieldContent Claim IDclaim_low_dose_adaptation Claim State- ment Chronic low-dose radiation (0.003â0.3 mGy/hr) induces a distinct adaptive phenotype characterized by meta- bolic reprogramming and proteostatic stress, contrast- ing with the hypertrophic senescence seen at higher doses. Process Labelsmetabolic_reprogramming, adaptation_recovery, pro- teostasis_upr Quantitative Evidence obs:obs_w6_d0.003, obs:obs_w9_d0.03, obs:obs_w7_d0.3 Retrieval Ref- erences jump:row_862582, jump:row_928254, lit:q1:1 Confidence Level High in Figure A6, the association remained positive for thedose_time mean Hit@10 summary (Pearsoní=0.661 with 95% bootstrap CI [â0.363, 1.000]; Spearman í= 0.700 with 95% CI [â0.875, 1.000]). Given the extremely small sample size, the confidence intervals are understandably wide. We retain this five-point analysis strictly as a low-resolution sanity check. Consistent with our approach in the main text, we do not interpret this as an independent biological validation, but rather as a basic verification that the modelâs scores co-vary positively with the expected morphological drift at the macro-dose level. Figure A6: Coarse dose-level comparison (í=5) evaluating the consistency between the LLM-external morphology drift and thedose_timemean Hit@10 summary. While a positive trend is visible, the small sample size limits statistical confi- dence.