Paper deep dive
Diagnosing and narrowing the simulation-to-real gap in powder X-ray diffraction with a wet-dry agentic loop
Shaoguang Wang, Weiyu Guo, Ben Fei, Xiaohong Shao, Zhihui Wang, Wanli Ouyang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/25/2026, 6:16:05 AM
Summary
The paper diagnoses the simulation-to-real gap in powder X-ray diffraction (PXRD) analysis, finding it is structural (peak-position drift) rather than additive (noise). It introduces Xtalyst, a multi-agent system that uses real-data fine-tuning, peak-aligned reranking, and conformal recalibration to narrow this gap, demonstrating improved phase identification and refinement on real measured spectra compared to synthetic-trained models.
Entities (10)
Relation Signals (7)
Powder X-ray diffraction → suffersfrom → Simulation-to-Real Gap
confidence 97% · Deep-learning analyzers excel on simulated patterns and degrade on measured ones. This simulation-to-real gap is structural
Peak-aligned reranking → mitigates → Simulation-to-Real Gap
confidence 95% · correcting a small peak-position drift more than doubles median retrieval correlation... peak-aligned reranking... narrow what remains
Xtalyst → uses → Peak-aligned reranking
confidence 95% · Xtalyst integrates these in an agent-orchestrated system... Real-spectrum fine-tuning, peak-aligned reranking, and recalibration narrow what remains
Conformal Calibration → requires → real-spectrum fine-tuning
confidence 93% · A conformal calibration built on synthetic anchors under-covers measured spectra... and recalibrating on real spectra recovers near-nominal coverage
Xtalyst → queries → Materials Project
confidence 92% · The phase-identification module performs a tiered retrieval ladder against the Materials Project (MP) database
Xtalyst → uses → PyWPEM
confidence 90% · It integrates established components for property prediction (CHGNet (45)) and refinement (PyWPEM (46))
Xtalyst → uses → CHGNet
confidence 90% · It integrates established components for property prediction (CHGNet (45))
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Powder X-ray diffraction (PXRD) is the routine probe of crystalline matter, yet its analysis is the rate-limiting step as laboratories automate acquisition. Deep-learning analyzers excel on simulated patterns and degrade on measured ones. This simulation-to-real gap is structural, not additive: synthetic denoising gives no measurable lift on real spectra, whereas correcting a small peak-position drift more than doubles median retrieval correlation. Real-spectrum fine-tuning, peak-aligned reranking, and recalibration narrow what remains and restore the coverage synthetic anchors lose. Xtalyst integrates these in an agent-orchestrated system spanning phase identification, refinement, and calibrated property prediction. On a frozen held-out partition (n=534) each module measured on both splits reproduces its development finding -- including the synthetic-anchor under-coverage, whose magnitude differs between the two pools -- while held-out refinement converges and preserves symmetry without reaching profile-quality fits, and on a diffractometer its wet-dry recommend-rescan-reanalyze loop flips a blinded silicon standard to a gated PASS and changes which minor phase is resolved on a multi-metal alloy.
Tags
Links
- Source: https://arxiv.org/abs/2608.22400v1
- Canonical: https://arxiv.org/abs/2608.22400v1
Trouble viewing inline? Open PDF directly →
Full Text
192,908 characters extracted from source content.
Expand or collapse full text
Diagnosing and narrowing the simulation-to-real gap in powder X-ray diffraction with a wet–dry agentic loop Shaoguang Wang, 1 Weiyu Guo, 1 Ben Fei, 2 Xiaohong Shao, 3 Zhihui Wang, 4,5 Wanli Ouyang 2 1 The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China. 2 The Chinese University of Hong Kong, Hong Kong SAR, China. 3 Suzhou National Laboratory, Suzhou, China. 4 Shenzhen Loop Area Institute, Shenzhen, China. 5 Dalian University of Technology, Dalian, China. Powder X-ray diffraction (PXRD) is the routine probe of crystalline matter, yet its analysis is the rate-limiting step as laboratories automate acquisition. Deep-learning analyzers excel on simulated patterns and degrade on measured ones. This simulation-to-real gap is structural, not additive: synthetic denoising gives no measurable lift on real spectra, whereas correcting a small peak-position drift more than doubles median retrieval correlation. Real-spectrum fine-tuning, peak-aligned reranking, and recalibration narrow what remains and restore the coverage synthetic anchors lose. Xtalyst integrates these in an agent-orchestrated system spanning phase identification, refinement, and calibrated property pre- diction. On a frozen held-out partition (풏=534) each module measured on both splits reproduces its development finding—including the synthetic-anchor 1 arXiv:2608.22400v1 [cond-mat.mtrl-sci] 23 Aug 2026 under-coverage, whose magnitude differs between the two pools—while held-out refinement converges and preserves symmetry without reaching profile-quality fits, and on a diffractometer its wet–dry recommend–rescan–reanalyze loop flips a blinded silicon standard to a gated pass and changes which minor phase is resolved on a multi-metal alloy. Powder X-ray diffraction (PXRD) is the predominant routine probe of crystalline matter, sup- plying phase identity, lattice parameters, microstructure, and, when coupled with full-profile re- finement (1, 2), quantitative phase composition for materials chemistry, metallurgy, mineralogy, pharmaceuticals, and energy storage. Yet the value locked in a powder pattern is released only through an analytical workflow that begins with phase identification, proceeds to Rietveld refine- ment of lattice parameters and atomic positions (as implemented in established packages such as GSAS-I (3), FullProf (4), and TOPAS (5)), and ends in downstream property estimation. This workflow has historically been operator-intensive, with expert time dominating both interpretation and quality control (6). Laboratory automation has accelerated the acquisition of powder patterns far faster than their interpretation, so that analysis, still manual, time-consuming, error-prone, and difficult to scale, is now widely identified as the rate-limiting human step in high-throughput and autonomous-laboratory campaigns (7, 8). The bottleneck is acute on multi-phase samples, on broader-chemistry minerals where chemical-system (chemsys)-restricted libraries do not pin down a unique phase, and on low-symmetry triclinic or monoclinic cells whose refinement loss landscape is shallow and populated by local minima. An automated PXRD analysis that is deployable on real measured spectra, calibrated with quantified uncertainty, and traceable end-to-end would remove this bottleneck, turning a laboratory scan, optionally accompanied by a composition or crystal-system prior, into a vetted structure-and- property answer. The obstacle to that goal is not any single algorithm but the reliability of the whole analysis chain on measured, rather than simulated, data. Deep learning has rapidly entered every stage of this workflow. End-to-end convolutional and attention models now address phase identification on single and multi-phase patterns (7– 13); crystal-system and space-group classifiers operate on stick patterns or full profiles (14–17); and a growing body of generative models attempts direct PXRD-to-structure mappings (18–25). Substantial progress on simulated test splits is well documented: correlations against synthetic 2 targets routinely exceed 0.9, and per-class accuracy on synthetic crystal-system tests exceeds 90% for the most common systems. Performance on measured spectra decays sharply. On a representative subset of the recently released opXRD corpus of measured minerals (26), the median peak-aligned Pearson correlation (a pattern-to-pattern match scored after correcting a small rigid 2휃 peak- position offset) between top-1 retrieval candidates and the observed pattern is 0.21 before any of the interventions reported here, and a synthetic-trained crystal-system classifier (a one-dimensional convolutional neural network trained on Materials Project stick patterns, our own pre-fine-tuning checkpoint) scores 27.5% on real RRUFF spectra (27) against 44.1% on clean simulated patterns rendered from held-out structures through the same code path (푛=399), its own simulated ceiling; the > 90% values quoted above are reported by larger models on their own simulated splits and are not reproduced here (Fig. 1). This simulation-to-real gap, not any individual model’s accuracy, is the dominant source of unreliability in a deployed PXRD deep-learning pipeline, yet it is rarely quantified end-to-end on real measured spectra. This work makes three contributions, each established on real measured spectra: two concern the nature of the simulation-to-real gap, and the third is the end-to-end system that operationalizes them. First, we show that the gap is structural rather than additive. A one-dimensional denoiser that performs as designed on its synthetic task yields no measurable lift on real opXRD spectra (statistically equivalent to zero within a±0.05 correlation margin), whereas correcting a small rigid peak-position drift is the single largest correctable component and more than doubles the median retrieval correlation (the full intervention sequence is summarized in fig. S1). This finding redirects effort away from noise removal and toward alignment and real-data adaptation. Second, we find that the uncertainty layer inherits the same gap. A conformal calibration built on synthetic anchors under-covers measured spectra, on held-out as well as development data, and recalibrating on real spectra recovers near-nominal coverage on a disjoint development calibration. This calibration failure on real spectra, to our knowledge not previously quantified end-to-end in a deployed PXRD pipeline, is the most transferable of our findings: it applies to any spectroscopic pipeline that ships synthetic-calibrated conformal intervals. Third, turning these findings into a deployable capability, we build Xtalyst, an end-to-end multi-agent PXRD pipeline that—to our knowledge, the first to do so—closes the loop from a raw measured spectrum through phase identification, full- profile refinement, quantitative phase analysis, and calibrated property prediction under cross-stage 3 reliability gates, with every design choice reported against the quantified negative result it displaced (the system and its newly introduced components are detailed below). The mismatch between simulated training distributions and measured deployment distributions is not unique to PXRD. Spectroscopy domains adjacent to ours have named the phenomenon ex- plicitly and built it into the design loop: radiation spectroscopy reports a synthetic-trained source classifier improving from 75% to 96% after a 64-spectrum real-data fine-tune (28); foundation models for stellar spectra fine-tune on real observational data to bridge a synthetic gap (29); magnetic-resonance spectroscopy narrows a simulation-to-real gap through physics-informed aug- mentation of the simulated training distribution (30); reconstructive spectroscopy reports the same trend (31); and seismic imaging deploys explicit synthetic-to-real domain adaptation (32). The remedy that recurs across these domains is real-data adaptation, whether by fine-tuning or by aug- mentation, of synthetic-pre-trained models. PXRD has lacked an equivalent systematic study that quantifies the gap on real measured spectra, identifies its dominant cause, and demonstrates what intervention most narrows it. Three further gaps limit the PXRD deep-learning evidence base. First, most reported systems address a single task, whether crystal-system classification (15–17), phase identification (9, 12, 13), or generative structure proposal (18, 19, 22, 24). End-to-end pipelines that combine identification, refinement, quantitative analysis, and property estimation in a single tested system are rare. The closest concurrent work, Dara (33), automates multiple-hypothesis phase identification by adjudi- cating candidate phase combinations against real refinement residuals rather than against a learned classifier, and demonstrates it at scale in an autonomous synthesis laboratory; it stops deliberately short of detailed structural refinement. It therefore supports, independently, the principle our profile gate applies—that a refinement residual rather than a classifier score should arbitrate whether a phase assignment is trustworthy—while sitting upstream of the refinement and property layers evaluated here. Second, calibrated reliability with formal coverage guarantees on real measured spectra is correspondingly rare. The dominant uncertainty mechanism reported in PXRD deep- learning work is Bayesian or Monte-Carlo-dropout ensembling (17); distribution-free conformal prediction, which wraps any predictor with a finite-sample coverage guarantee (34–37), is widely used in adjacent scientific domains but, to our knowledge, has not been calibrated on real-mineral PXRD with the explicit question of whether the calibration itself transfers across the synthetic-real 4 domain shift (38–40). Third, multi-agent and large-language-model (LLM)-orchestrated systems for chemistry and materials (41–43) have had limited uptake in PXRD, and we are aware of no published PXRD-specific multi-agent system that closes the loop from raw spectrum to evaluated structure and property. A concurrent agentic platform for electron microscopy, EMSeek (44), adopts a similar high-level design (autonomous microservices, provenance tracking, and a gated mixture- of-experts (MoE) predictor) for a different modality. Our primary contribution is the diagnosis of the simulation-to-real gap and the finding that the calibrated-uncertainty layer inherits it; our sys- tem contribution is distinct from EMSeek in both modality and design, comprising PXRD-specific chemistry-aware tiered retrieval, cross-stage reliability gates, and a recommend–rescan–reanalyze loop closed on a real diffractometer. Integrating these stages into a single multi-agent system, rather than running the individual tools independently, is a deliberate design choice with concrete value on real measured data. An integrated chain enforces end-to-end reliability that no isolated tool can: three separately reported cross-stage reliability gates, namely a symmetry gate (space-group preservation), a profile-quality gate (푅 푤푝 /푅 푝 ), and an energy-plausibility gate (formation energy within a conformal interval), are checked across stage boundaries, and a bounded replan loop reacts to a failed gate before the result is reported. We report these gates separately because they can disagree; a refinement can converge and preserve symmetry yet still fit the profile poorly. The chain also makes the analysis traceable: an append-only provenance trace records every tool call, validator decision, and rationale for a given spectrum, so a downstream conclusion can be audited back to its inputs. And it reduces operator burden by converting raw outputs into vetted, gated decisions and operator- facing recommendations rather than leaving cross-tool reconciliation to manual expert effort. These are reliability, provenance, and human-effort arguments. We do not optimize for or benchmark throughput; as an order-of-magnitude characterization only, end-to-end analysis of a single-phase pattern runs in minutes on a single CPU host, dominated almost entirely by the full-profile refinement (phase identification and formation-energy inference are sub-second to a few seconds each), so we make no state-of-the-art speed claim. We address these gaps with a system that operationalizes the two findings above. Xtalyst is a multi-agent end-to-end PXRD pipeline covering tiered phase identification, whole-pattern lattice- parameter refinement, quantitative per-phase abundance analysis, formation-energy prediction with 5 conformal uncertainty, and post-chain agents for polymorph disambiguation and experiment rec- ommendation. It integrates established components for property prediction (CHGNet (45)) and refinement (PyWPEM (46)), and augments the open intensity-only phase identifier CPICANN (12) with a peak-aligned reranker that lifts its development median Pearson by +74%; XtalCS, our real-RRUFF fine-tuned crystal-system classifier, lifts accuracy from 26.5% (synthetic pre-training) to 49.8% on the same development split (푛=313) and 49.4% on a frozen held-out split (푛=251). Throughout, where we build on existing components we cite and credit them; where we extend them, through peak-aligned reranking, real-spectrum fine-tuning, and real-spectrum conformal re- calibration, we report the extension and its measurable effect on the same frozen held-out data used for the rest of the evaluation. Each negative result we report is presented as the evidence for the production choice it determined, with sample sizes stated throughout; we deliberately do not pursue single-algorithm state of the art. Results The Xtalyst pipeline The results below follow a difficulty-escalating arc: we first characterize the simulation-to-real gap and its dominant cause, then show how the wet–dry loop narrows that gap on progressively harder samples—a single-phase standard, then multi-metal alloys where a single scan cannot resolve the coexisting phases that the loop recovers. Each step reports a concrete finding, and every production choice is stated against the quantified negative result it displaced; the component-level analyses that support these findings (uncertainty calibration, held-out benchmarking, and the discriminative- versus-generative comparison) are summarized here and detailed in the Supplementary Materials. Xtalyst is organized around three analysis modules—phase identification, full-profile refine- ment, and structure validation—together with a multi-agent control layer (Fig. 2 gives the full architecture, illustrated end-to-end on a Si standard measured on a Rigaku SmartLab). Throughout the paper we use complementary measured worked examples: this Si closed-loop standard (a clean, gated end-to-end pass on real instrument data) to demonstrate the full recommend–rescan– reanalyze loop; Sb 2 O 3 (a challenging gap case, high 푅 푤푝 , recorded mislabeling) to illustrate the 6 simulation-to-real problem and the advisory agents; and PbSO 4 anglesite as a further clean single- phase refinement example. The phase-identification module performs a tiered retrieval ladder against the Materials Project (MP) database (47) with peak-aligned Pearson correlation against the observed pattern as the candidate-quality score; the refinement module runs whole-pattern refinement of the retrieved candidate’s lattice parameters against the observed profile; and the validation module evaluates the refined structure with a machine-learning interatomic potential (MLIP; CHGNet (45)) and reports a conformal half-width around the predicted formation energy. A multi-agent layer plans configuration, performs post-chain polymorph disambiguation, exposes a (size/strain) microstructure agent, and emits operator-facing experiment recommendations. Of the constituent modules, the crystal-system classifier (XtalCS) and the conformal calibration are trained or fine-tuned on real-data anchors within this work; the peak-aligned reranker, the tiered fallback wiring, the automated peak-list (peak0) synthesis that lets the refinement engine run automatically on an arbitrary retrieved candidate without a hand-authored peak list, the non-negative least-squares (NNLS) per-phase abundance recovery and its multi-phase slot-to-Materials-Project resolution, and the multi-agent control are introduced by this work; and CHGNet, PyWPEM, MEGNet (48), MACE-MP-0 (49), CPICANN (12), and XDecomposer (50) (multi-phase decomposition) are external components integrated and credited as such. The evaluation rests on three open mineralogical corpora and a small collaborator subset. We ingest 2,696 spectra from four sources (1,277 from opXRD (26), 1,365 from RRUFF (27), 32 from the Simonnet 2024 quantitative phase-analysis round-robin (51), and 22 from a collaborator- supplied internal subset); one additional X-ray reflectivity pattern was identified and excluded as out of scope. We freeze a 534-spectrum held-out partition (255 opXRD + 273 RRUFF + 6 Simonnet, a frozen held-out manifest) read exactly three times across the entire study, each read fixed in advance: the headline held-out evaluation, a single conformal recalibration check, and one small quantitative-phase-analysis read-out on the held-out Simonnet mixtures. We never report any number as evaluated on 2,696; 2,696 is the ingest total. Per-module evaluation uses its own labeled subset (푛 stated alongside every number), and the end-to-end refinement evaluation is conducted at 푛 = 146 on the held-out partition—the subset of the 254 scored held-out opXRD spectra for which phase identification returned a candidate above the correlation floor, so convergence rates on it are conditional on that gate and are reported unconditionally as well. The 254 are the 255 7 held-out opXRD spectra less one whose measured range (6.7–51.1 ◦ ) does not cover the 10–80 ◦ window the benchmark resamples onto, excluded for comparability rather than for quality. Table S1 summarizes the scope and the development/held-out split. Per-claim sample sizes vary because each claim is measured on the split appropriate to it: the frozen held-out partition for the headline pipeline metrics, development pools for the diagnostic ablations, an external Materials Project benchmark for the interatomic-potential comparison, the Simonnet real-weighed anchor for quantitative phase analysis, and curated fixtures for the control- layer ablations. Table S2 in the Supplementary Materials maps every principal claim to its sample size and split, so that the varying 푛 are read as a deliberate evaluation design rather than an opportunistically pooled count. The simulation-to-real gap is structural, not additive Key result: the gap that breaks intensity-only methods on real spectra is dominated by a correctable peak-position drift, not by additive noise—so the right intervention is alignment, not denoising. The gap between simulated and measured spectra is large enough that it must dominate any deployment claim. On the opXRD development subset (푛=306), the median peak-aligned Pearson correlation between top-1 chemsys-exact retrieval candidates and the observed pattern is 0.146 without alignment (distinct from the 0.21 headline value of Fig. 1, which is the peak-aligned top-1 correlation on a smaller 푛=75 diagnostic subset). For scale, profile correlations near 0.9 are the simulated-pattern reference level: our own synthetic-degradation validation, described below, recovers 0.88–0.90 against simulated ground truth. We did not run the retrieval module on a synthetic query split, so 0.9 marks that reference level here, not a measurement of this module. We test two hypotheses for the cause of the gap (Fig. 1B and D). A natural first hypothe- sis is that the gap reflects additive degradations (background structure, instrumental broadening, counting noise) that a denoising preprocessor could absorb. A one-dimensional U-Net trained on a synthetic-degradation task and validated against synthetic ground truth moves the median top-1 profile correlation by 0.318 → 0.315 on real opXRD spectra (median paired Δ = −0.004, 푛=25). This change is not merely small: the paired change is statistically indistinguishable from zero (Wilcoxon signed-rank 푝 = 0.49) and is statistically equivalent to zero within a pre-specified±0.05 8 correlation margin (two one-sided tests, 푝 < 0.001), with a bootstrap 95% confidence interval (CI) on the median paired change of[−0.018,+0.015] that brackets zero; the denoiser in fact low- ers the correlation on 16 of 25 spectra. Two classical baselines (iterative-polynomial background subtraction and rolling-min subtraction) move it within ±0.002 of the raw value. The null is a conservative test rather than an artifact of an already-degraded starting point: the 푛=25 cohort is the cleaner subset with resolvable top-1 candidates, and the same network lifts its synthetic validation correlation 0.55–0.60→ 0.88–0.90 (+0.30), so the null reflects a genuine synthetic-to-real domain mismatch, not a failure of model capacity. Crucially, the alignment intervention below works on this same cohort under this same direct-Pearson metric: applying the±0.5 ◦ peak-alignment slide search to the raw (un-denoised) spectra of these 25 samples lifts the median correlation 0.318 → 0.504 (per-sample median Δ =+0.080, improving 22 of 25, Wilcoxon 푝 = 8× 10 −7 ), whereas denoising the same 25 moves nothing (9 of 25, 푝 = 0.49). The additive and structural hypotheses are thus adjudicated head-to-head on one matched cohort, removing any confound from the differing sample or metric used for the larger 푛=306 alignment sweep below. A second hypothesis is that the gap is structural, dominated by peak-position shifts that arise when an MP candidate lattice, relaxed by density functional theory (DFT), does not match the measured lattice exactly. A peak-aligned Pearson score with a±1.0 ◦ shift search moves that median 0.146 → 0.385 (푛=306, +0.24). Of the 306 development samples, 276 (90%) improve by more than 0.01 and 230 (75%) by more than 0.05; on the chemsys-exact-hit (hybrid) stratum the median lifts 0.205→ 0.482 (median per-sample+0.131; median absolute shift applied 0.14 ◦ ; Fig. 3D), and on the intensity-only fallback stratum it lifts 0.001→ 0.259 (per-sample+0.206; median absolute shift 0.58 ◦ ). The fallback stratum, whose candidates tend to carry larger systematic 2휃 offsets, benefits the most. Taken together, the denoiser zero-lift and the peak-align effectiveness establish that the simulation-to-real gap is structural rather than additive: a noise-removal preprocessor cannot close it on real data, even when it performs as expected on the synthetic task it was trained for. Within the structural causes, a small rigid peak-position drift between the simulated candidate lattice and the measured pattern is the single largest correctable component: peak-alignment recovers+0.24 of the median, roughly a third of the gap to the synthetic regime. A substantial structural residual remains after alignment (aligned median 0.385, still well below the∼0.9 synthetic level); by elimination this 9 residual is dominated by intensity redistribution (texture, preferred orientation, site occupancies) and anisotropic, non-rigid lattice mismatch that a single rigid shift cannot absorb. We test the non-rigid part directly. Replacing the single rigid offset with a two-parameter angle-dependent cor- rection Δ2휃(휃) = 푎+ 푏 tan휃 (a zero-shift term 푎 and a specimen-displacement/strain term 푏, fitted per spectrum) raises the aligned median from 0.354 to 0.377 and improves 303 of 304 real spectra individually (median per-sample +0.012, Wilcoxon 푝 ≈ 2× 10 −51 ; the tan휃 term is engaged in 95% of spectra, median|푏| = 0.41 ◦ ). Here the 0.354 baseline is the rigid-offset median recomputed on the same 푛=304 angle-fit samples under the identical per-spectrum 푎+ 푏 tan휃 procedure with 푏 constrained to zero, so it is directly comparable to the two-parameter 0.377; it sits slightly below the 0.385 reported above because that value comes from a wider±1.0 ◦ grid search rather than this per-spectrum least-squares offset fit, not from the negligible 푛=306-versus-304 difference. That an angle-dependent model beats the rigid one on essentially every spectrum confirms the residual drift is systematically non-rigid rather than a single global offset, though its modest median size shows peak alignment—rigid or angle-dependent—is not the whole gap. We bridge the remaining residual downstream with real-spectrum fine-tuning and physics-based refinement rather than claim it away. Real-spectrum fine-tuning narrows the gap on single-phase samples Key result: two lightweight, real-data interventions—fine-tuning a synthetic-pre-trained classifier on measured spectra, and peak-aligned reranking of intensity-only candidates—recover much of the lost accuracy on single-phase minerals without retraining any external model. Crystal-system classification. The synthetic-pre-trained crystal-system classifier (a 1-D convo- lutional neural network, CNN, trained on stick patterns simulated from 2,800 Materials Project structures, 400 per system, of which 2,401 are used for training and 399 held out) reaches 27.5% on RRUFF measured spectra (푛=400 development sample; Wilson 95% CI [23.4, 32.1]; Fig. 3C), and 26.5% on the smaller RRUFF development test set (푛=313) on which the remaining rungs of this ladder are measured. Class-balanced retraining does not move that aggregate (25.9%, 푛=313); what it moves is the worst-class floor, from 0 of 24 Tetragonal spectra correct to a min-per-class accuracy of 9.8%, which is why the balanced checkpoint is used as the base for fine-tuning and not reported as an accuracy gain. The production checkpoint, XtalCS, fine-tunes the balanced model 10 on a real-RRUFF development pool (730 spectra: FT-train 584, FT-val 146, disjoint from that test set) with min-per-class checkpoint selection, lifting accuracy to 49.8% on the same test set (푛=313, Wilson[44.3, 55.3]) and 49.4% on the frozen held-out partition (푛=251, Wilson[43.3, 55.5]; one RRUFF sample, rruffR060379, was inadvertently in both the fine-tuning and held-out sets and is excluded from this held-out evaluation, see Materials and Methods). The held-out figure is sta- tistically indistinguishable from the development figure, indicating clean generalization. Per-class held-out accuracy ranges from 81% on Cubic and 60% on Triclinic to 39% on Monoclinic and 37% on Hexagonal (fig. S2C), and the low-symmetry pair dominates the off-diagonal: 16 of the 50 Monoclinic errors are read as Triclinic (as many as are read as Orthorhombic), and true Monoclinic supplies 16 of the 21 spurious Triclinic predictions; peak-level attribution (fig. S3) shows the clas- sifier keying on genuine reflections in a correct case and scattering across the crowded low-angle region in this confusion. We report these residual weaknesses as quantified, structural limits of the 7-way task at the current real-data FT pool size rather than as upper-bound claims (the Discussion); a follow-up class-weighting study confirms that loss-shaping can lift Monoclinic and Hexagonal individually but only by sacrificing accuracy elsewhere and the overall ceiling stays at∼ 49.8%. Peak-aligned reranking of intensity-only candidates. CPICANN (12) returns an intensity-only top-퐾 candidate set with no pattern-correlation score. We rerank its top-6 by peak-aligned Pearson against the observed spectrum. On the development opXRD subset where phase identification falls through to CPICANN (푛=26), the median candidate Pearson lifts 0.216 → 0.377 (paired comparison on 푛=17 both-scored: median Δ=+ 0.103, mean Δ=+ 0.128, improving on 13 of 17 with the remaining 4 tied and none worse; Wilcoxon signed-rank and exact sign test on the 13 nonzero pairs both give 푝 = 0.0002); the MP-resolvability of the selected candidate (whether it maps to a retrievable Materials Project structure) lifts from 65% to 100%. We also tested a hard crystal-system filter (using XtalCS predictions to restrict the top-퐾 ); at ∼ 50% classifier accuracy the filter removes correct candidates more often than incorrect ones and degrades median correlation to 0.141. The peak-aligned reranker is adopted; the hard crystal-system filter is reported and rejected. 11 Tiered retrieval identifies phases from intensity alone Key result: a three-tier retrieval ladder identifies a phase from the measured pattern alone, reach- ing an MP-resolvable candidate for every development input while gracefully degrading from composition-aided to intensity-only when no prior is supplied. The only required input is the measured intensity–2휃 scan (with wavelength and angular range); a composition or crystal-system prior is an optional accelerator, not a requirement. Phase identification routes every input through a three-tier fallback chain (fig. S4): a chemsys-exact hybrid retrieval against Materials Project, a Jaccard-ranked nearest-chemistry retrieval, and the intensity-only CPICANN candidate set aug- mented with peak-aligned reranking. When a chemical system is supplied it activates the first two tiers; with spectrum alone the chain falls through to the intensity-only tier, which is why identifi- cation quality is composition-dependent. Peak-aligned scoring is applied uniformly across all tiers. We use coverage here in the retrieval sense, the fraction of inputs for which the chain returns at least one MP-resolvable candidate structure (distinct from the conformal coverage of the uncertainty layer below). On the development opXRD pool (푛=510), the chain attains 100% candidate coverage (every input reaches at least one MP-resolvable candidate) with median peak-aligned Pearson 0.594 on chemsys-exact hits and 0.167 on the intensity-only fallback; tier-2 nearest-chemistry retrieval, when it returns a candidate, has a median of 0.460 on the development no-MP-match subset (푛=64 of 127 no-MP-match samples). The held-out result reproduces this ladder cleanly: median peak- aligned Pearson is 0.609 (bootstrap 95% CI[0.555, 0.654]) on the 푛=212 samples where the chain returns at least one MP candidate (of 254 held-out samples reaching phase identification), and 0.643 ([0.593, 0.684]) on the chemsys-exact-hit substratum (푛=176) versus 0.391 ([0.279, 0.485]) on the nearest-chemistry substratum (푛=36). Held-out and development medians agree within sampling uncertainty: the held-out chemsys-exact median 0.643 and the development chemsys-exact me- dian 0.594 (bootstrap CI[0.566, 0.638]) are not significantly different (Mann–Whitney 푝 = 0.32). Resolving this held-out median by chemistry (fig. S5) shows retrieval is strongest for refractory transition-metal oxides and weakest for light, disorder-prone chemistries, with the asymmetry that the strong end rests on one to three held-out materials per element while the weak end is well sampled (Mg 0.409, 푛=34; Ag 0.370, 푛=7). 12 The residual is explicit and located. On the held-out partition, 42 of 254 samples (16.5%) are broader-chemistry minerals that exit the chain without any MP candidate at any tier, neither chemsys-exact nor Jaccard-near nor intensity-only resolution to an MP structure. This subset is the dominant cohort outside Xtalyst’s current addressable scope; we report it as a quantified limit, not as an artifact of the chain wiring. The absolute median of 0.6 on the chemsys-exact stratum also retains the residual simulation-to-real ceiling: peak-alignment more than doubles the correlation in relative terms but does not bring it into the synthetic near-saturation regime. Real-spectrum recalibration restores coverage the synthetic calibration loses Key result: the simulation-to-real gap reaches the uncertainty layer too—a synthetic-anchor con- formal calibration under-covers real spectra—and recalibrating on real measurements restores valid coverage. We treat the calibration layer with the same simulation-to-real lens applied to the predictors. Split conformal calibration of the formation-energy predictor (CHGNet residuals) on synthetic Materials Project anchors gives a half-width of 89.8 meV/atom at nominal coverage 1− 훼 = 0.90 (Materials and Methods). When this synthetic-anchor calibration is applied to a real-opXRD evaluation pool (푛=97 refined structures), empirical coverage at 훼=0.10 collapses to 0.546 (Wilson 95% CI [0.447, 0.642]), far below the nominal 0.90, which the interval excludes. That pool is every refined structure in the development set, including the refinement failures and symmetry-gate mismatches; on the narrower verdict-pass subset from which the real calibration set below is drawn, the same synthetic half-width covers 0.781 (50/64, Wilson[0.666, 0.865], nominal 0.90 still excluded). We give both, because part of the distance between 0.546 and the recalibrated coverage is that verdict-pass filter rather than the recalibration itself. This is a calibration failure, not a predictor failure: the upstream point estimator is unchanged. We recalibrate the conformal layer on real measured spectra, using a development-only cali- bration set that is disjoint from the frozen held-out partition: the 49 development opXRD refined structures with DFT ground truth. (An earlier 64-structure real-spectrum set inadvertently included 15 samples that were later frozen into the held-out partition; we exclude those and report only the leak-free development (푛=49) set here, so the calibration set and any held-out evaluation are strictly disjoint.) A leave-one-out calibration on these 49 structures, in which each test residual 13 is scored against a conformal quantile computed from the other 48 so that no point calibrates on itself, restores near-nominal coverage on the development domain: 0.959 at 훼=0.05, 0.918 at 훼=0.10 (Clopper–Pearson [0.804, 0.977], nominal 0.90 included), 0.816 at 훼=0.20, and 0.714 at 훼=0.30; every measured nominal level lies inside the exact binomial interval of the corresponding measured coverage (the 훼=0.10 level is plotted in fig. S4E). The gap also reproduces on held-out: applying the synthetic-anchor calibration (the Materials Project DFT anchors, which are disjoint from the measured spectra) to the held-out refined-structure CHGNet residuals undercovers (81.1% empirical at 훼=0.10, Wilson[73.4, 87.0], 푛=127; nominal 0.90 excluded), with no upstream re-run (fig. S4F and G). Transferring the leak-free real-spectrum calibration (the same 49 development structures, verified disjoint from the held-out partition) to the 127 held-out CHGNet residuals then recovers valid coverage: the empirical held-out coverage is at or above every nominal level—95.3% at 훼=0.10 (Wilson [90.1, 97.8], 푛=127), 90.6% at 훼=0.20, and 96.9% at 훼=0.05—so the recali- brated interval is conservative on held-out (it never under-covers), in contrast to the synthetic-anchor calibration that excludes nominal. We state this as conservative rather than near-nominal: because the development residuals are larger than the held-out residuals, a development-calibrated interval slightly over-covers on held-out, which is the safe direction for a reliability gate but not a tight one. To our knowledge this is the first end-to-end measurement that a synthetic-anchor calibration under- covers real measured spectra in a deployed PXRD pipeline, and that real-spectrum recalibration restores valid (at-or-above-nominal) coverage on a disjoint held-out partition. We do not over-claim the recovery. It rests on a 49-structure real calibration set; while every measured nominal level is contained in its exact binomial interval, the point estimates carry the wide intervals of that sample size, and at 훼=0.01 the set is too small to estimate the corresponding half- width reliably. The headline statement is therefore conditional on the calibration source matching the deployment domain, a statement that matches the rest of the simulation-to-real diagnosis. The full chain holds up on a frozen held-out benchmark The end-to-end held-out test ran on 푛=146 refinement-eligible held-out samples (held-out opXRD that pass the chemsys-exact-hit gate of phase identification with peak-aligned Pearson> 0.4; fig. S6). Of these, 127 (87.0%, Wilson 95% CI[80.6, 91.5]) converged through the full chain (retrieval, full- 14 profile refinement, and CHGNet evaluation with a conformal half-width). That 87.0% is conditional on the eligibility gate, and we give the unconditional rate beside it so the two are not read as the same quantity: over all 254 held-out opXRD spectra, including the 108 for which phase identification never cleared the correlation floor and which therefore never reached refinement, the chain returns a refined structure for 127/254 = 50.0% (Wilson[43.9, 56.1]). The conditional rate answers whether refinement converges once a phase has been identified; the unconditional rate answers how often the whole chain succeeds on an arbitrary measured pattern, and it is limited by phase identification rather than by refinement. On the 127 converged samples, the refined space group matches the input space group in 127/127 = 100% of cases (Wilson 95% lower bound 0.971): when PyWPEM converges it does not break the space-group symmetry assumed at retrieval. The formation-energy gate reports 113 pass / 6 marginal / 8 fail on the same 127 samples. The remaining 19/146 = 13.0% are chain failed: PyWPEM aborts inside its M-step with a numerical breakdown on extreme- angle low-symmetry cells (mostly triclinic), a failure mode we quantify and locate rather than discard. Table S3 lists one converged held-out refinement per crystal system, spanning 푅 푤푝 from 13.2 to 96.8, each with its input space group preserved. Here 푅 푤푝 and 푅 푝 are the weighted and unweighted profile residuals (the percentage misfit between the calculated and observed diffraction profiles, lower being better); throughout, a normalized 푅 푤푝 above∼20 is treated as an unreliable fit by the profile gate (Materials and Methods). Those seven entries are the best fit in each crystal system rather than the typical one, so we give the distribution they are drawn from. Over all 127 converged held-out refinements the normalized 푅 푤푝 has median 116.7 (interquartile range 74.0 to 138.4, full range 13.2 to 215.5) and the scale-invariant 푅 푝 has median 92.4; only 2 of 127 clear the 푅 푤푝 ≤ 20 profile gate and 20 of 127 sit at or below 50. The held-out benchmark therefore establishes that refinement converges on 87.0% of eligible spectra, that convergence never breaks the retrieved space group, and that the predicted formation energy survives the re-celling—and it does not establish profile-quality fits on this cohort. That follows from what the module refines: the six lattice parameters move and the atomic coordinates, occupancies and displacement parameters stay at the retrieved candidate’s values (Materials and Methods), so half the cohort barely moves at all—the largest relative change in any of 푎, 푏, 푐 is under 0.1% for 67 of the 127 samples, median 0.091%—and a real mineral pattern carrying accessory phases, texture and a composition offset from the database candidate 15 cannot be driven to a low residual by re-celling alone. The single-phase laboratory samples below, where the phase is known and well crystallized, are the regime in which this chain reaches gated passes (푅 푤푝 7.1 to 19.8); the held-out mineral cohort is not that regime, and we report it as located rather than solved. The MLIP choice rests on a five-model comparison spanning the strongest universal interatomic potentials currently available (45, 48, 49, 52, 53). On a single 307-structure benchmark drawn from Materials Project—the same set for every model, so the comparison is not confounded by sample- size differences—the mean absolute errors (MAEs) against DFT formation energies are: CHGNet 43.5 meV/atom, MatterSim v1.0.0-1M 189.6 (53), SevenNet-0 215.4 (52), MACE-MP-0 218.7, MEGNet 348.0, and a simple-mean three-model ensemble 181.2; configuration-specific MAE values are in Materials and Methods. On the full benchmark the bootstrap 95% MAE intervals are well separated—CHGNet [39, 48] against MatterSim [172, 208], SevenNet [196, 236], and MACE-MP-0 [199, 239], so CHGNet’s lead is decisive rather than a small-sample artifact (paired Wilcoxon signed-rank, CHGNet versus MatterSim and versus SevenNet, 푝 < 10 −30 each). On this Materials-Project-referenced formation-energy plausibility task CHGNet is the best-performing predictor among the models tested, leading by 4.4–8×; because the benchmark energies are MP- DFT values and CHGNet is trained on MP-adjacent data, we read this as the best matched predictor for our MP-referenced gate rather than as a domain-independent ranking of interatomic potentials. Its top ranking among universal potentials on the independent Matbench Discovery benchmark (54) is consistent external evidence. A learned mixture-of-experts blend overCHGNet, MEGNet, MACE-MP-0 reaches a best gated MAE of 59.7 meV/atom on the 푛=307 benchmark, 37% worse than single CHGNet; the per-structure absolute errors confirm the gap is not a mean artifact (paired Wilcoxon signed-rank, CHGNet versus the gated blend, 푝 = 1.6× 10 −7 ). The gate’s composite difficulty signal does correlate with the gated error (푟=0.47 at the best regularizer, 0.87 at heavier regularization), which is why we retain the gate as an opt-in uncertainty signal but not as a point predictor (Materials and Methods). The pipeline reproduces development performance on the frozen held-out set where it is strong— crystal-system top-1 49.4% (푛=251, versus 49.8% development), phase-identification median peak- aligned Pearson 0.609 (푛=212, versus 0.594), space-group preservation 127/127 on converged refinements, and real-spectrum conformal coverage 95.3% at 훼=0.10 on held-out (푛=127)—and 16 exposes the residual gaps where it is not: conformal calibration under the synthetic-anchor recipe under-covers (81.1%), the broader-chemistry no-candidate stratum (16.5%), and non-convergence on extreme-angle low-symmetry minerals. Table S4 gives the full per-module development-versus- held-out breakdown with every 푛 and 95% confidence interval. Interval convention: proportions are reported with Wilson score intervals throughout, except the leave-one-out conformal coverage on the 49 real-spectrum calibration structures, which is reported with the Clopper–Pearson exact binomial interval emitted by the conformal module itself. Intervals on medians and on mean absolute errors are percentile bootstraps with 200,000 resamples and a fixed per-quantity seed, recomputed by docs/research/paperbootstrapcis.py into a committed artifact so that every interval quoted here is reproducible to the digits shown. On a real diffractometer, the recommended re-scan flips a single-phase stan- dard from fail to pass Key result: on physical hardware, the recommendation—not merely the act of re-measuring—is what carries a single-phase silicon standard across the reliability gate, the simplest rung of the wet–dry loop. Beyond the archival held-out benchmark, we close the recommend–rescan–reanalyze loop on physical hardware, running the same two-scan cycle on three real samples that span three crystal systems (table S5). All measurements are on laboratory Cu K훼 diffractometers outside the development corpora; for each sample the pipeline reads a fast coarse survey, the recommender emits a data-driven fine-scan window, a human operator re-measures at those settings, and the chain re-analyzes the fine scan. The decisive result is that the recommended re-scan is load-bearing (Fig. 3A): on a blinded silicon standard (identified as Si, Materials Project mp-149,퐹푑 ̄ 3푚, #227) the coarse survey refines only to 푅 wp = 22.2 and is flagged unreliable by the profile gate, whereas the fine scan acquired at the recommended window flips the same sample to a gated pass (푅 wp = 16.4, 푅 p = 8.8), with refined lattice 푎 = 5.43198 ̊ A—reproducing the 5.43189 ̊ A cell supplied to the refinement to+17 ppm, and+0.02% from the tabulated Si value 5.4309 ̊ A—space group preserved (227 → 227), and a CHGNet formation energy of +0.9 meV/atom whose 90% split-conformal interval brackets the true value of zero for elemental Si. The phase identity, space group, and energy gate are correct on both scans; only the fit quality crosses the reliability threshold after the 17 recommended re-scan, so the recommendation, not merely the act of re-measuring, is what produces the passing analysis. This is a single-instrument, human-in-the-loop demonstration, in which the operator executes and submits the recommended re-scan, rather than an autonomous campaign or a multi-instrument benchmark; its value is that the recommend–rescan–reanalyze cycle runs on real laboratory hardware and the reliability gates behave correctly on raw measured counts. The other two samples map the operating envelope rather than adding further passes. A corun- dum (훼-Al 2 O 3 ) specimen is identified with the correct space group (푅 ̄ 3푐) and passes the energy gate on both scans, and the recommended fine scan improves the fit (푅 wp 186.5→ 154.4), but the low counts keep it above the profile-reliability threshold on both passes, so the pipeline correctly withholds a pass. A multi-metal alloy is a harder test still, and the case in which the re-scan changes which phases are resolved and not only how well they fit (Fig. 4). The residual-adjudicated phase counter rejects the single-phase hypothesis on both passes, so the phase count is not what changes: on the coarse survey it pairs TiFe 2 with an FeNi 3 candidate (single-phase 푅 wp = 29.0, two-phase 24.5), but the non-negative least-squares estimator—a separate fit of the measured pattern against pure-phase references, which it renders at a 0.15 ◦ width that a 1 ◦ grid therefore samples poorly— assigns that second phase a 0% intensity share, so no minor phase is resolved, and the pass is flagged unreliable with the diagnostic that a true phase may be missing from the candidate database. On the recommended fine re-scan the counter again rejects one phase (single-phase 푅 wp = 26.3) and now pairs TiFe 2 with a face-centered-cubic Fe candidate carrying a 10.7% least-squares phase fraction—on the per-mass-normalized basis in force for that run, not the peak-normalized intensity share the coarse pass reports—converging to 푅 wp = 19.8, though the database-absent solid-solution constituent keeps the overall verdict partial. The best single-phase fit leaves a systematic residual that a second phase reduces but does not remove (푅 wp 26.3 → 19.8; Fig. 4B), and the per-phase gates report the outcome honestly—TiFe 2 passes while the Fe phase is flagged for review (Fig. 4D). Across the three samples the loop demonstrates a genuine coarse-fail-to-fine-pass flip on the stan- dard, a fit-quality gain that does not yet clear the gate under low counts, and a change in which minor phase is resolved driven by the recommended re-scan—and in every case the gates report the outcome, including non-passing ones, rather than over-stating reliability. A complementary benefit of the survey-first design is that it targets where expensive scan time is spent. The coarse survey sweeps the full angular range quickly—on the multi-metal alloy loop 18 above, 78 points over 3–80 ◦ at a 1 ◦ step—and the recommender then confines the slow, high-count fine acquisition to the narrower window it identifies as informative (1,901 points over 39–77 ◦ at a 0.02 ◦ step), rather than acquiring the entire range at fine resolution. Quantitatively, acquiring the full 3–80 ◦ range at the fine 0.02 ◦ step would demand ≈3,850 points; by restricting the slow pass to the informative 39–77 ◦ window the loop collects 1,901, roughly half the fine-resolution points, and skips the low-information low-angle and high-angle tails entirely. Because the coarse and fine passes deliberately differ in both range and sampling density, we read this as a targeting of acquisition effort rather than a controlled beam-time measurement, which we do not claim; the point is that the survey-first loop directs high-count counting to the angular region the coarse pass flags as diagnostic, instead of paying fine-resolution cost across the whole range. Scaling to multi-metal alloys: full multi-phase decomposition with per-phase verdicts Key result: on the hardest multi-metal alloys, the full chain decomposes several coexisting phases from a single measured scan and returns a calibrated per-phase verdict, passing the phases it can certify and flagging the rest rather than over-interpreting them. The closed-loop samples above establish the recommend–rescan–reanalyze cycle but reach only a two-phase alloy at their most complex. To exercise the multi-phase decomposition that motivates this work, we trace the full chain on a real three-phase Ti-15Nb alloy scan and a Ni-base GH4169 superalloy scan (Fig. 5), both laboratory Cu K훼 measurements outside the development corpora: the Ti-15Nb pattern is a worked case distributed with PyWPEM (46) and the GH4169 scan was measured for this work. On the Ti-15Nb pattern the whole-profile refinement resolves three coexisting metal phases, a 훽-Ti (BCC) matrix and 훼-Ti (HCP) with a metastable 훼 ′ -HCP variant, to a good fit (푅 wp = 9.0, 푅 p = 5.0) with per-phase abundances of 5.5, 48.3, and 46.2% from PyWPEM’s decomposed-intensity estimator, which weights each phase’s Lorentz–polarization-corrected decomposed intensity by its crystal density (a proxy for phase abundance: neither a weight fraction nor a structure-factor-weighted quantification; see Materials and Methods) (Fig. 5A to D). The reliability gates are reported per phase and honestly: the 훼-Ti HCP phase is symmetry-preserving (191 → 191) yet still carries an energy-gate warning (|Δ퐸 form | = 0.089 eV/atom), whereas the retrieved 훽-Ti cubic polymorph 19 (225 → 221, |Δ퐸 form | = 0.167 eV/atom) and the 훼 ′ variant, which has no Materials Project entry, are flagged for review rather than passed silently (Fig. 5E). The GH4169 superalloy is a supplied-hypothesis case rather than a blind one: the grade was known when the run was made, so the 훾-Ni plus 훿-Ni 3 Nb pair was handed to the multi-phase chain, which then retrieved, refined, and gated both phases on its own. It validates the industrially decisive precipitate—훿-Ni 3 Nb keeps 푃푚푛 (59 → 59) with |Δ퐸 form | = 5× 10 −5 eV/atom, a per-phase pass—and flags the 훾 matrix, a face-centered-cubic Ni–Cr–Fe solid solution with no Materials Project entry: Stage 1 retrieves a body-centered-cubic Ni entry for it instead, whose symmetry gate duly fails (229→ 221, |Δ퐸 form | = 0.148 eV/atom), and substituting the textbook face-centered-cubic Ni makes the fit worse (푅 wp 35.9→ 42.3), the signature of a solid-solution lattice that no database entry reproduces. The whole-pattern residual accordingly stays high (푅 wp = 35.9, 푅 p = 16.2) and the overall verdict is partial (Fig. 5F). These cases show the multi-phase decomposition, quantitative analysis, and per-phase reliability gates operating together on real alloy measurements, with the failing verdicts as informative as the passes. Together the three rungs trace the same loop up a difficulty gradient—a single-phase standard carried across the gate, then coexisting alloy phases a coarse survey cannot separate, then a multi-metal decomposition with per-phase verdicts—each step reporting what it resolves and, just as plainly, what it cannot. On real spectra, retrieval-plus-physics beats direct generation Key result: on the hardest broader-chemistry cohort, discriminative retrieval with physics-based refinement outperforms generative structure models that map a pattern straight to a candidate, so the production path is the retrieval-plus-refinement one. A natural question for an end-to-end PXRD system is whether the retrieval + refinement path could be supplanted by a generative model that maps a measured pattern directly to a candidate crystal structure. We compare the two paths on the broader-chemistry opXRD subset where the truth formula has no exact MP entry, the cohort where phase identification falls through to the intensity-only fallback and where, by construction, the generative models cannot rely on a retrieved chemsys-exact structure either. On an expanded 36-sample matched test pool (each method scored on the same broader- chemistry no-MP-match samples), the median peak-aligned Pearson of the top-1 candidate is 0.459 20 for Jaccard-ranked nearest-chemistry retrieval (the production tier-2 component), versus 0.193 for MatterGen (24) conditioned on the chemical system and 0.141 for deCIFer (22) with constrained chemistry-and-atom-mask decoding; retrieval is at least as good as both generators on 26 of the 36 materials. On the original 12-sample pool, two further deCIFer ablations score lower still: 0.137 for a fine-tuned deCIFer that received 2,000 additional synthetic training pairs, against 0.217 for the same base checkpoint without them, and 0.086 for unconstrained decoding. The extra training data therefore did not help, but we do not read an effect size out of it: two recorded runs of that one unconstrained configuration, at sampling temperature 1.0, span 0.086 to 0.217, a wider gap than the one the extra data produced. What is invariant across all four configurations is that none recovers the truth formula (0/12), which locates the bottleneck in the frozen decoder rather than in the amount of conditioning data. (These 12-sample medians follow that earlier ablation’s convention of taking the median over the samples with positive correlation; the 36-sample medians above are over all 36.) MatterGen returns the truth formula in 2 of 12 cases (17%) and deCIFer in 0 of 12, and its single best instance reaches 푟=0.959 (CuFeSSn), level with the highest retrieval score anywhere in the 36-sample pool (0.956); those two maxima are different materials drawn from pools that do not overlap, so they bound each method’s best case rather than comparing one sample. The generative path is genuinely improving but has not yet overtaken retrieval+physics on the cohort that exposes the simulation-to-real challenge most acutely (Fig. 6). A second, parallel, comparison is on quantitative phase analysis, the established Rietveld- based route to phase fractions (55, 56). On the Simonnet 2024 real-weighed-truth anchor set, our production non-negative least-squares estimator against peak-normalized pure-phase stick patterns (an intensity share; see Materials and Methods), evaluated on all development mixtures (the six held-out Simonnet patterns excluded), achieves an unweighted mean absolute error against the weighed truth of 0.195 at 퐾=2 (푛=10), 0.111 at 퐾=3 (푛=10), and 0.119 at 퐾=4 (푛=3); it improves on a uniform-share trivial baseline at 퐾=2 (0.195 versus 0.265) and 퐾=3 (0.111 versus 0.133), while at 퐾=4 it is worse than that baseline (0.119 versus 0.075, 푛=3), a negative we carry rather than average away. A head-only fine-tuned 퐾=2 CNN reaches 0.149 under leave-one-out over the twelve 퐾=2 mixtures, competitive with non-negative least-squares at two phases; but at 퐾 ≥ 3 that CNN architecture becomes structurally inapplicable (its binary-pair semantics do not generalize to ternary or quaternary mixtures) and, forced through pairwise decomposition, underperforms even 21 the trivial baseline. We therefore use non-negative least-squares as the primary estimator at all 퐾 ∈ 2, 3, 4, reserving the 퐾=2 CNN for two-phase mixtures, where it is competitive. Across the two comparisons, the same pattern recurs: the discriminative-and-physics path (chemistry-aware retrieval, physics-informed peak-aligned scoring, non-negative least-squares with explicit non-negativity, CHGNet as the trained single-point predictor with conformal half- widths from real-spectrum calibration) outperforms the generative-and-learning path on the real-data evidence base we tested, without claiming the former is theoretically superior. Each com- ponent is simply the strongest tool actually measured on the cohort where the simulation-to-real challenge is most exposed. Orchestration and advisory agents The control layer around the physical chain is a convenience and reliability layer, not a new scientific capability: an LLM planner schedules per-input configuration, three post-chain agents turn outputs into operator-facing advice, and a reliability layer wraps the loop (architecture in Materials and Methods). The planner’s authority is bounded to orchestration knobs—retrieval mode, top-퐾 , the refinement iteration cap, the 2휃 fitting window, and single- versus multi-phase routing—each range-checked against a fixed schema; no agent supplies a physical quantity, so candidate scores, refined parameters, mass fractions, energies, and gate verdicts are computed by the deterministic engines alone. Every plan is JSON-schema-validated with a deterministic rule planner as fallback, so pipeline availability never depends on an LLM, and the planner is opt-in and off by default: every held-out and benchmark number reported above was produced with it switched off. Keeping the LLM out of the reported numbers does not by itself make a run bit-reproducible: the refinement engine carries its own unseeded pseudo-random parameter selection on monoclinic cells, quantified in Materials and Methods. We nonetheless quantify the layer with the on/off discipline used for every other module, on development and curated data only (never the frozen held-out partition). The planner is silent on ordinary inputs and load-bearing on hard ones: on the two-phase Mn 2 O 3 /RuO 2 fixture, planner-on output is identical to planner-off in 12 of 13 headline fields (the thirteenth differs by 1.2× 10 −7 eV/atom, subprocess noise), at the cost of one∼1.7-s call; across eight feature-level inputs all three backends (DeepSeek-V3, on-prem Qwen2.5-7B, rule fallback) 22 emit 100% schema-valid plans, and on three edge cases engineered to need a non-default deviation the rule/Qwen/ DeepSeek planners deviate on 0/2/3 (one deviation correctly truncates a biased in- situ tail). The three advisory agents (polymorph, size/strain, recommender) report actual outputs on real spectra (fig. S7): the polymorph agent reaches 49.1% top-1 crystal-system accuracy (54/110, Wilson [39.9, 58.3]) and recovers the correct compound on 91% (31/34, [77.0, 97.0]) of the polymorph-ambiguous subset, with a discriminative confidence flag (confident verdicts correct 63.6% vs 34.5% for ambiguous; 푧=3.05, 푝=0.002); the size/strain agent returns a Williamson–Hall microstrain trend on the Fe–Mn series (휌=0.87, 푝=0.0025; martensite only, on four reflections per strain level, with per-level 푅 2 of 0.00–0.21, so the trend across levels is the result and no single level is) with a non-strippable low-absolute-reliability caveat; and the recommender emits at least one ranked next-experiment action on all 93/93 development runs. The agents abstain when appropriate—with the formula withheld, the polymorph agent returns ambiguous with zero candidates and the size/strain agent declines a single-spectrum input. An optional deep-research step composes a literature-grounded report (fig. S8), also LLM-generated and advisory. The reliability layer (Guardian bounded replan over the gates, verdict reproduction on 10/10 of the development chains that carry a stored canonical verdict, an eleventh skipped for a missing artifact, plus a Ti-15Nb replan-to-pass; Scribe append-only provenance; Phase-Guard detection-limit checks that convert a 퐾=3 refinement’s unsupported 20.3–46.5% HCP abundance into a quantified 3.4–11.4% limit of quantification, both on the same intensity-share basis) is summarized in fig. S9 and described in Materials and Methods. The recommender’s downstream usefulness is advisory and not benchmarked, and absolute crystallite size/strain is reported as a trend only. Table S6 consolidates the component-wise ablation evidence across the study’s eight dimen- sions (planner, post-chain agents, phase-identification fallback chain, simulation-to-real bridge, uncertainty method, interatomic potential, training strategy, and discriminative-versus-generative comparison), with the sample size of every comparison stated. Each production choice corresponds to a measured increment on real data, and each rejected alternative is kept as a quantified negative. 23 Discussion Implications and limitations The central result of this work is methodological rather than a single performance number: across three open mineralogical corpora the dominant obstacle to deploying a powder-diffraction deep- learning pipeline on measured spectra is a simulation-to-real gap, and that gap is best addressed by locating its cause before attempting to remove it. Our diagnosis shows the gap is structural rather than additive, and identifies a small rigid peak-position drift between simulated candidate lattices and measured patterns as its single largest correctable component, with a substantial structural residual (intensity redistribution and anisotropic lattice mismatch) that alignment does not remove and that we bridge downstream. The practical consequence is concrete: a denoising preprocessor, the obvious first response to noisy measured data, yields no lift on real opXRD spectra even when it performs as designed on its synthetic training task, whereas a peak-alignment search, a far cheaper intervention, more than doubles the median retrieval correlation while recovering about a third of the gap. A pipeline budget spent on denoising is, on this evidence, misallocated; a budget spent on alignment and on real-data fine-tuning is not. We regard this reorientation as transferable beyond the specific modules reported here. A further facet of the domain shift appears in the refinement quality metric itself. The weighted profile residual 푅 wp , whose absolute value has long been recognized as a fragile fitness criterion that depends on data conditioning as much as structural correctness (57), is sensitive to the absolute intensity scale: synthetic and database spectra are intensity-normalized, whereas raw laboratory counts are not, so high-count measured patterns are systematically penalized. The sensitivity enters through the weighting. Our refinement engine weights each point by 1/max(퐼 obs , 1) rather than by 1/퐼 obs ; with a true 1/퐼 obs weight both sums in 푅 wp would be homogeneous of degree one in the intensity scale and the ratio would be exactly invariant. The floor breaks that invariance in one direction only: on a peak-normalized pattern most points fall below unity (3,603 of 3,801 on the Si scan below) and their misfit is divided by 1 instead of by a smaller number, which under-weights it, whereas on the raw-count scale the floor almost never binds. On a real Si scan, refined twice from the same file against the same reference structure over the same angular window, 푅 wp falls from 37.8 to 15.4 under peak-normalization alone. The refined lattice parameter (5.43189 ̊ A) and 24 the space group (227) are identical across the pair and the unweighted 푅 p agrees to 8.79, and rescaling the recorded profile by hand, without refitting, reproduces the same movement in 푅 wp at bit-identical 푅 p : what moves is the metric, not the structure. The two refinements are not, however, bit-identical, and the third digit of 푅 p (8.797 against 8.786) shows it, because the expectation- maximization convergence test compares a fixed absolute tolerance against a log-likelihood that itself scales linearly with intensity, so the two runs halt at marginally different points. We therefore normalize input intensities to a common peak scale before refinement so that gate thresholds calibrated on normalized development data transfer to raw-count measurements. The corollary for anyone comparing against their own refinement is that an 푅 wp reported here is peak-normalized and a raw-count 푅 wp for the same fit will be substantially higher. The simulation-to-real gap is thus inherited not only by the predictors and the uncertainty layer but by the metric used to gate them. The reliability analysis extends the same domain-shift lens to the uncertainty layer, where it produces what we believe is the most broadly relevant observation of the study. Conformal prediction is increasingly adopted in scientific machine learning precisely because it offers distribution-free coverage guarantees, but those guarantees are conditional on the calibration set being drawn from the deployment distribution. We find that a conformal calibration built from clean density-functional- theory anchors under-covers measured spectra on development data—in plain terms, an interval advertised as 90%-reliable brackets the truth only about 55% of the time on real spectra until it is recalibrated—and, critically, that this under-coverage reproduces on held-out evaluation with a confidence interval that excludes the nominal level. The same conformal machinery recalibrated on a disjoint set of real measured spectra recovers near-nominal coverage on the development domain. The uncertainty layer, in other words, inherits the simulation-to-real gap of the predictors it wraps; a coverage guarantee certified on synthetic anchors is not a coverage guarantee on measured data. Any spectroscopic pipeline that ships synthetic-calibrated conformal intervals should expect the same effect and should calibrate on real measurements before quoting coverage. We have been deliberate in reporting where the pipeline does not work, and those boundaries are as informative as the successes, since each is the flip side of a design decision it justifies. The crystal- system classifier reaches roughly fifty percent absolute accuracy on a seven-way task; this is double the synthetic-only baseline and reproduces cleanly on held-out data, but half of single-spectrum predictions remain incorrect, with Monoclinic and Hexagonal the weakest classes, and a follow-up 25 loss-weighting study confirmed that the ceiling is set by the size of the real fine-tuning pool rather than by the training objective. The classifier is therefore used as a soft prior for downstream ranking, not as a hard filter. Phase-identification retrieval returns no candidate at any tier for roughly one in six held-out broader-chemistry minerals; these are genuinely outside the addressable scope of a Materials-Project-backed retrieval ladder rather than an artifact of the chain wiring, and they mark the chemistry frontier where a generative route may eventually be required, even though every generator we tested currently underperforms retrieval on exactly this cohort. Automated full-profile refinement does not converge on roughly thirteen percent of complex minerals, almost all extreme-angle low-symmetry cells where the external engine breaks down numerically; we log these rather than silently dropping them, and when refinement does converge it never violates the space group. Quantitative per-phase abundance validation rests on a small real-weighed-truth anchor set and is reported as small-sample throughout; the abundances themselves are intensity shares rather than weight fractions, so the weighed-truth agreement we report is agreement between an intensity share and a mass share, and it is correspondingly tighter for similar-scattering mixtures than for contrasting ones. At 퐾=4 the estimator does not improve on a uniform-share baseline on this anchor, which we report as a limit rather than as a result in either direction, given 푛=3 mixtures. Finally, the real-spectrum conformal recovery rests on a modest development-only calibration set (푛=49, disjoint from held-out); its coverage intervals are correspondingly wide and the smallest 훼 we report is limited by that set’s size. Transferring that calibration to the held-out partition yields conservative coverage (95.3% at 훼=0.10), which we therefore read as a coverage check rather than a tight, fully powered independent recalibration benchmark. Each limitation is reported with its sample size beside the corresponding result, and each implies a specific next step. Those next steps follow directly from the located limits, which we consolidate as an operating envelope (table S7). The single most consequential improvement is more real labeled measure- ment: a larger real RRUFF fine-tuning pool to lift the weak crystal-system classes, and a larger real calibration set to tighten conformal coverage at small 훼 and make real-spectrum recalibration the unconditional default rather than a domain-matched option. Broadening the structure database beyond Materials Project (for example to the Crystallography Open Database (58) or large model- discovered structure repositories (59)), or coupling retrieval with a generative proposer specifically on the broader-chemistry no-candidate cohort, would attack the phase-identification blind spot 26 where it actually occurs. Extending the evaluation to neutron diffraction and to in-situ or operando measurement series would test whether the diagnosis and the bridging interventions transfer to ac- quisition regimes with different degradation structure. A further promising direction is to close the loop with autonomous experimentation: beyond analyzing a pattern, the pipeline’s recommender agent already emits concrete next-experiment advice (for example a slower re-scan, a finer step, or an extended 2휃 range), and these suggestions could in principle drive an autonomous diffraction platform that can scan, analyze, judge data quality, and adaptively re-scan without a human in the loop, in the spirit of recent self-driving materials laboratories and machine-learning-guided adaptive diffraction (7,8,60). We demonstrate the human-in-the-loop form of this loop as a single-instrument proof of concept: the recommender proposes scan parameters, an operator re-measures on a Rigaku SmartLab, and the pipeline re-analyzes to a gated PASS on a Si standard (Results), while the fully autonomous, human-out-of-the-loop coupling and cross-instrument generalization remain future work. We deliberately make no claim of state-of-the-art performance on any single task; the con- tribution is Xtalyst as a reliability-first PXRD workflow—an integrated, multi-agent, end-to-end pipeline whose simulation-to-real behavior is, to our knowledge, the first to be diagnosed, bridged where possible, and calibrated on real measured spectra, with the residual gaps quantified and lo- cated so that the path forward is explicit. That a synthetic-to-real shift can erode a coverage guarantee is unsurprising; our contribution, to our knowledge a first for a real-mineral PXRD pipeline, is mea- suring that under-coverage on real spectra and showing that real-measurement recalibration restores near-nominal coverage. We establish this on a modest calibration set (larger-scale confirmation is fu- ture work), but the implication already reaches beyond diffraction to any spectroscopic pipeline that ships synthetic-calibrated uncertainty. On a real-data evidence base, we prioritize failures that are measured, located, and corrected in a working pipeline over a higher single-task benchmark number. 27 Real (opXRD, n = 75) 0.0 0.2 0.4 0.6 0.8 1.0 Top-1 peak-aligned Pearson r median 0.21 synthetic test split: r > 0.9 (reported) RawBG sub. U-NetRawPeak align 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Real opXRD median Pearson 0.318 0.320 0.315 0.146 0.385 Additive test (n = 25) Structural test (n = 306) denoiser null (−0.003) align +0.24 20253035404550 2θ (deg) Intensity (a.u.) Simulated (mp-2136) Measured 2030405060 2θ (deg) −0.4 −0.2 0.0 0.2 0.4 Peak drift Δ2θ (deg) mean ≈ 0 (no rigid shift) RMS 0.19°, anisotropic A The Gap ExistsB The Gap Is Structural, Not Additive C Simulated vs Measured (Sb 2 O 3 ) D Per-Peak Structural Drift Figure 1: The simulation-to-real gap, established (A) and diagnosed as structural (B to D). (A) Top-1 retrieval correlation (peak-aligned Pearson 푟 ): synthetic test-split regime (푟 > 0.9, reference band) versus real opXRD spectra (median 0.21, 푛=75). (B) Two candidate causes adjudicated on one matched cohort of 25 real opXRD spectra (direct top-1 Pearson). Additive test: raw 0.318, background subtraction ±0.002, 1-D U-Net denoiser 0.315 (median paired Δ=−0.004, Wilcoxon 푝=0.49, equivalent to zero within±0.05; yet the same net gains+0.30 synthetically). Structural test on the same 25: a peak-alignment slide search lifts the median 0.318 → 0.504 (Δ=+0.080, 22/25 improved, Wilcoxon 푝=8× 10 −7 ); it lifts the larger 푛=306 pool 0.146→ 0.385 (+0.24). The gap is structural, not additive. (C) A measured pattern (orange; Sb 2 O 3 ) on its correct-phase MP candidate (mp-2136, blue): right phase, yet intensities redistribute and peaks shift. (D) Per-peak drift Δ2휃 (measured− candidate), same example: mean≈ 0 (no rigid shift) but root-mean-square 0.19 ◦ , peaks up to∼0.4 ◦ , indicating anisotropic lattice mismatch. 28 Sample Unknown Sample Diffractometer Coarse Pattern Fast Survey Scan Coarse Survey XtalCS + CPICANN Candidate (Si, mp-149) recommended fine scan 2θ 24–100° / 0.02° slow / high count Fine Pattern 3801 pts / 0.02° / slow sharp, high S/N Phase Identification Si Fd-3m (#227) Refinement Rwp 16.4 Rp 8.8 Property & UQ Gate PASS ΔE 0.9 meV/atom Output Si cubic · Fd-3m space group 227→227 ΔE 0.9 meV/atom Rp 8.8 / Rwp 16.4 Opt-In LLM Planner plans the route & schema-checks every step Cross-Cutting Reliability Layer Size/Strain Phase-Guard Cs-Hint Polymorph scan 1scan 2 automated closed loop advisor re-issues a fine scan LLM Agent Method Stage Advice C Quantitative evidence Measured Coarse vs. Fine Scan Fine-Scan whole-pattern fit (PyWPEM) Si Fd-3m (#227) Rwp 16.4 Rp 8.8 Fine-Scan Verdict Candidate Si mp-149 Profile Rwp 16.4 Lattice +0.02% Energy±UQ +0.95 ± 154.6 Verdict PASS Case-Study Deep Research literature-grounded advice Optional Advice post-hoc, not in metrics B Overall Framework Prior Deep-Learning PXRDXtalyst Point Estimate No Uncertainty, No Gate ? Point Estimate & Conformal Interval Distribution-free, Covers the Truth Gated PASS A What Xtalyst adds: every prediction ships calibrated and gated Figure 2: The Xtalyst system: an end-to-end PXRD pipeline that ships every prediction calibrated and gated. (A) Prior deep-learning PXRD looks near-solved on simulation (median peak-aligned Pearson 푟 > 0.9) yet collapses on real spectra (median 푟 ≈ 0.21, 푛=75 opXRD minerals). (B) The pipeline, end to end: a fast coarse survey on a laboratory diffractometer is screened by the crystal-system classifier (XtalCS) and the identifier CPICANN, which propose a candidate and a recommended fine-scan window; the re-measured fine pattern flows through phase identification, whole-pattern lattice refinement (PyWPEM), and property prediction with split-conformal uncertainty (CHGNet), under a cross-cutting reliability layer (symmetry, profile-quality and energy gates, plus the four auxiliary checks the panel names). The optional large-language-model steps (planner, Deep-Research advisory, post-hoc agents) are marked separately and enter no reported result or gate verdict. (C) An illustrative worked trace on a blinded silicon standard; Fig. 3 develops this single-phase case quantitatively. 29 A The wet--dry loop carries a single-phase standard across the gate (real Si) coarse survey Phase IDRefineGate Si Fd ̄ 3m (#227) R wp 22.2 FAIL · unreliable recommend fine window · re-scan fine scan Phase IDRefineGate Si Fd ̄ 3m 227 →227 R wp 16.4 gated PASS B Full-profile fit of the fine scan (PyWPEM) C Fine-tuningD Peak alignment E Calibrated verdict space group 227 →227 PASS |ΔE form | 0.0009 eVPASS profile R wp 16.4PASS lattice fidelity a = 5.43198 vs 5.43189 Å (+17 ppm) CHGNet energy interval 0.94 ± 154.6 meV/atom brackets 0 30405060708090100 2θ (deg) 0 50 100 intensity (norm.) gated PASS: R wp = 16.4, R p = 8.8 measuredPyWPEM fitdifference synthbal.FTheld 0 20 40 60 XtalCS acc. (%) 26.5 25.9 49.8 49.4 rawaligned 0.0 0.2 0.4 0.6 0.8 1.0 retrieval r 0.21 0.48 synthetic r > 0.9 Figure 3: The wet–dry loop, and how real-spectrum learning narrows the simulation-to-real gap on a single-phase standard. (A) The loop on a blinded silicon standard: a fast coarse survey flows through the three analysis stages under a cross-stage reliability gate; the phase is correctly identified (Si, 퐹푑 ̄ 3푚) but the profile gate fails (푅 푤푝 = 22.2), the advisor recommends a fine-scan window, and the re-measured fine scan returns a gated pass (푅 푤푝 = 16.4, 227→ 227). (B) The passing full-profile fit (measured, calculated, difference; 푅 푤푝 = 16.4, 푅 푝 = 8.8). (C) Real-spectrum fine- tuning: XtalCS crystal-system accuracy 26.5% (synthetic-only)→ 25.9% (class-balanced)→ 49.8% (fine-tuned) on one RRUFF development test set (푛=313), and 49.4% on the frozen held-out partition (푛=251; detail in fig. S2). Class balancing raises the worst-class floor rather than the aggregate, so the middle bar is flat by construction. (D) Peak-aligned reranking lifts the median retrieval correlation 0.21→ 0.48 on real opXRD spectra. (E) The calibrated verdict: all three gates pass, the refined lattice reproduces the cell supplied to the refinement to+17 ppm (+0.02% from the tabulated Si reference), and the CHGNet formation energy (+0.94 meV/atom against a true value of zero) carries a conformal interval that brackets that true value. The interval is quoted at the half-width of the leak-free real-spectrum recalibration, ±154.6 meV/atom at 훼=0.10 (푛=49), which is the calibration we adopt for measured patterns; the synthetic-anchor calibration would report±89.8 on the same prediction, and its over-confidence on real spectra is the deficit quantified in table S6, row E. All values are from recorded runs (Materials and Methods). 30 A The re-scan changes which minor phase is resolved (real multi-metal alloy) coarse survey Phase IDRefineGate 2 phases TiFe 2 + FeNi 3 FeNi 3 share 0% R wp 24.5 · unreliable recommend fine window · re-scan fine scan Phase IDRefineGate 2 phases TiFe 2 + Fe Fe share 10.7% R wp 19.8 · partial B A single-phase fit leaves a systematic residual (PyWPEM) C Intensity share (NNLS) TiFe 2 89.3% Fe 10.7% D Per-phase reliability gates phase SG in →out |ΔE| gate TiFe 2 (mp-2454) 194 →194≈ 0 PASS Fe (mp-150) 225 →221 1.61review 4045505560657075 2θ (deg) −50 0 50 100 intensity (norm.) TiFe 2 reflections single-phase R wp = 26.3 → two-phase R wp = 19.8 (C, D) measured single-phase TiFe 2 fit residual % of intensity Figure 4: The wet–dry loop changes which minor phase is resolved. Real multi-metal alloy, traced through the loop twice. (A) On the fast coarse survey (78 points at 1 ◦ ) the residual-adjudicated counter rejects the one-phase model and pairs TiFe 2 with an FeNi 3 candidate (푅 푤푝 29.0→ 24.5), but the intensity-share estimator gives that second phase 0%, so no minor phase is actually resolved and the pass is flagged unreliable; the advisor recommends a fine-scan window, and on the re-scan (1,901 points at 0.02 ◦ ) the counter again resolves two phases, now TiFe 2 plus a face-centered-cubic Fe candidate carrying 10.7% (푅 푤푝 = 19.8, partial); the two structures resolved on the fine pass are shown inline. (B) Why the second phase is added: the best single-phase (TiFe 2 -only) fit leaves a systematic residual (푅 푤푝 = 26.3), so the counter rejects the one-phase model and a two-phase fit converges to 푅 푤푝 = 19.8 (measured, calculated, and residual profiles; TiFe 2 reflection ticks). (C) Per-phase abundances from the production non-negative least-squares estimator: TiFe 2 89.3, Fe 10.7%, on the per-mass-normalized basis in force for this run; the coarse pass’s 0% is a peak-normalized intensity share, so the two passes’ fractions are not on one basis, and neither is a weight fraction corrected for absorption. (D) Per-phase gates: TiFe 2 passes (194 → 194, |Δ퐸 form | ≈ 0), while the database-absent solid-solution Fe phase (225 → 221, |Δ퐸 form | = 1.61) is flagged for review rather than passed silently. All values are from recorded runs (Materials and Methods). 31 405060708090 2θ (deg) 0.0 0.5 1.0 intensity (norm.) R wp = 9.0, R p = 5.0, 3 coexisting phases measured WPEM fit difference 405060708090 2θ (deg) per-phase profile β-Ti (BCC) α-Ti (HCP) α ′ -Ti (HCP var.) β-Ti α-Ti 6% 48% 46% phase abundance density-weighted intensity estimate, not scale-factor QPA phase SG in →out|ΔE| gate β-Ti BCC (mp-6985) 225 →221 0.167review α-Ti HCP (mp-72) 191 →191 0.089warn α ′ -Ti (not in MP) N/AN/Areview Each phase carries its own verdict (reported separately). 20406080100 2θ (deg) 0.0 0.5 1.0 intensity (norm.) δ-Ni 3 Nb: Pmmn 59 →59, |ΔE| ≈ 0, per-phase pass γ solid solution: BCC Ni retrieved, 229 →221, flagged supplied γ+δ hypothesis (grade known); whole-pattern R wp = 35.9, R p = 16.2, overall verdict partial A Ti-15Nb Alloy, Whole-Pattern Fit (PyWPEM) B Deconvolved Phase ContributionsC Refined Structures D Per-Phase AbundanceE Per-Phase Structure-Validation Gates F GH4169 Superalloy, Per-Phase Gates on a Supplied γ+δ Hypothesis Figure 5: Multi-phase resolution of real metal alloys from single laboratory scans. (A) Whole-pattern fit of the Ti-15Nb alloy (measured, PyWPEM calculated, and difference; 푅 wp = 9.0, 푅 p = 5.0). (B) Deconvolved per-phase contributions for the 훽-Ti (BCC), 훼-Ti (HCP), and metastable 훼 ′ -Ti (HCP variant) phases. (C) Refined 훽-Ti and 훼-Ti structures, rendered from the refined CIFs. (D) Per-phase abundances from PyWPEM’s decomposed-intensity estimator (5.5, 48.3, 46.2%): decomposed intensity weighted by crystal density (Materials and Methods), a proxy that is neither a weight fraction nor a scale-factor Rietveld quantitative-phase analysis. (E) Per-phase structure-validation gates (space-group preservation and|Δ퐸 form |), each phase with its own verdict. (F) GH4169 Ni-base superalloy scan, run from a supplied 훾+훿 hypothesis (the grade was known): 훿-Ni 3 Nb refines with 푃푚푛 preserved (59 → 59) and |Δ퐸 form | = 5× 10 −5 eV/atom (a per-phase pass), while the 훾 Ni–Cr–Fe solid solution is flagged for review (229→ 221, 0.148 eV/atom); 푅 wp = 35.9 and the overall verdict is partial. All panels are real archived measurements and pipeline outputs. 32 Retrieval (ours) Matter- Gen deCIFer 0.0 0.1 0.2 0.3 0.4 0.5 Median Pearson r 0.459 0.193 0.141 Held-out broader-chemistry material (sorted by retrieval) 0.0 0.2 0.4 0.6 0.8 1.0 Peak-aligned Pearson r Retrieval (this work) MatterGen deCIFer (chem-mask) 15202530354045505560 2θ (deg) Norm. intensity (offset) Measured (opXRD) retrieval top-1: r = 0.83 MatterGen: r = 0.23 deCIFer: r = 0.19 A Aggregate Median (n = 36)B Per-Material (Retrieval ≥ Both on 26/36) C Worked Example: Generated Patterns Miss the Peaks Figure 6: Discriminative retrieval versus generative structure models on a matched 36-sample broader-chemistry no-MP-match pool (peak-aligned Pearson 푟 ; all methods scored on the same samples). (A) Aggregate median: Jaccard nearest-chemistry retrieval (this work) 0.459 versus MatterGen 0.193 and deCIFer 0.141. (B) Per-material breakdown sorted by retrieval; retrieval is at least as good as both generators on 26 of the 36 materials (dotted lines: retrieval and MatterGen medians). (C) Worked example (opXRD sample opxrdhkust000520): the best MatterGen (푟 = 0.23) and deCIFer (푟 = 0.19) generated patterns miss the measured peaks, whereas the retrieval top-1 candidate scores 푟 = 0.83. Two further deCIFer ablations on the original 12-sample pool (data-scale-up 0.137; unconstrained decoding 0.086 and 0.217 in two recorded runs) are noted in the text, with the median convention they use. 33 Materials and Methods Datasets and the held-out protocol We ingest 2,696 powder X-ray diffraction patterns from four sources: 1,277 measured patterns from the opXRD open mineralogical corpus (26) (the HKUST(Guangzhou) subset), 1,365 measured pat- terns from RRUFF (27), 32 measured patterns from the Simonnet et al. quantitative phase-analysis round-robin (51) with real-weighed ground-truth mass fractions, and 22 collaborator-supplied in- ternal patterns, one of which was identified during ingest as an X-ray reflectivity measurement and excluded as out of scope. All measured patterns are interpolated onto a common 2휃 grid (step 0.02 ◦ , range 10–80 ◦ ). We partition the ingest into a frozen held-out set (푛 = 534: 255 opXRD, 273 RRUFF, 6 Simonnet) and a development pool (푛 = 2,696− 534− 22 = 2,140, less the collaborator subset). The held-out partition is fixed in advance and is read exactly three times across the entire study, each read fixed before the data were touched: once for the headline held-out evaluation; once for a single conformal recalibration check on the existing refined structures (no upstream phase-identification or refinement re-run); and once for a small quantitative-phase-analysis read-out on the held-out Simonnet mixtures (5 of the 6 admit a 퐾 -phase decomposition; table S4). After the third read the held-out partition is retired and is not consulted again in any later work. Per-module evaluation 푛 is reported alongside every quantitative number; the figure 2,696 is the ingest total, and is never used on its own as an evaluation sample size. Crystal-system classifier (XtalCS) XtalCS is a one-dimensional convolutional neural network (six Conv1d blocks with stride-2 down- sampling, BatchNorm and SiLU activations, an adaptive-average-pool head over the final feature map, followed by two fully-connected layers and a 7-way softmax, with 1,641,351 trainable pa- rameters; input shape(퐵, 1, 3584)) trained in two stages. The first stage is synthetic pre-training on stick patterns simulated from 2,800 Materials Project structures (47), balanced as 400 structures per crystal system and split into 2,401 training and 399 held-out structures, with online degrada- tion augmentations; on clean (non-degraded) renders of the 399 held-out structures the pre-trained checkpoint reaches 44.1%, and on the degraded renders it reaches 40.4%, which bounds what the 34 same architecture achieves before any real spectra are seen. The second stage is a real-data fine-tune on the RRUFF development pool of 730 measured spectra (FT-train 584, FT-val 146). After exclud- ing one RRUFF sample (rruffR060379) that was inadvertently present in both the fine-tuning pool and the held-out partition, the fine-tuning pool has no overlap with the held-out 273 or with the development test set of 313; the held-out crystal-system evaluation excludes this sample (푛=251). The fine-tune uses the Adam optimizer with learning rate 10 −4 , batch size 49 (7× 7), cosine an- nealing, and mini-batches mixing real and synthetic samples at a 50% ratio. Checkpoint selection is by minimum per-class accuracy on FT-val. The production checkpoint (step 1,100 of 1,500) is the one reported as XtalCS throughout this work; intermediate synthetic-only and class-balanced retraining variants are reported as ablations in the corresponding Results section. Denoiser (synthetic-degradation U-Net) The denoiser evaluated as the additive-hypothesis test is a one-dimensional U-Net (four en- coder/decoder levels, base width 32 channels doubling to 256, on the common 3,584-point 2휃 grid). It is trained purely on synthetic (clean, degraded) pairs: clean patterns are stick patterns rendered from Materials Project structures, and the degradation model draws, per sample, five gap-source corruptions—amorphous background humps (1–3 humps at 5–45% of the maximum peak), peak broadening (FWHM 0.20–0.60 ◦ ), texture/intensity redistribution (log-normal 휎 0.3– 0.8), contaminant peaks (probability 0.5, 5–30% relative intensity), and combined white and 1/푓 counting noise (0.5–4%). The network is trained to regress the clean pattern from the degraded one (Adam, learning rate 10 −3 , cosine annealing, batch 16, 3,000 steps, seed 42); the extrinsic evalua- tion reports the top-1 profile Pearson correlation between the denoised real spectrum and its top-1 candidate simulation. This module is a diagnostic only and is not part of the production pipeline. Phase identification: tiered retrieval with peak-aligned reranking of candidate sets Phase identification operates as a three-tier fallback chain. Tier 1 performs a chemsys-exact hybrid retrieval against Materials Project: the candidate pool is first restricted to materials whose chemical system exactly matches the input formula’s elements (a composition constraint), and the surviving candidates are then ranked by peak-aligned Pearson correlation against the observed pattern (a pattern-similarity score). We call this tier hybrid because it combines the composition constraint 35 with the pattern-similarity reranking rather than relying on either alone. Tier 2, invoked when Tier 1 returns no candidate, performs a Jaccard-ranked nearest-chemistry retrieval that broadens the candidate pool to materials whose element set has high Jaccard overlap with the input. Tier 3, invoked when Tier 2 also returns no candidate, runs the open-source CPICANN intensity-only deep- learning identifier (12) and augments its top-퐾 output with peak-aligned reranking introduced in this work. We distinguish two senses of the term throughout: peak-aligned Pearson is the candidate- quality metric (the Pearson correlation between observed and simulated patterns after a small rigid 2휃 alignment, ±0.5 ◦ on the 10–80 ◦ scans of the development set; the window is grid-relative, see below), whereas peak-aligned reranking (equivalently, peak-alignment) is the intervention that reorders candidate lists by that metric. The peak-aligned scoring function is a Pearson correlation between the observed pattern and a forward-simulated candidate pattern, evaluated over a slide search of±25 points of the resampling grid; the highest-correlation shift defines the candidate’s score. Both the increment and the reach of that search are set by the submitted scan, not fixed in degrees: the pattern is resampled onto 3,500 points spanning the submitted 2휃 range, so the increment is(2휃 max − 2휃 min )/3499 and the window is 25 times that. For the 10–80 ◦ scans of the development set this is 0.020 ◦ and±0.500 ◦ , the values we quote as the production setting throughout; a narrower re-scan is searched more finely over a pro- portionally narrower window (a 60–80 ◦ window gives 0.006 ◦ and±0.143 ◦ ). The structural-cause diagnostic of Fig. 1B instead used a wider±1.0 ◦ search to characterize the full peak-shift distribu- tion (per-stratum median absolute shifts 0.14 ◦ and 0.58 ◦ , means 0.27 ◦ and 0.54 ◦ ), which is why a reported shift can exceed the production window. The shift the search settles on is reported back to the submitter as an instrument-facing 2휃 offset, per phase and as one median for the sample, carrying two caveats it cannot be read without: it is quantized to the grid increment, and 38 of the 304 de- velopment spectra (12.5%) return the largest shift the window admits, where the true displacement may be larger and the reported figure is a lower bound. Forward simulation for the reranker uses the pymatgen XRDCalculator at a single wavelength (1.5418 ̊ A, the unresolved Cu K훼 average) with each reflection broadened by a Gaussian of FWHM = 0.15 ◦ and the pattern normalized to unit maximum; it carries no K훼 1 /K훼 2 doublet, Debye–Waller factor, or asymmetry (the Pseudo- Voigt/Caglioti 푈/푉/푊 profile and the 1.540593/1.544414 ̊ A doublet enter only downstream, in the PyWPEM full-profile refinement, not in this retrieval score). On the Tier-3 path, we resolve the 36 CPICANN candidate’s reduced formula to a Materials Project structure, forward-simulate, and se- lect the highest-scoring among the top-퐾 = 6 candidates. The hard crystal-system filter we tested as a candidate intervention degraded median correlation at the current XtalCS accuracy and is not used. Coordinate-consistent whole-pattern refinement The refinement module performs whole-pattern profile refinement using the PyWPEM engine (46) integrated via a coordinate-consistent wrapper that we introduce in this work. PyWPEM decomposes the whole pattern as a physics-constrained probabilistic mixture solved by expectation– maximization with Bragg consistency imposed, which its authors position as an alternative to the conventional least-squares Rietveld formulation rather than an implementation of it (46). The refined quantities returned to the chain are the six lattice parameters per phase. Atomic coordinates are carried through from the retrieved candidate unchanged and no atomic displacement parameters are refined, so an exported refined Crystallographic Information File (CIF) should be read as a re-celled candidate structure rather than an independently solved one; the structural degrees of freedom a conventional Rietveld refinement would release (fractional coordinates, site occupancies, isotropic or anisotropic displacement parameters, preferred orientation) are held fixed here. The wrapper resolves axis permutations between an MP retrieval cell and an author-supplied cell when present (detecting the permutation of the 푎푏푐 triple by lattice-parameter ratio matching), promotes a primitive cell to its conventional counterpart through the pymatgen (61) SpacegroupAnalyzer with symprec= 0.01, synthesizes a starting peak list (peak0) at the candidate’s HKL positions over the configured 2휃 window; PyWPEM otherwise expects a hand-authored per-phase peak list, so this synthesis is what lets the engine run automatically on an arbitrary retrieved candidate, and invokes PyWPEM with the production configuration subset number = 9, lowbound = 15, upbound = 55, bta = 0.85, itermax = 150, asyC = 0, and a background fit with bacnum = 1000 background points, selected by an FFT low-pass followed by a Savitzky–Golay filter and fitted with a degree-6 polynomial. Because the background is fitted as part of the refinement, patterns enter as measured; a pattern whose background has already been subtracted lies outside the configuration under which every number here was obtained. The refined CIF is exported; the post-refinement check computes the input-versus-refined space-group number via pymatgen and flags any symmetry 37 change as a reliability-gate failure. The refinement subprocess is invoked from a dedicated conda environment (pyxplore) to isolate dependency conflicts with the property-prediction stage. Reported cell precision and run-to-run repeatability. The refinement returns a point estimate of the six cell parameters and no covariance matrix, so no estimated standard deviation is computed for any refined cell parameter and none is reported; the digits we print are the engine’s numerical precision, not a statement of how well a parameter is determined. Repeatability, unlike uncertainty, is measurable, and it is crystal-system dependent. Across the deployment’s recorded jobs we identified eight sets of repeat runs (23 runs total), grouping runs only where the submitted intensity file was byte-identical and the wavelength, angular window, iteration cap, reference structure and starting cell all matched; the background-fit variance is identical within every group, which localizes any divergence to the Bragg refinement rather than to preprocessing. Six sets (14 runs; cubic, hexagonal and trigonal materials) returned bit-identical cell parameters on every repeat. Two sets (nine runs, both a monoclinic Al 2 Si 2 O 9 specimen) did not: 푎, 푏 and 푐 spanned up to 0.0027 ̊ A and 훽 up to 0.007 ◦ across repeats, with only the first two decimals of푎, 푏,푐 common to all runs. The mechanism is specific rather than general: the Bragg-step routine selects which cell parameter to update by an unseeded pseudo-random draw, and does so in the monoclinic branch alone, while every other crystal system follows a fixed update order. Because the recorded repeat runs reuse working directories, and a shared directory could in principle explain agreement without reproducibility, we confirmed the bit-identical result on a purpose-built control: one cubic specimen (Si) refined three times from three separate, freshly created working directories returned identical cell parameters and an identical background-fit variance. We therefore describe the chain as free of LLM influence on reported numbers, which it is, rather than as bit-reproducible, which holds for the crystal systems we have replicate runs for but not for monoclinic cells. Two limits on the measurement: it covers five materials and four of the seven crystal systems, so an untested system is untested rather than shown stable, and repeatability bounds nothing about accuracy, which remains governed by zero-shift, sample-height and wavelength calibration. Profile-quality gate and intensity normalization. Because the weighted profile residual 푅 푤푝 is sensitive to the absolute intensity scale (Discussion), raw laboratory counts are peak-normalized 38 to a common maximum before refinement so that residual thresholds calibrated on normalized development data transfer to raw-count measurements. The rescaling is uniform, so it preserves peak positions and relative intensities; 푅 푝 is scale-invariant by construction, and on the Si pair reported in the Discussion the refined lattice and space group are identical across scales while 푅 푝 agrees to three significant figures. Reported 푅 푤푝 values are therefore peak-normalized figures and are not directly comparable to a raw-count 푅 푤푝 for the same fit. Two distinct residual criteria are used and should not be conflated. The profile-reliability threshold used for the reported reliability verdicts and the recommender’s high-푅 푤푝 trigger flags a normalized 푅 푤푝 above∼20 as an unreliable fit (so, for example, the silicon standard’s normalized 푅 푤푝 = 16.4 is reported as a reliable pass). Separately, the orchestrator carries a stricter internal Boolean rwp gatepass at 푅 푤푝 < 8, a conservative auto-accept guard used only in the multi-phase routing logic; it is intentionally tighter than the reliability threshold and is not the criterion behind the reported single-phase passes. Per-phase abundance quantification Quantitative phase analysis uses a non-negative least-squares estimator against pure-phase reference patterns. For each input pattern with an identified set of 퐾 phases, we forward-simulate the stick pattern of each phase using the same Pseudo-Voigt model as phase identification, scale each per- phase pattern to a common peak height, stack them into a design matrix 퐴 ∈ R 푚×퐾 over the common 2휃 grid, and solve min 푤≥0 ∥퐴푤− 푦∥ 2 2 where 푦 is the observed pattern. The returned weights are renormalized to sum to unity. What those weights measure follows from how the columns are scaled, and we state it explicitly because it is easy to over-read. A column scale is not identifiable from a least-squares fit: multiply- ing column 푘 of 퐴 by any 푠 > 0 leaves the residual unchanged and simply divides 푤 푘 by 푠. Scaling every column to a common peak height therefore discards the relative amplitude between phases, and what the renormalized weights carry is each phase’s share of the scattered intensity, not of the specimen mass. We report them as per-phase abundances on that basis throughout, including in the figures. Dividing the stick heights by the unit-cell mass instead—with one common factor across the library, so the relative per-mass amplitudes survive—gives coefficients that do carry mass-share semantics, and we implemented and measured that variant on the same weighed-truth anchor: it is indistinguishable from the shipped basis on accuracy (paired Wilcoxon 푝 = 0.70; 39 bootstrap 95% CI on the per-prediction difference [−0.011,+0.022], spanning zero, 푛=76 phase predictions), which is expected, since without an absorption and preferred-orientation correction the per-mass columns are not yet a weight-fraction basis either. We therefore report the intensity share the estimator actually computes rather than the weight fraction it does not. This estimator is non-learning, has no 퐾 -specific training, and operates identically at 퐾 ∈ 2, 3, 4. A 퐾 = 2 CNN we evaluated in development collapses below a uniform-share baseline at 퐾 ≥ 3 and is retained only as an opt-in 퐾 = 2 alternative. The estimator omits a microabsorption correction and the per-phase 푍·푀·푉 (formula units× molar mass× cell volume) factor, unlike a full Rietveld quantitative-phase analysis, which fits scale factors against structure-factor-weighted intensities and can incorporate a microabsorption correction (55, 62). Reading an intensity share as a weight fraction is therefore safe only where the phases scatter comparably per unit mass, and it over-represents a strong scatterer where they do not. It is most reliable for same-element or similar-scattering polymorph mixtures; on development mixtures with strong scattering contrast or trace (≲ 5%) phases the recovered fractions can bias toward equal shares and are reported as approximate. Two of the reported case studies predate this estimator becoming the production default and are labeled with the route that produced them. The TiFe 2 /Fe fractions (Fig. 4C) are from the non- negative least-squares estimator above. The three-phase Ti-15Nb fractions (Fig. 5D) are PyWPEM’s own MassFraction estimate route, which weights each phase by the mean Lorentz–polarization- corrected intensity of its 푛 strongest decomposed reflections times its crystal density, without structure factors; it is a decomposed-intensity proxy, and like the non-negative least-squares route it is not a scale-factor Rietveld quantitative-phase analysis. Neither number should be read as a structure-factor-weighted phase quantification. Formation-energy prediction and conformal uncertainty The formation energy of the refined structure is predicted by CHGNet (45). Two different energy differences appear in this work and are not the same quantity. The reliability gate acts on the refinement response |Δ퐸 form | = |퐸 form (refined)− 퐸 form (retrieved)|, with both terms predicted by the same model on the two structures: it asks whether re-celling moved the predicted energy, its thresholds are < 0.05 eV/atom passing and < 0.10 a warning, and every per-sample |Δ퐸 form | quoted in Results is this quantity. The conformal interval instead acts on the prediction residual 40 | ˆ푒 − 푒| between the model prediction and the density-functional-theory reference, which is also what the 43.5 meV/atom benchmark mean absolute error measures. The two are independent—a sample can have a small refinement response and a large prediction residual, or the reverse— and fig. S8 reports one sample where they are 1.86 and 16.9 meV/atom respectively. The choice of CHGNet over alternative machine-learning interatomic potentials was made on the basis of a five-MLIP benchmark on a 307-structure subset of Materials Project with density-functional- theory formation-energy reference values (whose own accuracy against experiment is∼few-tens of meV/atom (63)); CHGNet achieved the lowest mean absolute error on the 307-structure benchmark (43.5 meV/atom). Every model was scored on this same 307-structure set: the nearest competitors were MatterSim v1.0.0-1M (189.6 (53)), SevenNet-0 (215.4 (52)), MACE-MP-0 (218.7 (49)) and MEGNet (348.0 (48)). A stratified 49-structure subset gives the same ordering (CHGNet 42.1, MatterSim 152.3, SevenNet 167.3); on that subset the three MACE-MP-0 size variants score 171.6 (small), 168.6 (medium) and 173.6 (large), so model capacity within a family does not close the gap. The full benchmark, being harder and more chemically diverse, raises every absolute error but leaves CHGNet’s lead decisive (bootstrap 95% MAE intervals non-overlapping: CHGNet [39, 48] vs. MatterSim[172, 208] and SevenNet[196, 236]; paired Wilcoxon 푝 < 10 −30 ). SevenNet was run in an isolated environment so its dependency stack did not perturb the other models. The comparison is not exhaustive: other universal potentials, among them M3GNet (64) and NequIP (65), were not benchmarked here, so the claim is that CHGNet is the best of the five models tested on this MP-referenced task, not that it is optimal among all published potentials. A learned mixture-of- experts blend overMEGNet, MACE-MP-0, CHGNet gave a best gated MAE of 59.7 on the same 307-structure benchmark, 37% worse than single CHGNet; the mixture is retained as an opt-in difficulty signal only. The gate is a single softmax-normalized linear layer (21 parameters) over six composition features (transition-metal, oxygen, and halogen indicators; chalcogen fraction; unique- element count; and maximum atomic number), producing a per-sample convex combination of the three experts; it is trained by full-batch gradient descent (mean-squared error to the DFT truth, 퐿 2 weight decay) and emits a composite difficulty signal (gate entropy plus expert disagreement) that correlates with the gated error. The gate’s 23 calibration anchors overlap the 307-structure benchmark by five structures; because the MoE underperforms the pretrained CHGNet, this overlap 41 can only flatter the weaker model and does not affect the conclusion (CHGNet itself is pretrained, not fitted on the anchors, so its 43.5 meV/atom benchmark MAE is unaffected). Uncertainty quantification uses split conformal prediction (34, 37): given a calibration set of pairs ( ˆ푒 푖 ,푒 푖 ) where ˆ푒 푖 is the predicted formation energy and 푒 푖 is the DFT truth, the half-width at nominal coverage 1 − 훼 is the empirical ⌈(푛 + 1)(1 − 훼)⌉-th order statistic of | ˆ푒 푖 − 푒 푖 |. The synthetic-anchor calibration is scored on the production CHGNet residuals against Materials Project DFT anchors; its half-widths at 훼 = 0.05, 0.10, 0.20 are 109.4, 89.8, 60.7 meV/atom re- spectively, measured at the 118-structure stage of the benchmark set, and it is the 89.8 meV/atom (훼=0.10) width against which every synthetic-anchor coverage figure reported here was scored. The benchmark set was later extended to 307 structures, where the same split conformal gives 123.2, 95.3, 70.1 meV/atom; re-scoring both coverage measurements against those wider intervals leaves the conclusions unchanged—real-spectrum coverage at 훼=0.10 is 53/97 either way, and held-out coverage moves from 81.1% to 82.7% with nominal 0.90 outside the interval in both cases—so we report the widths the recorded runs actually used. Separately, an 푛=23 MEGNet- scored calibration file predates the production switch to CHGNet and retains MEGNet predictions, giving much wider half-widths of 746.4, 684.2, 538.7 meV/atom at the same 훼; it remains what the pipeline’s opt-in conformal call reads at run time, and it is reported here only as a legacy diagnostic baseline. The real-spectrum recalibration set of 푛 = 49 development refined structures with DFT ground truth (disjoint from the held-out partition) is scored on the production CHGNet residuals, and its half-widths at the same 훼 are 235.8, 154.6, 68.0 meV/atom respectively; this recalibration is the default for the real-mineral domain and is demonstrated in the corresponding Results section to recover near-nominal coverage on the disjoint development domain (the frozen held-out partition is deliberately not re-used for calibration). Multi-agent orchestration A multi-agent control layer schedules the per-input configuration of the chain and runs three post-refinement agents. The configuration planner is a large-language-model agent (DeepSeek- V3 as the primary backend, with an on-prem Qwen2.5-7B fallback) that ingests input-feature summaries and emits a JSON-schema-validated configuration override; a rule planner serves as a deterministic fallback when both LLM backends are unavailable. The override is restricted to 42 six fields with fixed admissible ranges—retrieval mode (chemistry-aware hybrid or intensity-only), top-퐾 ∈ [1, 20], refinement iteration cap ∈ [50, 700], 2휃 window lower bound ∈ [10, 40] ◦ and upper bound ∈ [40, 80] ◦ , and a Boolean multi-phase force—and any missing, out-of-range, or unparseable field rejects the whole plan in favor of the defaults, so the planner can narrow or widen a search but cannot introduce a physical value. The planner is opt-in and off by default; every archival benchmark, ablation, and held-out number reported in this paper was produced with it disabled, and the agent layer itself is characterized separately on development and curated fixtures. The closed-loop demonstration likewise runs with the planner off; its scan windows come from the deterministic recommender described below, not from an LLM. The post-chain agents are: a polymorph agent that ranks candidate Materials Project polymorphs of the identified chemistry by a weighted score combining spectral correlation, composition match, hull stability, and crystal-system agreement, emitting a confident verdict when the top-to-runner-up margin exceeds a fixed threshold and ambiguous otherwise; a size/strain agent that, when a multi-spectrum series is provided, estimates trends in crystallite size and microstrain via Williamson–Hall analysis (66). This agent operates on the measured patterns rather than on the refined profile: it fits Gaussian profiles to detected peaks and assigns them to phases by proximity to reference reflections computed from fixed lattice parameters, so the implementation is specific to the Fe–Mn series reported here, and generalizing it would mean taking those reference reflections from the Stage-2 result. The third agent is a recommender that emits operator-facing experiment suggestions (for example, advising a LaB 6 or Si instrument standard when none is present in the run metadata). All three are deterministic Python services with documented schemas; the LLM planner is the only LLM-driven component. The recommender also produces the coarse-to-fine re-scan window that drives the closed-loop demonstration. From a fast coarse survey it detects significant reflections with a robust relative- intensity floor at 12% of the strongest peak (a candidate peak must additionally span at least a minimum physical width, FWHM ≥ 0.04 ◦ , floored at two grid points, to reject single-point noise spikes), then proposes a fine-scan window bracketing the detected reflections with a 4 ◦ low-angle and 12 ◦ high-angle margin, subject to instrument bounds (2휃 ∈ [5, 120] ◦ , never starving the high angle below 70 ◦ ); a real reflection within 5 ◦ of a scan edge flags the survey as truncated, and the fine scan is recommended at a slower speed and finer step over that window. 43 Reliability layer (Guardian, Scribe, Phase-Guard) A reliability layer wraps the agent loop; the three names denote components introduced in this work and are not claimed as standardized terms. Guardian is a bounded observe–validate–replan loop over the pipeline’s reliability gates (space-group preservation, formation-energy plausibility with |Δ퐸 form | < 0.05 eV/atom passing and < 0.10 a warning, and the profile-quality 푅 푤푝 check): each step’s verdict yields one of accept, retry, replan, or abort, replans are hard-capped at three (after which the run aborts with reason replanlimitexceeded), and a replan may inject a known metastable-polymorph space group when exactly one phase fails the symmetry gate (for example Im ̄ 3m, no. 229, for a 훽-Ti candidate). Tool exceptions abort with a classified reason (for example an extreme-angle PyWPEM index-error is tagged as a narrow-range Stage-2 failure) rather than retrying indefinitely; in an expected-plan mode with replanning disabled the wrapper reproduces the chain’s own verdict without mutating any tool output. Scribe serializes every tool call, output verdict, and validator decision (with redacted file paths and blob-size limits) into an append-only per-job markdown and JSON provenance trace. Phase-Guard applies a closed-form detection-limit check to low-content phase claims: from the background standard deviation 휎 bg in a reflection-free window it sets a single-peak limit of detection 3휎 bg and limit of quantification 9휎 bg (relative to a peak-normalized reference pattern), with a multi-peak limit scaled by 1/ √ 푛 peaks , and flags any fitted phase fraction whose supporting reflection lies below this limit as fit-residual absorption rather than a detected phase. Use of large language models. Large language models enter the pipeline in two advisory roles only: the configuration-planning agent described above (DeepSeek-V3 via its cloud API as the primary backend, Qwen2.5-7B-Instruct on-premise as the fallback, and a deterministic rule planner when neither is available), and an optional literature-grounded deep-research report (large-language- model-generated, with web search) that summarizes a completed analysis for the operator. Every planner-proposed configuration is parsed and validated against a fixed JSON schema before use; the deep-research report is purely descriptive. No language-model output enters the physical phase- identification, refinement, and validation computations or any reported quantitative result, which are produced by the deterministic retrieval, refinement, phase-abundance, and interatomic-potential modules. No large language model is listed as an author of this work. 44 Compute. Learned components (the crystal-system classifier, the denoiser, and the interatomic- potential inference) run on NVIDIA A800 GPUs; retrieval, refinement, and phase-abundance analysis run on CPU. Consistent with our deliberate no-speed-claim stance, we report per-stage wall-clock only as an order-of-magnitude characterization from the recorded runtime benchmark: phase-identification retrieval is 0.7–3.4 s per sample, formation-energy inference is sub-second (CHGNet median ∼0.2 s), and the whole-pattern refinement dominates end-to-end time (median ∼175 s, up to tens of minutes for the hardest multi-phase patterns), so a single-phase analysis completes in a few minutes on a single host. 45 References and Notes 1. H. M. Rietveld, A Profile Refinement Method for Nuclear and Magnetic Structures. Journal of Applied Crystallography 2, 65–71 (1969), doi:10.1107/S0021889869006558. 2. A. Le Bail, Whole Powder Pattern Decomposition Methods and Applications: A Retrospection. Powder Diffraction 20 (4), 316–326 (2005), doi:10.1154/1.2135315. 3. B. H. Toby, R. B. Von Dreele, GSAS-I: The Genesis of a Modern Open-Source All Purpose Crystallography Software Package. Journal of Applied Crystallography 46 (2), 544–549 (2013), doi:10.1107/S0021889813003531. 4. J. Rodr ́ ıguez-Carvajal, Recent Advances in Magnetic Structure Determination by Neutron Powder Diffraction. Physica B: Condensed Matter 192 (1–2), 55–69 (1993), doi:10.1016/ 0921-4526(93)90108-I. 5. A. A. Coelho, TOPAS and TOPAS-Academic: An Optimization Program Integrating Computer Algebra and Crystallographic Objects Written in C++. Journal of Applied Crystallography 51 (1), 210–218 (2018), doi:10.1107/S1600576718000183. 6. C. F. Holder, R. E. Schaak, Tutorial on Powder X-ray Diffraction for Characterizing Nanoscale Materials. ACS Nano 13 (7), 7359–7365 (2019), doi:10.1021/acsnano.9b05157. 7. P. M. Maffettone, et al., Crystallography Companion Agent for High-Throughput Ma- terials Discovery. Nature Computational Science 1 (4), 290–297 (2021), doi:10.1038/ s43588-021-00059-2. 8. N. J. Szymanski, et al., Adaptively Driven X-ray Diffraction Guided by Machine Learning for Autonomous Phase Identification. npj Computational Materials 9 (1), 31 (2023), doi: 10.1038/s41524-023-00984-y. 9. J.-W. Lee, W. B. Park, J. H. Lee, S. P. Singh, K.-S. Sohn, A Deep-Learning Technique for Phase Identification in Multiphase Inorganic Compounds Using Synthetic XRD Powder Patterns. Nature Communications 11 (1), 86 (2020), doi:10.1038/s41467-019-13749-3. 46 10. J. Schuetzke, A. Benedix, R. Mikut, M. Reischl, Enhancing Deep-Learning Training for Phase Identification in Powder X-ray Diffractograms. IUCrJ 8 (3), 408–420 (2021), doi:10.1107/ S2052252521002402. 11. H. Dong, et al., A Deep Convolutional Neural Network for Real-Time Full Profile Analysis of Big Powder Diffraction Data. npj Computational Materials 7 (1), 74 (2021), doi:10.1038/ s41524-021-00542-4. 12. S. Zhang, et al., Crystallographic Phase Identifier of a Convolutional self-Attention Neural Network (CPICANN) on Powder Diffraction Patterns. IUCrJ 11 (4), 634–642 (2024), doi: 10.1107/S2052252524005323. 13. B. Cao, et al., XQueryer: An Intelligent Crystal Structure Identifier for Powder X-Ray Diffrac- tion. National Science Review 12 (12), nwaf421 (2025), doi:10.1093/nsr/nwaf421. 14. W. B. Park, et al., Classification of Crystal Structure Using a Convolutional Neural Network. IUCrJ 4 (4), 486–494 (2017), doi:10.1107/S205225251700714X. 15. L. C. O. Tiong, J. Kim, S. S. Han, D. Kim, Identification of Crystal Symmetry from Noisy Diffraction Patterns by a Shape Analysis and Deep Learning Approach. npj Computational Materials 6 (1), 196 (2020), doi:10.1038/s41524-020-00466-5. 16. F. Oviedo, et al., Fast and Interpretable Classification of Small X-ray Diffraction Datasets Using Data Augmentation and Deep Neural Networks. npj Computational Materials 5 (1), 60 (2019), doi:10.1038/s41524-019-0196-x. 17. R. Xing, et al., Interpretable X-ray Diffraction Spectra Analysis Using Confidence Evaluated Deep Learning Enhanced by Template Element Replacement. npj Computational Materials 11 (1), 281 (2025), doi:10.1038/s41524-025-01743-x. 18. Q. Lai, et al., End-to-End Crystal Structure Prediction from Powder X-Ray Diffraction. Ad- vanced Science 12 (8), 2410722 (2025), doi:10.1002/advs.202410722. 19. G. Guo, et al., Towards End-to-End Structure Determination from X-ray Diffraction Data Using Deep Learning. npj Computational Materials 10 (1), 209 (2024), doi:10.1038/ s41524-024-01401-8. 47 20. T. Xie, X. Fu, O.-E. Ganea, R. Barzilay, T. Jaakkola, Crystal Diffusion Variational Autoencoder for Periodic Material Generation, in International Conference on Learning Representations (ICLR) (2022), arXiv:2110.06197. 21. R. Jiao, et al., Crystal Structure Prediction by Joint Equivariant Diffusion, in Advances in Neural Information Processing Systems (NeurIPS), vol. 36 (2023), p. 17464–17497, doi: 10.52202/075280-0767, diffCSP; arXiv:2309.04475. 22. F. L. Johansen, et al., deCIFer: Crystal Structure Prediction from Powder Diffraction Data Using Autoregressive Language Models. arXiv preprint arXiv:2502.02189 (2025), accepted to Transactions on Machine Learning Research (2026); OpenReview LftFQ35l47. 23. L. M. Antunes, K. T. Butler, R. Grau-Crespo, Crystal Structure Generation with Au- toregressive Large Language Modeling. Nature Communications 15, 10570 (2024), doi: 10.1038/s41467-024-54639-7. 24. C. Zeni, et al., A Generative Model for Inorganic Materials Design. Nature 639 (8055), 624–632 (2025), published Nature version of MatterGen (preprint arXiv:2312.03687)., doi: 10.1038/s41586-025-08628-5. 25. Q. Li, et al., Powder Diffraction Crystal Structure Determination Using Generative Models. Nature Communications 16 (1), 7428 (2025), doi:10.1038/s41467-025-62708-8. 26. D. Hollarek, et al., opXRD: Open Experimental Powder X-ray Diffraction Database. Ad- vanced Intelligent Discovery 2 (2), e202500044 (2025), preprint arXiv:2503.05577 (2025); dataset on Zenodo (DOI 10.5281/zenodo.15298026); print issue April 2026., doi:10.1002/aidi. 202500044. 27. B. Lafuente, R. T. Downs, H. Yang, N. Stone, The Power of Databases: The RRUFF Project, in Highlights in Mineralogical Crystallography, T. Armbruster, R. M. Danisi, Eds. (Walter de Gruyter GmbH, Berlin, M ̈unchen, Boston), p. 1–30 (2015), doi:10.1515/9783110417104-003. 28. P. Lalor, H. Adams, A. Hagen, Sim-to-Real Supervised Domain Adaptation for Radioisotope Identification. Nuclear Instruments and Methods in Physics Research Section A: Acceler- 48 ators, Spectrometers, Detectors and Associated Equipment 1083, 171159 (2026), preprint arXiv:2412.07069 (2024)., doi:10.1016/j.nima.2025.171159. 29. N. Koblischke, J. Bovy, SpectraFM: Tuning into Stellar Foundation Models. arXiv preprint arXiv:2411.04750 (2024), accepted at the NeurIPS 2024 Workshop on Foundation Models for Science., doi:10.48550/arXiv.2411.04750. 30. Z. Ma, S. M. Shermer, O. Karakus, F. C. Langbein, The Sim-to-Real Gap in MRS Quantifi- cation: A Systematic Deep Learning Validation for GABA. arXiv preprint arXiv:2602.20289 (2026), doi:10.48550/arXiv.2602.20289. 31. J. Chen, P. Li, Y. Wang, P.-C. Ku, Q. Qu, Sim2Real in Reconstructive Spectroscopy: Deep Learning with Augmented Device-Informed Data Simulation. APL Machine Learning 2 (3), 036106 (2024), doi:10.1063/5.0209339. 32. T. Alkhalifah, H. Wang, O. Ovcharenko, MLReal: Bridging the Gap Between Training on Synthetic Data and Real Data Applications in Machine Learning. Artificial Intelligence in Geosciences 3, 101–114 (2022), doi:10.1016/j.aiig.2022.09.002. 33. Y. Fei, M. J. McDermott, C. L. Rom, S. Wang, G. Ceder, Dara: Automated Multiple-Hypothesis Phase Identification and Refinement from Powder X-Ray Diffraction. Chemistry of Materials 38 (3), 1364–1376 (2026), doi:10.1021/acs.chemmater.5c02820. 34. V. Vovk, A. Gammerman, G. Shafer, Algorithmic Learning in a Random World (Springer) (2005), doi:10.1007/b106715. 35. J. Lei, M. G’Sell, A. Rinaldo, R. J. Tibshirani, L. Wasserman, Distribution-Free Predictive Inference for Regression. Journal of the American Statistical Association 113 (523), 1094– 1111 (2018), doi:10.1080/01621459.2017.1307116. 36. Y. Romano, E. Patterson, E. Cand ` es, Conformalized Quantile Regression, in Advances in Neural Information Processing Systems (NeurIPS), vol. 32 (2019), arXiv:1905.03222. 37. A. N. Angelopoulos, S. Bates, Conformal Prediction: A Gentle Introduction. Foundations and Trends in Machine Learning 16 (4), 494–591 (2023), doi:10.1561/2200000101. 49 38. R. J. Tibshirani, R. Foygel Barber, E. J. Cand ` es, A. Ramdas, Conformal Prediction Un- der Covariate Shift, in Advances in Neural Information Processing Systems, vol. 32 (2019), arXiv:1904.06019. 39. I. Gibbs, E. Cand ` es, Adaptive Conformal Inference Under Distribution Shift, in Advances in Neural Information Processing Systems (NeurIPS), vol. 34 (2021), p. 1660–1672, arXiv:2106.00170. 40. R. F. Barber, E. J. Cand ` es, A. Ramdas, R. J. Tibshirani, Conformal Prediction Beyond Ex- changeability. The Annals of Statistics 51 (2), 816–845 (2023), doi:10.1214/23-AOS2276. 41. A. M. Bran, et al., Augmenting Large Language Models with Chemistry Tools. Nature Machine Intelligence 6 (5), 525–535 (2024), doi:10.1038/s42256-024-00832-8. 42. D. A. Boiko, R. MacKnight, B. Kline, G. Gomes, Autonomous Chemical Research with Large Language Models. Nature 624 (7992), 570–578 (2023), doi:10.1038/s41586-023-06792-0. 43. A. Ghafarollahi, M. J. Buehler, SciAgents: Automating Scientific Discovery Through Bioin- spired Multi-Agent Intelligent Graph Reasoning. Advanced Materials 37, 2413523 (2025), doi:10.1002/adma.202413523. 44. G. Chen, W. Yuan, F. You, Bridging electron microscopy and materials analysis with an autonomous agentic platform. Science Advances 12 (14), eaed0583 (2026), doi:10.1126/sciadv. aed0583. 45. B. Deng, et al., CHGNet as a Pretrained Universal Neural Network Potential for Charge- Informed Atomistic Modelling. Nature Machine Intelligence 5 (9), 1031–1041 (2023), doi: 10.1038/s42256-023-00716-3. 46. B. Cao, et al., AI-Driven Structure Refinement of X-Ray Diffraction. arXiv preprint arXiv:2602.16372 (2026), doi:10.48550/arXiv.2602.16372. 47. A. Jain, et al., Commentary: The Materials Project: A Materials Genome Approach to Accel- erating Materials Innovation. APL Materials 1 (1), 011002 (2013), doi:10.1063/1.4812323. 50 48. C. Chen, W. Ye, Y. Zuo, C. Zheng, S. P. Ong, Graph Networks as a Universal Machine Learning Framework for Molecules and Crystals. Chemistry of Materials 31 (9), 3564–3572 (2019), doi:10.1021/acs.chemmater.9b01294. 49. I. Batatia, et al., A Foundation Model for Atomistic Materials Chemistry. The Journal of Chemical Physics 163 (18), 184110 (2025), published journal version of MACE-MP-0 (preprint arXiv:2401.00096, 2024). Author list truncated after the first twelve; full list in the published record., doi:10.1063/5.0297006. 50. H. Gao, B. Cao, Y. Su, T.-Y. Zhang, Q. Liu, XDecomposer: Learning Prior-Free Set De- composition for Multiphase X-Ray Diffraction. arXiv preprint arXiv:2605.05866 (2026), doi: 10.48550/arXiv.2605.05866. 51. T. Simonnet, et al., Phase Quantification Using Deep Neural Network Processing of XRD Patterns. IUCrJ 11 (5), 859–870 (2024), doi:10.1107/S2052252524006766. 52. Y. Park, J. Kim, S. Hwang, S. Han, Scalable Parallel Algorithm for Graph Neural Network Interatomic Potentials in Molecular Dynamics Simulations. Journal of Chemical Theory and Computation 20 (11), 4857–4868 (2024), doi:10.1021/acs.jctc.4c00190. 53. H. Yang, et al., MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures (2024), arXiv:2405.04967. 54. J. Riebesell, et al., A Framework to Evaluate Machine Learning Crystal Stability Predic- tions. Nature Machine Intelligence 7 (6), 836–847 (2025), matbench Discovery; preprint arXiv:2308.14920 (2023)., doi:10.1038/s42256-025-01055-1. 55. D. L. Bish, S. A. Howard, Quantitative Phase Analysis Using the Rietveld Method. Journal of Applied Crystallography 21 (2), 86–91 (1988), doi:10.1107/S0021889887009415. 56. N. V. Y. Scarlett, I. C. Madsen, Quantification of Phases with Partial or No Known Crystal Structures. Powder Diffraction 21 (4), 278–284 (2006), doi:10.1154/1.2362855. 57. B. H. Toby, R Factors in Rietveld Analysis: How Good Is Good Enough? Powder Diffraction 21 (1), 67–70 (2006), doi:10.1154/1.2179804. 51 58. S. Graˇzulis, et al., Crystallography Open Database – An Open-Access Collection of Crys- tal Structures. Journal of Applied Crystallography 42 (4), 726–729 (2009), doi:10.1107/ S0021889809016690. 59. A. Merchant, et al., Scaling Deep Learning for Materials Discovery. Nature 624 (7990), 80–85 (2023), doi:10.1038/s41586-023-06735-9. 60. N. J. Szymanski, et al., An Autonomous Laboratory for the Accelerated Synthesis of Inorganic Materials. Nature 624 (7990), 86–91 (2023), doi:10.1038/s41586-023-06734-w. 61. S. P. Ong, et al., Python Materials Genomics (pymatgen): A robust, open-source python library for materials analysis. Computational Materials Science 68, 314–319 (2013), doi: 10.1016/j.commatsci.2012.10.028. 62. R. A. Young, Introduction to the Rietveld Method, in The Rietveld Method, R. A. Young, Ed., no. 5 in IUCr Monographs on Crystallography (Oxford University Press), p. 1–38 (1993), doi:10.1093/oso/9780198555773.003.0001. 63. S. Kirklin, et al., The Open Quantum Materials Database (OQMD): Assessing the Accuracy of DFT Formation Energies. npj Computational Materials 1 (1), 15010 (2015), doi:10.1038/ npjcompumats.2015.10. 64. C. Chen, S. P. Ong, A Universal Graph Deep Learning Interatomic Potential for the Periodic Ta- ble. Nature Computational Science 2 (11), 718–728 (2022), doi:10.1038/s43588-022-00349-3. 65. S. Batzner, et al., E(3)-Equivariant Graph Neural Networks for Data-Efficient and Ac- curate Interatomic Potentials. Nature Communications 13, 2453 (2022), doi:10.1038/ s41467-022-29939-5. 66. G. K. Williamson, W. H. Hall, X-Ray Line Broadening from Filed Aluminium and Wolfram. Acta Metallurgica 1 (1), 22–31 (1953), doi:10.1016/0001-6160(53)90006-6. 52 Acknowledgments Xtalyst, its multi-agent control layer, and the components introduced here are our own; the pipeline integrates external engines and public data resources, whose developers and maintainers we ac- knowledge (all cited in the text): the PyWPEM, CHGNet, MEGNet, MACE-MP-0, CPICANN, and XDecomposer packages, and the opXRD, RRUFF, and Materials Project databases. Funding: This work was supported by institutional funds from The Hong Kong University of Science and Technology (Guangzhou), the Shenzhen Loop Area Institute, and The Chinese University of Hong Kong. No specific external grant funding was received for this work. Author contributions: S.W. designed and implemented the pipeline, conducted the experiments, and drafted the manuscript. W.G. and B.F. contributed to the methodology and the analysis. X.S. and Z.W. contributed to data curation and validation. W.O. supervised the project and acquired funding. All authors reviewed and approved the final manuscript. Competing interests: The authors declare that they have no competing interests. Data and materials availability: The three open corpora analyzed in this study are publicly available from their original sources: the opXRD measured-pattern corpus (26), the RRUFF miner- alogical database (27), and the Simonnet et al. quantitative phase-analysis round-robin dataset (51). Reference crystal structures and density-functional-theory formation energies were obtained from the Materials Project (47). Two further groups of measured patterns analyzed here lie outside those corpora and are not deposited in a public repository: the laboratory Cu K훼 scans acquired for this work (the three coarse/fine closed-loop pairs of table S5 and the GH4169 superalloy scan of Fig. 5F), and the collaborator-supplied internal patterns, which include the Fe–Mn strain series and serve only as a cross-instrument check, excluded from every reported development and held-out number. The Ti-15Nb three-phase pattern of Fig. 5A is a worked case distributed with the PyWPEM package (46). Third-party components integrated by the pipeline (PyWPEM, CHGNet, MEGNet, MACE-MP-0, CPICANN, and XDecomposer) are obtained from their original public repositories as cited and are not redistributed here. 53 List of Supplementary Materials Figs. S1 to S9 Tables S1 to S7 54 Supplementary Materials This appendix contains: Figs. S1 to S9 Tables S1 to S7 Each is reproduced exactly as analyzed; panel labels and captions match the main-text narrative. 55 baseline Synthetic-trained retrieval no correction applied to the real patterns 0.146 median Pearson r vs measured (n = 306) rejected 1-D U-Net denoiser additive-noise hypothesis; no lift on real data no change 0.318 → 0.315 median Pearson r, matched cohort (n = 25) adopted Peak-position alignment structural drift, the largest correctable term 0.146 → 0.385 median Pearson r vs measured (n = 306) adopted Real-spectrum fine-tuning (XtalCS) crystal-system prior bridged onto real spectra 26.5% → 49.8% top-1 accuracy, dev set (n = 313) adopted Chemistry-aware tiered retrieval phase identification with peak-aligned reranking 0.609 median Pearson r, held-out (212 of 254 hits) adopted Whole-pattern lattice refinement six cell parameters per phase, under gates 127/146 · 127/127 converged · space group kept (n = 146) adopted Conformal recalibration on real spectra reliability recovery under sim-to-real shift nominal 0.90 81.1% → 95.3% coverage at α = 0.10, held-out (n = 127) 00.51 before → afterlevel as reported Each row is a different quantity on its own 0–1 range, measured on its own cohort: the rows do not add up. Figure S1: The simulation-to-real intervention ladder. Each row is one intervention, in the order established in the main text, with its own metric on that metric’s 0–1 range: an open marker joined to a filled one where a before→after pair was recorded; one diamond per reported level otherwise, so the whole-pattern refinement row carries two, its convergence rate and then its space-group-preservation rate; and an open marker ringing a filled one where before and after coincide at this scale, as on the denoiser row, which the figure marks “no change”. The rows are separate tracks rather than one axis: they are different quantities measured on different cohorts and share only the 0–1 bound, as the figure’s footer states. On the coverage row the dashed tick marks the nominal 0.90, where the ideal is the nominal level rather than 1, so 95.3% is conservative rather than better. The denoiser is the one rejected step—its paired median difference is −0.004 (95% CI [−0.018,+0.015], Wilcoxon 푝=0.49, matched 푛=25 cohort)—whereas correcting the structural peak-position drift is the largest single correctable component (per-sample median lift+0.152 on 푛=306). Every value printed is a recorded number also reported in the main text or table S4. 56 Synth-only pre-train Balanced retrain XtalCS (dev) XtalCS (held-out) 0 10 20 30 40 50 60 Crystal-system accuracy (%) 26.5% 25.9% 49.8% 49.4% Cub n=32 Tet n=18 Hex n=46 Trg n=0 Ort n=48 Mon n=82 Tcl n=25 0 20 40 60 80 Per-class accuracy (%) 81 56 37 N/A 50 39 60 agg 49.4% CubTetHexTrgOrtMonTcl Predicted Cub Tet Hex Trg Ort Mon Tcl True 26312 31014 7517683 22224162 486163216 32515 0255075100 Crystal-system accuracy (%) Orthorhombic 10 →43 Monoclinic 17 →43 Triclinic 71 →45 Tetragonal 17 →46 Hexagonal 22 →48 Cubic 49 →87 Aggregate 26 →50 class-balanced base real fine-tuned Simulated PXRD (forward model) Synthetic pre-train Class- balanced retrain Real-RRUFF fine-tune (730 spectra) CS prior sim → real 0.0 0.2 0.4 0.6 0.8 1.0 Row-normalized A Simulation-to-Real Training Bridge B Accuracy Across Training StagesC Per-Class Held-Out Accuracy D Confusion Matrix (Held-Out)E Real-Spectrum Fine-Tuning Lifts Each Class Figure S2: Real-RRUFF fine-tuning of the crystal-system classifier (XtalCS). (A) The simulation-to-real training bridge; the 730 real-RRUFF fine-tuning spectra are drawn from a 1,365-spectrum corpus. (B) Accuracy at three training stages on the same RRUFF development test set (푛=313); balancing lifts the worst-class floor from 0 of 24 Tetragonal to 9.8% min-per-class rather than the aggregate, and the frozen held-out value matches (푛=251, Wilson [43.3, 55.5] against[44.3, 55.3]). (C) Held-out per-class accuracy; Trigonal is N/A (unmeasured, not zero). Majority- class baseline 32.7%, balanced macro-average 53.8%; the classifier is a soft prior, not a hard filter. (D) Held-out confusion matrix; Monoclinic/Triclinic dominates the off-diagonal (16 of 50 Monoclinic errors, 16 of 21 spurious Triclinic predictions). (E) Per-class effect of real-spectrum fine-tuning, both sides scored on the same 푛=313 test set; every system improves except Triclinic, which trades head-class advantage for the weak classes, and the worst per-class value moves 9.8%→ 42.6%. The 146-spectrum checkpoint-selection split moves the aggregate 24.7%→ 63.0% and the worst class 8.3%→ 51.1%; these select the checkpoint and are not test accuracy. 57 1020304050607080 2θ (deg) truth Hexagonal → pred Hexagonal A Correct: attribution tracks the diagnostic peaks measured attribution (toward prediction) 1020304050607080 2θ (deg) truth Monoclinic → pred Triclinic B Confusion: Monoclinic read as Triclinic (the dominant error) measured attribution (toward prediction) Figure S3: Peak-level attribution for the crystal-system classifier (XtalCS). Integrated gradients (64 steps, zero- intensity baseline) computed for the production XtalCS checkpoint on real held-out RRUFF spectra, attributing the predicted-class logit back to each 2휃 channel and mapped onto the measured pattern (measured, dark; attribution toward the prediction, color). (A) A correct high-confidence case (Hexagonal): the attribution peaks coincide with the measured diffraction peaks (∼50% of the positive attribution mass falls on peak regions above 10% of maximum intensity, the remainder spread thinly over background), indicating the classifier keys on genuine reflections rather than background or instrument artifacts. (B) The dominant confusion (a Monoclinic pattern read as Triclinic; true Monoclinic supplies 16 of the 21 spurious Triclinic predictions on held-out, reported in the main-text Results and detailed in the held-out confusion matrix, fig. S2D): the positive attribution mass is enriched in the crowded low-angle region, where low-symmetry cells overlap: 13–30 ◦ carries 43% of it (against 30% for the correct case in A) while spanning 24% of the scanned range, and 69% of the mass still lands on peak regions, so the error is not background- driven. This is consistent with the two systems being hard to separate from peak positions alone. No spectrum in the held-out crystal-system label set carries a Trigonal label (fig. S2C), so no Trigonal attribution case is shown here; Trigonal structures do occur among the refined candidates (table S3). Attribution is reported as computed; no channel is selected or smoothed. 58 Tier 1 Hybrid Tier 2 Nearest Tier 3 CPICANN 0.0 0.2 0.4 0.6 Median peak-aligned Pearson 0.59 0.46 0.17 0.00.20.40.6 Before rerank (Pearson) 0.0 0.2 0.4 0.6 After rerank (Pearson) −0.40.00.4 Δ2θ ( ∘ ) 0.0 0.9 slide search improved (13/17) tie (4) y = x HybridNearest chemistry Blind spot −0.2 0.0 0.2 0.4 0.6 0.8 1.0 Top-1 peak-aligned Pearson n = 176n = 36 42/254 = 16.5% no candidate Synthetic calib. (n=97) Real-spec. recalib. (n=49) Nominal 0.90 0.0 0.2 0.4 0.6 0.8 1.0 Coverage at α = 0.10 0.546 0.918 0.900 0.70.80.91.0 Nominal coverage 0.6 0.7 0.8 0.9 1.0 Empirical coverage (n=127) 149 0 154 k = 45 half-width vs rank y = x Synthetic (held-out) α=0.30α=0.20α=0.10α=0.05 50 60 70 80 90 100 Empirical coverage (%) Nominal Synthetic (held-out) Measured pattern Tier 1 Chemsys-exact hybrid MP retrieval Tier 2 Jaccard nearest- chemistry Tier 3 CPICANN intensity-only + peak-align rerank no MP match still none Ranked MP candidate → Stage-2 lattice refinement on hit 0 20 40 60 80 100 Coverage (%) 75 50 100 A Three-Tier Retrieval Ladder B Per-Tier CoverageC Peak-Aligned RerankingD Held-Out by Tier E Synthetic vs Real CPF Coverage vs NominalG Per-α Coverage (n=127) Figure S4: The reliability layer: identification coverage and calibrated uncertainty. (A) The three-tier retrieval ladder: a pattern escalates only when its tier returns no MP-resolvable candidate, and any hit routes to Stage-2 refinement. (B) Per-tier coverage and median peak-aligned Pearson (opXRD development, 푛=510): 0.594 hybrid, 0.460 nearest- chemistry (coverage 64/127), 0.167 intensity-only. (C) Peak-aligned reranking of the intensity-only top-퐾 set: the푛=17 paired before/after correlations (13/17 improved, 4 tied, none worse; medianΔ=+ 0.103); inset, the±0.5 ◦ slide search. (D) Per-tier held-out correlations; 42/254 (16.5%) return no MP-resolvable candidate at any tier (broader-chemistry blind spot; this coverage is candidate return, not conformal coverage). (E) Synthetic vs real-spectrum conformal coverage at 훼=0.10 (development): synthetic-anchor calibration (±89.8 meV/atom) under-covers on a real-opXRD pool; real-spectrum recalibration recovers 0.918 (Clopper–Pearson[0.804, 0.977]). (F) Held-out empirical vs nominal coverage (푛=127) under synthetic-anchor calibration, below 푦=푥 at all 훼; inset, the split-conformal half-width as the ⌈(푛+1)(1−훼)⌉=45th order statistic of the 푛=49 real-spectrum |residual| set. (G) Held-out synthetic-anchor coverage at the four nominal levels it prints, with Wilson 95% intervals (81.1% at 훼=0.10). 59 H 0.59 He Li 0.59 Be 0.72 B 0.55 C 0.64 N 0.53 O 0.64 F 0.63 Ne Na 0.58 Mg 0.41 Al 0.50 Si 0.57 P 0.47 S 0.58 Cl 0.58 Ar K 0.47 Ca 0.65 ScTi 0.53 V 0.72 Cr 0.50 Mn 0.47 Fe 0.47 Co 0.48 Ni 0.45 Cu 0.56 Zn 0.67 GaGe 0.40 As 0.52 Se 0.81 BrKr RbSr 0.71 Y 0.84 Zr 0.63 Nb 0.84 MoTcRuRhPd Ag 0.37 Cd 0.60 In 0.60 Sn 0.76 Sb 0.60 Te 0.76 IXe CsBa 0.73 HfTa 0.88 W 0.92 ReOsIrPt 0.79 Au Hg 0.45 Tl 0.33 Pb 0.49 Bi 0.70 PoAtRn FrRaRfDb Sg BhHsMt LaCe 0.48 Pr 0.16 NdPmSmEuGdTbDyHoErTmYb 0.84 Lu AcThPaU 0.45 NpPu * ** * ** 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Median peak-aligned Pearson r Figure S5: Chemical-coverage map of held-out phase-identification retrieval. Each element is colored by the median peak-aligned Pearson correlation over the held-out materials whose reference formula contains it (the 푛=212 of 254 held-out samples returning a candidate; the same set behind the reported median of 0.609). The strongest cells are simple refractory transition-metal oxides, but they are thin: Ta 0.875 (푛=2), Y 0.845 (푛=3), Zr 0.631 (푛=3), and the two highest medians of all—W 0.924 and Nb 0.845—rest on one material each and are hatched for that reason. The strongest cells with 푛≥8 are Ba 0.735 (푛=9), Be 0.723 (푛=8) and V 0.722 (푛=8). The weak end is the better-sampled one: of every cell with 푛≥5 the two lowest are the light and disorder-prone chemistries Mg 0.409 (푛=34) and Ag 0.370 (푛=7). Elements seen in fewer than two held-out materials are hatched and elements absent from the held-out set are gray. Every cell is a median over recorded runs; no value is fitted or smoothed. 60 −4−3−2−10 DFT E form (eV/atom) −4 −3 −2 −1 0 CHGNet E form (eV/atom) MAE 43.5 meV/atom n = 307 0150300 |ΔE| (meV/atom) 50100200500 MAE vs DFT (meV/atom, log) MEGNet MACE-MP-0 SevenNet MatterSim Mean ens. MoE gate CHGNet 348 219 215 190 181 60 44 127/127 space groups preserved ■ pass 113 ■ warn 6 ■ fail 8 534 Held-out spectra 512 Tier-eligible 96% 254 Scored opXRD opXRD only 146 Refinement-eligible 57% 127 Converged refinements 87% 152025303540455055 2θ (deg) Intensity (norm.) Observed Fitted Difference A Property Predictor AccuracyB Model Selection (Single 307-set)C Structural Gates (E form ) D End-to-End Held-Out Funnel E Converged Fit: ZrSiO 4 Figure S6: End-to-end held-out evaluation. (A) Per-structure agreement between the production CHGNet formation- energy predictor and the Materials Project DFT reference across the 307-structure benchmark, colored by absolute error (mean absolute error 43.5 meV/atom); the points track the parity line over the full energy range. (B) Model-selection benchmark: mean absolute error against DFT for the strongest available machine-learning interatomic potentials, all scored on the same 307-structure set (CHGNet 43.5, mixture-of-experts gate 59.7, simple-mean ensemble 181.2, MatterSim 189.6, SevenNet-0 215.4, MACE-MP-0 218.7, MEGNet 348.0 meV/atom), on a logarithmic axis because the values span an order of magnitude. (C) Structure-validation gate outcomes on the 127 converged held-out refinements: the refined space group matches the input in all 127 cases (100%), and the formation-energy gate splits 113 pass, 6 marginal, 8 fail. (D) The frozen held-out evaluation funnel, from 534 held-out spectra through 512 tier-eligible (the 22 excluded are 13 whose scan range is clipped and 9 RRUFF entries with no crystal-system label) and the 254 scored opXRD spectra, to 146 refinement-eligible and the 127 converged refinements that complete the end-to-end assessment. The step from 512 to 254 is a change of population rather than an attrition: refinement eligibility is assessed on the opXRD portion alone, because the eligible RRUFF spectra are the crystal-system labels and the six Simonnet mixtures are the separate quantitative-phase-analysis read-out, so neither is a refinement target. The annotated rates are retention only for the three steps that are attrition. (E) A representative converged held-out refinement (ZrSiO 4 zircon, mp-4820, 푅 wp = 30.4) with observed, fitted, and difference profiles and the refined crystal structure rendered inline. All values are from recorded runs on real data. 61 20304050 2θ (deg) 0.0 0.2 0.4 0.6 0.8 1.0 Norm. intensity Sb 2 O 3 opXRD measurement 0.00.20.40.60.81.01.2 Polymorph score mp-2136·56 mp-1205345·14 mp-1047300·63 mp-1044869·1 mp-1398749·194 0.978 0.411 0.400 0.392 0.378 verdict: confident (margin 0.57) 051015 Applied strain (%) −2 −1 0 W-H strain ε (10 −3 , uncal.) BCC phase Spearman ρ=0.867 (p=0.0025) abs. reliability: low On gate-fail (R wp 67, CP tier C): 5 follow-ups priorityrecommended action 0.90 polymorph_check 0.56 composition_assay 0.54 refine_full_budget 0.50 longer_scan 0.40 standard_sample Production run, verbatim recommendation “Measure a LaB6/Si line-broadening standard on the same instrument — unlocks absolute crystallite size/strain instead of trend-only.” A Worked-Example SpectrumB Polymorph Disambiguation C Size/Strain Agent (Fe-Mn Series)D Recommender Agent Figure S7: Advisory agents on real spectra. (A) The measured pattern of the agent case-study sample (opXRD, Sb 2 O 3 ). (B) Polymorph-disambiguation agent output: 40 same-chemistry (O–Sb) candidates scored by spectral corre- lation, composition match, hull stability, and crystal-system agreement. The top-ranked entity, Sb 2 O 3 (mp-2136, space group 56), wins at margin 0.566 over the runner-up Sb 2 O 3 (mp-1205345, space group 14); the remaining displayed can- didates (mp-1047300, mp-1044869, mp-1398749) are SbO 2 polymorphs. The result overturns the recorded SbO 2 soft label (verdict confident). (C) Size/strain agent on the nine-spectrum Fe–Mn series: the per-level Williamson–Hall BCC strain against applied strain (Spearman 휌 = 0.87, 푝 = 0.0025), reported trend-only with a hard-coded low-absolute- reliability caveat. The plotted휀 is the uncalibrated Williamson–Hall slope—the series carries no instrument-broadening standard, so 휀 is negative at the lowest levels and only its ordering is quotable. The trend is martensite-only at every strain level: the strongest austenite reflection, FCC(111) at 43.6 ◦ , lies within 1 ◦ of BCC(110) at 44.6 ◦ , is discarded as an unresolvable doublet, and leaves austenite below the three points a Williamson–Hall line requires. The four martensite reflections per level give per-level 푅 2 of 0.00–0.21, so the ordering across levels is the result rather than any individual level’s fit. (D) Recommender agent: five prioritized next-experiment recommendations on the sample’s recorded failure profile, and the verbatim production recommendation emitted in the multi-phase showcase run. All four panels are unedited agent outputs from the recorded runs; the reliability layer that gates them is fig. S9. 62 A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) A Verdict receipt PASS PbSO4 anglesite mp-3472 Pnma (#62) · Orthorhombic Deterministic metrics Rp = 3.02 Rwp = 7.12 |ΔE| CHGNet vs DFT = 16.9 meV/atom predictor CHGNet v0.3.0 Structure-validation gates space group Pnma → Pnma |ΔE| within 10–50 meV band lattice < 0.2% vs literature B Assessment 1 Phase-ID confidence (high) Orthorhombic PbSO4 in space group Pnma (#62), matching the known mineral anglesite. Stage-1 candidate mp-3472 is the correct structural prototype; refined lattice a=8.486, b=5.403, c=6.964 Å agree with published anglesite. No competing polymorph under ambient conditions. 2 Refinement quality Rp = 3.02%, Rwp = 7.12%. For a well-crystallised inorganic powder on a laboratory diffractometer, Rwp below 10% is good; 7.1% is typical for a routine anglesite pattern. The Rp/Rwp gap points to mild preferred orientation along (001) and minor grinding strain. 3 Physical plausibility Refined cell volume ~319.5 ų is consistent with the density of PbSO4. |ΔEform| = 16.9 meV/atom (CHGNet vs DFT) falls in the 10–50 meV/atom band — acceptable agreement, typical for GGA-level functionals, with no structural anomaly and lattice constants within 1% of published baselines. C Next-step recommendations A Primary · low cost · half-day Re-run XRD at slower scan (0.5°/min, 0.01° step) over 20–60° 2θ to resolve subtle peak asymmetry; Rwp 7.1% could drop to ~5%. B Secondary · medium · 1–2 days SEM-EDX to confirm Pb:S:O stoichiometry and check trace PbO/PbS impurity (~1 at% limit), validating phase purity. C Investigative · higher cost If precipitation-synthesised, anneal 400°C/4 h in air to improve crystallinity; re-collect XRD to verify lattice unchanged. D Literature [1] Electronic structure of PbCr(1-x)SxO4 solid solution (RSC) [2] mp-623984: PbSO4 (Pnma, 62) — Materials Project [3] arXiv:1411.6528 [cond-mat.mtrl-sci] (2014) [4] Accuracy of DFT-computed formation energies (OSTI) [5] Crystal lattice structures: orthorhombic space groups [6] DFT of high-pressure polymorphs of organic crystals (IJMS) Figure S8: Deep-research advisory output on a worked example (PbSO 4 anglesite; production CHGNet pre- dictor). After the deterministic chain returns a verdict, an optional large-language-model step composes a literature- grounded advisory report. (A) Verdict receipt: the gated pass, the phase identity, the deterministic metrics (푅 푝 , 푅 푤푝 , and the CHGNet-versus-DFT prediction residual, 16.9 meV/atom here), and the three structure-validation gates. The energy gate acts on the refinement response, a different quantity (1.86 meV/atom here; Materials and Methods). (B) As- sessment in three parts: phase-identification confidence (here high), refinement quality, and a plausibility check against reference lattice parameters. (C) Tiered next-experiment recommendations, primary to investigative by cost, with effort estimates. (D) The ranked list of web-retrieved citations. Text is verbatim from one real run; as with the configuration planner, this report is LLM-generated and enters neither the phase-identification, refinement and validation computa- tions nor any reported quantitative result. 63 chain verdict gates: SG · ΔE form · R wp PASSreplan acceptfail ≤ 3 × abort cap hit replan action, e.g. inject known SG (Ti →229) 039ctrl Fe-Mn series: strain (%) and control 0 10 20 30 40 50 60 HCP abundance (% of intensity) detection limit (LOQ) observed HCP(101) K=3 fit HCP A Guardian Replan Loop B Phase-Guard Detection Limit Figure S9: The reliability layer. (A) The Guardian bounded-replan loop (schematic): the chain verdict is checked against the space-group, formation-energy, and 푅 푤푝 gates; a passing verdict is accepted, a failing one triggers a replan (for example injecting a known metastable-polymorph space group, Ti→229) that loops back through the gates up to a hard cap of three, after which the job aborts with a classified reason. (B) Phase-Guard detection-limit check on the ten recorded Fe–Mn patterns (the nine strain levels plus the unstrained control): a 퐾=3 refinement reports a 20.3–46.5% HCP intensity share (red), but the clean-window limit of quantification (3.4–11.4%, dashed) and the observed HCP(101) excess (green,≤ 0.14휎 above background) are both far below it, exposing the fitted HCP fraction as residual absorption rather than a detected phase. Both percentages are decomposed-intensity shares, not weight fractions (Materials and Methods). (A) is a schematic of the implemented loop; (B) is the recorded detection-limit analysis. 64 Table S1: Data scope. Held-out is frozen as a fixed manifest and read exactly three times across the entire study, each read fixed in advance: the headline held-out evaluation, one conformal recalibration check, and one quantitative-phase- analysis read-out on the held-out Simonnet mixtures. Per-module evaluation 푛 varies by module and is reported in full alongside every individual number used throughout the study. The three public corpora (opXRD, RRUFF, and the Simonnet 2024 quantitative-phase-analysis set) are cited in the main-text reference list. Corpus푛Role opXRD measured1,277 Phase ID, retrieval, conformal recalibration, refinement RRUFF measured1,365 Crystal-system fine-tuning and test Simonnet 2024 quantitative phase analysis 32 Mass-fraction real-weighed truth Collaborator (internal)22 Cross-instrument check (21 powder patterns; the remaining one is an X-ray reflectivity mea- surement, excluded as out of scope) Total ingest2,696 n/a Held-out (frozen)534 255 opXRD + 273 RRUFF + 6 Simonnet End-to-end held-out146 Refinement-eligible subset of held-out 65 Table S2: Evaluation design: each principal claim with the sample size and the data split on which it is measured. “Held-out” is the frozen held-out partition; “Dev” is a development pool; “MP benchmark” is a Materials Project DFT- energy benchmark; “QPA anchor” is the Simonnet 2024 real-weighed-truth set; “Fixture/case” are curated real-spectrum fixtures and case studies. Claims are grouped into headline pipeline performance (frozen held-out), diagnostic ablations, and control-layer/quantitative blocks; every subset is kept disjoint from the frozen held-out partition. This consolidates the evaluation design referenced from the main text. Claim푛Split Headline pipeline claims (frozen held-out) Crystal-system accuracy (XtalCS, 49.4%)251Held-out, leak-excl. (development 49.8%, 푛=313) Phase-identification retrieval median Pearson (0.609) 212 of 254Held-out(development0.594, 푛=383) Space-group preservation (100%)127 of 127Held-out, converged Refinement convergence (87.0%)127 of 146Held-out (refinement-eligible) Conformal: synthetic undercovers held-out (81.1% at 훼=0.10) 127Held-out (synthetic calib.) Diagnostic ablations (development / benchmark sets) Conformal: real recalib recovers (91.8% at훼=0.10) 49Development (leak-free calibration, disjoint from held-out) Gap is structural: denoiser null vs peak-align (+0.24) 25 / 306Dev (opXRD) CHGNet best of five MLIPs (43.5 meV/atom)307MP benchmark Discriminative > generative (0.459 vs 0.193 vs 0.141) 36Dev (broader-chemistry, matched) Control-layer and quantitative (fixtures / anchors) Phase-abundance NNLS (퐾=2 MAE 0.195)10QPA anchor (development); held- out read-out 5 of 6 Planner value (on/off; three backends)1 fixture+ 8 inputsFixture/case Advisory agents emit real payloads1 sample× 5 configsCase (real spectra) 66 Table S3: Representative converged held-out refinements, one per crystal system (lowest available 푅 푤푝 per system shown). The crystal system is assigned from the refined space-group number. Every listed refinement converged and preserves its input space group (the symmetry gate); the triclinic entry (푅 푤푝 = 96.8) converges and preserves symmetry but has a poor profile fit and would not clear a profile-quality threshold. Space-group preservation and profile quality are reported as separate gates (Materials and Methods). These are per-system minima and not typical values: over all 127 converged held-out refinements the median 푅 푤푝 is 116.7 and 2 of 127 clear the 푅 푤푝 ≤ 20 gate (main text). Sample (opXRD)MP id푅 푤푝 Space group opxrdhkust000495 mp-1227996 13.2 214 (Cubic) opxrdhkust001128 mp-1226487 13.9 44 (Orthorhombic) opxrdhkust000105 mp-1196550 23.1 12 (Monoclinic) opxrd hkust000730 mp-395327.8 167 (Trigonal) opxrdhkust000640 mp-482030.4 141 (Tetragonal) opxrdhkust000254 mp-56026542.0 179 (Hexagonal) opxrdhkust000700 mp-81868396.8 2 (Triclinic) 67 Table S4: Per-module development versus held-out performance. CS = crystal system; ID = phase identification; CP = conformal prediction; SG = space group. Confidence intervals are Wilson 95%, except the 푛=49 real-recalibration coverage row, which carries the Clopper–Pearson exact interval emitted by the conformal module. The pipeline reproduces development performance on the frozen held-out set where it is strong and exposes the residual gaps where it is not; these numbers are summarized in the main text. ModuleMetricDevelopmentHeld-outComment CS classifier (XtalCS)accuracy49.8% (푛=313)49.4%(푛=251,Wilson [43.3, 55.5]) clean gen. (leak- excl.) Phase ID, hybrid hitsmedian Pearson 0.594 (푛=383)0.643 (푛=176)matches develop- ment Phase ID, all hitsmedian Pearson n/a0.609 (푛=212/254)n/a Phase ID, no candidategap raten/a42/254 = 16.5%broader-chem blind Refinementconver- gence raten/a127/146 = 87.0%strong Refinementprofile residual median 푅 푤푝 n/a116.7(IQR74.0–138.4, 푛=127;2/127clearthe 푅 푤푝 ≤20 gate) lattice-only refinement; con- vergence,not profile-quality passes SG preservation on con- verged raten/a127/127 = 100%strong CP 훼=0.10 (synthetic)coverage54.6% (푛=97)81.1%(푛=127,Wilson [73.4, 87.0]) held-out: nominal excluded CP훼=0.10 (real recalib) coverage91.8%(푛=49, Clopper–Pearson [80.4, 97.7]) 95.3%(푛=127,Wilson [90.1, 97.8]) dev.near- nominal;held- out conservative (≥ nominal) Phase-abundance NNLS MAEvs weighed 0.195(퐾=2, 푛=10 dev.) 0.128 (5 of the 6 held-out Si- monnet mixtures admit a 퐾 - phase decomposition; all 퐾 ) small-푛; intensity share vs weighed truth 68 Table S5: Real-instrument closed loop across three samples and three crystal systems. For each sample the pipeline reads a fast coarse survey, the recommender emits a data-driven fine-scan window, the operator re-measures, and the chain re-analyzes the fine scan; the coarse and fine 푅 푤푝 are the whole-profile residuals on the two passes, and the profile-reliability gate is graded pass (푅 푤푝 ≤ 20), warn (20 to 50) and fail (> 50), with any grade other than pass making the overall verdict unreliable. All scans are laboratory Cu K훼 measurements outside the development corpora. The silicon standard is the load-bearing case: its coarse survey trips the profile gate (warn) and the recommended fine scan flips it to a gated pass, with identity, space group, and energy gate correct on both passes. The corundum and multi-metal samples map the operating envelope: the fine scan improves the fit (and, for the alloy, changes which minor phase is resolved: the counter accepts two phases on both passes, but the second one is assigned no weight at all on the coarse pass—0% of the peak-normalized intensity share—against 10.7% on the fine pass, which predates that basis and is therefore a per-mass-normalized least-squares fraction), but low counts / a database-absent solid-solution constituent keep the overall verdict below pass. SampleCrystal systemCoarse 푅 푤푝 Fine 푅 푤푝 Fine verdict Si standard (blinded) Cubic (퐹푑 ̄ 3푚)22.2 (warn)16.4PASS (gate flip) 훼-Al 2 O 3 (corundum) Trigonal (푅 ̄ 3푐)186.5154.4unreliable:low counts(ID/SG/Δ퐸 ok) Multi-metal alloyMetal (multi-phase) 24.5 (퐾=2) 19.8 (퐾=2) partial:DB-absent solid solution 69 Table S6: Component-wise ablation summary across the eight dimensions of the study. Every number is from a recorded run on real data with 푛 stated; rejected alternatives are quantified negatives, not omissions. Correlations are median peak-aligned Pearson; mean absolute errors (MAEs) are in meV/atom; FT = fine-tuning; UQ = uncertainty quantification. DimensionAblationKey result (푛)Production decision A. PlannerLLMplanneron/off; rule vs DeepSeek-V3 vs Qwen2.5-7B 12/13 fields bit-identical, +1.7 s (푛=1 fixture); schema-valid 8/8 all backends; edge-case deviations rule 0/3, Qwen 2/3, DeepSeek 3/3 (푛=8) LLMplannerwith schema validation and rule fallback B.Post-chain agents polymorphdevelopment benchmark; recommender coverage; size/strain trend polymorph crystal-system top-1 49.1% (54/110), confident 63.6% vs ambiguous 34.5%; recommender≥1 action on 93/93 development runs; size/strain Spearman 휌=0.87 (푛=9) post-hoc only; identifi- cation, refinement, and validation outputs never altered C.Phase- identification fallback tiers wired incrementally; peak-align on/off coverage 0.74 → 1.00 (푛=101); hybrid median 0.214 (75 hits in the 푛=101 slice) → 0.594 (383 hits in the 푛=510 cohort), and 0.214 → 0.603 on that identical 푛=101 slice re-measured with alignment on; peak-align medians 0.146 → 0.385 (푛=306, per-sample median lift +0.152); tier-2 standalone 0.460 (푛=64) full three-tier ladder, peak-align default D. Simulation- to-real bridge synthetic-only vs real-FT; denoiser vs alignment XtalCS 26.5% →49.8% (푛=313); matched cohort (푛=25): denoiser 0.318→ 0.315 (푝=0.49) vs peak-align 0.318 → 0.504 (푝=8×10 −7 ); peak-align 0.146 → 0.385 (푛=306) real-spectrumfine- tuning + alignment; denoiser rejected E. UQ methodsplit vs Mondrian; syn- thetic vs real calibration equal coverage at every 훼, Mondrian in- tervals wider at 훼≤0.20 and narrower at 훼=0.30 (푛=23); synthetic-on-real 0.546 (푛=97, every refined structure; 0.781 on the 푛=64 verdict-pass subset) vs real re- calibration 0.918 (푛=49 leak-free devel- opment) split CP with real- spectrum calibration 70 DimensionAblationKey result (푛)Production decision F. MLIPfive MLIPs on one shared benchmark; MoE gate vs single model CHGNet 43.5 vs MoE 59.7, simple- mean ensemble 181.2, MatterSim 189.6, SevenNet-0 215.4, MACE-MP-0 218.7, MEGNet 348.0 (all 푛=307); MACE small/medium/large 171.6/168.6/173.6 (푛=49) single CHGNet; MoE as opt-in difficulty sig- nal G.Training strategy synthetic-only vs class- balanced vs real-FT with min-per-class selection aggregate 26.5% → 25.9% → 49.8% (푛=313); min-per-class 0 → 9.8% → 42.6% on the same set, and 8.3% → 51.1% on FT-val (푛=146, checkpoint se- lection only) balanced retrain, then real-FT (= XtalCS) H. Disc. vs gen. retrievalvs MatterGen/deCIFer; NNLS vs CNN retrieval 0.459 > MatterGen 0.193 > deCIFer 0.141 (푛=36 matched); CNN be- low uniform baseline at 퐾 ≥ 3 discriminative+ physics path 71 Table S7: Operating envelope: located limits and the next step each implies. Each row pairs a quantified limitation (reported with its sample size beside the corresponding result in the main text) with the specific measurement or method change that would lift it. The limits are stated as boundaries of the current addressable scope, not as failures concealed; each is the direct source of a prioritized next step. ComponentLocated limit (with 푛)Next step it implies Crystal-systemprior (XtalCS) 49.4% top-1 on held-out (푛=251); Monoclinic and Hexagonal weakest; ceiling set by real-FT pool size, not the objective Enlarge the real RRUFF fine-tuning pool for the weak classes; used as a soft prior, not a hard filter Phase-identification re- trieval no MP-resolvable candidate at any tier for 42/254 (16.5%) held-out broader-chemistry minerals Extend the reference library / chem- istry frontier; a generative route may eventually be required (though generators currently underperform retrieval here) Full-profile refinementnon-convergence on ∼13% of complex minerals, almost all extreme-angle low-symmetry cells (ex- ternal engine breaks down); logged, never silently dropped; on the 127 that do converge the pro- file residual stays high (median 푅 푤푝 116.7; 2/127 clear the gate) because only the lattice is refined Robustify the refinement engine on low-symmetry cells and release coordinates, occupancies and dis- placement parameters so the resid- ual can come down; convergence never violates the space group Mass-fraction quantifi- cation validated on a small real-weighed-truth anchor set; reported small-sample throughout Acquire a larger quantitative phase- analysis anchor set with weighed ground truth Conformal uncertaintyreal-spectrum recovery rests on a modest devel- opment calibration set (푛=49, disjoint from held- out); intervals wide, smallest 훼 limited by set size; held-out coverage conservative (95.3% at훼=0.10) Enlarge the real calibration set to tighten intervals and make real- spectrum recalibration the uncondi- tional default Instrument loopsingle-instrument, human-in-the-loop demonstra- tion on real laboratory hardware Multi-instrument and autonomous- campaign extension (future work) 72