Paper deep dive
Towards single-shot coherent imaging via overlap-free ptychography
Oliver Hoidn, Albert Vong, Aashwin Mishra, Steven Henke, Matthew Seaberg
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 1:05:39 PM
Summary
The paper introduces PtychoPINN, a physics-constrained, self-supervised deep learning framework for coherent diffractive imaging (CDI) and ptychography. It extends previous work to enable overlap-free, single-shot reconstructions in Fresnel CDI geometry while accelerating conventional multi-shot ptychography. By coupling a differentiable forward model with a Poisson photon-counting likelihood, the framework treats real-space overlap as a tunable parameter. Results demonstrate that PtychoPINN achieves high reconstruction quality (SSIM 0.904 for single-shot) with significantly lower data requirements (1,024 images vs 16,384 for supervised baselines) and higher throughput (40x faster than least-squares methods), validated on experimental data from the Advanced Photon Source (APS) and Linac Coherent Light Source (LCLS).
Entities (10)
Relation Signals (9)
PtychoPINN → enables → overlap-free, single-shot reconstructions
confidence 95% · extend PtychoPINN ... to deliver overlap-free, single-shot reconstructions in a Fresnel coherent diffraction imaging (CDI) geometry
PtychoPINN → uses → Poisson photon-counting likelihood
confidence 95% · The framework couples a differentiable forward model of coherent scattering with a Poisson photon-counting likelihood
PtychoPINN → validatedon → Advanced Photon Source
confidence 95% · These results, validated on experimental data from the Advanced Photon Source
PtychoPINN → validatedon → Linac Coherent Light Source
confidence 95% · and the Linac Coherent Light Source
PtychoPINN → achieves → SSIM 0.904
confidence 90% · overlap-free single-shot reconstruction with an experimental probe reaches amplitude structural similarity (SSIM) 0.904
PtychoPINN → outperforms → Supervised Baseline
confidence 90% · PtychoPINN achieves higher SSIM with only 1,024 images ... against a data-saturated supervised model
PtychoPINN → outperforms → LSQ-ML
confidence 90% · Per-graphics processing unit (GPU) throughput is approximately 40x that of least-squares maximum-likelihood (LSQ-ML)
PtychoPINN → uses →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Ptychographic imaging at synchrotron and XFEL sources requires dense overlapping scans, limiting throughput and increasing dose. Extending coherent diffractive imaging to overlap-free operation on extended samples remains an open problem. Here, we extend PtychoPINN (O. Hoidn \emph{et al.}, \emph{Scientific Reports} \textbf{13}, 22789, 2023) to deliver \emph{overlap-free, single-shot} reconstructions in a Fresnel coherent diffraction imaging (CDI) geometry while also accelerating conventional multi-shot ptychography. The framework couples a differentiable forward model of coherent scattering with a Poisson photon-counting likelihood; real-space overlap enters as a tunable parameter via coordinate-based grouping rather than a hard requirement. On synthetic benchmarks, reconstructions remain accurate at low counts ($\sim\!10^4$ photons/frame), and overlap-free single-shot reconstruction with an experimental probe reaches amplitude structural similarity (SSIM) 0.904, compared with 0.968 for overlap-constrained reconstruction. Against a data-saturated supervised model with the same backbone (16,384 training images), PtychoPINN achieves higher SSIM with only 1,024 images and generalizes to unseen illumination profiles. Per-graphics processing unit (GPU) throughput is approximately $40\times$ that of least-squares maximum-likelihood (LSQ-ML) reconstruction at matched $128\times128$ resolution. These results, validated on experimental data from the Advanced Photon Source and the Linac Coherent Light Source, unify single-exposure Fresnel CDI and overlapped ptychography within one framework, supporting dose-efficient, high-throughput imaging at modern light sources.
Tags
Links
- Source: https://arxiv.org/abs/2602.21361v3
- Canonical: https://arxiv.org/abs/2602.21361v3
Trouble viewing inline? Open PDF directly →
Full Text
38,205 characters extracted from source content.
Expand or collapse full text
Towards single-shot coherent imaging via overlap-free ptychography O. Hoidn 1,* A. Vong 2 S. Henke 2 A. Mishra 1 and M. Seaberg 1 1SLAC National Accelerator Laboratory, Menlo Park, California 94025, USA 2Argonne National Laboratory, Lemont, Illinois 60439, USA *ohoidn@slac.stanford.edu †journal: opticajournal†articletype: Research Article abstract* Ptychographic imaging at synchrotron and XFEL sources requires dense overlapping scans, limiting throughput and increasing dose. Extending coherent diffractive imaging to overlap-free operation on extended samples remains an open problem. Here, we extend PtychoPINN (O. Hoidn et al., Scientific Reports 13, 22789, 2023) to deliver overlap-free, single-shot reconstructions in a Fresnel coherent diffraction imaging (CDI) geometry while also accelerating conventional multi-shot ptychography. The framework couples a differentiable forward model of coherent scattering with a Poisson photon-counting likelihood; real-space overlap enters as a tunable parameter via coordinate-based grouping rather than a hard requirement. On synthetic benchmarks, reconstructions remain accurate at low counts (∼104 \!10^4 photons/frame), and overlap-free single-shot reconstruction with an experimental probe reaches amplitude structural similarity (SSIM) 0.904, compared with 0.968 for overlap-constrained reconstruction. Against a data-saturated supervised model with the same backbone (16,384 training images), PtychoPINN achieves higher SSIM with only 1,024 images and generalizes to unseen illumination profiles. Per-graphics processing unit (GPU) throughput is approximately 40×40× that of least-squares maximum-likelihood (LSQ-ML) reconstruction at matched 128×128128× 128 resolution. These results, validated on experimental data from the Advanced Photon Source and the Linac Coherent Light Source, unify single-exposure Fresnel CDI and overlapped ptychography within one framework, supporting dose-efficient, high-throughput imaging at modern light sources. 1 Introduction Modern light sources, such as fourth-generation synchrotrons and X-ray Free-Electron Lasers (XFELs), generate coherent diffraction data far faster than images can be reconstructed [1]. This growing gap between acquisition and analysis precludes real-time feedback and on-the-fly experimental steering, both essential for maximizing the scientific output of these facilities. Ptychographic coherent diffraction imaging (CDI) is a cornerstone x-ray nanoscale imaging technique [2], but the computational reconstruction of real-space images from diffraction faces some limitations and tradeoffs. First, classical iterative algorithms like the Ptychographic Iterative Engine (PIE) require ∼ 60–70% scan overlap for robust convergence and process only ∼ 0.1–1 diffraction patterns per second on standard hardware [3, 4]; even graphics processing unit (GPU)-accelerated solvers struggle to keep pace with high-repetition-rate sources [5, 6]. Supervised machine learning (ML) approaches have been introduced to accelerate reconstruction by moving from iterative optimization-based procedures to single-shot inference using trained models. These approaches can accelerate inference but are often limited by poor generalization and the need for large labeled training sets generated by iterative solvers. [7, 6] Moreover, single-frame supervised methods cannot exploit overlap redundancy, failing outright when overlap constraints are required. In short, neither conventional methods nor direct supervised inversion unifies speed, resolution, and flexible handling of real-space constraints. Beyond supervised direct inversion techniques, recent developments in machine learning-based phase retrieval for ptychography include hybrid physics-learning approaches (e.g., deep-prior regularization and learned accelerators within iterative solvers) [8, 9], implicit neural representations including sinusoidal representation networks (SIREN)-style parameterizations [10], learned probe-position correction for large scan errors [11], and unrolled transformer-based ptychography networks [12]. Related unsupervised physics-aware inversion has also been demonstrated for 3D Bragg CDI (AutoPhaseNN) [13]. Within ptychography and extended-sample CDI, no single prior approach has jointly demonstrated reusable pre-trained inference, label-free training, and operation without strict overlap constraints. In this context, we address several limitations of prior approaches with a physics-constrained, self-supervised framework: a trainable inverse-mapping network is composed with a differentiable forward simulator of coherent scattering, and the full system is optimized end-to-end as an autoencoder using diffraction-domain losses (Poisson photon-counting likelihood [14, 15]). A key property of this formulation is that real-space redundancy is treated as a configurable parameter rather than a hard requirement. Specifically, the number of simultaneously reconstructed coherent scattering shots can be dialed to match the acquisition regime, including the overlap-free setting. In practice, when a curved or defocused probe provides sufficient phase diversity, the diffraction-domain likelihood alone can anchor reconstruction and spatial redundancy can be reduced to zero. This is the principle underlying Fresnel CDI [16, 17]. We use “single-shot” throughout in the limited sense of a single diffraction measurement with a structured probe (without lateral scanning, beam multiplexing [18, 19], or downstream modulators [20, 21]). Our previous work [22] demonstrated this physics-constrained approach on synthetic data; here, we extend it to realistic probes, arbitrary scan geometries, and single-shot reconstruction. We evaluate the model under both typical and non-ideal conditions, including low photon dose and large position jitter, and demonstrate good performance on experimental data from the Advanced Photon Source (APS) and the Linac Coherent Light Source (LCLS). Specifically, we demonstrate (i) self-supervised reconstruction of experimental data (APS, LCLS) at ∼6.1×103 6.1× 10^3 diffraction patterns/s; (i) overlap-free, single-shot reconstruction in Fresnel CDI geometry; (i) dose-efficient imaging via Poisson likelihood at ∼104 10^4 photons/frame; and (iv) an order-of-magnitude improvement in data efficiency over a supervised baseline with the same network architecture. In this study all reconstructions are performed in overlap-free single-shot mode, except in explicitly labeled overlap ablations. (a) Idealized — CDI (b) Idealized — Ptycho (c) Semi-synthetic — CDI (d) Semi-synthetic — Ptycho Figure 1: Reconstruction comparison across probe types and acquisition modes. Rows: idealized probe (Gaussian-smoothed disk, uniform phase) vs semi-synthetic (experimental probe, synthetic object). Columns: single-shot CDI vs overlapped ptychography. 2 Methods and Architecture 2.1 Formulation and Forward Model We learn an inverse map G:X→YG:X\!→\!Y from diffraction space to real space and optimize it by composing with a differentiable forward model F:Y→XF:Y\!→\!X. The overall autoencoder is F∘GF G, trained to match measured diffraction statistics without ground-truth images. Data model and notation. Each training sample comprises CgC_g diffraction amplitude images xkk=1Cg\x_k\_k=1^C_g acquired at probe coordinates r→kk=1Cg\ r_k\_k=1^C_g. The network G(x,r)G(x,r) outputs CgC_g complex object patches Okk=1Cg\O_k\_k=1^C_g on an N×N× N grid. In the expressions that follow, Δr→[⋅]T_ r[·] denotes real-space translation by Δr→ r; Pad[⋅]Pad[·] zero-pads to a canvas large enough to contain all translated patches; PadN/4[⋅]Pad_N/4[·] embeds a central N/2×N/2N/2× N/2 tile into an N×N× N grid; CropN[⋅]Crop_N[·] center-crops to N×N× N; 1 is an all-ones array of appropriate size; and ⊙ denotes the elementwise (Hadamard) product. Constraint map (FcF_c): translation-aware merging. To enforce overlap consistency, per-patch reconstructions are merged in a translation-aligned frame: Oregion(r→)=∑k=1Cg−r→k[Pad(Ok)]∑k=1Cg−r→k[Pad()]+ϵ,ϵ=10−3. O_region( r)\;=\; _k=1^C_g\;T_- r_k\! [Pad\! (O_k ) ] _k=1^C_g\;T_- r_k\! [Pad\! (1 ) ]+ε, ε=10^-3. (1) This "translational pooling” applies to arbitrary scan geometries. Coordinate-aware grouping. Training groups are formed locally by nearest-neighbor sampling. For each anchor r→i r_i, let K(r→i)N_K( r_i) be its K nearest distinct neighbors. A group i,jG_i,j draws Cg−1C_g-1 neighbors uniformly without replacement: i,j=r→i∪Si,j,Si,j⊂K(r→i),|Si,j|=Cg−1,G_i,j=\ r_i\∪ S_i,j, S_i,j _K( r_i),\;|S_i,j|=C_g-1, repeated nsamplesn_samples times per anchor. If duplicate neighbor sets are disallowed, the effective number of distinct groups per anchor is neff=min(nsamples,(KCg−1)),n_eff\;=\; \! (n_samples,\, KC_g-1 ), so the total number of training examples is Nscan×neffN_scan× n_eff, with the combinatorial upper bound Nscan(KCg−1)N_scan KC_g-1. Choosing nsamples>1n_samples>1 augments the dataset through combinatorial re-grouping while preserving local spatial consistency. Coordinates within each group are expressed in a stable local frame by re-centering to the group centroid r→global=1Cg∑k=1Cgr→k,r→krel=r→k−r→global. r_global= 1C_g _k=1^C_g r_k, r^\,rel_k= r_k- r_global. Diffraction map (FdF_d): coherent scattering. Given OregionO_region, the kkth translated object patch and exit wave are Ok′(r→) O _k( r) =CropN[r→krel(Oregion)], =Crop_N\! [T_ r^\,rel_k\! (O_region ) ], (2) Ψk _k =ℱOk′(r→)⋅P(r→), =F\! \O _k( r)· P( r) \, (3) where P(r→)P( r) is the (estimated) probe and ℱF is the 2D Fourier transform. Predicted detector-plane amplitudes include a global intensity scale eαloge _ that links normalized network outputs to physical photon counts: A^k=|Ψk|eαlog. A_k\;=\;| _k|\;e _ . (4) 2.2 Data Preprocessing A dataset consists of diffraction images from one or more objects measured with a fixed probe illumination P. After grouping images into samples of CgC_g diffraction patterns each (Section 2.1), we normalize the raw diffraction amplitudes to ensure favorable neural net activation magnitudes during training: xk=xk′⋅(N/2)2⟨∑i,j|xij′|2⟩, x_k\;=\;x _k· (N/2)^2 _i,j|x _ij|^2 , (5) where x′x denotes raw measurements and the average is over all images in the dataset. This choice ensures order-unity activations in the neural network: by Parseval’s theorem, unit-amplitude real-space objects produce diffraction power of approximately N2/4N^2/4, so this normalization maps experimental amplitude images to internal activations of order unity. Additionally, we introduce a trainable scalar αlog _ that converts between the dimensionless internal model activations and absolute per-pixel integrated amplitudes. The final, scaled, network input is xin=x⋅e−αlogx_in=x· e^- _ . 2.3 Neural Network Architecture The inverse map G follows an encoder–decoder design (as in [22]; see also [23] for a PyTorch implementation and novel training procedures), conditioned on xkk=1Cg\x_k\_k=1^C_g and r→krelk=1Cg\ r^\,rel_k\_k=1^C_g, and outputs complex patches Okk=1Cg\O_k\_k=1^C_g. To respect oversampling while avoiding probe truncation artifacts, the decoder allocates most capacity to the central, well-posed region and a lightweight continuation to the periphery. Handling extended probes. Convolutional neural network (CNN) architectures are limited to modest dimensions (N≤128N≤ 128) because convolutional receptive fields capture long-range interactions only inefficiently in this Fourier inversion setting, and we must furthermore restrict high-resolution reconstruction to the central N/2×N/2N/2× N/2 region to satisfy oversampling conditions [24]. Probes with extended tails force inefficient use of this limited number of pixels because the real-space area brightly illuminated by the probe is small compared to the total probe area that must be represented to avoid truncation artifacts from non-zero amplitude at the edge of the real-space grid. Consequently, given the modest magnitude of N, fully inscribing the probe—tails included—within the central N/2×N/2N/2× N/2 pixels may require too much binning. This causes a dilemma: one must choose between truncation artifacts (and possible lack of convergence due to the associated physical inconsistency) and violation of the diffraction-space oversampling condition for coherent imaging. We resolve this by reconstructing the object in high resolution in the central N/2×N/2N/2× N/2 region of the real-space grid and low resolution in the periphery. Presuming the absence of high spatial frequency components in the probe tail, extending the probe times object reconstruction into the periphery does not compromise well-posedness of the inverse problem. Concretely, we split the penultimate decoder layer’s channels into a majority set for the central region and the remaining 4 channels to coarsely reconstruct the periphery: Oamp O_amp =PadN/4(σA(Conv(HAcentral)))+σA(ConvUp(HAborder))⊙Mborder, =Pad_N/4\! ( _A(Conv(H^central_A)) )\;+\; _A(ConvUp(H^border_A)) M_border, (6) Ophase O_phase =PadN/4(πtanh(Conv(Hϕcentral)))+πtanh(ConvUp(Hϕborder))⊙Mborder, =Pad_N/4\! (π (Conv(H^central_φ)) )\;+\;π (ConvUp(H^border_φ)) M_border, (7) Ok O_k =Oamp⋅exp(iOphase), =O_amp· \! (i\,O_phase ), (8) where H⋅centralH^central_\·\ targets the central region, H⋅borderH^border_\·\ (the last 4 channels) produces a low-resolution continuation, and MborderM_border is a binary mask that isolates the boundary contributions to the outer region. This modification avoids artifacts from truncation of the exit wave and enables stable reconstruction with experimentally-realistic probes. 2.4 Training Objective and Optimization Poisson negative log-likelihood (NLL). The training procedure optimizes the inverse map G using a negative log-likelihood loss under Poisson statistics: ℒPoiss=−∑k,i,jlogfPoiss(Nkij;λkij)=∑k,i,j(λkij−Nkijlogλkij), _Poiss\;=\;- _k,i,j f_Poiss(N_kij; _kij)\;=\; _k,i,j ( _kij-N_kij\, _kij ), (9) where Nkij=|xkij′|2N_kij=|x _kij|^2 is the measured photon count and λkij=|A^kij|2 _kij=| A_kij|^2 is the predicted count. Since the network operates on normalized inputs (Eq. 5) for numerical stability, a scale parameter eαloge _ bridges normalized and physical units. When the mean photon flux NphotonsN_photons is known, we initialize: eαlog←2NphotonsN. e _ \;←\; 2 N_photonsN. (10) This ensures predicted intensities match measurement statistics. The parameter eαloge _ may be fixed or learned (see Table 3); learning it can absorb modest calibration errors. Amplitude loss for unknown counts. For datasets lacking absolute photon counts, we resort to mean absolute error (MAE) on normalized amplitudes: ℒMAE=∑k,i,j|xkij−A^kije−αlog|.L_MAE= _k,i,j |x_kij- A_kije^- _ |. In the results reported here we do not use any real-space loss; training is driven solely by the diffraction-domain losses (Poisson NLL or MAE). Implementation notes. All operators in FcF_c and FdF_d are differentiable and implemented with padding-aware translations and fast Fourier transform (FFT)-based diffraction. Batching is performed over groups i,jG_i,j; nearest-neighbor sampling with nsamples>1n_samples>1 provides dataset augmentation while preserving local spatial consistency. Default architectural and training hyperparameters are summarized in Table 3. 2.5 Supervised Baseline The supervised baseline uses the same encoder-decoder backbone and input representation as PtychoPINN (cf. PtychoNN [7]). It is trained with direct real-space supervision on paired diffraction/reference-reconstruction data, without enforcing the differentiable forward model in the training loss. Data splits, normalization, and scan-coordinate conditioning are matched to the PtychoPINN runs so the comparison isolates training paradigm rather than architecture. 2.6 Datasets and Evaluation Protocol We evaluate on an APS Velociprobe Siemens-star dataset, an LCLS X-ray Pump-Probe (XPP) test pattern dataset (hereafter, LCLS XPP dataset), a synthetic Siemens-star dataset simulated from APS Siemens-star reconstructions (ground truth for Table 1), and a synthetic line-pattern dataset of randomly oriented high-aspect-ratio features from [22] (used for the overlap ablation in Table 2). APS and LCLS experiments are run in single-shot mode (one diffraction frame per group), except where overlap ablations are explicitly labeled. For the Siemens-star experiments, we use a spatial holdout: the top half of the scan is used for training and the bottom half for testing. For out-of-distribution transfer, models trained on APS data are evaluated on LCLS data without retraining, with beamline-specific forward parameters (probe/geometry) substituted at inference. 3 Results We report results on the APS Velociprobe Siemens-star data, the LCLS XPP dataset, the synthetic Siemens-star dataset, and the synthetic line-pattern dataset; see Methods for dataset definitions and evaluation protocol. 3.1 Reconstruction Quality Figure 2 compares reconstructions on the APS Siemens-star data across two sampling budgets (512 and 8192 diffraction patterns), using a spatial holdout where the top half of the scan is used for training and the bottom half for testing. At 8192 patterns (Fig. 2b), the supervised baseline reconstructs training-region data well but degrades on held-out positions, whereas PtychoPINN maintains consistent quality across both. At 512 patterns, this train–test gap widens further for the supervised baseline. On the synthetic Siemens-star dataset (simulated from APS Siemens-star reconstructions), PtychoPINN also attains higher phase fidelity than the supervised baseline; see Table 1. Table 1: Reconstruction quality metrics at maximum training set size (16,384 images): peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM). Values shown are mean ± standard deviation across 5 trials. Best values per dataset are highlighted in green. PSNR (dB) SSIM Dataset Method Amplitude Phase Amplitude Phase synthetic Siemens-star Supervised baseline 84.83±0.2384.83± 0.23 68.62±0.0268.62± 0.02 0.930±0.0020.930± 0.002 0.912±0.0030.912± 0.003 PtychoPINN 85.53±0.0285.53± 0.02 70.54±0.0670.54± 0.06 0.955±0.0010.955± 0.001 0.962±0.0010.962± 0.001 (a) 512 diffraction patterns of the Siemens star test pattern. (b) 8192 diffraction patterns of the Siemens star test pattern. Figure 2: Comparison of reconstruction quality with different numbers of diffraction patterns. 3.2 Overlap-Free Reconstruction In overlap-free operation, we set the group size to a single diffraction frame (Cg=1C_g=1), removing overlap-based real-space consistency. Reconstruction then relies entirely on the diffraction likelihood and the known probe structure (defocused probe/Fresnel geometry). Figure 1 illustrates this single-frame mode compared with multi-position ptychography. Quantitative comparisons across overlap and probe-structuring variants on a synthetic line-pattern dataset are summarized in Table 2 (overlap-free Cg=1C_g=1 vs overlap Cg=4C_g=4). With an experimental probe, removing overlap reduces amplitude SSIM by 0.064 (0.968 to 0.904) and PSNR by 4.14 dB (73.03 to 68.89). On the accepted same-generated-split, support-constrained direct-stitch comparator, PtychoPINN reached SSIM 0.943263 and PSNR 70.738232 dB, whereas PyNX HIO/ER reached SSIM 0.005343 and PSNR 38.934707 dB. Table 2: Synthetic line-pattern amplitude reconstruction metrics on the test split. Ground truth is the simulated object (amplitude only; the object has constant zero phase). Case PSNR (dB) SSIM overlap-free (Cg=1C_g=1) + idealized probe 60.67 0.620 overlap-free (Cg=1C_g=1) + experimental probe 68.89 0.904 overlap (Cg=4C_g=4) + idealized probe 71.34 0.952 overlap (Cg=4C_g=4) + experimental probe 73.03 0.968 Same-split experimental-probe benchmark PtychoPINN (Cg=1C_g=1, same split) 70.74 0.943 PyNX HIO/ER (τ=0.05τ=0.05, known-probe support) 38.93 0.005 The first four rows report the historical overlap/probe ablation values. The same-split benchmark block uses the accepted generated-split comparator bundle; PyNX HIO/ER uses PyNX 2024.1 with a known-probe support prior, |P|≥0.05max|P||P|≥ 0.05 |P|, and direct support-anchored evaluation without oracle shift, twin-image, orientation, or phase-sign alignment. 3.3 Photon-Limited Performance Figure 3 compares resolution using the 50% Fourier ring correlation criterion (FRC50) as a function of photon dose for Poisson NLL versus MAE training objectives. At low dose (∼104 10^4 photons/frame), the Poisson NLL achieves comparable resolution to MAE at roughly 10×10× higher dose, corresponding to an order-of-magnitude improvement in dose efficiency. This advantage arises because the Poisson likelihood correctly models photon-counting noise, preserving sensitivity to the low-count, high-q components that carry fine spatial detail but are overwhelmed by bright-pixel residuals under an amplitude MAE. MAE objective Poisson NLL objective 10910^9 photons 10410^4 photons (a) Reconstruction comparison at 10910^9 and 10410^4 photons for MAE versus Poisson NLL objectives (left: representative diffraction patterns). (b) Resolution (FRC50) as a function of on-sample photon dose. Figure 3: Photon-limited performance for two self-supervised PtychoPINN variants trained with mean absolute error (MAE) and Poisson negative log likelihood (NLL) reconstruction penalties. 3.4 Data Efficiency Figure 4 illustrates the reconstruction quality (phase SSIM) as a function of dataset size. PtychoPINN maintains high fidelity (SSIM >0.85>0.85) from as few as 1024 diffraction patterns. In contrast, the supervised baseline degrades rapidly below 2048 samples. At small training set sizes, PtychoPINN achieves comparable quality using roughly an order of magnitude less training data. This suggests that the physical constraints enforced by the training procedure act as an effective prior for this inverse-imaging task. Figure 4: Structural similarity of PtychoPINN and the supervised baseline as a function of training set size. 3.5 Out-of-distribution Generalization Figure 5 compares an in-distribution LCLS control (train LCLS XPP, test LCLS XPP) with an out-of-distribution transfer setting (train APS, test LCLS XPP). Under APS→\!→\!LCLS shift, the supervised baseline largely collapses, whereas PtychoPINN preserves edge structure, albeit with visible phase artifacts. The reference column is an extended ptychographic iterative engine (ePIE) reconstruction of the LCLS data. Out-of-Distribution (Train A → Test B)In-Distribution (Train B → Test B)Reference (ePIE)PtychoPINNSupervised AAPS-2-ID; BLCLS XPP; AAPS-2-ID; BLCLS XPP; Figure 5: Comparison of methods for an in-distribution LCLS control (train LCLS XPP, test LCLS XPP) and out-of-distribution transfer (train APS, test LCLS XPP). The reference column shows an ePIE reconstruction of the LCLS data. 3.6 Computational Performance PtychoPINN processes approximately 6.1k diffraction patterns per second at 64×6464× 64 image resolution and 2.6k patterns per second at 128×128128× 128 in single-GPU inference measurements, excluding stitching/reassembly time. As a high-performance conventional baseline, we benchmarked LSQ-ML with pty-chi [25] at 128×128128× 128 (batch size 96) and measured 1.444 s per epoch over 10,304 frames. Assuming 100 iterations for convergence, this corresponds to 71.36 frames/s (10,304/(100×1.444)) (10,304/(100× 1.444) ). At matched 128×128128× 128 resolution, PtychoPINN therefore provides an approximately 40×40× throughput advantage over LSQ-ML. 4 Discussion Overlap-free reconstruction Table 2 reveals a clear interaction between probe structure and overlap. With the idealized (flat-phase) probe, removing overlap (Cg=1C_g=1) drops amplitude SSIM from 0.952 to 0.620; with the experimental (curved) probe, the same change yields 0.968 to 0.904. Probe curvature largely compensates for the loss of overlap-based redundancy, consistent with the expected role of structured phase diversity in Fresnel CDI. In the same-generated-split direct-stitch comparison, the PtychoPINN overlap-free experimental-probe reconstruction also outperformed the PyNX HIO/ER reference. This is a scoped comparison to one support-constrained CDI baseline with a known-probe support prior, not a broad claim about every classical CDI method. These trends indicate that overlap and probe diversity are partially substitutable sources of constraint, but their interaction warrants further study across a broader range of probe geometries. Making overlap a tunable parameter rather than a hard requirement has concrete implications for experimental design. Scans can use fewer positions, less overlap, or—in the Fresnel regime—no scanning at all, reducing acquisition time and total dose. The framework is also more tolerant of position jitter than overlap-dependent methods, since the reconstruction does not rely on precise inter-frame registration to enforce real-space consistency. Together, these properties are particularly relevant for dynamic or radiation-sensitive samples at high-rate sources, where overlap requirements, scan precision, and photon budget are simultaneous constraints. Diffraction-space supervision The forward model provides dense physical constraint per measurement: each diffraction pattern encodes the full exit-wave amplitude, and the Poisson NLL correctly weights every detector pixel—including the low-count, high-q pixels where fine spatial detail resides (Fig. 3). By contrast, real-space supervision constrains the network against a single reference reconstruction that already carries the ambiguities intrinsic to the inverse problem, such as global phase offsets. Because these nuisance parameters are not uniquely determined by the data, a supervised network can overfit to them, which likely explains both the supervised baseline’s train–test gap on held-out scan positions (Fig. 2) and its collapse under cross-facility transfer (Fig. 5). Data efficiency follows from the same mechanism: the forward-model constraint is far more informative per sample than a pixel-wise real-space loss, so the network converges with roughly an order of magnitude fewer training patterns (Fig. 4). Open problems The main methodological limitation is the fixed-probe assumption: the current formulation uses a pre-estimated probe and fixed scan coordinates during training, so it does not correct probe drift or position errors. A direct extension is to jointly refine probe and position parameters within the same self-supervised loop. The framework is modular: inverse backbone, differentiable forward model, and loss are separable components. This design should allow further speedups from mixed precision and architecture-level optimization without changing the architecture or training procedure. At higher resolution, the dominant scaling bottleneck is the CNN inverse backbone. Replacing it with a Fourier neural operator (FNO) backbone is a likely next step, because global spectral mixing is expected to scale better with image size N and improve high-resolution reconstruction quality. The same modular structure should also simplify adaptation to other coherent imaging geometries, including Bragg CDI. 5 Conclusions We presented an extended PtychoPINN framework that unifies overlap-free single-shot Fresnel coherent diffraction imaging and overlapped ptychography within a single self-supervised formulation. The method combines a differentiable coherent-scattering forward model with diffraction-domain training losses and supports arbitrary scan geometries through coordinate-aware grouping. Across APS and LCLS experiments, we measured approximately 6.1×1036.1× 10^3 diffraction patterns/s at 64×6464× 64 and 2.6×1032.6× 10^3 at 128×128128× 128 in single-GPU inference. In overlap ablations on synthetic line-pattern data with an experimental probe, overlap-free reconstruction reached amplitude SSIM 0.904 versus 0.968 for overlap-constrained reconstruction. In photon-limited regimes, Poisson NLL training improved dose efficiency by roughly an order of magnitude relative to MAE at comparable FRC50. Relative to a supervised baseline with the same backbone, the method maintained high quality with substantially fewer training samples. Future work will focus on joint probe/position refinement and higher-capacity inverse-mapping neural network backbones for large-image reconstructions. Appendix A: Key Configuration Parameters These parameters control critical aspects of the reconstruction process and should be tuned based on experimental conditions and computational constraints. Table 3: Model parameters, default code values, and settings used for the APS/LCLS experiments in this paper Parameter Default Description N 64 Patch dimension (pixels) C_g 1 Patterns per group (code default: 4) K 4 Nearest neighbors for scan-position grouping pad_object True Restrict object to N/2×N/2N/2× N/2 for oversampling probe.mask False Apply circular mask to probe gaussian_smoothing_sigma 0.0 Gaussian smoothing σ applied to probe illumination intensity_scale.trainable True Whether αlog _ is optimized during training n_filters_scale 2 Network width multiplier amp_activation sigmoid Amplitude decoder activation offset 4 Scan step size (pixels) d 3-5 Encoder depth (resolution-dependent) Table 4: Symbol definitions Symbol Type / Structure Description x′x Set of CgC_g real images Raw diffraction patterns for one sample x Set of CgC_g real images Normalized diffraction patterns for one sample r→k r_k 2D Position Vector Absolute scan position for the k-th image within a sample r→global r_global 2D Position Vector Centroid of a solution region (group of scans) r→krel r^\,rel_k 2D Offset Vector Relative scan offset within a solution region eαloge _ Scalar (trainable or fixed) Log-intensity scale parameter NphotonsN_photons Scalar Target average total photons per diffraction pattern P(r→)P( r) N×N× N complex array Effective probe function OkO_k N×N× N complex array k-th object patch decoded by the network G OregionO_region M×M× M complex array Merged object representation for a solution region Ok′O _k N×N× N complex array Object patch extracted from OregionO_region for forward model Ψk _k N×N× N complex array Predicted complex wavefield at the detector A^k A_k N×N× N real array Predicted final diffraction amplitude for one patch λijk _ijk Scalar Poisson rate parameter for a single pixel N: patch dimension, CgC_g: patches per group, M: merged region size Funding This work was supported by the U.S. Department of Energy, Laboratory Directed Research and Development program at SLAC National Accelerator Laboratory, under Contract No. DE-AC02-76SF00515. It was also supported by the U.S. Department of Energy (DOE) Office of Science-Basic Energy Sciences award Collaborative Machine Learning Platform for Scientific Discovery 2.0. Use of the Advanced Photon Source was supported by the U. S. Department of Energy, Office of Science, Office of Basic Energy Sciences, under Contract No. DE-AC02-06CH11357. Disclosures The authors declare no conflicts of interest. Data availability Data and code supporting this study are available from the corresponding author upon reasonable request. The PtychoPINN source code is available at https://github.com/hoidn/PtychoPINN. References [1] SLAC National Accelerator Laboratory, “LCLS-I-HE: Design and Performance,” https://lcls.slac.stanford.edu/lcls-i-he/design-and-performance (2023). Accessed: 2025-08-14. [2] M. Guizar-Sicairos and P. Thibault, “Ptychography: A solution to the phase problem,” Today 74, 42–48 (2021). [3] O. Bunk, M. Dierolf, S. Kynde, et al., “Influence of the overlap parameter on the convergence of the ptychographical iterative engine,” 108, 481–487 (2008). [4] A. M. Maiden and J. M. Rodenburg, “An improved ptychographical phase retrieval algorithm for diffractive imaging,” 109, 1256–1262 (2009). [5] S. Marchesini, H. Krishnan, B. J. Daurer, et al., “Sharp: a distributed gpu-based ptychographic solver,” of Applied Crystallography 49, 1245–1252 (2016). [6] A. V. Babu, T. Zhou, S. Kandel, et al., “Deep learning at the edge enables real-time streaming ptychographic imaging,” Communications 14, 7059 (2023). [7] M. J. Cherukara, T. Zhou, Y. S. G. Nashed, et al., “Ai-enabled high-resolution scanning coherent diffraction imaging,” Physics Letters 117, 044103 (2020). [8] C. A. Metzler, P. Schniter, A. Veeraraghavan, and R. G. Baraniuk, “prdeep: Robust phase retrieval with a flexible deep network,” in Proceedings of the 35th International Conference on Machine Learning, vol. 80 of Proceedings of Machine Learning Research (2018), p. 3501–3510. [9] A. R. C. McCray, S. M. Ribet, G. Varnavides, and C. Ophus, “Accelerating iterative ptychography with an integrated neural network,” of Microscopy 300, 180–190 (2025). [10] V. Sitzmann, J. N. P. Martel, A. W. Bergman, et al., “Implicit neural representations with periodic activation functions,” in Advances in Neural Information Processing Systems, vol. 33 (2020), p. 7462–7473. [11] M. Du, T. Zhou, J. Deng, et al., “Predicting ptychography probe positions using single-shot phase retrieval neural network,” Express 32, 36757–36780 (2024). [12] W. Gan, Q. Zhai, M. T. McCann, et al., “Ptychodv: Vision transformer-based deep unrolling network for ptychographic image reconstruction,” Open Journal of Signal Processing 5, 539–547 (2024). [13] Y. Yao, H. Chan, S. K. R. S. Sankaranarayanan, et al., “Autophasenn: unsupervised physics-aware deep learning of 3d nanoscale bragg coherent diffraction imaging,” Computational Materials 8, 124 (2022). [14] P. Thibault and M. Guizar-Sicairos, “Maximum-likelihood refinement for coherent diffractive imaging,” Journal of Physics 14, 063004 (2012). [15] J. P. Seifert, Z. Chen, M.-J. Yoon, et al., “Maximum-likelihood ptychography in the presence of poisson–gaussian noise,” Letters 48, 4897–4900 (2023). [16] G. J. Williams, H. M. Quiney, B. B. Dhal, et al., “Fresnel coherent diffractive imaging,” Review Letters 97, 025506 (2006). [17] M. Stockmar, P. Cloetens, I. Zanette, et al., “Near-field ptychography: phase retrieval for inline holography using a structured illumination,” Reports 3, 1927 (2013). [18] P. Sidorenko and O. Cohen, “Single-shot ptychography,” 3, 9–14 (2016). [19] K. Kharitonov, M. Mehrjoo, M. Ruiz-Lopez, et al., “Single-shot ptychography at a soft x-ray free-electron laser,” Reports 12, 14430 (2022). [20] F. Zhang, I. Peterson, J. Vila-Comamala, et al., “Phase retrieval by coherent modulation imaging,” Communications 7, 13367 (2016). [21] X. Dong, X. Pan, C. Liu, and J. Zhu, “Single shot multi-wavelength phase retrieval with coherent modulation imaging,” Letters 43, 1762–1765 (2018). [22] O. Hoidn, A. A. Mishra, and A. Mehta, “Physics constrained unsupervised deep learning for rapid, high resolution scanning coherent diffraction reconstruction,” Reports 13, 22789 (2023). [23] A. Vong, S. Henke, O. Hoidn, et al., “Towards generalizable deep ptychography neural networks,” abs/2509.25104, 1–1 (2025). [24] J. Miao, P. Charalambous, J. Kirz, and D. Sayre, “Extending the methodology of x-ray crystallography to allow imaging of micrometre-sized non-crystalline specimens,” 400, 342–344 (1999). [25] M. Du, H. Ruth, S. Henke, et al., “Pty-chi: A pytorch-based modern ptychographic data analysis package,” abs/2510.20929, 1–1 (2025).