Paper deep dive
Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response
Timothy C. Pearce, David J. T. Smith, Alec Dobney, Alessia Freddo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/17/2026, 5:23:23 AM
Summary
The paper introduces CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a machine-learning framework that uses meteorological data to predict fugitive hydrogen sulphide (H2S) and methane emissions from landfills. By identifying causal drivers (wind direction, wind speed, atmospheric pressure) across multiple timescales using information-theoretic methods, CAIRN enables proactive public health alerts. The system was validated against sensor data and independent community odour complaints, demonstrating its ability to estimate community impact in real-time.
Entities (9)
Relation Signals (7)
CAIRN → predicts → Hydrogen Sulphide
confidence 95% · Trained to predict gas measurements, CAIRN operates using only routine weather variables
CAIRN → validatesagainst → Community Odour Complaints
confidence 93% · tracks an independent record of community odour complaints at lag zero
Atmospheric Pressure → drives → Hydrogen Sulphide Exposure
confidence 92% · barometric pumping modulates the source flux as multi-hour soil–atmosphere pressure gradients drive gas out of the waste mass
Wind Speed → drives → Hydrogen Sulphide Exposure
confidence 92% · mechanical dilution at sub-hourly scales sets the instantaneous concentration through the inverse-wind-speed dependence
Wind Direction → drives → Hydrogen Sulphide Exposure
confidence 92% · wind direction... dominant at every scale up to 6 h... advection sets which receptor is exposed
CAIRN → generates → Tiered Alert
confidence 90% · Combining four such nowcasters produces a site-level, tiered alert aligned with WHO odour guidance
CAIRN → predicts → Methane
confidence 90% · framework transfers unchanged to a second monitoring station and to co-emitted methane
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. Here we show that the meteorological drivers of elevated hydrogen sulphide (HS) at a long-monitored European landfill, and the timescales over which they act, can be identified directly from routine monitoring data. We introduce CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a machine-learning framework whose internal memory is matched to these measured timescales: a fast component tracking hour-scale wind-borne transport and a slow component tracking multi-hour weather changes. Trained to predict gas measurements, CAIRN operates using only routine weather variables and the calendar, without hand-engineered features. Its behaviour is consistent with the identified transport mechanisms, and the framework transfers unchanged to a second monitoring station and to co-emitted methane. Combining four such nowcasters produces a site-level, tiered alert aligned with WHO odour guidance that closely reproduces the alert generated by a direct sensor network and tracks an independent record of community odour complaints. Weather-driven nowcasting can therefore estimate community impact as an emission episode unfolds, providing public-health authorities with a validated, graded trigger for intervention and enabling exposure to be reduced during events rather than after them.
Tags
Links
- Source: https://arxiv.org/abs/2608.14254v1
- Canonical: https://arxiv.org/abs/2608.14254v1
Trouble viewing inline? Open PDF directly →
Full Text
303,509 characters extracted from source content.
Expand or collapse full text
Fugitive emissions from waste sites increasingly expose communities to toxic and odorous gases, yet public-health responses remain largely retrospective, with episodes investigated only after residents have been exposed. Here we show that the meteorological drivers of elevated hydrogen sulphide (H2S) at a long-monitored European landfill, and the timescales over which they act, can be identified directly from routine monitoring data. We introduce CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a machine-learning framework whose internal memory is matched to these measured timescales: a fast component tracking hour-scale wind-borne transport and a slow component tracking multi-hour weather changes. Trained to predict gas measurements, CAIRN operates using only routine weather variables and the calendar, without hand-engineered features. Its behaviour is consistent with the identified transport mechanisms, and the framework transfers unchanged to a second monitoring station and to co-emitted methane. Combining four such nowcasters produces a site-level, tiered alert aligned with WHO odour guidance that closely reproduces the alert generated by a direct sensor network and tracks an independent record of community odour complaints. Weather-driven nowcasting can therefore estimate community impact as an emission episode unfolds, providing public-health authorities with a validated, graded trigger for intervention and enabling exposure to be reduced during events rather than after them. Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response Timothy C. Pearce Email: t.c.pearce@le.ac.uk Affiliation: Biomedical Engineering Research Group, School of Engineering, University of Leicester, Leicester, United Kingdom David J. T. Smith Email: toby.smith@ukhsa.gov.uk Affiliation: Environmental Hazards and Emergency Department, UKHSA, Nottingham, United Kingdom Alec Dobney Email: alec.dobney@ukhsa.gov.uk Affiliation: Environmental Hazards and Emergency Department, UKHSA, Nottingham, United Kingdom Alessia Freddo Email: alessia.freddo@ukhsa.gov.uk Affiliation: Environmental Hazards and Emergency Department, UKHSA, Nottingham, United Kingdom Introduction Communities living adjacent to regulated industrial sources of fugitive odorous and toxic gases (for example, landfills, sewage works, composting facilities, intensive livestock operations and biogas plants) bear a disproportionate burden of respiratory, neurological and mental-health morbidity that current response frameworks manage primarily through retrospective intervention, once impacts in surrounding populations are already documented (89; 100). The scale of the underlying challenge is growing: municipal solid waste alone is projected to nearly double globally by 2050 (60), and landfill methane emissions are substantially under-reported worldwide (102). Moreover, evidence indicates that odour-emitting facilities are more likely to be located in less privileged communities characterised by lower income, lower levels of education attainment, and higher proportions of minority populations (25; 72). This unequal distribution of exposure contributes to environmental health inequalities, as the odour-related impacts on health and wellbeing are compounded by a reduced capacity to influence decision-making. Consequently, although concerns surrounding odorous sites are often framed as a sustainability or nuisance issue, the problem is more fundamentally one of environmental justice, resulting in persistent and unequal health burdens. Among the gases these odorous industrial sites release, hydrogen sulphide (H2S) is a sentinel: generated when sulphate-reducing bacteria metabolise sulphate in decomposing waste under anaerobic conditions, it carries a characteristic “rotten-egg” odour detectable at concentrations far below those of toxicological concern (1). Its health burden is correspondingly graded. Acute, high-level exposure is overtly toxic, but the dominant impact on communities near regulated sites can be chronic – persistent malodour that degrades wellbeing through sleep disturbance, headache, irritation and psychological stress (93; 54). Odour is widely recognised as a significant environmental nuisance with measurable population-level impacts. Across Europe, environmental odours are recognised as a major source of public concern and are among the leading causes of environmental complaints to regulatory authorities, second only to noise (33). The prevalence of odour annoyance varies substantially across contexts, with relatively low levels reported in general population surveys and markedly higher levels (often exceeding 40–50%) in populations residing near agricultural, industrial, or waste-related emission sources (10; 104). The regulatory threshold for intervention in the UK is defined under Part I of the Environmental Protection Act 1990, where odour may constitute a statutory nuisance if it is “prejudicial to health or a nuisance.” Local authorities have a duty to investigate complaints and, where satisfied that a statutory nuisance exists, must serve an abatement notice requiring the responsible party to stop or mitigate emissions (31). However, determining nuisance relies on professional judgement and witnessed evidence, meaning intermittent odours, one-off events, or highly sensitive receptors may fall into “grey areas” where impacts are real but do not meet the statutory threshold. In these situations, the Environmental Permitting Regulations (EPR) provide an important complementary control. Environmental permits often include conditions to manage emissions such as odour, and breaches can be enforced by regulators independently of the statutory nuisance regime. For permitted facilities, even impacts below the statutory nuisance threshold may still breach permit conditions and require action, highlighting the role of EPR alongside statutory nuisance controls (24). Assessment of odour impacts in the UK is structured around the FIDOL framework, comprising Frequency, Intensity, Duration, Offensiveness and Location, which provides a practical basis for evaluating how odours affect receptors. FIDOL is embedded in the Institute of Air Quality Management (IAQM) guidance on odour assessment, widely applied in planning and regulatory decision-making. It recognises that odour impact is not solely concentration-based but depends on the interaction of these perceptual and contextual factors. This approach helps standardise subjective assessments, but it can still introduce variability and uncertainty, particularly where monitoring data are limited and reliance is placed on human sensory judgement. Operationally, statutory nuisance cases are managed in the UK through a multi‑stakeholder system to support joint risk assessment, communication, and response. Together, this layered system integrates environmental regulation, public health science, and emergency planning. Within this framework, air quality guideline levels for H2S, such as WHO 24-hour recommendations for health protection (150μgm−3150\, \,m^-3) and 30 minute odour annoyance guideline (7μgm−37\, \,m^-3), are used to assess exposures to the gas and whether odour episodes are likely to affect health, wellbeing or quality of life. The physical apparatus for that knowledge is dispersion modelling, which represents the atmosphere through steady-state, well-mixed approximations and requires site-specific stability and emission-flux calibrations Three gaps stand between this monitoring infrastructure and a community-protective intervention, and they are evidential rather than merely computational. First, the precise links are missing: we lack an empirical, data-derived account of how routine meteorological variables map onto fugitive emission and exposure, as opposed to a mechanistic account assumed in advance. Second, the drivers are multiscale: exposure is governed by processes acting from minutes to days that interact non-linearly, and the canonical timescales of those processes have not been recovered from observation and used to structure a predictive model. Third, and most consequential for public health, the validation is missing: no learnt classifier of fugitive emission exposure has been tested against an independent record of community wellbeing – the very outcome that fugitive-emission regulation exists to protect (34). The first two gaps share a common root in the physics of the boundary layer. At a monitoring receptor, H2S concentration is set by a hierarchy of meteorological controls operating across distinct timescales: advection sets which receptor lies downwind; mechanical dilution at sub-hourly scales sets the instantaneous concentration through an inverse dependence on wind speed (85); barometric pumping modulates the source flux as multi-hour soil–atmosphere pressure gradients drive gas out of the waste mass (106; 86; 35); and the diurnal growth and collapse of the mixing layer governs the depth of the volume into which emissions are diluted (84). Each process leaves a signature on a different temporal horizon, so the meteorology–exposure relationship is inherently multiscale and compounding. This is precisely what existing approaches fail to capture organically. Engineered-feature machine-learning pipelines, which now dominate air-quality classification (74; 90), hard-code their memory horizons – derivative windows, lag depths, stagnation intervals – in advance, tying performance to an analyst’s prior choices rather than to the physics itself. Structured state-space models (49) can in principle learn long-memory representations directly from raw meteorology, but their timescale priors are initialised data-agnostically; resolving this would let a model autonomously capture the compounding physics of weather over time without constant human recalibration, yet whether the canonical boundary-layer horizons can be recovered from data and used as architectural priors has not been demonstrated for fugitive emissions. The third gap is of a different kind. Uncovering the mathematical links between weather and emission is only half of an implementation problem: a model validated solely against the sensor signal it was trained on remains a technical exercise, blind to whether its outputs track the real-world outcome (i.e. human exposure and the distress it causes). Demonstrating real-world effectiveness, as distinct from in-sample efficacy, requires confronting model predictions with an external, independently generated record of community impact. Such records exist, in the form of daily odour complaints lodged by affected residents, but they have not been used to externally validate a learnt exposure model operating in an uncontrolled, real-world setting. Here we address all three gaps using post-calibration-adjustment monitoring (28) from a long-observed European landfill site associated with persistent odour-related impacts and community concern. For the first two, information-theoretic causal inference (91) recovers the missing links empirically, identifying wind direction, wind speed and atmospheric pressure as the causal core of receptor exposure and ranking their influence across ten temporal scales spanning fifteen minutes to ninety-six hours – the multiscale structure made explicit, from observation rather than assumption. We then introduce CAIRN (Causal-Anchored Inference for Receptor Nowcasting), a dual-pathway state-space classifier whose memory kernels are anchored on these data-derived scales. Supplied only with raw meteorology, CAIRN matches the high-exposure detection that engineered pipelines reach only through multi-stage hand-coded derivative, stagnation and recirculation features, and it transfers without modification across two receptor geometries and to co-emitted methane. For the third gap, we couple the model’s output directly to lived community experience. Fusing four per-channel classifiers under a Bayesian log-odds accumulator yields a four-tier alert system, aligned with WHO odour-annoyance guidance, that agrees with a deterministic raw-sensor ground truth and – critically – tracks an independent record of daily community odour complaints at lag zero (21). This external validation moves the contribution from a better algorithm to a framework whose internal exposure signal demonstrably reflects real community outcomes in a complex, uncontrolled environment. Trained on a single year of routine meteorology already collected continuously at a regulated UK landfill, the CAIRN architecture bridges multiscale physical drivers and lived human experience, establishing a validated evidential basis for real-time community protection at regulated emission sources, to be tested across the source diversity that such settings present. Results Causal hierarchy of meteorological drivers We analysed a year of post-calibration-adjustment 2024 monitoring data from the principal receptor site (MMF9, n=34,015n=34,015 valid 15-minute records of co-located H2S, CH4 and four meteorological channels) together with concurrent records from two further downwind sites (MMF1, MMF2; data and processing in Methods, “Data and study sites”). H2S monitoring at this site is exceptional rather than routine – it was instituted in response to the community-impact emergency described in the Introduction, not as a standard regulatory requirement – and we use the resulting record to derive a model whose inference-time inputs depend only on the four routinely collected meteorological channels. Two complementary analyses establish the meteorological control on receptor-level exposure: a multi-method environmental characterisation of the source–pathway–receptor system (Fig. 1), and an information-theoretic ranking of the candidate causal drivers across ten temporal scales (Fig. 2; Methods, “Causal analysis”). Conditional probability function (CPF) rose plots at the three co-located stations (Fig. 1a) converge on a common source area: at the principal receptor the exceedance probability peaks in the westerly sector (peak bearing 265∘265 , CPF ≈0.37≈ 0.37, peak/opposing-sector ratio 4.09×4.09×; circular–linear correlation r=0.149r=0.149, p<10−3p<10^-3; spike-conditional Fisher’s odds ratio 1.431.43, p=2.9×10−4p=2.9× 10^-4). The modest circular–linear r reflects sectoral rather than sinusoidal dependence on bearing; the peak/opposing-sector and spike-conditional odds ratios are the stronger directional evidence. This anisotropy is consistent with a continuous ground-level fugitive source upwind of the receptor and is the basis on which a receptor-modelling triangulation (5), whose intersection falls inside the permitted landfill boundary, narrows the emitting area to within a single upwind sector. Bivariate analysis of simultaneously measured H2S and CH4 (Fig. 1b) yields a strong positive correlation (r=0.830r=0.830, p<10−15p<10^-15; Spearman ρ=0.683ρ=0.683; power-law exponent b=1.95b=1.95) and a Jaccard spike co-occurrence index J=0.654J=0.654 (χ2=21,298χ^2=21,298, p<10−15p<10^-15). PM10 negative controls give r=−0.028r=-0.028, ruling out mechanical dust resuspension as a shared pathway, while a moderate NOx correlation (r=0.377r=0.377) is accounted for by joint nocturnal trapping (a 6-hour diurnal phase offset between the two species; Supplementary Note 2). Together these results identify the source as anaerobically generated landfill gas (1). The remaining four panels resolve the boundary-layer mechanisms through which meteorology modulates exposure. The temperature–H2S phase loop (Fig. 1c) traces a strong counter-clockwise hysteresis (normalised area A∗=0.369A^*=0.369; cooling/warming-limb slopes βcool=−2.83 _cool=-2.83, βwarm=−2.45 _warm=-2.45), demonstrating that the recent thermal history of the boundary layer, not the instantaneous temperature, carries information about the surface-layer mixing state. The mean diurnal profile (Fig. 1d) peaks at 02:00 LST (9.69μgm−39.69\, \,m^-3 mean; 51.6μgm−351.6\, \,m^-3 at P95) and reaches an early-afternoon minimum across 12:00–15:00 LST (1.31μgm−31.31\, \,m^-3), giving peak/trough ratios of 7.4×7.4× in the mean and 12.3×12.3× at the upper tail (Kruskal–Wallis H=338.5H=338.5, p=7×10−58p=7× 10^-58). Aligning every calendar day to astronomical sunrise (Fig. 1e) collapses 365 independent trajectories onto a single reproducible dispersal curve: a pre-sunrise baseline of 9.49μgm−39.49\, \,m^-3 falls to 1.30μgm−31.30\, \,m^-3 within four hours (7.3×7.3× reduction, one-sided Mann–Whitney p<10−15p<10^-15). Bivariate regression (Fig. 1f) confirms all three theoretically predicted negative associations (WS r=−0.122r=-0.122; TEMP r=−0.173r=-0.173; dT/dtdT/dt r=−0.143r=-0.143; all p<10−3p<10^-3); concentration-spike events cluster in the low-wind, low-temperature, negative-tendency regime corresponding to Pasquill class E–F stable boundary layers (85; 94). Pressure is excluded from this bivariate panel because its causally significant horizon is multi-hour and not resolvable in instantaneous 15-minute scatter; the pressure-tendency response is established below by the multiscale transfer-entropy analysis (Fig. 2a). Multiscale transfer entropy across seven candidate drivers and ten temporal scales (15 min to 96 h, the nine estimable of which are shown in Fig. 2a; Methods, “Causal analysis”) ranks these drivers quantitatively. Among the four directly measured meteorological variables, three form a robust core of directed information flow, significant after Benjamini–Hochberg correction across a single family of all 70 driver–scale tests: atmospheric pressure (ETE=0.005ETE=0.005–0.0580.058 nats, significant at every scale from 15 min to 12 h, its share of total positive ETE rising from 11.5%11.5\% at 15 min to 31.4%31.4\% at 12 h, the range over which barometric responses to passing synoptic systems become resolvable), wind direction (0.0210.021–0.0580.058 nats, dominant at every scale up to 6 h, 3838–50%50\% of total ETE, and not significant beyond), and wind speed (0.0110.011–0.0290.029 nats, significant at τ≤3τ≤ 3 h and again at 12 h). The scale-by-scale composition of that total is shown in Fig. 2b. The core is robust to the dependence assumption: under a correction valid under arbitrary dependence between the tests, wind direction retains all six scales and pressure all seven, and 18 of the 19 cells belonging to these three drivers survive, the exception being wind speed at 12 h (Fig. 2a, filled versus open symbols; Methods). Temperature enters through its rate of change rather than its level. The raw variable reaches significance at the 15 min scale only and under Benjamini–Hochberg alone, and is not interpreted; the six-hour temperature tendency is significant under both corrections (ETE=0.022ETE=0.022, q=0.009q=0.009 and q=0.042q=0.042 respectively) and is the only cell of the temperature family to survive the arbitrary-dependence correction. The tendency variables are deterministic transforms of drivers already in the family and are not counted as separate discoveries, but the contrast between the pair is informative: at six hours the rate of cooling carries more directed information than the temperature itself (0.0220.022 against 0.0180.018 nats). Six hours is the interval over which nocturnal radiative cooling establishes the surface inversion, and the result corroborates, by an independent route, the thermal-history dependence that the phase-loop hysteresis of Fig. 1c shows. The analysis does not localise where in the source–pathway–receptor chain a driver acts, so a thermal signature consistent with inversion trapping is equally consistent with temperature modulating the rate at which the waste mass generates gas; the two are not separable here. Directed information from every driver falls away beyond 12 h: no cell is significant at 24 or 48 h, both of which are estimable, and 96 h is not estimable (n=89n=89 blocks). That fall-away is a property of the test rather than an absence of coupling: the surrogate null widens approximately twentyfold as coarse-graining reduces the sample from 34,01534,015 blocks at 15 min to 180180 at 48 h, so a given effective transfer entropy is progressively harder to distinguish from chance at coarser scales. Raw ETE is consequently not comparable across scales. The coarse scales are reported as underpowered at this record length rather than as uncoupled. The picture that emerges is a five-phase exposure lifecycle. H2S and CH4 are co-emitted from the surface (Fig. 1b) under barometric modulation (Fig. 2a, raw pressure, whose share of directed information rises with aggregation scale and peaks at 12 h); advection sets which receptor is exposed (Fig. 1a; wind direction dominant at every scale up to 6 h and not significant beyond in Fig. 2a); mechanical dilution at sub-hourly scales sets the instantaneous concentration through the inverse-wind-speed dependence of the Gaussian plume (Fig. 1f); under stable nocturnal boundary layers a shallow inversion traps emissions over successive hours (Fig. 1c–d), recorded as the strong counter-clockwise thermal hysteresis and, independently, as the six-hour temperature tendency that is the one thermal cell to survive the arbitrary-dependence correction in Fig. 2a; and post-sunrise convective mixing erodes the inversion and entrains the accumulated reservoir into a much larger volume (Fig. 1e). The two analyses agree: the variables that information-theoretic causality identifies as significant at a given timescale are the variables that classical dispersion physics identifies as operative on that timescale. We note that the pressure signal will enter the engineered-feature classifier in the next subsection through its derivative dP/dtdP/dt and the stagnation index rather than as raw pressure: the engineered features already encode the synoptic information that pressure carries, a point that the ALE analysis confirms below. Engineered-feature benchmark and physical interpretability To establish whether the causal hierarchy can be operationalised by a model that uses it explicitly, and to provide a strong baseline against which a learnt representation can be tested, we trained four seasonal XGBoost classifiers on a feature set that encodes the causal hierarchy explicitly: causal Butterworth derivatives of T, P and WS, each fitted over a strictly backward window [t−6h,t][t-6\,h,\,t] and verified causal by truncation invariance (values computed on a truncated series are identical to those computed on the full series at every retained timestamp). The backward window preserves a 6 h span while placing its effective centre at t−3ht-3\,h rather than at t, so individual feature attributions shift in both directions and not only through the removal of lookahead; 2- and 6-hour stagnation indices; a recirculation index; cyclic time and direction encodings; and four annual Fourier harmonics (Methods, “XGBoost seasonal nowcaster”). Three ordinal classes follow World Health Organization guideline thresholds (Low <2<2, Medium 22–77, High ≥7μgm−3≥ 7\, \,m^-3). Aggregated across the calendar year 2024 (n=33,617n=33,617 evaluation samples), the classifier attains a sample-weighted accuracy of 79.9%79.9\%, with per-season accuracy of 76.476.4–83.6%83.6\% and per-season High-class recall of 0.5830.583–0.7490.749 (Fig. 3a; Supplementary Tables 2–4). Hyperopt-validation and seasonal-evaluation F1 agree within |ΔF1|≤0.009| _1|≤ 0.009 across all four seasons, confirming that the temporally isolated holdout prevents hyperparameter leakage (15). The strictly causal derivative formulation carries no lookahead, which brings hyperparameter selection and seasonal evaluation into closer agreement, which is independent evidence that the lookahead was real. A 3-hour majority-vote ensemble (Fig. 3d) suppresses classification noise; rolling accuracy is sustained across each seasonal window (Fig. 3c), and High-class predictions co-locate in time with daily community odour complaints (Fig. 3b)–the first indication that the classifier captures an operationally meaningful exposure signal. Accumulated local effects (ALE) computed on held-out validation data (Fig. 4a; Methods, “Ensemble classifier-tree seasonal nowcaster”) identify a physical hierarchy that overlaps with the one the causal analysis recovered from the data alone. The leading features are the diurnal cosine hour_coshour\_cos (class-mean importance I¯=0.138 I=0.138), the 2-hour stagnation index stagnation_2hstagnation\_2h (I¯=0.045 I=0.045) and wind speed (I¯=0.029 I=0.029); the temperature tendency dT/dtdT/dt, which led this ranking before the derivative correction, falls to twelfth (I¯=0.001 I=0.001). The dependence curves (Fig. 4b) are physically interpretable: wind speed enters with the negative slope of the inverse-wind-speed dilution law, and stagnation contributes monotonically positively. Notably, raw atmospheric pressure has zero ALE importance despite being a top-tier MSTE driver: the 2-hour stagnation index already encodes the synoptic information that pressure carries, and the raw value adds no further signal. The benchmark thus recovers the same meteorological drivers that the causal analysis identifies. It does so, however, chiefly through raw wind channels and a diurnal encoding rather than through the hand-coded derivative kernels: of the engineered features, only the stagnation index carries appreciable attribution. Detailed seasonal interpretation of each feature is given in Supplementary Note 3. Physics-anchored learnt nowcasting To test whether the same physics is discoverable from raw meteorology alone, we constructed CAIRN, a dual-pathway diagonal structured state space (S4D) model whose only structural prior is a channel-wise split into a fast lane (kernels initialised at a 1-hour timescale, matching the wind-driven advective band of Fig. 2a) and a slow lane (initialised at a 6-hour timescale, which lies within the band of scales over which pressure remains causally significant in Fig. 2a) (49). The input is restricted to four raw meteorological channels and minimal calendar encodings (Fig. 5a); no derivative, lag, rolling statistic, stagnation index or recirculation index is supplied (Methods, “CAIRN: structured state-space dual-pathway nowcaster”). Strict causality is preserved at every layer: the SSM kernel is one-sided, the convolution is explicitly truncated, and the readout is the final-timestep slice. We evaluate the model under a weekly expanding-window walk-forward protocol (13 retrains over the autumn 2024 quarter), the conservative choice because autumn was the worst season for the engineered benchmark (Macro F1=0.608_1=0.608) and the season in which long-memory drivers carry the largest predictive load. A six-condition initialisation ablation, holding optimiser, loss, features and walk-forward protocol fixed (Extended Data Table ED1), isolates the architectural contribution. Sweeping the slow-lane anchor across two orders of magnitude (τslow⋆∈6,48,147τ _slow∈\6,48,147\ h – respectively a scale lying within the causally significant pressure band identified by the MSTE analysis, an intermediate sub-synoptic value lying outside it, and the anchor of the preceding production configuration) identifies the MSTE-aligned anchor as the sole condition that separates from the field. Each condition was run three times from independent random seeds; Extended Data Table ED1 reports mean ± s.d., and the pooled within-condition standard deviation is 0.013 in F1-High (12 d.f.). The 6 h anchor attains 0.57±0.010.57± 0.01 and outperforms every alternative initialisation on a paired comparison across the 13 weekly folds (two-sided Wilcoxon signed-rank, Benjamini–Hochberg-corrected q≤0.048q≤ 0.048 in all five comparisons). The optimum lies inside the causally significant pressure band (τ≤12τ≤ 12 h; Fig. 2a), whereas both larger anchors lie outside it. Two limitations qualify this. The sweep did not test anchors between 6 and 48 h, so the optimum is identified within the tested set rather than located within the band, and the band is now known to extend to 12 h, which was not tested — and 12 h is the scale at which pressure carries its largest share of directed information (Fig. 2a), so the causal analysis itself identifies an untested anchor as its strongest candidate. Because 12 h lies both inside the band and within the learnt kernels’ reachable-memory regime, testing it is also the experiment that would separate the band-membership and kernel-reachability explanations, which the present sweep cannot; it is the natural next test. Independently, the learnt slow-lane kernels accumulate 99% of their sensitivity mass by 27–45 h and retain only 0.04–0.23% of peak sensitivity at 48 h, so the two larger anchors initialise memory the model cannot use. Band membership and kernel reachability are therefore confounded across the tested anchors: both explanations predict the observed ordering, and this ablation cannot separate them. Only the unbounded condition, in which the upper Δ clamp is lifted and the optimiser is free to migrate the slow lane, falls clearly below every bounded condition (0.48±0.010.48± 0.01; 10 of 13 folds against the best bounded alternative, Wilcoxon p=0.011p=0.011). Because that condition differs from the 147 h baseline only in the removal of the clamp, bounding the slow lane is consequential independently of where within the bound it is anchored. Within the bound, the conditions resolve into two tiers rather than a graded ordering. The 147 h, 48 h, overlap and random conditions are mutually indistinguishable (0.500.50–0.520.52; one-way ANOVA over three seeds each, F=1.75F=1.75, p=0.23p=0.23), so their apparent ranking within any single run is not interpretable. The 147 h condition alone moves from second place to fifth between one seed and the three-seed mean. Two caveats constrain the interpretation. First, random initialisation (0.52±0.010.52± 0.01) outperforms both the 48 h and 147 h anchors, so these data do not establish physics anchoring as a general mechanism; they support the narrower claim that the MSTE-derived 6 h anchor is superior to every alternative tested, including no anchor at all. Second, the comparison against random initialisation is the weakest of the five: the 6 h anchor wins on 9 of 13 folds, which reaches significance on the signed-rank test (q=0.048q=0.048) but not on a two-sided sign test (p=0.267p=0.267); under two protocols that remove the checkpoint selection — scoring every fold at the final retained epoch, and at a common fixed epoch — the same comparison improves to q≤0.013q≤ 0.013. The skill retained under random initialisation indicates that the channel-wise lane partition itself provides a structural benefit even without anchored sampling. A complementary lane-zeroing ablation (Supplementary Methods, “Walk-forward validation”) does not resolve the two pathways. Under the kernel-path mask of Eq. (S55) the paired per-fold difference between them is −0.002-0.002 (fast leads on 8 of 13 folds; sign test p=0.58p=0.58), the ordering reverses between aggregation rules, and the interaction term exceeds either marginal contribution while ranging from −0.468-0.468 to +0.474+0.474 across folds. The lanes therefore contribute jointly rather than separably, and the marginal values reported in Supplementary Methods, “Walk-forward validation” should be read as a decomposition that does not resolve rather than as an attribution of High-class skill to either lane. The trained model reproduces phenomenology it was not given directly. Hour-of-day and 10∘ wind-sector aggregates of the predicted High-class rate (Fig. 5b) reproduce the empirical nocturnal accumulation peak at 02:00 LST. The directional response tracks the empirical one closely (sector-profile correlation r=0.93r=0.93) and concentrates predicted High probability in the source sector, which carries 0.2910.291 against 0.0770.077 outside it – an enrichment of 3.8×3.8×. The predicted peak sector is nonetheless displaced approximately 34∘34 clockwise of the empirical peak, so the model recovers the source direction as a broad sector rather than resolving its precise bearing. Per-feature, per-lane sensitivities (Fig. 5c) show qualitatively distinct integration kernels: the fast lane decays smoothly from k=0k=0 over the first ∼ 5 h with wind variables retained longest, while the slow lane retains appreciable sensitivity across the whole displayed window. That extent reflects the input-sensitivity construction, not the learnt kernel timescales, which move only marginally during training (slow-lane mean 7.55→7.567.55→ 7.56 h under the MSTE-aligned configuration). Most consequentially, the marginal response of the predicted High probability to the unprovided pressure tendency dP/dtdP/dt (Fig. 5d) and temperature tendency dT/dtdT/dt (Fig. 5e) is non-monotonic and physically interpretable in both cases, with the empirical High-class frequency tracking the model curve within bootstrap confidence: from raw P and T alone, the model reproduces marginal responses that engineered pipelines obtain only by hand-coding explicit derivatives. The corresponding lane-resolved association is weak (ρs=+0.14 _s=+0.14) and is reported as consistency with the joint synoptic signature rather than as evidence that the signature is encoded in any particular lane. Walk-forward performance and community-impact validation Spike detection is the operational target of CAIRN: training uses focal cross-entropy with explicit High-class weighting to drive rare-event sensitivity (Methods, “CAIRN: structured state-space dual-pathway nowcaster”). Across the 13 weekly walk-forward retrains over Oct–Dec 2024 (n=8,040n=8,040 evaluation timesteps, of which 872 are High-class; Methods, “Walk-forward validation”), CAIRN leads the engineered XGBoost benchmark on High-class detection under a matched protocol in which both models select their stopping point on the same inner temporal probe: F1-High 0.5330.533 against 0.5010.501, a gap of +0.032+0.032 (+6.3%+6.3\% relative). The paired per-fold difference does not reach significance (two-sided Wilcoxon p=0.110p=0.110; CAIRN leads in 8 of 13 weeks), although the paired High-class decision over all 8,040 predictions does (McNemar, 344 against 292, p=0.043p=0.043, uncorrected for the eight metrics compared). Both figures pool the confusion counts over all 8,040 evaluation timesteps before computing the metric. The alternative convention, an unweighted mean of the thirteen weekly values, gives 0.5280.528 against 0.4650.465 and a larger gap of +0.063+0.063; the pooled rule is reported because it matches the operational question the alert system poses and because it is the more conservative of the two here. The difference between the rules is itself informative: High-class support ranges from 8 to 150 samples across the thirteen weeks, so an unweighted mean gives a week with eight positives the same influence as one with a hundred and fifty. Recall-High is 0.6180.618 against 0.5750.575 (+0.044+0.044) and precision-High 0.4680.468 against 0.4450.445 (+0.024+0.024), so the detection advantage holds on both components. The probabilistic terms do not follow it. Brier-High marginally favours CAIRN (0.0750.075 against 0.0770.077), but the ranked probability score and the multiclass log-loss both favour the benchmark, the latter substantially (0.1390.139 against 0.1020.102, and 0.9000.900 against 0.5370.537). CAIRN’s advantage is therefore in discrimination rather than calibration: it separates High-class events more sharply while returning a less well-calibrated posterior. This is the expected consequence of the protocol correction. The published configuration selected each weekly checkpoint on High-class F1 measured on the evaluation block itself, a criterion that optimises the decision boundary and not the probability estimate; removing it leaves a better-ranked and less well-calibrated model (full per-week breakdown in Extended Data Table ED2 and Supplementary Table 5). Both figures are lower than the values reported in the first version of this work, in which the two models were compared under different selection protocols: CAIRN selected its checkpoint on the evaluation block while the benchmark used a fixed iteration budget. Correcting both to a common inner probe reduces the gap from +0.086+0.086 to +0.032+0.032. The benchmark is essentially unaffected by the correction in its own right (selection-bias removal −0.004-0.004 at a 7-day block, +0.001+0.001 at 28 days), so the change is attributable to CAIRN’s protocol rather than to a change in the benchmark. Discrimination is higher than the benchmark (AUC 0.8980.898 vs 0.8770.877; Fig. 6b, left), although the per-fold paired difference does not reach significance (CAIRN leads in 8 of 13 weeks, sign test p=0.59p=0.59) and the claim is therefore not made across the operating range. CAIRN wins F1-High in 8 of 13 weeks. CAIRN detects 38 more High-class events than the benchmark while producing 14 fewer false alarms (612 against 626). Because both models decide by the argmax of the three-class posterior, neither has a free operating point, so this is not a threshold trade: the ordering of events by predicted probability is better, which the higher AUC confirms independently across 8,0408,040 timesteps. Rolling 24-hour metrics (Fig. 6a) show the gain is sustained rather than concentrated in single weeks, and the spike-detection track at the top of the panel shows that residual false negatives concentrate near the High-class threshold rather than on the largest plumes. Comparing CAIRN with the engineered benchmark varies two things at once, the input representation and the model, so a third arm isolates them. An identically configured tree ensemble was trained on CAIRN’s own sixteen raw inputs, none of which carries a lag, rolling window or derivative; it is therefore memoryless by construction, and it reaches F1-High 0.4630.463. Against that common floor, the twenty-four-feature engineered set, whose Butterworth derivatives and stagnation indices supply memory by hand, recovers +0.040+0.040 in pooled F1-High, while the state-space architecture, given only the raw sixteen, recovers +0.073+0.073. Both increments reach significance on the paired per-fold test that the endpoint comparison fails (two-sided Wilcoxon p=0.027p=0.027 and p=0.022p=0.022; each leads in 10 of 13 folds), and the learnt route recovers roughly 1.81.8 times as much. Hand-coded and learnt memory therefore recover the same quantity by different means, and only one of the two required an analyst to specify the physics in advance. The ordering holds on every High-class metric and on false alarms, which fall monotonically across the three arms (717, 626, 612); it does not extend to overall accuracy or multiclass log-loss, where the single-channel model remains last for the reason given above. Both comparisons are sensitive to fold composition, and unequally so. Two folds have an inner probe containing no High-class sample, and between them they carry 201 of the 872 High-class samples, including the largest single week. Excluding them reduces the endpoint difference from +0.032+0.032 to +0.012+0.012, while the architecture increment falls only from +0.073+0.073 to +0.042+0.042 and the engineering increment from +0.040+0.040 to +0.030+0.030; the ordering survives every jackknife (Supplementary Table S18). The decomposition against the memoryless floor is therefore the more stable of the two comparisons, which is why the architectural claim is made against that floor rather than against the engineered benchmark alone. A second model with an architecturally identical specification, retrained without modification on co-emitted CH4 under the same walk-forward protocol, attains a daily-mean predicted High probability that correlates with that of the H2S model at r=0.777r=0.777 (R2=0.604R^2=0.604, slope =1.00=1.00, n=88n=88 days, p=5.5×10−19p=5.5× 10^-19; Fig. 6d, right) and a Brier-High score (0.07240.0724) within 0.0030.003 of the H2S model (Supplementary Table S14). This comparison is cross-protocol: the H2S series is the matched inner-probe model and the CH4 series is not, since the cross-species arm was not re-trained (Methods, “Walk-forward validation”). The comparison is therefore between protocols as well as between species, and the difference is not attributable to the models alone. The near-unit slope and small intercept indicate that the same physics-anchored architecture transfers to a second co-emitted species despite differing absolute concentration scales, providing the structural pre-condition for a multi-species ensemble whose redundancy rests on two independent measurement chains, though not on independent labels: the CH4 class boundaries are derived from the H2S boundaries by ordinary least squares (H2S=5.55CH4−8.29H_2S=5.55\,CH_4-8.29), so the two targets are a linear transform of one another and the redundancy is instrumental rather than definitional for two co-emitted gases sharing a common source (Supplementary Table 6). The final test couples the model output to an independent record of community wellbeing. The lagged Pearson cross-correlogram between the day-mean predicted High probability and the daily community odour complaint count over the same Oct–Dec 2024 window (Fig. 6c) shows a single pronounced peak at lag 0 for both species: H2S r=0.585r=0.585 (R2=0.342R^2=0.342, p=2.2×10−9p=2.2× 10^-9, n=88n=88 days) and CH4 r=0.517r=0.517 (R2=0.267R^2=0.267, p=2.5×10−7p=2.5× 10^-7); the correlation falls to roughly half its peak within ±1± 1 day on either side of the lag-zero maximum (Fig. 6c). This single-channel correlation must be read against what naive meteorology alone achieves on the same days. A single predictor, the negated daily mean temperature, reaches r=0.563r=0.563 when fitted out of sample under blocked cross-validation, and across four independent runs of the identical protocol the model attains 0.581±0.0280.581± 0.028. The margin over that floor, +0.018+0.018, is within run-to-run variation (one-sided t=1.28t=1.28, p=0.145p=0.145). Correlation with complaints at a single receptor is therefore not, on its own, evidence that the model captures a community-impact signal beyond ambient seasonal meteorology. Two consequences follow. First, the lag structure is informative even where the magnitude is not: the absence of a dominant positive lag indicates that the strictly causal architecture is behaving as a nowcaster of present community exposure, not a forecaster of future complaints. Second, the signal strengthens substantially under multi-channel fusion: it is the fused result reported below, rather than the single-receptor correlation reported here, which carries the community-impact claim. Multi-site network tier alerting and public-health framing A natural deployment of CAIRN is a multi-receptor public-health alert system in which independently trained per-channel classifiers fuse into a single tiered indicator at the timescale of community impact. This deployment tests three properties that go beyond the single-receptor analyses above: (i) whether the CAIRN architecture transfers without retuning to a second receptor geometry; (i) whether classification skill is maintained when each constituent classifier is driven by the meteorology of its own receptor rather than by the principal MMF9 record; and (i) whether the fused output maps onto an operationally interpretable tiered-alert framework matched to the public-health response landscape of Supplementary Note 1. MMF1 is excluded from this analysis: it was relocated to Maria’s Way mid-record and the post-move record is too short to support an independent walk-forward retrain. Four constituent CAIRN nowcasters – H2S and CH4 at MMF9 and MMF2, each retrained on the co-located meteorology of its own receptor under the protocol of Sec. Walk-forward validation and metrics – feed a Bayesian log-odds accumulator (Sec. Multi-site Bayesian alert-tier classifier, Eq. (3)). The predicted tier is constructed purely from meteorology: no raw target-species concentration enters the predicted pipeline at inference time. Per-channel debouncing matches the WHO 30-minute averaging convention, and the channel-state likelihood ratios are discounted below each classifier’s bare precision-prevalence ratio to absorb the inter-channel dependence implicated by the shared advection bearing and synoptic barometric modulation of Figs. 1a and 2a (Supplementary Note 6, “Bayesian fusion of multi-site, multi-species evidence”; Supplementary Table S3). The continuous posterior is mapped to four sequential ordinal tiers at cut-points (θ1,θ2,θ3)=(0.15,0.50,0.92)( _1, _2, _3)=(0.15,0.50,0.92) aligned with the Supplementary Note 1 public-health response: Tier 1 at the FIDOL nuisance-assessment regime; Tier 2 at multi-channel corroboration; Tier 3 at the near-maximum attainable posterior under coordinated four-channel meteorological evidence (analytical ceiling Pmax≈0.95P_ ≈ 0.95). Across 13 weekly walk-forward folds spanning Jan–Mar 2025 (n=8,536n=8,536 synchronous 15-min timesteps; Fig. 7a–c), the fused predicted tier agrees with the deterministic raw-concentration ground-truth tier at quadratic-weighted Cohen’s kappa κw=0.709 _w=0.709 (block-bootstrap 95% CI 0.557,0.8520.557,0.852), strict accuracy 79.4%79.4\% (Extended Data Table ED3). Two reference points are required to read those figures. Because the tier distribution is dominated by Tier 0 (normal background), a majority-class predictor attains strict accuracy 80.0%80.0\%; the accuracy figure alone therefore does not demonstrate skill, and the agreement claim rests on κw _w, which is chance-corrected. The unweighted κ is 0.4620.462: the quadratic weighting materially improves the reported agreement because most errors are off-by-one tier, which is the operationally tolerable failure mode but should be stated explicitly. Two points concern selection, at different levels. The hysteresis counts, likelihood ratios and tier cut-points that parameterise the fusion were selected by grid search on nine of the thirteen weeks (Supplementary Note 6). Refitting these parameters leave-one-week-out, so that every timestep is scored out of sample with respect to them, gives κw=0.703 _w=0.703 against 0.7090.709 and a complaint correlation of 0.7210.721 against 0.7290.729: the fusion-parameter optimism is 0.0060.006 in agreement and 0.0090.009 in correlation (Extended Data Table ED3, Supplementary Table S2). The insensitivity is structural rather than fortunate: the accumulator consumes latched binary channel states after debouncing, so the fused output has coarse leverage over the parameters being tuned. The four constituent channels separately retain the original checkpoint-selection rule and were not re-trained under the inner temporal probe adopted for the H2S nowcaster of Fig. 6a,b, so each carries the single-channel optimism measured in Methods (0.0310.031 in F1-High); the leave-one-week-out result above indicates that the latching stage attenuates parameter-level differences, though the checkpoint-level passthrough is not separately measured. Performance is concentrated at the extremes: Tier 0 and Tier 3 are recovered (F1=0.91_1=0.91 and 0.570.57, the latter roughly half of ground-truth Tier 3 support from meteorology alone) while the intermediate tiers are not (F1≈0.35_1≈ 0.35; Extended Data Table ED3). Tier 2 is the operationally consequential intermediate state, and recalibration of the constituent nowcasters would be required before this layer of the alert system is deployable for statutory-nuisance confirmation (Supplementary Note 6, Discussion). Three properties of the deployed configuration follow directly from its arithmetic and bound how it can be used. Predicted Tier 3 requires all four channels active, because the largest three-channel posterior is 0.9030.903 against θ3=0.92 _3=0.92; a single channel missing therefore caps the posterior at that value and makes Tier 3 unreachable, whereas the ground-truth tier requires only three latched channels, so the two scales are not aligned at the top. Both H2S receptors active with both CH4 channels quiet gives P=0.40P=0.40 and reaches Tier 1 only. These follow from the likelihood-ratio set and prior of Supplementary Table S1 and account for the intermediate-tier recalls above rather than leaving them unexplained. A same-architecture vote-count baseline returns κwvote=0.639 _w^vote=0.639; the Bayesian fusion delivers Δκw=+0.07 _w=+0.07 over baseline, and under a paired block bootstrap, which resamples the same blocks for both estimators so that the shared fold-to-fold variation cancels, the improvement is significant (95% CI 0.025,0.1090.025,0.109; p=0.003p=0.003). An unpaired bootstrap, which resamples the two κ values independently, inflates the interval by approximately a factor of three and is not the appropriate test for two estimators evaluated on identical timesteps. The fusion is retained on both statistical and operational grounds: beyond the agreement gain it produces a calibrated continuous PtP_t that the vote-count baseline cannot deliver. The predicted tier nonetheless coincides with the vote count minus one on 94.2%94.2\% of timesteps, so the fusion should be understood as a calibrated refinement of channel counting rather than as a categorically different decision rule. External validation against an independent record of daily community odour complaints (Fig. 7d) – not used in training, hysteresis or LR selection – confirms the operational relevance. The daily-mean CAIRN tier correlates with daily complaint count at Pearson r=0.729r=0.729 (95% CI 0.450.45–0.840.84 by a 7-day moving-block bootstrap, which respects the strong day-to-day autocorrelation of both series (lag-1 +0.63+0.63 and +0.59+0.59); Spearman ρ=0.617ρ=0.617, n=89n=89 days; days on which no complaint record was returned are excluded rather than treated as zero counts; Fig. 7e). Under the leave-one-week-out refit the same correlation is 0.7210.721, and it sits 0.0720.072 below the deterministic ground-truth reference (r=0.793r=0.793); a paired 7-day moving-block bootstrap on the difference, in which the same resampled blocks are used for both series so that shared day-to-day variation cancels, gives Δr=+0.072 r=+0.072 (95% percentile CI −0.303-0.303 to +0.127+0.127; B=4,000B=4,000). The paired difference is not statistically distinguishable from zero, though the interval is wide and asymmetric because a small number of high-count complaint days dominate the resampled blocks. Measured against the best naive meteorological predictor on the same 89 days (daily stagnation fraction, r=0.475r=0.475), the fused tier improves on ambient meteorology by +0.255+0.255 (paired 7-day moving-block bootstrap, 95% CI +0.098+0.098 to +0.377+0.377, P(Δ≤0)=5×10−4P( ≤ 0)=5× 10^-4; the same resampled blocks are used for both series so that shared seasonal variation cancels). The corresponding single-channel margin over its own floor is +0.018+0.018 and does not reach significance. It is this multi-channel result that supports the community-impact claim: fusing four independently trained channels recovers a signal that no single channel separates from seasonal meteorology. The lagged cross-correlogram peaks at lag 0 (Fig. 7f, Supplementary Table S8). The lag-0 bar is the unique global maximum of a broad lag-zero-centred structure; its location supports interpretation as a nowcaster rather than a delayed reflector of past events or a forecaster of future ones.. The architecture transfers without modification across two receptor geometries and two chemical species, and recovers the community-impact signal at near-ground-truth quality despite the modest within-tier mismatch on the non-extreme classes. Figure 8 maps the same complaint record in space: weekly postcode-level counts over January–March 2025 are greatest in the postcodes nearest the landfill and decline with distance. The temporal agreement of Fig. 7e–f and this spatial gradient are independent expressions of one exposure signal, consistent with plume transport and dilution from the site rather than with reporting-driven variability. Discussion Systems used to monitor fugitive landfill emissions are inherently retrospective: exposure is typically characterised only after community impact has occurred, reflecting the source–pathway–receptor framework that underpins chemical-incident response. This posture understates both the scale and distribution of impact. Satellite inversions indicate that landfill methane emissions exceed reported inventories (20; 78; 102), associated health and wellbeing burdens are increasingly recognised (50; 52; 89), and these burdens fall disproportionately on less advantaged populations (25; 72). The results presented here demonstrate that receptor-scale exposure can be inferred directly from meteorological conditions, represented in a causal nowcaster, and validated against independent records of community impact. Together, these findings support a shift from retrospective characterisation towards anticipatory response. The controls exerted by meteorology on receptor hydrogen sulphide have previously been inferred from dispersion theory and partial field evidence (61; 80). Here, they are resolved directly from observations. Multiscale transfer entropy identifies a consistent hierarchy of drivers, with wind direction, wind speed and atmospheric pressure forming the dominant pathways of information flow (Fig. 2). That hierarchy is not an artefact of either the estimator settings or the particular twelve months: it survives a six-configuration estimator ablation (Supplementary Note 5) and reproduces independently on two disjoint half-years, with the ordering identical in both halves and in the full-year analysis and the sign of the effective transfer entropy agreeing in 19 of 21 driver–scale cells. Pressure carries an increasing share of directed information as the aggregation scale coarsens, peaking at 12 h. That is the horizon on which both of the synoptic mechanisms identified above operate, barometric pumping of the source flux and subsidence-driven suppression of the depth into which emissions are diluted, and directed information alone does not separate them. The hierarchy is a statement about control rather than about where in the chain that control is exerted. The dominant reading is that meteorology sets the conditions under which emissions reach a receptor, but a contribution from temperature- or pressure-dependent emission rates at the source is not excluded by these data. This interpretation is strengthened by agreement across independent approaches: the same structure emerges from system characterisation, causal inference and model attribution (Figs. 1, 2, 4). While each method has distinct limitations, convergence across them supports a physical interpretation of the meteorology–exposure relationship rather than an artefact of any single analysis. These insights extend to model design. Air-quality prediction systems typically impose temporal structure through engineered lags and aggregation windows (73; 55). In contrast, CAIRN links its temporal receptive field to the coupling timescales the causal analysis measures. The tested anchor inside the causally significant pressure band outperformed every tested anchor outside it (Extended Data Table ED1). Bounding the slow pathway to physically plausible timescales is consequential in its own right: releasing the bound and allowing the optimiser to migrate the lane freely costs 0.040.04 in F1-High on 10 of 13 folds (p=0.011p=0.011, Extended Data Table ED1). The marginal contributions of the two lanes are not separately identifiable (Supplementary Methods, “Walk-forward validation”). This alignment allows the model to reconstruct responses to pressure and temperature tendencies directly from raw meteorological inputs (Fig. 5), and it leads the engineered benchmark on every high-exposure metric while producing fewer false alarms (612 against 626 across 8,0408,040 timesteps; Extended Data Table ED2). The individual margins are modest and none reaches significance across the thirteen folds, though the paired High-class decision does (McNemar p=0.043p=0.043, uncorrected for multiplicity). Encoding temporal structure through measured dynamics thus reduces reliance on site-specific feature engineering. Model performance is further corroborated by validation against an independent record of community odour complaints. Evaluation of odour models rarely extends beyond comparison with sensor data or olfactometric measurements, and complaint data are typically used for operational support rather than for validation against a health-relevant endpoint (74; 88; 11). Here, complaints withheld from all stages of model development provide an external benchmark. That independence is procedural rather than physical: the complaint record is independent of the model-fitting process, but not of the process generating the sensor labels, since complaints and concentrations are both downstream of the same emission and dispersion events. The comparison therefore tests whether a meteorology-only signal tracks a human-reported endpoint, not whether two causally unrelated series agree. The model’s exposure probability co-varies with complaint counts at r=0.585r=0.585 (n=88n=88 days), with a cross-correlation peak at zero lag, with no lag beyond ±1± 1 day exceeding half the peak; one day is the temporal resolution of the complaint record (Fig. 6). This indicates that the model captures the same exposure signal experienced by the community. At a single channel that margin over a naive meteorological floor is small and within run-to-run variation; the claim rests on the fused tier reported below, not on this correlation alone. The zero-lag relationship is central to interpretation. A delayed peak would indicate reconstruction of exposure after it has been reported, while a leading relationship would imply the use of future information. Instead, the observed alignment places a quantified exposure indicator at the time of occurrence. This coincides with the point at which statutory-nuisance response operates: Environmental Health Officer assessments, Odour Management Plan audits and escalation through Local Resilience Forum structures currently rely on complaint-driven evidence (Supplementary Note 1). A real-time indicator therefore provides the situational awareness required for anticipatory and proportionate intervention without reliance on forecast information. A practical limitation arises from the dependence on receptor H2S measurements during training. Although inference relies only on meteorology, the model requires periodic retraining to reflect evolving site conditions, including changes in waste composition, gas extraction performance and boundary-layer dynamics (Fig. 6). Continuous H2S monitoring of this type is uncommon and resource-intensive, constraining deployment. The close correspondence between model output and complaint records suggests a potential alternative. Because the fused tier tracks daily complaint counts at r=0.721r=0.721 under the leave-one-week-out refit, against r=0.793r=0.793 for the deterministic raw-sensor ground truth scored the same way (n=89n=89 days; Fig. 7e), and therefore recovers most of the correlation a direct-sensor installation achieves from meteorology alone (paired Δr=+0.072 r=+0.072, 95% CI −0.303-0.303 to +0.127+0.127), complaint data may provide a viable supervisory signal. A nowcaster trained directly on real-time complaint counts could learn the meteorology–impact relationship without specialised instrumentation, while retaining adaptive retraining to capture non-stationary system behaviour. The initial use of H2S as a training target remains essential in this context. A physical, threshold-referenced quantity anchors model interpretation and allows complaints to function as an independent validation signal. The finding that these two signals encode comparable information is therefore a result of this study, not an assumption, and it is this result that motivates future exploration of complaint-based training. The causal structure of the nowcaster also provides a pathway to forecasting. Because predictions depend solely on meteorological inputs, substituting forecast for observed weather would allow forward projection without altering model architecture. Recent advances in numerical and machine-learning-based weather prediction systems provide the necessary inputs at appropriate spatial and temporal resolution (63; 87; 65). Whether predictive skill is retained under forecast uncertainty remains an open question, but the present work establishes the nowcasting basis required to address it. Taken together, these components support construction of a tiered exposure indicator derived from Bayesian fusion of meteorologically trained channels (Fig. 7). The resulting ordinal scale aligns with the FIDOL/WHO odour-annoyance framework used in UK statutory assessment (79). It reproduces sensor-derived classifications with substantial agreement (κw=0.709 _w=0.709, of which nine of the thirteen weeks are in sample with respect to the fusion parameters; 0.7030.703 under a leave-one-week-out refit) and closely tracks community complaint patterns. This tiered system integrates model outputs into an operational decision framework rather than constituting a separate analytical result. Evidence for the effectiveness of short-term alerting is mixed, with benefits strongest when alerts are coupled to required mitigation measures (71; 22). Nonetheless, the results establish a key prerequisite for such systems: meteorology-only models can recover receptor-scale exposure conditions and align closely with independently recorded community impact. These findings have direct implications for public health practice, regulation and environmental policy, because they enable a shift from retrospective, complaint-driven response to anticipatory, evidence-based management of odour exposure. The framework presented here integrates machine learning within the established source–pathway–receptor paradigm, extending it from a tool for post hoc risk assessment to one that supports prospective, data-driven intervention. By linking meteorology-derived exposure predictions to health-relevant values and community impact, it provides a basis for coordinated action across public-health responders, regulators, operators and affected communities. For public-health systems, a tiered, real-time exposure indicator enables proportionate intervention aligned with existing response frameworks (Supplementary Note 1). Classification-based outputs mapped to WHO guideline values translate model predictions into operationally meaningful risk categories, facilitating decision-making and risk communication. The spatial gradient in complaint activity observed in Fig. 8, greatest in the postcodes closest to the site and declining with distance, indicates that odour impacts are not uniformly distributed across the surrounding population, although the mapped counts are not population-normalised and therefore locate where impact is reported rather than per-capita risk. Proximity to the source and the atmospheric conditions associated with elevated exposure together allow public-health responses to be targeted in space and in time, supporting the protection of susceptible populations, including individuals with asthma or chronic obstructive pulmonary disease, who may experience disproportionate adverse effects during odour episodes. (93; 71). By identifying elevated exposure at the time it occurs, the system enables proactive targeting of interventions and tailored communication strategies. More broadly, combining predictive exposure with complaint data and meteorological context operationalises the four key determinants of odour impact, namely frequency, intensity, duration and offensiveness, providing a more comprehensive basis for public-health decision-making. As observed in other air-quality contexts, the benefits of such systems depend on the actions they trigger, with measurable health gains strongest where alerts are coupled to implemented control measures rather than advisory guidance alone (22). For affected communities, the availability of a predictive, transparent exposure signal addresses a central feature of odour-related harm: uncertainty. Persistent exposure to malodour is associated not only with physical symptoms but also with psychological stress and reduced wellbeing, particularly where exposure is unpredictable and difficult to manage (34). A tiered communication system provides advance indication of likely exposure conditions, enabling residents to anticipate and adapt behaviour accordingly. This may reduce perceived and actual impacts, improve community resilience, and support trust in regulatory processes when information is delivered clearly and consistently. By moving beyond complaint-driven response, in which residents act only after harm has occurred, the framework also promotes a more equitable distribution of protection, particularly for populations that experience disproportionate environmental burdens (25). For regulators and operators, the availability of a real-time and predictive exposure signal supports a transition toward risk-based, anticipatory management. Regulatory oversight can be targeted to periods and conditions associated with elevated risk, rather than applied uniformly or triggered only after complaints are received, improving both efficiency and proportionality (79). The identified causal drivers provide directly actionable insight for site management. Elevated exposure is associated with specific meteorological conditions, including low wind speed, falling temperature and stable boundary-layer regimes, allowing operational activities with a direct bearing on emissions, waste handling and capping among them,to be scheduled away from high-risk periods. This anticipatory approach has the potential to reduce emissions, minimise community disturbance, improve regulatory compliance and lower operational costs. Integration with meteorological forecasting systems would further extend this capability, enabling forward planning of mitigation actions. At the level of environmental policy, linking exposure prediction to health-based thresholds strengthens the evidentiary basis for regulatory decision-making. Interventions can be evaluated against quantified wellbeing outcomes, supporting both effectiveness and proportionality and addressing a key barrier to the adoption of stricter or more innovative measures. The use of general meteorological and temporal predictors, rather than site-specific operational inputs, enhances the potential transferability of the approach across different industrial contexts. This suggests applicability beyond landfill settings, to other sources of fugitive emissions such as waste management, wastewater and industrial facilities, although validation across such settings remains an important next step. More broadly, the framework illustrates how exposure prediction, source management and health protection can be integrated within a single operational system. By aligning scientific inference with regulatory practice and community experience, it provides a pathway for translating advances in environmental data science into tangible public-health and policy outcomes, supporting the development of more responsive, predictive and health-informed environmental governance. Methods Data and study sites Continuous ambient air-quality monitoring data were collected at three downwind monitoring stations (denoted MMF9, MMF1, MMF2) and one supplementary receptor (denoted MMF1A) throughout the calendar year 2024 (1 January–31 December). Records at 15-minute sampling interval yielded 34,60934,609, 23,10423,104, 28,33128,331 and 11,37411,374 observations with valid H2S respectively, against a complete 2024 grid of 35,136 15-minute intervals. The two shorter records are not the result of instrument downtime: MMF1 ceased operation on 31 August 2024 and the supplementary receptor began on 2 September 2024, a relocation rather than an outage, and the two are therefore not concurrent. Each station recorded H2S (μgm−3 \,m^-3), CH4 (ppm), wind direction (WD, deg), wind speed (WS, m s-1), air temperature (TEMP, ∘C) and barometric pressure (P, hPa); the temperature tendency dT/dtdT/dt (∘C hr-1) was computed as a first-order finite difference at the 15-minute sampling interval (TEMP and T are used interchangeably hereafter). The receptor-modelling analysis in Fig. 1a uses all three downwind stations; per-panel chemical and boundary-layer analyses (Fig. 1b–f, Fig. 2) and all classifier training and evaluation use the principal receptor MMF9, which provides the longest contiguous valid record during the 2024 evaluation window. The 2024 record uses post-calibration-adjustment data following the independent reanalysis of the H2S analyser by the regulator (28); no pre-2024 H2S observation enters any evaluation set. Training windows for the walk-forward protocol begin on 1 September 2023, so pre-adjustment data contribute to model fitting but never to any reported metric. Pasquill–Gifford stability classes were approximated using measured wind speed combined with a solar-insolation proxy from time-of-day and time-of-year (pyranometer data unavailable). All analyses were conducted in Python 3.10 (NumPy, SciPy, pandas, scikit-learn, hyperopt, PyTorch and Matplotlib). Environmental characterisation The six panels of Fig. 1 use standard statistical procedures, specified panel by panel in Supplementary Methods, “Environmental characterisation”: conditional probability function (CPF) rose plots over ns=36n_s=36 ten-degree wind-direction sectors at an H2S exceedance threshold Cthr=5μgm−3C_thr=5\, \,m^-3 (5) (a); Pearson and Spearman correlation, a log–log OLS power-law fit and a Jaccard spike co-occurrence index tested by χ2χ^2 for the H2S–CH4 pair (b); the closed diurnal (TEMP, H2S) phase loop, whose enclosed area is computed by the Shoelace formula and normalised by the mean ranges, with a hysteresis-asymmetry index from cooling/warming-limb slopes (c); the mean diurnal profile with hourly variation tested by Kruskal–Wallis (d); the sunrise-aligned composite over all 365 days, pre- versus post-sunrise concentrations compared by a one-sided Mann–Whitney U test (e); and bivariate OLS regressions of H2S on wind speed, temperature and dT/dtdT/dt with two-tailed t-tests on the slopes (f). the quantitative summary across all panels is in Supplementary Note 2 and Supplementary Table 1. Causal analysis Multiscale transfer entropy (MSTE) quantifies directed nonlinear information flow from each candidate driver to H2S across ten temporal scales τ∈15min, 30min, 1h, 2h, 3h, 6h, 12h, 24h, 48h, 96hτ∈\15\,min,\,30\,min,\,1\,h,\,2\,h,\,3\,h,\,6\,h,\,12\,h,\,24\,h,\,48\,h,\,96\,h\, following the framework of 91. For each driver–scale combination we (i) coarse-grain by non-overlapping block averaging at aggregation factor m=τ/(15min)m=τ/(15\,min); (i) compute conditional transfer entropy TX→Y|=I(Yt+1;Xt∣Yt(h),t)T_X→ Y =I\! (Y_t+1;\,X_t Y_t^(h),Z_t ) (92) using the Frenzel–Pompe k-nearest-neighbour estimator (37) conditioned on a single greedily selected confounder; (i) apply the rank-uniform marginal transform, tie-breaking jitter, Theiler exclusion window W=4W=4 (95) and an adaptive neighbour count k, all specified in Supplementary Methods, “Multiscale transfer entropy”; (iv) assess significance using B=2,000B=2,000 circular-shift surrogates that preserve the marginal distribution, periodicity and autocorrelation of the source while breaking its temporal alignment with the target; (v) compute effective transfer entropy by subtracting the surrogate-mean baseline; (vi) apply Benjamini–Hochberg false-discovery-rate correction at α=0.05α=0.05 across a single family of all 70 driver–scale tests, with surrogate p-values computed as p=(1+#b:TEb≥TEobs)/(B+1)p=(1+\#\b:TE_b _obs\)/(B+1) so that p=0p=0 is unattainable and the correction is well posed (7). To test whether the directed information reflects shared diurnal periodicity rather than transport, the analysis was repeated with the circular-shift offsets restricted to integer multiples of 24 h. This null preserves diurnal co-oscillation between driver and target while destroying their day-to-day correspondence, and is an exact permutation test over the 365 distinct day shifts the record admits. The directed information is essentially unchanged: the median proportion of effective transfer entropy attributable to shared diurnal structure is 0.0270.027 across the four directly measured drivers, and no cell exceeds 0.750.75. This test is uninformative at aggregation scales of 24 h and coarser, where every circular shift is already a day multiple and the day-shift null coincides with the ordinary one; those scales are reported as not tested for diurnal confounding rather than as unconfounded. The scale grid, driver set, embedding depths, conditioning strategy and surrogate scheme were fixed before the reported analysis was run and are reported in full; the estimator-configuration comparison of Supplementary Note 5 is exploratory and is reported as a robustness check rather than as a set of confirmatory tests. Only the single pre-specified configuration enters the Benjamini–Hochberg family. Benjamini–Hochberg controls the false discovery rate under independence or positive regression dependence. These tests are dependent by construction: the scales are nested coarse-grainings of a single series, and the three tendency variables are deterministic transforms of drivers already in the family. Positive regression dependence is therefore plausible but not demonstrated. The correction was therefore repeated as a sensitivity analysis under Benjamini–Yekutieli, which is valid under arbitrary dependence at the cost of a harmonic-sum penalty. Benjamini–Hochberg remains the reported procedure; cells surviving both corrections are drawn with filled symbols in Fig. 2 and cells significant under Benjamini–Hochberg alone with open symbols, so the dependence assumption can be read directly off the figure rather than taken on trust. A scale is reported as estimable only if it retains n≥100n≥ 100 coarse-grained blocks and the adaptive neighbour count satisfies k≤n/10k≤ n/10; a scale failing either test is reported as “not estimable” rather than “not significant”. Under this rule the 96 h scale (n=89n=89) is not estimable, and the coarsest estimable scale is 48 h. Blocks were retained only when at least 50% of their nominal sub-samples survived. The history embedding depth h is set adaptively and capped at h≤2h≤ 2 to keep the conditioning dimension within the regime where the KSG estimator retains adequate sensitivity. Pseudo-code for the full algorithm and the Frenzel–Pompe KSG-CMI subroutine is given in Supplementary Methods, “Multiscale transfer entropy”. A six-configuration ablation study confirms that the reported causal hierarchy is robust to estimator hyperparameter choice (Supplementary Note 5; Supplementary Tables 7–9). That ablation varies the estimator settings; it does not vary the sample. Stability with respect to the sample was assessed separately by split-half resampling on January–June and July–December 2024 with every estimator parameter held fixed: the ordering of the three directly measured drivers is identical in both halves and in the full-year analysis, and the sign of the effective transfer entropy agrees in 19 of 21 driver–scale cells. Per-half significance retention, the two cells that disagree, and the limitations of the test are reported in Supplementary Note 5, “Split-half stability of the recovered hierarchy”. Ensemble classifier-tree seasonal nowcaster Four independent XGBoost classifiers (17) were trained, one per calendar quarter of 2024 (Winter Q1, Spring Q2, Summer Q3, Autumn Q4), each using an expanding-window protocol rooted at the start of the monitoring record (September 2023) and extending to the boundary of the target evaluation quarter. Three ordinal classes use WHO guideline values, which include the odour detection and the odour annoyance guidance levels for H2H_2S: Low <2<2, Medium 22–77, High ≥7μgm−3≥ 7\, \,m^-3. The feature set encodes the causal hierarchy explicitly: causal Butterworth derivatives of T, P and WS (82); 22-h and 66-h stagnation indices (fraction of 15-minute intervals with WS<1.5ms−1WS<1.5\,m\,s^-1); a scalar/vector recirculation index; cyclic time and wind-direction encodings; and four annual Fourier seasonality harmonics. The resolved set contains 24 features and includes no lagged or rolling-window terms. One of the 24, the station identifier, is constant in the single-receptor configuration reported here and carries no information; it is retained only for compatibility with the multi-site pipeline. Correcting the derivative window to be strictly causal changed the attribution of the classifier substantially while leaving its skill essentially unaltered (Macro F1 0.608→0.6120.608→ 0.612), indicating that the engineered temperature tendency was load-bearing for interpretation rather than for predictive performance. Class imbalance (∼9.5:1 9.5:1 Low:High) is corrected by inverse prior-probability weighting via scikit-learn’s compute_sample_weight with class_weight=’balanced’ (26); SMOTE (16) and scaled-exponent weighting alternatives were rejected because the former produces physically implausible feature combinations (interpolation between temporally distant samples) and the latter degraded Low-class precision below acceptable levels. Bayesian hyperparameter optimisation uses the Tree-structured Parzen Estimator (8) over an eight-dimensional search space (max_depth, learning rate, n_estimators, subsample, colsample_bytree, α, λ, min_child_weight) with up to 40 trials and patience-10 early stopping. To prevent leakage between hyperparameter selection and seasonal evaluation, each season uses a temporally isolated one-week hyperopt holdout consumed before the season boundary, separated from the evaluation quarter by a one-week gap (15). The final model uses the multi:softprob objective with mlogloss as the training metric. The benchmark reported in Extended Data Table ED2 was trained for a fixed 200 boosting rounds with early stopping disabled, so no evaluation-window data influenced its stopping point. A 3-hour majority-vote ensemble (12 consecutive 15-minute samples per block) post-processes predictions to align with the 30-minute WHO averaging convention; ties are broken in favour of the highest-prior class (Low) for conservative false-alarm control. Model interpretability uses Accumulated Local Effects (ALE) computed on held-out validation data independently per seasonal classifier with PyALE at K=20K=20 quantile-based grid points (3); the full ALE derivation, the per-class importance metric, and the multiclass aggregation procedure are given in Supplementary Methods, “Accumulated Local Effects”. Hyperparameter values, per-season performance, sample counts and the hyperopt-vs-evaluation comparison are in Supplementary Tables 2–4; per-season ALE interpretation is in Supplementary Note 3. CAIRN: structured state-space dual-pathway nowcaster The CAIRN architecture removes the engineered-feature inductive bias and learns meteorological memory directly from a minimal sixteen-dimensional input: four raw meteorological channels (WD,WS,T,PWD,\,WS,\,T,\,P) augmented only by the sin/cos decomposition of WD (resolving the 0/360∘0/360 angular discontinuity), an hour-of-day sin/cos pair, and four annual Fourier seasonality harmonics represented as sin/cos pairs (eight channels). Raw WD is retained alongside its sin/cos decomposition, so the 0/360∘0/360 discontinuity the decomposition removes is reintroduced on one of the sixteen channels. This was not intended and is reported for completeness; the decomposed pair carries the directional information, and the ablation of Extended Data Table ED1 holds this input set fixed across all conditions, so no reported comparison is affected. No derivative, lag, rolling statistic, stagnation index or recirculation index is supplied. The model is a stack of nlayers=3n_layers=3 pre-norm S4D blocks (49; 47) of width H=128H=128 and state size N=256N=256, applied to a sequence of length L=384L=384 steps (96 h at 15-minute sampling interval). Each S4D layer convolves its input against a length-L, one-sided convolution kernel ¯=(¯,¯¯,…,¯L−1¯)∈ℝL, K\;=\; (C B,\;C A B,\;…,\;C A^\,L-1 B )\;∈\;R^L, (1) which is parameterised in S4D form with diagonal A, complex eigenvalues λn=−exp(logAnreal)+iπ(n−1) _n=- ( A^real_n)+i\,π(n-1), and learnable timescale Δ . The full continuous-time SSM, zero-order-hold discretisation, complex-Vandermonde diagonal-kernel construction, and S4D block specification (LayerNorm, gated-linear-unit fusion, FFN, DropPath, residual; 105; 23; 53; 56) are given in Supplementary Methods, “S4D state space model”. The architectural innovation is a channel-wise split into a fast lane (large decay magnitude |Afastreal|=10.0|A^real_fast|=10.0, short-Δ window) covering sub-hour to few-hour timescales, and a slow lane (small decay magnitude |Aslowreal|=0.5|A^real_slow|=0.5, long-Δ window) covering half-day to multi-day timescales, with 50/50 channel split (f=0.5f=0.5). Per-channel Δ values within each lane are sampled in log-space around a physics-anchored centre that biases coverage onto the canonical horizons identified by the MSTE analysis (Fig. 2a): a fast anchor τfast⋆=1τ _fast=1 h matching the wind-driven advective band, and a slow anchor τslow⋆=6τ _slow=6 h – a scale within the band over which atmospheric pressure remains causally significant (15 min to 12 h; Fig. 2a). Given the 15-minute sampling base Δtbase=0.25 t_base=0.25 h, Δlane⋆=Δtbase|Alanereal|⋅τlane⋆,logΔlane(h)∼(logΔlane⋆,s2), _lane\;=\; t_base|A^real_lane|·τ _lane, ^(h)_lane\; \;N\! ( _lane,\,s^2 ), (2) clamped to log-uniform per-lane bounds with log-space spread s=0.5s=0.5. Strict causality is preserved at every layer: the SSM kernel is one-sided, the FFT convolution is explicitly truncated to length L, the readout is the final-timestep slice, and no bidirectional, attention-based or global-pool operation is used. The training objective is focal cross-entropy (68) with focusing parameter γ=2.62γ=2.62 and per-class weights (αLow,αMed,αHigh)=(1.00,11.26,14.07)( _Low, _Med, _High)=(1.00,11.26,14.07) imposed manually rather than derived from inverse frequencies, complemented by inverse-prior balanced minibatch sampling. The optimiser is AdamW (70) with five differential learning-rate groups, a six-epoch warmup during which logAreal A^real and AimagA^imag are frozen, and a ReduceLROnPlateau scheduler monitoring validation High-class F1 with patience-2 reduction and floor 10−510^-5. Full optimiser configuration is given in Supplementary Methods, “Optimiser schedule”. Walk-forward validation and metrics CAIRN models are evaluated under a weekly expanding-window walk-forward protocol covering 1 October–31 December 2024 (13 non-overlapping weekly blocks; n=8,040n=8,040 evaluation timesteps). For each weekly block, training data comprise all observations up to the previous Sunday with the training cutoff retreated by L−1=383L-1=383 steps (≈96≈ 96 h) before the first validation timestamp, reserving a causal context window for the input convolution that contains no training labels. The engineered benchmark applies no such reserve: it is a pointwise classifier with no context window and therefore no corresponding leakage mode, so it trains to the week boundary and sees approximately four additional days per fold. The matched comparison retains this difference, since removing it would handicap the benchmark for a risk it does not carry. A fresh StandardScaler is fitted on training data alone and applied to the validation week and its context reserve. A fresh model is trained from scratch (no carry-over of weights, optimiser state or learning-rate schedule) for up to 60 epochs with patience-7 early stopping on an inner temporal probe: the seven days immediately preceding each evaluation week are withheld from training and used solely to select the epoch at which training stops. This departs from the protocol used in the first version of this work, in which early stopping, best-epoch restoration and checkpoint selection all keyed on the evaluation block itself. Re-running all 13 folds with the inner probe, together with a matched control that withholds the same quantity of training data but retains the original selection rule and so isolates the effect of the reduced training set, shows that the original procedure inflated F1-High by 0.0310.031 (12 of 13 folds negative; sign test p=0.0017p=0.0017, Wilcoxon p=0.0085p=0.0085). This protocol governs the H2S nowcaster evaluated over Oct–Dec 2024, which is the model reported in Fig. 6a,b, Extended Data Table ED2 and Supplementary Table 5. The cross-species CH4 model and the four channels fused in Fig. 7 were trained before the correction and retain the original selection rule; re-training them was not undertaken. Their reported values therefore carry the same selection optimism, of the order of the 0.0310.031 measured above at the single-channel level, although the effect on a fused ordinal tier is not separately quantified and the hysteresis and log-odds accumulation operate on latched activations rather than on raw posteriors. The affected quantities are identified where they appear (Supplementary Table 6; Fig. 6c,d; Fig. 7). The inflation is not a consequence of class sparsity: it is uncorrelated with the number of High-class samples in the evaluation week (r=−0.19r=-0.19, p=0.54p=0.54), and arises instead from selecting the maximum of a volatile epoch-wise metric (mean absolute epoch-to-epoch change 0.0480.048) over approximately 18 epochs. Probe adjacency matters. A 28-day probe ending four days before the evaluation week doubles the apparent effect relative to the adjacent 7-day probe, because a month-old block selects a checkpoint suited to conditions that have since changed. At this site’s event rate a seven-day block occasionally contains no High-class exceedance at all: evaluation weeks span 8–150 High samples and probe blocks 0–180. Any weekly-blocked protocol monitoring a rare-class metric at this event rate is therefore operating on a selection signal that is intermittently near-degenerate, and folds in which the probe contains no positive class are reported separately. In the matched comparison against the engineered benchmark, two of the thirteen inner probes contain no High-class sample and four more contain fewer than thirty; in those folds the stopping point is selected on a metric computed over a block containing none, or almost none, of the class it measures. Where the probe holds no High-class sample the monitored F1-High evaluates to zero at every epoch rather than to an undefined value, so no epoch registers an improvement, the early-stopping counter runs from the first post-warmup epoch and the weights restored are that epoch’s; the ReduceLROnPlateau scheduler receives the same zero and reduces the learning rate on its plateau schedule. The scheduler falls back to validation loss only when the monitored metric is genuinely undefined, which arises when the validation block is empty, not when the class is merely absent. The alternative, a 28-day probe, removes the degeneracy but places the selection block up to a month before the evaluation week, which at this site’s rate of synoptic change is its own bias. Both were run; the adjacent 7-day probe is reported because its selection block immediately precedes the period being predicted. The consequence is visible in one week of the benchmark comparison, whose inner block contains no High-class sample and whose F1 is correspondingly degraded despite the evaluation week itself holding 150 positives; this is a property of the selection protocol rather than of the model, and is a further reason for pooling rather than weighting weeks equally. Period-wide aggregates in Extended Data Table ED2 are computed by pooling the confusion counts over all 8,040 evaluation timesteps and deriving one metric from the pooled counts; the per-week table (Supplementary Table 5) reports the unweighted mean of the 13 per-week values. The two conventions give 0.533 and 0.528 respectively and are stated separately wherever both appear. Rows with any missing feature are dropped listwise; labels are forward-filled over gaps of at most 30 minutes; channels absent at a given timestep contribute their prior rather than an imputed probability. Under the reported protocol the production configuration reproduces to ±0.006± 0.006 F1-High across three independent random seeds, and the pooled within-condition standard deviation across all six ablation conditions is 0.0130.013 (12 d.f.). The metric suite spans hard classification (per-class and macro F1, and conditional-rolling F1/precision/recall and false-alarm rate over a 4-h moving window) and probabilistic calibration (Brier-High, 12; binary ranked probability score, 32; log-loss). Their mathematical definitions, the per-week pipeline enumeration and the no-leakage controls (label, feature, model, hyperparameter and retrain) are given in Supplementary Methods, “Walk-forward validation”. Per-week breakdown is reported in Supplementary Table 5; period-wide aggregates are in Extended Data Table ED2. Cross-species CH4 nowcaster and complaint validation A second CAIRN model with architecturally identical specification was trained on co-emitted CH4 at the principal receptor under the same weekly walk-forward protocol, retaining the original checkpoint-selection rule; it was not re-trained under the inner temporal probe, and its reported values are optimistic to the extent quantified in “Walk-forward validation”. CH4 class boundaries were derived from the H2S thresholds via an ordinary least-squares regression H2S=5.55CH4−8.29H_2S=5.55\,CH_4-8.29 fitted on the principal receptor record (giving Low <1.854<1.854, Medium 1.8541.854–2.7562.756, High ≥2.756≥ 2.756 ppm). External validation against community wellbeing uses an independent record of daily community odour-complaint counts (n=88n=88 days, Oct–Dec 2024, excluding days lacking complaint data). Daily-mean predicted High-class probabilities for both species were computed by averaging p^High p_High within calendar-day windows; lagged Pearson cross-correlations between p^High p_High and the daily complaint count were evaluated at lags −7-7 to +7+7 days, with the lag-0 correlation reported as the primary endpoint. The complaint record is not used during training and is unavailable to the classifier at inference time. Per-week performance for both species is in Supplementary Table 6; the underlying co-located 15-minute time series is shown in Supplementary Figure 1. Multi-site Bayesian alert-tier classifier A joint analysis fuses four CAIRN nowcasters – H2S and CH4 at receptors MMF9 and MMF2, each trained independently under Sec. CAIRN: structured state-space dual-pathway nowcaster and the walk-forward protocol of Sec. Walk-forward validation and metrics – into a continuous network-event probability and an ordinal alert tier τt∈0,1,2,3 _t∈\0,1,2,3\ at the native 15-minute sampling interval. The predicted tier is constructed purely from meteorology at inference time: no raw target-species concentration is read by any stage of the predicted pipeline. Each channel’s per-step High-class probability is thresholded at p⋆=0.5p =0.5 and passed through an asymmetric (Non,Noff)=(3,2)(N_on,N_off)=(3,2) latch (45-minute onset, 30-minute clearance – one full WHO 30-minute averaging interval plus a 15-minute safety margin, and one full interval, respectively). The four latched states ℓc,t _c,t then contribute additive log-likelihood-ratio updates to a prior log-odds at base-rate π=0.10π=0.10: logOt=logπ1−π+∑c∈logLRc,t,Pt=Ot/(1+Ot), O_t\;=\; \! π1-π\;+\; _c _c,t, P_t\;=\;O_t/(1+O_t), (3) with channel-state likelihood ratios calibrated below each nowcaster’s bare precision-prevalence ratio to absorb inter-channel chemical and meteorological dependence (Supplementary Note 6, Supplementary Table S3). H2S-quiet states contribute mild negative log-LR (silence at a direct receptor is informative); CH4-quiet and all missing-data states contribute logLR=0 =0 (neutral). The continuous posterior is mapped to four sequential ordinal tiers by hard thresholding alone – no raw-sensor override on the predicted side – at cut-points (θ1,θ2,θ3)=(0.15,0.50,0.92)( _1, _2, _3)=(0.15,0.50,0.92) aligned with the FIDOL public-health framework of Supplementary Note 1: Tier 1 at the FIDOL nuisance-assessment regime, Tier 2 at multi-channel corroboration, and Tier 3 at near-maximum attainable posterior (analytical ceiling Pmax≈0.949P_ ≈ 0.949). Robustness of the parameter choice is established by a ±25%± 25\% sensitivity analysis on the channel likelihood ratios (maximum |Δκw|=0.009| _w|=0.009; Supplementary Note 6, Supplementary Table S4). Performance is assessed against a deterministic ground-truth tier τtGTτ^GT_t constructed from raw concentrations by a channel-vote rule on the same debounced threshold exceedances (Supplementary Note 6, Eq. (S9)). The two pipelines share their per-channel hysteresis but differ in the multi-channel fusion stage (Bayesian log-odds versus hard channel-vote count) and in the target signal used as the per-channel input (model probability versus raw concentration). The primary statistic is the quadratic-weighted Cohen’s kappa (18; 64) on the synchronous Jan–Mar 2025 walk-forward window (n=8,536n=8,536 shared 15-min timesteps; 13 weekly retrains per constituent nowcaster). Confidence intervals are obtained by block bootstrap (B=1,000B=1,000 resamples, 24-hour block length) to accommodate within-episode autocorrelation. A vote-count baseline that replaces the Bayesian accumulator with a hard count of latched channels – equivalent to applying the ground-truth rule directly to the predicted activations – isolates the contribution of the soft-evidence fusion stage from the per-channel nowcasting stage. External validation against an independent record of daily community odour complaints is reported as a held-out endpoint that was not used in any stage of training, parameter selection or hysteresis design. Per-tier metrics with bootstrap CIs, the confusion matrix, the per-week kappa breakdown, the LR sensitivity analysis, and the complaint cross-correlation are in Supplementary Tables S4–S8. Code availability A reference implementation of the CAIRN nowcaster is publicly available at https://github.com/tcpearce/CAIRN. It contains the dual-pathway state-space model definition, the thirteen weekly walk-forward checkpoints for each species, the per-week standardisation parameters, the feature construction and labelling code, and an acceptance test suite that reproduces the per-timestep predictions reported here. Running python -m cairn.cli predict --species h2s --start 2024-10-01 --end 2024-12-30 regenerates the walk-forward results of Fig. 6 and Extended Data Table ED2 from the released checkpoints. The multiscale transfer-entropy estimator, the gradient-boosted benchmark and the scripts generating Figs. 1–7 are available from the corresponding author and will be deposited in the same archive on acceptance. Fig. 8 is not reproducible from the released data, since it rests on the community complaint record, which is withheld as personal data (Data availability). It is a descriptive map of the spatial distribution of that record and no quantitative result in this work is derived from it. A versioned snapshot with a persistent DOI will be deposited at Zenodo on acceptance. Data availability The 15-minute meteorological and gas-concentration record underlying the nowcasting results is released with the reference implementation above as a single Parquet file covering 1 September 2023 to 31 December 2024 at the principal receptor, comprising wind direction, wind speed, air temperature, barometric pressure, H2S and CH4. Community odour-complaint records are not released. They are personal data collected under a regulatory complaints process, are not required to reproduce any model result reported here, and are excluded from the released dataset and its file-level metadata. The complaint analyses of Figs. 6c–d, 7d–f and 8 therefore cannot be regenerated from the public release; the derived daily counts used in those panels are available from the corresponding author on reasonable request, subject to the information-governance conditions under which they were obtained. The monitoring site is identified by station code only. Author contributions • Conceptualization: A.F., T.C.P., D.J.T.S, A.D. • Methodology: T.C.P and A.F. • Investigation: T.C.P and A.F. • Visualization: T.C.P. and A.F. • Writing – original draft: T.C.P. and A.F. • Writing – review & editing: A.F., T.C.P., D.J.T.S, A.D. Funding acknowledgement This study is part-funded by the National Institute for Health and Care Research (NIHR) Health Protection Research Unit in Chemical Threats and Hazards. The views expressed are those of the author(s) and not necessarily those of the NIHR or the Department of Health and Social Care – T.C.P, D.J.T.S, A.D. and A.F (grant award NIHR207293). Acknowledgements The authors gratefully acknowledge Vince Jenner (Environmental Public Health Scientist, Environmental Hazards and Emergency Department, UK Health Security Agency) for valuable insights on the meteorological correlations associated with fugitive emissions. Sincere appreciation is also extended to Aphrodite Niggebrugge (Senior Geospatial Analyst, Geospatial Team, UK Health Security Agency) for the development of high-resolution GIS outputs, by which the visualisation of the research findings were effectively supported. Conflicting interests The authors declare no conflicting interests. References Agency for Toxic Substances and Disease Registry (2001) Agency for Toxic Substances and Disease Registry Landfill gas primer — an overview for environmental health professionals, chapter 2: landfill gas basics. Technical report U.S. Department of Health and Human Services, Atlanta, GA. Cited by: §S1, Introduction, Causal hierarchy of meteorological drivers. Agency for Toxic Substances and Disease Registry (2016) Agency for Toxic Substances and Disease Registry Hydrogen sulfide and carbonyl sulfide toxguide. Technical report U.S. Department of Health and Human Services, Public Health Service. Note: Accessed: 2026-07-03 External Links: Link Cited by: §S1. Apley and Zhu (2020) D.W. Apley and J. Zhu Visualizing the effects of predictor variables in black box supervised learning models. Journal of the Royal Statistical Society: Series B (Statistical Methodology) 82 (4), p. 1059–1086. External Links: Document Cited by: Definition., Accumulated Local Effects, Ensemble classifier-tree seasonal nowcaster. Asakura (2015) H. Asakura Sulfate and organic matter concentration in relation to hydrogen sulfide generation at inert solid waste landfill site — limit value for gypsum. Waste Management 43, p. 328–334. External Links: Document Cited by: §S1. Ashbaugh et al. (1985) L.L. Ashbaugh, W.C. Malm, and W.Z. Sadeh A residence time probability analysis of sulfur concentrations at grand canyon national park. Atmospheric Environment 19 (8), p. 1263–1270. External Links: Document Cited by: Conditional probability function (Fig. 1a)., Causal hierarchy of meteorological drivers, Environmental characterisation. Banerjee et al. (2025) S. Banerjee, K. Ghosh, M. De, U. Basu, and B. Basu Entropy-based analysis of urban pollutant–weather correlations. Note: arXiv:2508.17453; update venue/DOI on journal publication External Links: Document, 2508.17453, Document Cited by: §S4. Benjamini and Hochberg (1995) Y. Benjamini and Y. Hochberg Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B 57 (1), p. 289–300. External Links: Document Cited by: Surrogate testing., Causal analysis. Bergstra et al. (2011) J. Bergstra, R. Bardenet, Y. Bengio, and B. Kégl Algorithms for hyper-parameter optimization. In Advances in Neural Information Processing Systems 24 (NeurIPS), p. 2546–2554. Cited by: Ensemble classifier-tree seasonal nowcaster. Beychok (1994) M.R. Beychok Fundamentals of stack gas dispersion. Vol. 29, Milton R. Beychok, Irvine, CA. External Links: Document Cited by: §S1. Blanes-Vidal et al. (2012) V. Blanes-Vidal, E. S. Nadimi, T. Ellermann, H. V. Andersen, and P. Løfstrøm Perceived annoyance from environmental odors and association with atmospheric ammonia levels in non-urban residential communities: a cross-sectional study. Environmental Health 11 (1), p. 27. External Links: Document Cited by: Introduction. Brancher and De Melo Lisboa (2014) M. Brancher and H. De Melo Lisboa Odour impact assessment by community survey. Chemical Engineering Transactions 40, p. 139–144. External Links: Document Cited by: Discussion. Brier (1950) G.W. Brier Verification of forecasts expressed in terms of probability. Monthly Weather Review 78 (1), p. 1–3. External Links: Document Cited by: Metric definitions., Walk-forward validation and metrics. California Air Resources Board (2020) California Air Resources Board Estimation and comparison of methane, nitrous oxide, and trace volatile organic compound emissions and gas collection system efficiencies at california landfills. calpoly final report. Technical report CARB, Sacramento, CA. Cited by: §S4, §S4. Catena et al. (2022) A.M. Catena, J. Zhang, R. Commane, L.T. Murray, M.J. Schwab, E.M. Leibensperger, J. Marto, M.L. Smith, and J.J. Schwab Hydrogen sulfide emission properties from two large landfills in new york state. Atmosphere 13 (8), p. 1251. External Links: Document Cited by: §S4, §S4, §S4. Cawley and Talbot (2010) G.C. Cawley and N.L.C. Talbot On over-fitting in model selection and subsequent selection bias in performance evaluation. Journal of Machine Learning Research 11, p. 2079–2107. External Links: Link Cited by: Engineered-feature benchmark and physical interpretability, Ensemble classifier-tree seasonal nowcaster. Chawla et al. (2002) N.V. Chawla, K.W. Bowyer, L.O. Hall, and W.P. Kegelmeyer SMOTE: synthetic minority over-sampling technique. Journal of Artificial Intelligence Research 16, p. 321–357. External Links: Document Cited by: Ensemble classifier-tree seasonal nowcaster. Chen and Guestrin (2016) T. Chen and C. Guestrin XGBoost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, p. 785–794. External Links: Document Cited by: Ensemble classifier-tree seasonal nowcaster. Cohen (1968) J. Cohen Weighted kappa: nominal scale agreement provision for scaled disagreement or partial credit. Psychological Bulletin 70 (4), p. 213–220. External Links: Document Cited by: Multi-site Bayesian alert-tier classifier. Cui et al. (2019) Y. Cui, M. Jia, T.-Y. Lin, Y. Song, and S. Belongie Class-balanced loss based on effective number of samples. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Piscataway, NJ, USA, p. 9268–9277. External Links: Document Cited by: Joint loss.. Cusworth et al. (2024) D. H. Cusworth, R. M. Duren, A. K. Ayasse, R. Jiorle, K. Howell, A. Aubrey, R. O. Green, M. L. Eastwood, J. W. Chapman, A. K. Thorpe, J. Heckler, G. P. Asner, M. L. Smith, E. Thoma, M. J. Krause, D. Heins, and S. Thorneloe Quantifying methane emissions from United States landfills. Science 383 (6690), p. 1499–1504. External Links: Document Cited by: Discussion. Dai et al. (2026a) Y. Dai, J. Qian, Y. Yang, B. Liu, S. Li, K. Zhang, Q. Xie, C. Tong, Y. Chen, A. R. MacKenzie, and Z. Shi Significant benefits of pollution alerts for cleaner air and better health. PNAS Nexus 5 (3), p. pgag054. External Links: Document Cited by: Introduction. Dai et al. (2026b) Y. Dai, J. Qian, Y. Yang, B. Liu, S. Li, K. Zhang, Q. Xie, C. Tong, Y. Chen, A. R. MacKenzie, and Z. Shi Significant benefits of pollution alerts for cleaner air and better health. PNAS Nexus 5 (3), p. pgag054. External Links: Document Cited by: Discussion, Discussion. Dauphin et al. (2017) Y.N. Dauphin, A. Fan, M. Auli, and D. Grangier Language modeling with gated convolutional networks. In Proceedings of the 34th International Conference on Machine Learning (ICML), Vol. PMLR 70, p. 933–941. Cited by: S4D block., S4D block., CAIRN: structured state-space dual-pathway nowcaster. Department for Environment, Food & Rural Affairs (2017) Department for Environment, Food & Rural Affairs Interaction between environmental permitting and local authorities’ statutory nuisance duties. Note: Accessed: 24 June 2026 External Links: Link Cited by: Introduction. deSouza et al. (2025) P. N. deSouza, A. Rees, E. Oscilowicz, B. Lawlor, W. Obermann, K. Dickinson, L. M. McKenzie, S. Magzamen, S. Miller, and M. L. Bell Evaluating the environmental justice dimensions of odor in denver, colorado. Journal of Exposure Science & Environmental Epidemiology 36 (1), p. 67–76. External Links: Document Cited by: Introduction, Discussion, Discussion. Elkan (2001) C. Elkan The foundations of cost-sensitive learning. In Proceedings of the 17th International Joint Conference on Artificial Intelligence (IJCAI), p. 973–978. Cited by: Ensemble classifier-tree seasonal nowcaster. ENDS Report (2024) ENDS Report Growing number of persistently poor permitting performers in 2024 — 15 permits in band f over the last two years. External Links: Link Cited by: §S1. Environment Agency (2024) Environment Agency Walleys quarry landfill site, silverdale: hydrogen sulphide monitoring data adjustment statement. Environment Agency. Note: Statement on calibration issues and adjustment of historic hydrogen sulphide monitoring data External Links: Link Cited by: §S1, §S4, Introduction, Data and study sites. Environment Agency (2025) Environment Agency Environment agency chief regulator’s report 2024–25: supporting evidence. External Links: Link Cited by: §S1. Environmental Protection Act (1990a) Environmental Protection Act Environmental protection act 1990, c. 43. External Links: Link Cited by: §S1. Environmental Protection Act (1990b) Environmental Protection Act Section 79. External Links: Link Cited by: §S1, §S1, Introduction. Epstein (1969) E.S. Epstein A scoring system for probability forecasts of ranked categories. Journal of Applied Meteorology 8 (6), p. 985–987. External Links: Document Cited by: Metric definitions., Walk-forward validation and metrics. European Parliament Intergroup on Climate Change, Biodiversity and Sustainable Development (2021) European Parliament Intergroup on Climate Change, Biodiversity and Sustainable Development Revisiting odour pollution in europe: final agenda. Note: https://ebcd.org/wp-content/uploads/2021/10/Final-Agenda-Revisiting-Odour-Pollution-in-Europe.pdfEvent held 28 October 2021, hosted by MEP Maria Spyraki; co-organized with MIO-ECSDE and D-NOSES project Cited by: Introduction. Eykelbosh et al. (2021) A. Eykelbosh, R. Maher, D. de Ferreyro Monticelli, A. Ramkairsingh, S. Henderson, A. Giang, and N. Zimmerman Elucidating the community health impacts of odours using citizen science and mobile monitoring. Environmental Health Review 64 (2), p. 24–27. External Links: Document Cited by: Introduction, Discussion. Forde et al. (2019) O.N. Forde, A.G. Cahill, R.D. Beckie, and K.U. Mayer Barometric-pumping controls fugitive gas emissions from a vadose zone natural gas release. Scientific Reports 9 (1), p. 14080. External Links: Document Cited by: §S4, Introduction. [36] T. Freeman and R. Cudmore Review of odour management in new zealand. Note: Aurora Pacific Limited Cited by: §S1. Frenzel and Pompe (2007) S. Frenzel and B. Pompe Partial mutual information for coupling analysis of multivariate time series. Physical Review Letters 99 (20), p. 204101. External Links: Document Cited by: Frenzel–Pompe KSG estimator., Causal analysis. Friedman (2001) J.H. Friedman Greedy function approximation: a gradient boosting machine. Annals of Statistics 29 (5), p. 1189–1232. External Links: Document Cited by: Accumulated Local Effects. Glorot and Bengio (2010) X. Glorot and Y. Bengio Understanding the difficulty of training deep feedforward neural networks. In Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics (AISTATS), Vol. PMLR 9, p. 249–256. Cited by: Diagonal (S4D) parameterisation.. Goodwell and Kumar (2017) A.E. Goodwell and P. Kumar Temporal information partitioning: characterizing synergy, uniqueness, and redundancy in interacting environmental variables. Water Resources Research 53 (7), p. 5920–5942. External Links: Document Cited by: §S4, Multiscale transfer entropy. Gostelow et al. (2001) P. Gostelow, S.A. Parsons, and R.M. Stuetz Odour measurements for sewage treatment works. Water Research 35 (3), p. 579–597. External Links: Document Cited by: §S1. [42] Gov.uk Hydrogen sulphide: toxicological overview. External Links: Link Cited by: §S1. Gov.uk (2015) Gov.uk Nuisance smells: how councils deal with complaints. External Links: Link Cited by: §S1. Gov.uk (2024) Gov.uk Statutory nuisances: how councils deal with complaints. External Links: Link Cited by: §S1. Gov.uk (2025) Gov.uk Preparation and planning for emergencies: responsibilities of responder agencies and others. Cited by: §S1. Gu and Dao (2024) A. Gu and T. Dao Mamba: linear-time sequence modeling with selective state spaces. In First Conference on Language Modeling (COLM), Cited by: §S4. Gu et al. (2022a) A. Gu, K. Goel, and C. Ré Efficiently modeling long sequences with structured state spaces. In International Conference on Learning Representations (ICLR), Cited by: §S4, Continuous-time SSM., Zero-order-hold discretisation., S4D state space model, CAIRN: structured state-space dual-pathway nowcaster. Gu et al. (2020) A. Gu, T. Dao, S. Ermon, A. Rudra, and C. Ré HiPPO: recurrent memory with optimal polynomial projections. In Advances in Neural Information Processing Systems, Vol. 33, Red Hook, NY, USA, p. 1474–1487. External Links: Link Cited by: S4D state space model. Gu et al. (2022b) A. Gu, A. Gupta, K. Goel, and C. Ré On the parameterization and initialization of diagonal state space models. External Links: Document Cited by: §S4, Introduction, Diagonal (S4D) parameterisation., S4D state space model, Physics-anchored learnt nowcasting, CAIRN: structured state-space dual-pathway nowcaster. Guadalupe-Fernández et al. (2021) V. Guadalupe-Fernández, M. De Sario, S. Vecchi, L. Bauleo, P. Michelozzi, M. Davoli, and C. Ancona Industrial odour pollution and human health: a systematic review and meta-analysis. Environmental Health 20 (1), p. 108. External Links: Document Cited by: Discussion. He and Garcia (2009) H. He and E.A. Garcia Learning from imbalanced data. IEEE Transactions on Knowledge and Data Engineering 21 (9), p. 1263–1284. External Links: Document Cited by: Joint loss.. Heaney et al. (2011) C. D. Heaney, S. Wing, R. L. Campbell, D. Caldwell, B. Hopkins, D. Richardson, and K. Yeatts Relation between malodor, ambient hydrogen sulfide, and health in a community bordering a landfill. Environmental Research 111 (6), p. 847–852. External Links: Document Cited by: Discussion. Hendrycks and Gimpel (2016) D. Hendrycks and K. Gimpel Gaussian error linear units (gelus). arXiv preprint arXiv:1606.08415. External Links: Document Cited by: S4D block., CAIRN: structured state-space dual-pathway nowcaster. Hirasawa et al. (2019) Y. Hirasawa, M. Shirasu, M. Okamoto, and K. Touhara Subjective unpleasantness of malodors induces a stress response. Psychoneuroendocrinology 106, p. 206–215. External Links: Document Cited by: §S1, Introduction. Houdou et al. (2024) A. Houdou, I. El Badisy, K. Khomsi, S. A. Abdala, F. Abdulla, H. Najmi, M. Obtel, L. Belyamani, A. Ibrahimi, and M. Khalis Interpretable machine learning approaches for forecasting and predicting air pollution: a systematic review. Aerosol and Air Quality Research 24 (1), p. 230151. External Links: Document Cited by: Discussion. Huang et al. (2016) G. Huang, Y. Sun, Z. Liu, D. Sedra, and K. Q. Weinberger Deep networks with stochastic depth. In Computer Vision – ECCV 2016, Lecture Notes in Computer Science, Vol. 9908, Cham, p. 646–661. External Links: Document Cited by: S4D block., CAIRN: structured state-space dual-pathway nowcaster. Hudon et al. (2000) G. Hudon, C. Guy, and J. Hermia Measurement of odor intensity by an electronic nose. Journal of the Air & Waste Management Association 50 (10), p. 1750–1758. External Links: Document Cited by: §S1. Jang and Townsend (2001) Y.C. Jang and T. Townsend Sulfate leaching from recovered construction and demolition debris fines. Advances in Environmental Research 5 (3), p. 203–217. External Links: Document Cited by: §S1. Kass and Raftery (1995) R.E. Kass and A.E. Raftery Bayes factors. Journal of the American Statistical Association 90 (430), p. 773–795. External Links: Document Cited by: §S6. Kaza et al. (2018) S. Kaza et al. What a waste 2.0: a global snapshot of solid waste management to 2050. World Bank, Washington, DC. External Links: Document Cited by: §S1, Introduction. Ko et al. (2015) J. H. Ko, Q. Xu, and Y. Jang Emissions and control of hydrogen sulfide at landfills: a review. Critical Reviews in Environmental Science and Technology 45 (19), p. 2043–2083. External Links: Document Cited by: Discussion. Kraskov et al. (2004) A. Kraskov, H. Stögbauer, and P. Grassberger Estimating mutual information. Physical Review E 69 (6), p. 066138. External Links: Document Cited by: Conditional transfer entropy and confounder selection., Frenzel–Pompe KSG estimator., Rank transform and adaptive k.. Lam et al. (2023) R. Lam, A. Sanchez-Gonzalez, M. Willson, P. Wirnsberger, M. Fortunato, F. Alet, S. Ravuri, T. Ewalds, Z. Eaton-Rosen, W. Hu, A. Merose, S. Hoyer, G. Holland, O. Vinyals, J. Stott, A. Pritzel, S. Mohamed, and P. Battaglia Learning skillful medium-range global weather forecasting. Science 382 (6677), p. 1416–1421. External Links: Document Cited by: Discussion. Landis and Koch (1977) J.R. Landis and G.G. Koch The measurement of observer agreement for categorical data. Biometrics 33 (1), p. 159–174. External Links: Document Cited by: §S6, Multi-site Bayesian alert-tier classifier. Lang et al. (2024) S. Lang, M. Alexe, M. Chantry, J. Dramsch, F. Pinault, B. Raoult, M. C. A. Clare, C. Lessig, M. Maier-Gerber, L. Magnusson, Z. Ben Bouallègue, A. Prieto Nemesio, P. D. Dueben, F. Pappenberger, and F. Rabier AIFS – ECMWF’s data-driven forecasting system. Note: ECMWF Artificial Intelligence Forecasting System; operational at ECMWF since February 2025 External Links: 2406.01465, Document Cited by: Discussion. Lee et al. (2006) S. Lee et al. Reduced sulfur compounds in gas from construction and demolition debris landfills. Waste Management 26 (5), p. 526–533. External Links: Document Cited by: §S1. Leonardos (1996) G. Leonardos Review of odour control regulations in the usa. In Odors, Indoor and Environmental Air, p. 73–84. Cited by: §S1. Lin et al. (2017) T. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), p. 2980–2988. External Links: Document Cited by: Joint loss., CAIRN: structured state-space dual-pathway nowcaster. Local Government Association (2024) Local Government Association Civil resilience and emergency planning. External Links: Link Cited by: §S1. Loshchilov and Hutter (2019) I. Loshchilov and F. Hutter Decoupled weight decay regularization. In International Conference on Learning Representations (ICLR), Cited by: Differential learning-rate groups., CAIRN: structured state-space dual-pathway nowcaster. Lyons et al. (2016) R. A. Lyons, S. E. Rodgers, S. Thomas, R. Bailey, H. Brunt, D. Thayer, J. Bidmead, B. A. Evans, P. Harold, M. Hooper, and H. Snooks Effects of an air pollution personal alert system on health service usage in a high-risk general population: a quasi-experimental study using linked data. Journal of Epidemiology and Community Health 70 (12), p. 1184–1190. External Links: Document Cited by: Discussion, Discussion. Martuzzi et al. (2010) M. Martuzzi, F. Mitis, and F. Forastiere Inequalities, inequities, environmental justice in waste management and health. European Journal of Public Health 20 (1), p. 21–26. External Links: Document Cited by: Introduction, Discussion. Méndez et al. (2023) M. Méndez, M. G. Merayo, and M. Núñez Machine learning algorithms to forecast air quality: a survey. Artificial Intelligence Review 56 (9), p. 10031–10066. External Links: Document Cited by: Discussion. Mulrow et al. (2020) J. Mulrow, N. Kshetry, D.A. Brose, K. Kumar, D. Jain, M. Shah, T.E. Kunetz, and L.R. Varshney Prediction of odor complaints at a large composite reservoir in a highly urbanized area: a machine learning approach. Water Environment Research 92 (3), p. 418–429. External Links: Document Cited by: §S4, Introduction, Discussion. Nastev et al. (2001) M. Nastev et al. Gas production and migration in landfills and geological materials. Journal of Contaminant Hydrology 52 (1–4), p. 187–211. External Links: Document Cited by: §S1. National Academies of Sciences (1979) National Academies of Sciences Odors from stationary and mobile sources. National Academies Press, Washington, DC. External Links: Document Cited by: §S1. National Research Council (US) (2010) National Research Council (US) Acute exposure guideline levels for selected airborne chemicals: volume 9. External Links: Document, Link Cited by: §S1, §S1. Nesser et al. (2024) H. Nesser, D. J. Jacob, J. D. Maasakkers, A. Lorente, Z. Chen, X. Lu, L. Shen, Z. Qu, M. P. Sulprizio, M. Winter, S. Ma, A. A. Bloom, J. R. Worden, R. N. Stavins, and C. A. Randles High-resolution US methane emissions inferred from an inversion of 2019 TROPOMI satellite data: contributions from individual states, urban areas, and landfills. Atmospheric Chemistry and Physics 24 (8), p. 5069–5091. External Links: Document Cited by: Discussion. Nicell (2009) J.A. Nicell Assessment and regulation of odour impacts. Atmospheric Environment 43 (1), p. 196–206. External Links: Document Cited by: §S1, §S1, §S1, §S4, Discussion, Discussion. Njoku et al. (2025) P. O. Njoku, J. N. Edokpayi, and R. Makungo Assessment of landfill gas dispersion and health risks using AERMOD and TROPOMI satellite data: a case study of the thohoyandou landfill, south africa. Atmosphere 16 (12), p. 1402. External Links: Document Cited by: Discussion. Oke (1987) T.R. Oke Boundary layer climates. 2nd edition, Routledge, London. External Links: Document Cited by: §S4. Oppenheim et al. (1999) A.V. Oppenheim, R.W. Schafer, and J.R. Buck Discrete-time signal processing. 2nd edition, Prentice Hall, Upper Saddle River, NJ. External Links: Document Cited by: Ensemble classifier-tree seasonal nowcaster. Parker et al. (2002) T. Parker, J. Dottridge, and S. Kelly Investigation of the composition and emissions of trace components in landfill gas. r&d technical report p1-438/tr. Technical report Environment Agency, Bristol. Cited by: §S1. Parolari et al. (2021) A. J. Parolari, J. Sizemore, and G. G. Katul Multiscale legacy responses of soil gas concentrations to soil moisture and temperature fluctuations. Journal of Geophysical Research: Biogeosciences 126 (2), p. e2020JG005865. External Links: Document Cited by: §S4, Introduction. Pasquill and Smith (1983) F. Pasquill and F.B. Smith Atmospheric diffusion. 3rd edition, Ellis Horwood, Chichester. Cited by: §S4, Introduction, Causal hierarchy of meteorological drivers. Poulsen et al. (2003) T.G. Poulsen, M. Christophersen, P. Moldrup, and P. Kjeldsen Relating landfill gas emissions to atmospheric pressure using numerical modelling and state-space analysis. Waste Management & Research 21 (4), p. 356–366. External Links: Document Cited by: §S4, Introduction. Price et al. (2025) I. Price, A. Sanchez-Gonzalez, F. Alet, T. R. Andersson, A. El-Kadi, D. Masters, T. Ewalds, J. Stott, S. Mohamed, P. Battaglia, R. Lam, and M. Willson Probabilistic weather forecasting with machine learning. Nature 637 (8044), p. 84–90. Note: GenCast; published online Dec 2024 External Links: Document Cited by: Discussion. Prudenza et al. (2023) S. Prudenza, C. Bax, and L. Capelli Implementation of an electronic nose for real-time identification of odour emission peaks at a wastewater treatment plant. Heliyon 9 (10), p. e20437. External Links: Document Cited by: Discussion. Quist and Johnston (2024) A.J.L. Quist and J.E. Johnston Malodors as environmental injustice: health symptoms in the aftermath of a hydrogen sulfide emergency in carson, california, usa. Journal of Exposure Science & Environmental Epidemiology 34 (6), p. 935–940. External Links: Document Cited by: §S4, Introduction, Discussion. Rajesh et al. (2025) M. Rajesh, R. Ganesh Babu, U. Moorthy, and S. Veerappampalayam Easwaramoorthy Machine learning-driven framework for real-time air quality assessment and predictive environmental health risk mapping. Scientific Reports 15 (1), p. 28801. External Links: Document Cited by: Introduction. Runge et al. (2019) J. Runge, S. Bathiany, E. Bollt, G. Camps-Valls, D. Coumou, E. Deyle, C. Glymour, M. Kretschmer, M.D. Mahecha, J. Muñoz-Marí, E.H. van Nes, J. Peters, R. Quax, M. Reichstein, M. Scheffer, B. Schölkopf, P. Spirtes, G. Sugihara, J. Sun, K. Zhang, and J. Zscheischler Inferring causation from time series in earth system sciences. Nature Communications 10 (1), p. 2553. External Links: Document Cited by: §S4, Introduction, Conditional transfer entropy and confounder selection., Theiler window., Surrogate testing., Multiscale transfer entropy, Causal analysis. Schreiber (2000) T. Schreiber Measuring information transfer. Physical Review Letters 85 (2), p. 461–464. External Links: Document Cited by: Conditional transfer entropy and confounder selection., Causal analysis. Shusterman (1992) D. Shusterman Critical review: the health significance of environmental odor pollution. Archives of Environmental Health 47 (1), p. 76–87. External Links: Document Cited by: §S1, Introduction, Discussion. Stull (1988) R.B. Stull An introduction to boundary layer meteorology. Kluwer Academic Publishers, Dordrecht. External Links: Document Cited by: §S4, Causal hierarchy of meteorological drivers. Theiler (1986) J. Theiler Spurious dimension from correlation algorithms applied to limited time-series data. Physical Review A 34 (3), p. 2427–2432. External Links: Document Cited by: Theiler window., Causal analysis. Townsend et al. (2004) T. Townsend et al. Heavy metals in recovered fines from construction and demolition debris recycling facilities in florida. Science of the Total Environment 332 (1-3), p. 1–11. External Links: Document Cited by: §S1. Townsend et al. (2005) T. Townsend et al. C&D waste landfill in florida: assessment of true impact and exploration of innovative control techniques. Technical report Florida Center for Solid & Hazardous Waste Management. Cited by: §S1. Turner (1994) D.B. Turner Workbook of atmospheric dispersion estimates: an introduction to dispersion modeling. 2nd edition, CRC Press, Boca Raton. External Links: Document Cited by: §S1, §S4. UK Health Security Agency (2023) UK Health Security Agency Environmental public health surveillance system (ephss): report for 2021–2023. External Links: Link Cited by: §S1, §S1. UK Health Security Agency (2025) UK Health Security Agency Health risk assessment of air quality monitoring results from march 2021 to january 2025: walleys quarry landfill site, silverdale, newcastle-under-lyme. Technical report UKHSA, London. Cited by: §S1, §S1, §S4, Introduction. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A.N. Gomez, L. Kaiser, and I. Polosukhin Attention is all you need. In Advances in Neural Information Processing Systems 30 (NeurIPS), Red Hook, NY, USA, p. 5998–6008. Cited by: S4D block.. Wang et al. (2024) Y. Wang, M. Fang, Z. Lou, H. He, Y. Guo, X. Pi, Y. Wang, K. Yin, and X. Fei Methane emissions from landfills differentially underestimated worldwide. Nature Sustainability 7 (4), p. 496–507. External Links: Document Cited by: Introduction, Discussion. World Health Organization (2000) World Health Organization Air quality guidelines for europe. Technical report WHO Regional Publications, European Series, No. 91, WHO Regional Office for Europe, Copenhagen. Cited by: §S6. Wroniszewska and Zwoździak (2020) A. Wroniszewska and J. Zwoździak Odor annoyance assessment by using logistic regression on an example of the municipal sector. Sustainability 12 (15), p. 6102. External Links: Document Cited by: Introduction. Xiong et al. (2020) R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T.-Y. Liu On layer normalization in the transformer architecture. In Proceedings of the 37th International Conference on Machine Learning (ICML), Vol. PMLR 119, p. 10524–10533. Cited by: S4D block., S4D block., CAIRN: structured state-space dual-pathway nowcaster. Xu et al. (2014) L. Xu, X. Lin, J. Amen, K. Welding, and D. McDermitt Impact of changes in barometric pressure on landfill methane emission. Global Biogeochemical Cycles 28 (7), p. 679–695. External Links: Document Cited by: §S4, Introduction. Figure 1: Caption continued on next page. Environmental characterisation of receptor-level hydrogen sulphide exposure. a, Conditional probability function (CPF) rose plots at the three downwind monitoring stations (peak bearing 265∘265 at the principal receptor; CPF threshold 5μgm−35\, \,m^-3); dashed polygon, source area; n=34,609n=34,609 15-minute observations at the principal receptor. b, Bivariate H2S–CH4 relationship at the principal receptor, linear (left) and log–log (right) scales; Pearson r=0.830r=0.830 (p<0.001p<0.001); no-intercept OLS linear slope β^0=4.47μgm−3ppm−1 β_0=4.47\, \,m^-3\,ppm^-1; power-law fit y=0.356x1.95y=0.356\,x^1.95 (exponent b=1.95b=1.95). c, Hourly-mean diurnal trajectory in (temperature, H2S) phase space; counter-clockwise loop with normalised area A∗=0.369A^*=0.369. d, Mean diurnal H2S concentration profile (line) with 95% confidence interval of the hourly mean (shading; ± 1.96 SEM from daily composites); peak 02:00 LST, flat early-afternoon minimum across 12:00–15:00 LST. e, Sunrise-aligned mean concentration profile; pre-sunrise baseline 9.5μgm−39.5\, \,m^-3 (dashed), fumigation window 0–4 h (shading); concentrations are means ± 95% CI across all available calendar days. f, Bivariate scatter of H2S against wind speed (yellow), temperature and temperature tendency (navy); curves show smoothed 90th-percentile spline fitted to 20 equal-width bins in x, capturing upper-tail dependence on each driver. Methods, “Environmental characterisation”. Full numerical detail in Supplementary Note 2 and Supplementary Table 1. Figure 2: Multiscale information flow from meteorology to hydrogen sulphide. a, Effective transfer entropy (ETE; nats) from each of seven candidate meteorological drivers to H2S, evaluated at ten temporal scales (τ=15τ=15 min to 9696 h), of which the nine estimable scales (τ=15τ=15 min to 4848 h) are displayed; row colours identify each driver (shared with panel b). Cell values give ETE in nats. Colour intensity encodes ETE relative to that cell’s own surrogate standard deviation, because raw ETE is not comparable across scales (Methods); asterisks denote significance after Benjamini–Hochberg false-discovery-rate correction across a single family of all 70 tests (padj∗∗≤0.025^**\,p_adj≤ 0.025; ∗ 0.025<padj<0.05^*\,0.025<p_adj<0.05). Filled symbols mark cells that additionally survive Benjamini–Yekutieli correction, which is valid under arbitrary dependence; open symbols mark cells significant under Benjamini–Hochberg alone. Of the 29 cells reaching significance, 20 are filled and 9 open; 18 of the 19 cells belonging to the three directly measured core drivers are filled, the exception being wind speed at 12 h. Seven of the nine open cells are meteorological tendency variables and the remaining two are wind speed at 12 h and temperature at 15 min. The seven cells of the 96 h row are retained in the correction family although not estimable and not displayed, so the figure shows 63 of the 70 tested cells while the multiplicity penalty is applied across all 70. The three tendency-variable rows are shown in grey: they are deterministic transforms of drivers already present in the family, retained so that the Benjamini–Hochberg correction is applied across all 70 tests, but not interpreted in the main text (Supplementary Note 5). b, Stacked-bar decomposition of total positive ETE by driver at each temporal scale. Segment fills and stacking order match panel a; numerals give the cumulative total per scale; asterisks follow the same FDR threshold as panel a and are shown only for segments tall enough to carry a marker. Scales at which no cell reaches significance are labelled n.s. Estimator and confounder selection: Methods, “Causal analysis” (B=2,000B=2,000 circular-shift surrogates). Sensitivity to estimator hyperparameters in Supplementary Note 5. Figure 3: Caption continued on next page. Three-class XGBoost classifier of receptor exposure. WHO-anchored class boundaries: Low <2<2, Medium 22–77, High ≥7μgm−3≥ 7\, \,m^-3. a, Seasonal expanding-window classification of 15-minute H2S concentrations across 2024; WHO thresholds (dashed). True positives (TP=1,914=1,914), false positives (FP=2,395=2,395) and false negatives (FN=1,147=1,147) for High exceedance are overlaid. b, Co-located daily community odour-complaint counts, overlaid in time with the High-class predictions in panel (a). c, 7-day rolling classification accuracy across the full 2024 validation period. d, 3-hour majority-vote ensemble (12 consecutive 15-minute predictions per block) aligning the model output with the 30-minute WHO averaging convention; per-sample ground truth and class probabilities also shown. Period-wide macro-averaged performance: 15-min F1=0.651F_1=0.651 (precision =0.621=0.621, recall =0.702=0.702); 3-h ensemble F1=0.660F_1=0.660 (precision =0.621=0.621, recall =0.731=0.731). Methods, “XGBoost seasonal nowcaster”. Per-season performance, hyperopt parameters, sample counts and hyperopt-vs-evaluation comparison in Supplementary Tables 2–4. Figure 4: Caption continued on next page. Accumulated local effects (ALE) interpretation of the XGBoost classifier. a, Class-averaged ALE feature importance grouped by season (Methods, “Ensemble classifier-tree seasonal nowcaster”; full derivation in Supplementary Methods, “Accumulated Local Effects”); features sorted by descending mean importance across seasons. The leading features are hour_coshour\_cos, stagnation_2hstagnation\_2h and wind speed; the wind channels coincide with the multiscale transfer entropy core of Fig. 2a, while the diurnal term reflects the nocturnal accumulation regime resolved in Fig. 1d,e. b, Class-averaged ALE dependence curves, one subplot per feature (23 features); curves show four seasons overlaid. Positive ALE shifts predictions towards Medium/High; negative ALE shifts towards Low. ALE computed on held-out validation data for each seasonal classifier. Per-season interpretation in Supplementary Note 3. Figure 5: Caption continued on next page. Dual-pathway state space architecture and learnt representations.a, Schematic of the dual-pathway diagonal S4 model (Methods, “CAIRN: structured state-space dual-pathway nowcaster”). The 16-feature input (four raw meteorological channels and circular/Fourier calendar encodings; no derivative, lag, rolling, stagnation or recirculation feature) is linearly projected to model width H=128H=128 and split channel-wise into equal fast and slow lanes, each processed by three pre-norm S4D blocks with a gated linear unit, the two lanes then combined by a single GLU fusion. Fast lane: 1 h advective anchor; slow lane: 6 h synoptic anchor (a scale within the causally significant pressure band of Fig. 2a). b, Empirical (ground-truth) versus model-predicted High-class rates against hour of day (left) and 10∘10 wind-direction sector (right); 13-week walk-forward validation (Oct–Dec 2024). c, Column-normalised per-feature, per-lane sensitivity over the most recent 10 h of the 96 h input window; fast lane (top) and slow lane (bottom), averaged across the 11 of 13 weekly checkpoints that emitted sensitivity telemetry. The slow lane’s extended response reflects the sensitivity construction rather than the learnt kernel timescales, which change only marginally during training, and should not be read as memory acquired by the model. d, Scatter of predicted p^High p_High against the 96-hour normalised barometric-pressure change ΔP96h/σP P_96\,h\,/\, _P, coloured by the concurrent 96-hour temperature change ΔT96h T_96\,h; red curve, Savitzky–Golay smoothed trend; Spearman rsr_s annotation; both quantities are derived from raw P, T traces and are not supplied as input features. e, Violin distributions of p^High p_High across five equal-frequency quintile bins of the 96-hour temperature change (very rapid cooling to rapid warming); horizontal tick, within-quintile mean; violin colour reflects empirical High-class rate in that quintile (darker == higher rate). Six-condition initialisation ablation in Extended Data Table ED1; lane-ablation in Supplementary Methods, “Walk-forward validation”. Figure 6: Caption continued on next page. Operational walk-forward performance and external community-impact validation. 13 weekly expanding-window retrains over Oct–Dec 2024 (Methods, “Walk-forward validation”); n=8,040n=8,040 evaluation timesteps. a, Top: 15-minute H2S concentration with WHO Medium (2μgm−32\, \,m^-3) and High (7μgm−37\, \,m^-3) thresholds; per-sample S4 outcomes (true positives, false positives, false negatives) overlaid; weekly retrain boundaries marked. Lower rows: predicted High-class probability and rolling 24 h recall, precision, F1, false-alarm rate and 4 h moving-average Brier-High, log-loss-High and binary RPS-High for both models. Horizontal dashed lines: period-wide means. b, Receiver-operating characteristic (left) and 10-bin reliability diagram (right) over the full validation window; S4-H2S AUC =0.898=0.898, XGBoost AUC =0.877=0.877. The paired per-fold difference is not significant (S4 leads in 8 of 13 weeks, sign test p=0.58p=0.58; Wilcoxon p=0.11p=0.11). c, Lagged Pearson cross-correlogram between day-mean p^High p_High and the daily community odour-complaint count (n=88n=88 days; lags −7-7 to +7+7 days; negative lag, model leads complaints). Top: H2S model; bottom: CH4 model. d, Daily-aggregate scatter relationships (n=88n=88 days): complaints versus H2S model probability (left) and CH4 model probability (middle); cross-model agreement p^HighH2S p_High^H_2S versus p^HighCH4 p_High^CH_4 (right; OLS trend extrapolated to p^=1 p=1); ordinary least-squares fits with 95% confidence intervals for the mean response. Both nowcasters share the production architecture of panel (a) and Methods (“CAIRN: structured state-space dual-pathway nowcaster”). Panels (a) and (b), and the H2S series in (c) and (d), use the matched inner-probe protocol reported throughout the Results. The CH4 series in (c) and (d) is the published protocol: the cross-species arm was not re-trained under the correction, so the CH4 correlation and the cross-model agreement panel are descriptive rather than like-for-like comparisons. The same asymmetry is noted in Supplementary Table 6. Period-wide aggregates in Extended Data Table ED2; weekly breakdown in Supplementary Tables 5–6. Figure 7: Caption continued on next page. Multi-site Bayesian tier classifier matches the deterministic raw-sensor ground truth and the community-impact signal as an emissions nowcaster. a, Per-channel binary High-class activation for the four independently trained CAIRN nowcasters (H2S and CH4 at receptors MMF9 and MMF2), each weekly retrained under the walk-forward protocol of Sec. Walk-forward validation and metrics, under the original checkpoint-selection rule, so the figures below carry the selection optimism quantified in Methods. Within each channel row, the upper strip shows the ground-truth raw-concentration exceedance and the lower strip shows the model’s hysteresis-debounced High-class prediction; vertical dashed lines mark the 13 weekly walk-forward folds. b, The fused continuous network-event probability PtP_t (Eq. (3)); coloured horizontal bands shade the four ordinal-tier zones and the three dashed lines mark the Tier 1 (θ1=0.15 _1=0.15), Tier 2 (θ2=0.50 _2=0.50) and Tier 3 (θ3=0.92 _3=0.92) cut-points. c, CAIRN-predicted (top strip) versus deterministic raw-sensor ground-truth (bottom strip) ordinal alert tier τt∈0,1,2,3 _t∈\0,1,2,3\. Quadratic-weighted Cohen’s κw=0.709 _w=0.709, strict element-wise accuracy 79.4% (n=8,536n=8,536 synchronous 15-min timesteps). d, Daily community odour-complaint counts (symlog y), grey bars on the shared calendar axis of panels (a)–(c). The complaint record is fully external – never used during training or hyperparameter selection. e, Daily-mean alert tier versus daily complaint count (n=89n=89 days carrying a complaint record, including the 19 on which the recorded count was zero). Dark navy circles and dashed line: CAIRN predicted tier versus complaints (Pearson r=0.729r=0.729, R2=0.532R^2=0.532). Semi-transparent blue circles and dotted line: deterministic raw-sensor ground-truth tier versus complaints (r=0.793r=0.793, R2=0.629R^2=0.629). The two ordinary least-squares fits have essentially identical slopes; the fused tier tracks community impact at a quality comparable to the ground truth itself. f, Lagged Pearson cross-correlation of daily-mean predicted tier with daily complaint count across a ±7± 7-day window. Bars: Pearson r at each lag, shaded by significance; stars denote thresholds (∗p<0.001^***\,p<0.001, p∗∗<0.01^**\,p<0.01, ∗p<0.05^*\,p<0.05). The lag-0 bar is the unique global maximum at r=0.729r=0.729, with approximately symmetric roll-off at ±1± 1 day (r≈0.55r≈ 0.55–0.580.58). Correlations at ±5± 5 to ±7± 7 days also reach significance, so the roll-off is not strictly monotonic and the peak is the global maximum of a broad lag-zero-centred structure rather than an isolated one. The lag-0 peak identifies a nowcaster rather than a lagged reflector of past events or a forecaster of future ones. Figure 8: Caption continued on next page. Spatial reconstruction of weekly odour-complaint activity around the landfill site. Each panel shows the postcode-level distribution of daily odour complaints aggregated by week between January and March 2025. Colour intensity is proportional to the number of reports received, with darker shading indicating greater complaint activity. A persistent concentration of complaints is observed in communities closest to the landfill, with reporting intensity progressively declining with distance from the source. The stable spatial gradient provides independent community-based validation of the exposure patterns inferred from H2S measurements and meteorologically driven dispersion processes. Supplementary Information Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Supplementary Notes 1–6, Supplementary Methods, Supplementary Tables and Supplementary Figures. Extended Data tables are numbered ED1–ED5, Supplementary tables S1 onwards. References are listed once, at the end of the main article. Meteorology-driven Causal Nowcasting of Fugitive Landfill Emissions Enables Proactive Public Health Response Supplementary Information Contents References Extended Data S1 Regulatory and public-health context S2 Extended environmental characterisation S3 Per-season interpretation of the XGBoost ALE feature hierarchy S4 Extended comparison with prior literature S5 MSTE estimator ablation study S6 Joint site analysis Supplementary Methods Supplementary Tables Supplementary Figures This document accompanies the main manuscript and is organised in two parts. Extended Data (Tables ED1–ED3 below) holds the peer-reviewed quantitative tables referenced from the main figures and Results text. Supplementary Information consists of six Notes, six Methods subsections, nine Tables and one Figure that hold the contextual, per-panel, sensitivity and full-procedural material referenced from the main text, figure captions and Methods. Numbering: Extended Data items are prefixed ED, Supplementary items S. Supplementary Note 1 provides the regulatory and public-health context that frames the case-study site. Supplementary Note 2 gives the full per-panel quantitative detail underlying Fig. 1 of the main text. Supplementary Note 3 reports the per-season interpretation of the Accumulated Local Effects (ALE) hierarchy of the XGBoost classifier. Supplementary Note 4 provides an extended comparison with prior literature on landfill hydrogen sulphide, machine-learning nowcasting, and community-impact validation; this material currently overlaps with the main Discussion and may be relocated here in a compressed revision. Supplementary Note 5 documents the multiscale transfer entropy (MSTE) estimator ablation that supports the robustness of the causal hierarchy reported in main Fig. 2. Supplementary Methods holds the full mathematical and procedural detail for the methods summarised in main Methods, organised in six subsections that correspond one-to-one with the six Supplementary-Methods cross-references in main Methods: “Environmental characterisation” (per-panel mathematical detail for Fig. 1); “Multiscale transfer entropy” (full procedural specification, KSG–Theiler estimator, and Algorithm 1 + KSG-CMI subroutine pseudo-code); “Accumulated Local Effects” (ALE derivation, multiclass aggregation and importance metric); “S4D state space model” (continuous-time SSM, ZOH discretisation, S4D diagonal kernel, block specification, end-to-end architecture and causality guarantees); “Optimiser schedule” (joint focal-cross-entropy/MSE loss, AdamW with five differential learning-rate groups, warmup-freeze and ReduceLROnPlateau scheduler); and “Walk-forward validation” (per-week pipeline, no-leakage guarantees, lane-ablation and feature-sensitivity attribution, and metric definitions). Supplementary Tables 1–10 and Supplementary Figure 1 hold the numerical detail referenced from the main text and from the Notes and Methods below. Extended Data Table ED1: Six-condition initialisation ablation isolating the architectural contribution of physics-anchored timescale priors. All conditions share the dual-pathway diagonal S4 architecture, optimiser, focal-cross-entropy loss, sixteen-feature input, walk-forward protocol (13 weekly retrains over Oct–Dec 2024) and hyperparameter set; they differ only in the initialisation of the SSM kernel parameters Δlane,Alanereal\ _lane,A^real_lane\ and the per-lane Δ bounds. The MSTE-aligned condition (2a, bold) is the production configuration used in the main text. Period-wide F1-High is the operational target metric for High-class spike detection. Each condition was run three times from independent random seeds (42/43/44); values are mean ± s.d. over those three runs, quoted to two decimal places because the third decimal lies below the run-to-run noise floor. The pooled within-condition standard deviation is 0.0130.013 (12 d.f.). Conditions 1a, 1b, 1c and 2b are mutually indistinguishable (one-way ANOVA F=1.75F=1.75, p=0.23p=0.23); only 2a and 3b separate from that tier. Paired per-fold comparisons of 2a against each alternative use the seed-mean value per fold across the three seeds, giving 13 paired differences. Benjamini–Hochberg-corrected q ranges from 0.0030.003 to 0.0480.048 on two-sided Wilcoxon signed-rank tests; the comparison against random initialisation (1b) is the weakest at q=0.048q=0.048, and does not reach significance on a two-sided sign test (9 of 13 folds, p=0.267p=0.267). Under two protocols that remove the checkpoint selection — scoring every fold at the final retained epoch (patience expiry) and at a common fixed epoch, the last two columns — the same comparison improves to q≤0.013q≤ 0.013 and every paired margin is preserved or widened; per-condition values are given in those columns rather than restated here. Headline values are reported in main Results, “Physics-anchored learnt nowcasting”. ID Name Init. scheme τfast⋆τ _fast τslow⋆τ _slow Hypothesis tested F1-High (published) Patience expiry Common epoch 1a Baseline physics-anchored, hard-clamped 1 h 147 h Hard-clamped reference, single Δ value 0.50±0.020.50± 0.02 0.41 0.41 1b Random log-uniform (no anchor) — — Anchor irrelevant beyond bounds 0.52±0.010.52± 0.01 0.43 0.43 1c Overlap physics-anchored, wide spread 3 h 6 h Lane separation matters 0.52±0.010.52± 0.01 0.44 0.45 2a MSTE-aligned physics-anchored 1 h 6 h Slow anchor at MSTE 6 h scale (production) 0.57±0.010.57± 0.01 0.50 0.52 2b Mid-scale physics-anchored 1 h 48 h Intermediate slow anchor 0.51±0.020.51± 0.02 0.44 0.44 3b Unbounded physics-anchored, free upper bound 1 h 147 h (free) Optimiser-free slow timescale 0.48±0.010.48± 0.01 0.40 0.41 Table ED2: Period-wide walk-forward operational performance of the dual-pathway S4 model versus the engineered XGBoost benchmark under a matched selection protocol. Pooled across the 13 weekly walk-forward retrains over Oct–Dec 2024 (n=8,040n=8,040 evaluation timesteps) on the case-study site; both models select their stopping point on the same 7-day inner temporal probe, and the decision rule is the argmax of the three-class posterior. The S4 model leads on every High-class detection metric, on discrimination and on false alarms, contributing 38 additional true positives while producing 14 fewer false positives. The individual margins are modest and none reaches significance on a paired per-fold test over the 13 folds (F1-High p=0.110p=0.110); the paired High-class decision over all 8,040 predictions does (McNemar, 344 against 292, p=0.043p=0.043). The probabilistic terms separate from the detection terms: Brier-High marginally favours the S4 model, while the ranked probability score and log-loss favour the benchmark, reflecting that removing evaluation-block checkpoint selection leaves a better-ranked but less well-calibrated posterior (Results, “Walk-forward performance”). Whole-distribution metrics are reported for completeness and are an operational target of neither model. Weekly per-fold breakdown in Supplementary Table 5. Metric XGBoost S4 dual-pathway Δ High-class detection metrics (argmax of the three-class posterior) F1-High 0.501 0.533 +0.032+0.032 (+6.3%+6.3\%) Recall-High 0.575 0.618 +0.044+0.044 (+7.6%+7.6\%) Precision-High 0.445 0.468 +0.024+0.024 (+5.3%+5.3\%) False-alarm rate 0.087 0.085 −0.002-0.002 Discrimination AUC (High class) 0.877 0.898 +0.021+0.021 Probabilistic metrics Brier score (High class) 0.077 0.075 −0.003-0.003 Ranked probability score 0.102 0.139 +0.038+0.038 Log-loss 0.537 0.900 +0.362+0.362 Whole-distribution metrics (operational target of neither model) Accuracy 0.787 0.694 −0.092-0.092 Macro-F1 0.615 0.576 −0.039-0.039 Operational tallies True positives, High class 501 539 +38+38 False positives, total 626 612 −14-14 False negatives, High class 371 333 −38-38 Weekly F1-High wins 5/13 8/13 — Total evaluation timesteps n=8,040n=8,040 Table ED3: Headline performance of the meteorology-only multi-site Bayesian alert-tier classifier. All metrics evaluated on the synchronous Jan–Mar 2025 walk-forward window (n=8,536n=8,536 shared-grid 15-min timesteps; 13 weekly retrains per constituent nowcaster; Methods, “Multi-site Bayesian alert-tier classifier”). The predicted tier is constructed purely from meteorology: no raw target-species concentration enters the predicted pipeline at inference time. Confidence intervals are block bootstrap (B=1,000B=1,000 resamples, 24-h block length) on the synchronous timeseries to accommodate within-episode autocorrelation. Per-tier P/R/F1 values are paired with bootstrap 95% CIs in Supplementary Table S6; the full confusion matrix is in Supplementary Table S5. The same-architecture vote-count baseline replaces the Bayesian accumulator with a hard count of latched per-channel activations and holds all other pipeline stages fixed. The Δκw _w interval and p-value are computed from a paired block bootstrap, in which the same resampled blocks are used for both estimators so that shared fold-to-fold variation cancels. An unpaired bootstrap, which resamples the two κ values independently, approximately triples the interval width and is not appropriate for two estimators evaluated on identical timesteps. The majority-class baseline is reported because the tier distribution is dominated by the Normal state, so strict accuracy alone does not evidence skill; the unweighted κ is reported alongside κw _w because quadratic weighting materially raises the reported agreement when most errors are off-by-one tier. External validation against community odour complaints is fully held out: the complaint record was not used in training, hysteresis selection or LR design. The leave-one-week-out figures quantify the fusion-parameter tuning optimism directly (Δκw=0.006 _w=0.006); the constrained-ordering sensitivity is reported in Supplementary Note 6. Statistic Value Headline tier-match against the deterministic ground truth Quadratic-weighted Cohen’s kappa κw _w 0.709 [0.557, 0.852] Quadratic-weighted κw _w, leave-one-week-out refit 0.703 Complaint correlation, leave-one-week-out tier +0.721+0.721 Unweighted Cohen’s kappa κ 0.462 Strict element-wise accuracy 0.794 Majority-class baseline accuracy 0.800 Per-tier F1 (point estimates; full P/R/F1 with CIs in Supp. Table S6) Tier 0 (normal background) F1=0.912_1=0.912 Tier 1 F1=0.341_1=0.341 Tier 2 F1=0.366_1=0.366 Tier 3 F1=0.571_1=0.571 Comparison to a same-architecture vote-count baseline Vote-count baseline κw _w 0.639 [0.481, 0.792] Bayesian-fusion improvement Δκw _w +0.070+0.070 [0.025, 0.109] One-sided paired-bootstrap p (H0H_0: Bayesian ≤ baseline) 0.003 Predicted tier == (vote count − 1-\,1), fraction of timesteps 0.942 Robustness ±25%± 25\% LR perturbation: max |Δκw|| _w| (Supp. Table S4) 0.009 ±25%± 25\% LR perturbation: mean |Δκw|| _w| 0.002 External validation against daily community odour complaints (n=89n=89 days) Pearson r (daily-mean predicted tier vs daily complaints) +0.729 [+0.449, +0.841] Coefficient of determination R2R^2 0.532 Ground-truth tier ceiling (Pearson r) +0.793 [+0.307, +0.879] Share of ground-truth explained variance (R2R^2 ratio) 83% Paired difference from the ceiling (moving block) +0.072+0.072 [−0.303-0.303, +0.127+0.127] Lagged cross-correlation peak lag 00, r=+0.729r=+0.729 Constituent walk-forward classifiers (period-aggregate F1-High) H2S @ MMF9 (broad-coverage canary) 0.456 H2S @ MMF2 (tight-coverage confirmer) 0.623 CH4 @ MMF9 0.702 CH4 @ MMF2 0.659 Window and protocol Synchronous timesteps n 8,536 Weekly walk-forward folds 13 (Jan–Mar 2025) Per-channel walk-forward retrains 52 (13×413× 4 channels) Network prior base rate π 0.10 Tier cut-points (θ1,θ2,θ3)( _1, _2, _3) (0.15, 0.50, 0.92)(0.15,\,0.50,\,0.92) S1 Regulatory and public-health context Waste, landfills and fugitive gas emissions Global municipal solid waste generation reached an estimated 2.01 billion tonnes per year, representing approximately 0.74 kg per person per day on average, and is projected to rise to 3.40 billion tonnes by 2050 (60). Even under modern engineered designs that incorporate phased cell construction, liners, leachate extraction and capping, the long-term anaerobic decomposition of organic waste continues to release methane (typically 45–60% by volume), carbon dioxide (40–60%), water vapour, and a complex trace fraction including volatile organic compounds and reduced sulphur species, principally hydrogen sulphide (H2S) (1; 83). Landfill gas migration is nominally controlled through extraction wells supported by monitoring networks, but cap permeability, extraction rate, and the integrity of sub-cell containment together determine whether emissions remain on-site or migrate laterally to nearby receptors (75). Where these controls are incomplete, fugitive emissions to air pose persistent environmental and public-health risks for surrounding communities. Recent regulatory data indicate that the share of poor-performing regulated industrial sites in the waste sector is at the highest level on record, with such sites disproportionately implicated in pollution incidents and fugitive emissions (27; 29). Hydrogen sulphide: toxicology and odour H2S is generated in waste cells by sulphate-reducing bacteria converting sulphate (SO2−4_4^2-) under anaerobic conditions (66; 96; 97); in mixed waste streams the dominant sulphate source is gypsum (CaSO4) from construction and demolition debris (58; 4). Its odour threshold is cited as between approximately 0.75μgm−30.75\, \,m^-3 and 11μgm−311\, \,m^-3 (2), making it one of the most readily identifiable environmental odours, but prolonged or high-concentration exposure causes olfactory fatigue that masks levels potentially dangerous to health (42). Acute inhalation of low concentrations irritates the eyes and respiratory tract, producing sore throat, cough and breathing difficulty; sustained exposure prolongs these effects (42). At high concentrations (∼ 140mgm−3140~mg\,m^-3 and above), exposure may cause rapid collapse, respiratory paralysis, cyanosis, seizures, coma, cardiac arrhythmia, and death within minutes (42; 77). Chronic-exposure evidence from polluted communities and workplaces points to respiratory, ocular and neurological symptoms; H2S is not classified as a carcinogen by IARC (42). The 30-minute World Health Organization odour-annoyance guideline value (7μgm−37\, \,m^-3) is the relevant benchmark for community wellbeing, several orders of magnitude below acute-exposure thresholds (77). Odorous nuisance and the FIDOL framework Odour responses are highly variable between individuals and populations (93). In communities exposed to persistent odorous emissions, prolonged exposure has been associated with annoyance, insomnia, loss of appetite, unease, depression, headaches, sensory irritation, nausea and respiratory symptoms (41; 76; 93). Sub-irritant odorant levels can also trigger acute symptoms via non-toxicological mechanisms including innate odour aversion, exacerbation of underlying conditions, and stress-induced illness (93; 54). Surveys in Europe and North America consistently rank odours as the leading source of public complaints to regulatory agencies, with 13–20% of the European population reporting being bothered by environmental odours (67; 79; 57). The standard analytic framework for odour-impact assessment in the United Kingdom is FIDOL: Frequency, Intensity, Duration, Offensiveness and Location (79; 36). The Institute of Air Quality Management’s guidance translates these factors into the planning and permitting process. Whether a particular odour incident reaches the threshold of statutory nuisance under Part I of the Environmental Protection Act 1990 is determined by the local authority on the basis of sensory assessment, complaint analysis and contextual factors (30; 31; 44). Multi-stakeholder regulatory response Statutory nuisance investigations involving regulated industrial sources draw on a multi-stakeholder structure. Local-authority Environmental Health Officers undertake frontline assessment. Where odours originate from permitted industrial sites including landfills, sewage treatment works and waste facilities, the environmental regulator audits operator compliance with site-specific Odour Management Plans and the application of Best Available Techniques (45; 43). Public-health agencies provide assessment of community health implications, applying a tiered interpretation that distinguishes odour-related nuisance, short-term health considerations and long-term exposure risks against air-quality guideline values (WHO 24-h health-protection levels and 30-min annoyance thresholds) and against Acute Exposure Guideline Levels (AEGLs) used for emergency planning (99; 77). Where impacts become severe, formal multi-agency coordination escalates through Local Resilience Forum structures established under the Civil Contingencies Act, including Strategic and Tactical Coordinating Groups and Scientific and Technical Advisory Cells (31; 69). Public-health risk assessment for chemical incidents adopts a source–pathway–receptor (SPR) model: the source (here, the landfill gas), the pathway (atmospheric dispersion modulated by meteorology and topography), and the receptor (the surrounding residential population). Communication and decision-making within this framework necessarily occur after the receptor has been exposed. Forward-looking exposure estimation through dispersion modelling exists but relies on simplifying assumptions about atmospheric stability, emission flux and surface roughness that often do not match the highly non-linear, time-varying conditions found at real sites (9; 79; 98). Case study The case-study record analysed in the main text was generated at a European municipal landfill subject to a multi-year regulatory and public-health response. The site, regulated under an Environmental Permit, accepted predominantly non-hazardous waste with one cell containing stable non-reactive hazardous waste (gypsum and asbestos). From early 2021 onwards, community complaints rose substantially and 24-hour mean H2S concentrations approached but generally remained below the ATSDR Intermediate Minimal Risk Level (30μgm−330\, \,m^-3, 14–364 days) (99; 100). Critically, however, concentrations frequently exceeded the WHO 30-minute odour-annoyance guideline value (7μgm−37\, \,m^-3), an exceedance known to adversely affect wellbeing and contribute to headaches, irritation and sleep disturbance (99). Public-health agencies operating monthly risk assessments concluded throughout that period that rapid and sustained reduction at the source was required to protect community health and wellbeing; emergency-department syndromic surveillance showed no clear increase in attendances for physical conditions, but mental-health and wellbeing impacts were sufficient to prompt commissioning of additional support services (99; 100). A multi-agency Strategic Coordinating Group and a Scientific and Technical Advisory Cell were convened during the response. Site closure was issued by the regulator at the end of 2024. The analyses in the main text use post-calibration-adjustment monitoring data from the calendar year 2024 (Methods, “Data and study sites”). The site is treated anonymously throughout this manuscript; references to the relevant regulatory documentation provide the audit trail without re-naming the site in the prose (28; 100). S2 Extended environmental characterisation This Note expands the per-panel analyses summarised in main Fig. 1 with the full numerical detail referenced from the main Results. All statistics use the principal receptor record at MMF9 (n=34,609n=34,609 15-minute observations with valid H2S across 2024, from a complete grid of 35,136) unless otherwise stated. Source attribution by Conditional Probability Function (Fig. 1a) The exceedance threshold for the CPF was Cthr=5μgm−3C_thr=5\, \,m^-3, chosen to lie between the odour-detection threshold and the WHO 30-minute guideline value. The principal receptor (MMF9) shows a peak in exceedance probability in the westerly sector at peak bearing 265∘265 , with CPF≈0.37CPF≈ 0.37. Sector-averaged H2S concentrations were 11.2μgm−311.2\, \,m^-3 in the W (Cross-Right) sector and 10.6μgm−310.6\, \,m^-3 in the NW (Right-Up) sector, compared with 1.1μgm−31.1\, \,m^-3 in the opposing E (Cross-Left) sector, giving a peak/opposing sector relative-strength ratio of 4.09×4.09×. The circular–linear correlation between wind direction and H2S concentration was r=0.149r=0.149 (p<10−3p<10^-3, n=34,609n=34,609). The odds of observing a concentration spike (defined as >95>95th percentile, 15.3μgm−315.3\, \,m^-3) when wind direction aligned with the source bearing was elevated relative to background: Fisher’s exact odds ratio 1.431.43, p=2.9×10−4p=2.9× 10^-4. CPF cones from all three monitoring stations (MMF9, MMF1, MMF2) converged on a common source area, providing spatial triangulation of the emitting region. Chemical fingerprint: H2S–CH4 co-emission (Fig. 1b) Bivariate analysis of simultaneously measured H2S and CH4 at the principal receptor yielded: • Pearson correlation r=0.830r=0.830 (p<10−15p<10^-15) • Spearman rank correlation ρ=0.683ρ=0.683 (p<10−15p<10^-15) • Power-law exponent on log–log scale b=1.95b=1.95 for the relationship H2S=a⋅CH4bH_2S=a·CH_4^b • Jaccard spike co-occurrence J=0.654J=0.654 on top-5% spike sets • Spike co-occurrence count: 1,386 events vs. 87.7 expected under independence (χ2=21,298χ^2=21,298, p<10−15p<10^-15) The super-linear power-law exponent (b>1b>1) indicates that ambient H2S accumulates disproportionately as methane generation rises. At source, this is consistent with enhanced sulphate-reduction activity in deeper, more anaerobic waste cells where electron-donor competition favours sulphate-reducing bacteria. At the receptor, the ambient ratio is also modulated by joint dilution under stable conditions, differential biological oxidation in the cover soil, and gas-extraction performance, so the reported exponent is a downwind co-occurrence statistic and does not map one-to-one onto subsurface microbial kinetics. Negative controls To rule out non-gas-phase explanations for the H2S signal, two particulate-matter species were tested as negative controls: • PM10: r=−0.028r=-0.028, p<10−7p<10^-7 (no positive association; rules out shared mechanical resuspension) • PM2.5: r=−0.009r=-0.009, p=0.084p=0.084 (not significantly correlated) A moderate positive correlation with NOx (r=0.377r=0.377) was observed but is fully accounted for by joint nocturnal trapping rather than a shared source: the diurnal NOx profile peaks at approximately 08:00 LST whereas H2S peaks at 02:00 LST, a 6-hour phase offset that is incompatible with a common emission pathway. Multi-site meta-analysis Pooled across the four monitoring sites, the H2S–CH4 correlation gives rpooled=0.80r_pooled=0.80 (95% CI: 0.740.74–0.840.84, n=4n=4 sites). Heterogeneity is high (I2=99.8%I^2=99.8\%), reflecting site-specific factors including distance from the source area and local wind channelling. The pooled correlation comfortably satisfies the “Strong attribution” criterion of r>0.7r>0.7 and J>0.5J>0.5. Temperature–H2S diurnal hysteresis (Fig. 1c) The hourly-mean diurnal trajectory traces a counter-clockwise loop in the (TEMP, H2S) phase plane: • Normalised loop area A∗=0.369A^*=0.369 (classified “strong” hysteresis under A∗>0.30A^*>0.30) • Cooling-limb (18:00–06:00 LST) regression slope βcool=−2.83 _cool=-2.83 • Warming-limb (06:00–14:00 LST) regression slope βwarm=−2.45 _warm=-2.45 • Hysteresis-asymmetry index Hindex=0.072H_index=0.072 The corresponding wind-speed–H2S phase loop has negligible normalised area (A∗=0.012A^*=0.012, classified “no hysteresis”), confirming that wind-speed dilution acts effectively without memory at the diurnal timescale, consistent with the quasi-steady-state assumption of Gaussian-plume formulations. Diurnal concentration profile (Fig. 1d) The mean diurnal profile yields: • Peak at 02:00 LST: mean 9.69μgm−39.69\, \,m^-3, P95 51.6μgm−351.6\, \,m^-3 • Trough at 14:00–15:00 LST: mean 1.31μgm−31.31\, \,m^-3, P95 4.2μgm−34.2\, \,m^-3 • Mean peak/trough ratio 7.4×7.4×; P95 ratio 12.3×12.3× • Hourly variation: Kruskal–Wallis H=338.5H=338.5, p=7.0×10−58p=7.0× 10^-58 Approximate Pasquill–Gifford stability classification (Methods, based on measured wind speed and a time-of-day/year solar insolation proxy in the absence of pyranometer data) yielded unstable conditions (class A–B) in 40.0%40.0\% of observations and stable conditions (class E–F) in 26.1%26.1\%. The seasonal cycle further amplifies the diurnal effect: the February monthly peak reaches 14.97μgm−314.97\, \,m^-3 versus an August minimum of 0.99μgm−30.99\, \,m^-3, a ratio of 15.1×15.1× (Kruskal–Wallis p<10−15p<10^-15). Sunrise-aligned inversion break-up (Fig. 1e) Aligning every calendar day to astronomical sunrise (computed via the astral library at the site latitude/longitude) and aggregating observations within Δt∈[−2,+6] t∈[-2,+6] hours around sunrise: • Pre-sunrise baseline mean (Δt∈[−2,0] t∈[-2,0] h, n=2,920n=2,920): 9.49μgm−39.49\, \,m^-3 • Maximum hourly concentration at Δt=−2 t=-2 h: 10.42μgm−310.42\, \,m^-3 • +1+1 h post-sunrise: 2.65μgm−32.65\, \,m^-3 • +2+2 h: 1.54μgm−31.54\, \,m^-3 • +4+4 h floor: 1.30μgm−31.30\, \,m^-3 • Pre-/post-sunrise reduction: 7.3×7.3× • One-sided Mann–Whitney U=1.15×107U=1.15× 10^7, p<10−15p<10^-15 The reproducibility of this pattern across all 365 calendar days identifies it as a robust climatologically driven phenomenon rather than an episodic feature. Bivariate mechanism scatter (Fig. 1f) Linear regressions of H2S against three meteorological drivers (n=34,609n=34,609, all p<10−3p<10^-3): • Wind speed: r=−0.122r=-0.122, slope =−1.35μgm−3=-1.35\, \,m^-3 per m s-1 • Inverse wind speed: r=0.141r=0.141, slope =10.75μgm−3=10.75\, \,m^-3 per (m s-1)-1, 95% CI 9.49.4–12.112.1 • Air temperature: r=−0.173r=-0.173, slope =−0.78μgm−3=-0.78\, \,m^-3 per ∘C • Temperature tendency dT/dtdT/dt: r=−0.143r=-0.143, slope =−6.33μgm−3=-6.33\, \,m^-3 per ∘C hr-1 Mean concentration in low-wind conditions (u<1ms−1u<1\,m\,s^-1) was 11.8μgm−311.8\, \,m^-3 versus 1.8μgm−31.8\, \,m^-3 at u>5ms−1u>5\,m\,s^-1, a factor of 6.5×6.5×. Spike events occurred under significantly lower wind speeds than non-spike samples (1.751.75 vs 3.67ms−13.67\,m\,s^-1, Mann–Whitney p<10−15p<10^-15). Peak cooling rates reached approximately −0.6∘Chr−1-0.6\, C\,hr^-1, corresponding to the late-evening period of maximum inversion formation. Although individual mechanism R2R^2 values are small (0.0150.015–0.0300.030), reflecting the limited explanatory power of any single linear driver in isolation, extreme concentration spikes (>100μgm−3>100\, \,m^-3) cluster strongly in the low-wind, low-temperature, negative-tendency region of predictor space corresponding to Pasquill class E–F stable boundary layers. Summary table reference The complete numerical summary across all six panels is given in Supplementary Table 1. S3 Per-season interpretation of the XGBoost ALE feature hierarchy This Note expands the seasonal interpretation summarised in main Fig. 4 and the Results section. The four seasonal classifiers expose a coherent seasonal narrative connecting model structure to atmospheric physics. Note on revision. The per-season interpretation below has been rewritten. In the first version of this work these four subsections rested on the engineered temperature tendency, which was then the leading feature in every season. That feature was computed over a centred window and carried approximately three hours of lookahead (Methods, “Ensemble classifier-tree seasonal nowcaster”); correcting the window to be strictly causal removes it from the leading group in all four seasons, and the interpretation changes with it. Importances below are class-mean values from the corrected classifiers. Winter (January–March) Winter attributions are the smallest of the four seasons in absolute terms, with no feature exceeding I¯=0.04 I=0.04. Diurnal timing leads (hour_coshour\_cos, I¯=0.037 I=0.037), followed by wind direction (WD_cosWD\_cos, I¯=0.012 I=0.012); the 2-hour stagnation index contributes little (I¯=0.001 I=0.001) and the temperature tendency is negligible (I¯<0.001 I<0.001, twentieth of twenty-four features). The flatness of the whole ranking is itself the finding: during short winter days the nocturnal inversion is close to the default state rather than an episodic departure from it, so no single meteorological discriminator separates exceedance from background as sharply as in the transitional seasons. Spring (April–June) Spring shows the most evenly distributed attribution. Diurnal timing and stagnation are almost equally weighted (hour_coshour\_cos, I¯=0.146 I=0.146; stagnation_2hstagnation\_2h, I¯=0.130 I=0.130), with both wind-direction components following (WD_cosWD\_cos 0.0640.064, WD_sinWD\_sin 0.0560.056). Stagnation reaching near-parity with diurnal timing is specific to this season: spring months are generally well ventilated, so a stagnation episode is a sharp departure from the seasonal norm rather than the prevailing condition, and carries correspondingly more information. The strong L1 regularisation selected by hyperopt (α=5.65α=5.65) is consistent with the need to choose among several comparably informative signals in a meteorologically variable transitional season. Summer (July–September) Summer is the only season in which a directly measured meteorological variable leads: wind speed is the strongest feature (I¯=0.058 I=0.058), ahead of diurnal timing (0.0340.034) and the 2-hour stagnation index (0.0130.013). This is the expected signature of a convective regime. Daytime mixing suppresses accumulation, so exceedances are predominantly wind-stagnation events against a low-concentration baseline, and mechanical dilution rather than stability timing becomes the discriminating quantity. The deep tree depth selected by hyperopt (max_depth = 7) allows the model to represent the non-linear interaction between convective mixing and wind variability. Autumn (October–December) Autumn carries by far the largest single attribution of any season: diurnal timing at I¯=0.333 I=0.333, roughly an order of magnitude above the next feature (stagnation_2hstagnation\_2h, 0.0370.037), with wind direction and wind speed following (0.0200.020, 0.0120.012). This is the quarter in which daylight shortens fastest and nocturnal stable boundary-layer duration expands correspondingly, so time of day becomes a close proxy for stability state. It is also the evaluation quarter for the walk-forward analysis of main Fig. 6, which is worth noting when comparing the two: the season in which the engineered benchmark leans hardest on a calendar encoding is the season on which it is benchmarked. Cross-seasonal consistency with the causal hierarchy Two patterns hold across the four seasons. Diurnal timing leads in Winter, Spring and Autumn, and wind speed leads in Summer, consistent with time of day acting as a proxy for stability state wherever a nocturnal inversion regime is present, and being displaced by mechanical dilution once convective mixing dominates. Wind speed and wind direction, two of the three drivers the multiscale transfer entropy analysis identifies as the causal core (main Fig. 2), appear in the leading group of every season. Raw atmospheric pressure has zero ALE importance in all four, its synoptic content being carried by the stagnation index. The convergence is partial and should be read as such. The largest single ALE effect in three of four seasons is a calendar encoding rather than a measured driver, and calendar terms are not in the causal analysis driver set at all. What the two methods agree on is that wind transport carries the exceedance signal; they do not agree on a ranking, and the earlier claim that the ALE hierarchy mirrors the transfer-entropy hierarchy is not supported. Why the two rankings differ. The difference is largely structural rather than substantive, and is predictable from how the two quantities are defined. First, most of the apparent disagreement is non-measurement. Sixteen of the twenty-three features entering the ALE analysis are not in the transfer-entropy driver set, so no transfer-entropy ranking exists for them. These are the calendar encodings, both stagnation indices, the recirculation index and the directional components other than WD_sinWD\_sin. Their absence from that hierarchy is not a contradicting verdict. Second, on the seven features the two analyses share, they agree on the leading pair: restricted to those seven, ALE ranks wind speed first (I¯=0.029 I=0.029) and WD_sinWD\_sin second (0.0170.017), which are two of the three drivers the transfer-entropy analysis identifies as its core. The remaining five, including all three tendency variables, fall below I¯=0.002 I=0.002. Third, the one genuine discrepancy among the shared features is atmospheric pressure, which the transfer-entropy analysis ranks in its core and to which ALE assigns exactly zero. The mechanism is stated above: the engineered feature set contains a stagnation index that already encodes pressure’s synoptic content, so the classifier has a cheaper route to the same information and never splits on the raw channel. This is a property of the feature set, not a disagreement about the atmosphere. Fourth, the diurnal difference follows directly from the estimators’ definitions. Transfer entropy is computed conditional on the target’s own recent history, so any structure in H2S predictable from its own past, which the diurnal cycle largely is, is removed by construction before a driver is credited. Accumulated local effects apply no such conditioning, so a calendar term is credited with the full diurnal pattern. The day-multiple surrogate analysis quantifies this directly: only 2.7%2.7\% of the effective transfer entropy is attributable to shared diurnal structure, precisely because the conditioning has already absorbed it. Two methods that agreed on every feature would not constitute independent evidence. These agree where they measure the same quantity, and differ where their assumptions differ, in the direction those assumptions predict. S4 Extended comparison with prior literature This Note expands the discussion in the main text by situating each of the principal contributions against the most directly comparable published work. References cited here are the high-impact, narrative- relevant subset; the full reference list is in the main bibliography. Causal hierarchy versus mobile-laboratory and cover-type surveys Most published landfill H2S investigations report days-to-weeks of co-located gas and meteorology. 14 reports a mobile-laboratory survey at multiple US sites, with H2S/CH4 ratios varying by three orders of magnitude between sites; 13 quantifies 82 distinct landfill gases across 31 cover types in California. The case-study record analysed here, regulated under continuous environmental and public-health oversight, provides the multi-season temporal coverage required to test causal hypotheses about meteorological drivers. The MSTE result (main Fig. 2) maps directly onto the four canonical mechanisms established for landfill H2S transport: advection sets which receptor is exposed; mechanical dilution sets instantaneous concentration via the inverse-wind-speed dependence of the Gaussian plume (98; 85); barometric pumping modulates source flux through the multi-hour soil–atmosphere pressure equilibration timescale (86; 106); and the diurnal cycle of boundary-layer mixing height controls receptor exposure through the depth of the dilution volume (94; 81). The MSTE result that dP/dtdP/dt becomes causally significant precisely at τ≥2τ≥ 2 h is the information-theoretic counterpart of the controlled vadose-zone experiments showing that sudden barometric drops can trigger 20-fold gas breakthroughs over <24<24 h (35). The strong counter-clockwise temperature–H2S hysteresis (A∗=0.369A^*=0.369) reproduces, at the receptor, the multiscale low-pass behaviour directly measured for soil–gas exchange by independent groups (84). The instantaneous weakness of temperature as a causal driver, contrasted with cross-city pollutant– meteorology coupling studies that report ambient temperature leading pollutants at sub-hourly lags (6), identifies the case-study site as a nocturnal-accumulation-dominated regime in which solar heating matters not as an instantaneous variable but as the integrated forcing dictating the timing of inversion erosion. The H2S–CH4 co-emission with super-linear power-law exponent b=1.95b=1.95 sits comfortably within the inter-site range reported by mobile-laboratory surveys (14) and is mechanistically consistent with sulphate-reducing-bacteria competition with methanogens in sulphate-enriched waste cells. Physics-anchored architectural prior versus data-agnostic SSM initialisation Most published applications of structured state space models to environmental and physical time series use uniform log-Δ initialisation (47; 49). Recent extensions handle missing values with masking and adaptive temporal prototypes, or introduce input-dependent selectivity through Mamba’s parallel-scan algorithm (46), but in every case the timescale prior is data-agnostic. The dual-pathway MSTE-anchored S4 nowcaster developed here makes a different choice by instantiating the SSM kernel timescales on the causally significant horizons identified by transfer entropy on the same record. The empirical justification is the six-condition ablation (main text and Extended Data Table ED1): swept across two orders of magnitude of slow-anchor timescales (6, 48, 147 h), the optimum lies within the band in which atmospheric pressure remains causally significant in the MSTE analysis, and outperforms every alternative initialisation tested (q≤0.048q≤ 0.048). The sweep did not test anchors between 6 and 48 h, so the optimum is identified within the tested set rather than located within the band. This quantitative bridge between information-theoretic causal analysis (91; 40) and learnable architectural priors has not previously been instantiated for environmental nowcasting. The model reproduces the marginal response curves for dP/dtdP/dt and dT/dtdT/dt (main Fig. 5d,e) without these derivatives ever appearing as input features, a behaviour that engineered XGBoost pipelines must obtain by hand-coding Butterworth derivatives. The lane-resolved association is weak (ρs=+0.14 _s=+0.14) and this is therefore reported as consistency with the joint synoptic signature rather than as evidence that the signature is encoded in either pathway individually. This addresses a long-standing limitation of regulatory dispersion models: AERMOD and other Gaussian frameworks assume steady-state, well-mixed conditions, and large-eddy simulation is operationally infeasible for real-time exposure classification, leaving learnt nowcasters as a viable third path. Calibration to community impact versus hard threshold detection Two prior strands of work bracket the community-impact validation reported in main Fig. 6c–d. 74 demonstrated that random-forest classifiers operating on combined H2S sensor, weather and operational features could predict odour complaints at a wastewater reservoir, but reported that H2S sensor data alone were insufficient. Commercial decision-support platforms now provide reverse-trajectory dispersion modelling and weather-driven operational dashboards for landfill operators, but these are positioned as engineering tools rather than calibrated to community-wellbeing endpoints. The result reported in the main text, that a meteorology-only causal nowcaster’s internal probability tracks an independent community-wellbeing endpoint at 15-minute resolution across a full quarter (r=0.585r=0.585 lag-0, n=88n=88 days), closes that gap by demonstrating that machine-learning probability outputs can be made directly interpretable in terms of the FIDOL framework that already underpins UK statutory-nuisance assessment (79). The Dominguez Channel emergency in Carson, California, where chronic H2S exposure produced documented surges in both somatic and psychological morbidity (89), and the documented public-health impacts at the case-study site (100), together establish that the operational target for such systems is not the regulatory acute-exposure threshold but the much lower 30-minute WHO odour-annoyance guideline value at which community impact is already substantial. Operational pathway and limitations Three concrete actions are required to convert physics-anchored nowcasting into operational community protection. First, the regulatory monitoring infrastructure must be independently auditable: the present study deliberately uses only the post-calibration-adjusted record (28), and no machine-learning system can compensate for systematic measurement error in its training labels. Second, the causal hierarchy is regime-specific and must be re-estimated, not assumed, at each new site: the three orders of magnitude variability in H2S/CH4 ratios across landfills (14) and the cover-type-mediated emission diversity documented in California (13) both argue against transferring the causal hierarchy without verification. Third, the model’s strict causality makes the conversion to genuine forecasting architecturally trivial (substitute numerical-weather-prediction meteorology at inference), but the operational pipeline is not yet built; the demonstrated 20–40% peak-pollutant reductions and measurable morbidity benefits achievable through well-calibrated short-term alerts in other settings define a realistic upper bound for fugitive landfill emissions if dynamic gas-extraction control is co-deployed with anticipatory meteorology. The longer-term scientific opportunity is to couple the model’s continuous probability output to syndromic surveillance and primary-care morbidity coding, moving the evidence base from complaint correlation to clinical endpoints. S5 MSTE estimator ablation study To assess the sensitivity of the multiscale transfer entropy results in main Fig. 2 to methodological choices in the KSG–Theiler estimator, a six-configuration ablation study was performed. Each configuration modifies one or more hyperparameters relative to the preceding configuration, isolating the contribution of each methodological refinement. The configurations are summarised in Supplementary Table 7. All six were evaluated on the identical non-stationarised 15-minute series under the bounded-fill policy of Supplementary Methods, “Temporal coarse-graining” using the same coarse-graining scales, Theiler window (W=4W=4), rank transform, and circular-shift surrogates (B=10B=10). The surrogate count for this ablation was previously stated as B=20B=20. Configuration 4 has since been matched cell-by-cell against the stored results of the production run (16 of 16 driver–scale cells agreeing to six significant figures), establishing that Configuration 4 is that run; and that run reproduces its surrogate mean exactly at B=10B=10 and at no other value. The ablation therefore ran at ten surrogates, not twenty. This ablation ran under the plain convention p=#b:TEb≥TEobs/Bp=\#\b:TE_b _obs\/B, so at B=10B=10 its resolution is Δp=0.1 p=0.1 and a marked cell is one in which zero of ten surrogates exceeded the observed value. That is a coarse screen and is reported as one: it cannot support false-discovery correction across a family of this size, and the significance underpinning main Fig. 2 comes from the production analysis at B=2,000B=2,000 under the add-one convention (Supplementary Methods, “Surrogate testing”), not from this table. The ablation compares point estimates and the stability of which cells clear a fixed screen across estimator settings; it is not an independent significance analysis. Supplementary Table 8 reports the effective transfer entropy (ETE) for all seven drivers across the six temporal scales under each configuration. Cells in which zero of ten surrogates exceeded the observed value are marked with an asterisk. Supplementary Table 9 gives the aggregate comparison across all 70 driver–scale combinations. Degeneracy at coarse scales. Configurations 3–5 differ only in the cap hmaxh_ placed on the history embedding depth (4, 2 and 1 respectively). The depth is set adaptively as h=max(1,min(hmax,⌊3h/τ⌋))h= (1, (h_ , 3\,h/τ )), so for τ≥2τ≥ 2 h the adaptive term is already 1 and no cap can bind: configurations 3, 4 and 5 then coincide exactly. The ablation therefore distinguishes six configurations at the 15 min and 1 h scales but only four at 2 h and coarser: 1\1\, 2\2\, 3,4,5\3,4,5\ and 6\6\. Of those four, two use a fixed k=10k=10 rather than the adaptive rule, so the number of genuinely independent estimator checks at the coarse scales that carry the architectural claim is smaller than the six-configuration design implies. This is stated because the identical ETE values in those cells would otherwise appear to indicate a failed or duplicated run; they do not, and no configuration went unrun. Three principal findings (i) Core drivers are robust to all methodological choices. Wind direction, wind speed and atmospheric pressure clear the surrogate screen under all six configurations at scales ≤3h≤ 3\,h, confirming that these causal relationships are genuine features of the data rather than artefacts of any particular estimator setting. Temperature clears the screen only at the 6h6\,h scale across all configurations except the baseline. The baseline additionally marks temperature at 15min15\,min and 30min30\,min, which at h=4h=4 is consistent with estimator bias from the high-dimensional conditioning space. This ablation and the production analysis of Fig. 2 are not directly comparable on marginal cells: the ablation ran at B=10B=10, whose 0.10.1 resolution cannot adjudicate a cell whose raw p is of order 0.020.02, whereas the production run at B=2,000B=2,000 resolves it. Where the two differ on temperature, the production analysis is the reported result. (i) Bivariate transfer entropy (Configuration 6) inflates effect sizes and false-positive rates. Without conditioning on confounders, the bivariate estimator conflates genuine causal influence with shared meteorological forcing. The most evident case is dP/dtdP/dt, which appears strongly significant at all scales under Configuration 6 (ETE=0.6ETE=0.6–19.5×10−3nats19.5× 10^-3\,nats) but only at ≥2h≥ 2\,h under conditional configurations, consistent with barometric pumping operating on synoptic rather than sub-hourly timescales. This demonstrates the necessity of conditional transfer entropy with at least one confounder for physically interpretable causal attribution in this multivariate system. (i) History embedding depth is the dominant sensitivity axis. The transition from h≤4h≤ 4 (Configurations 1–3) to h≤2h≤ 2 (Configuration 4) at fine scales (τ≤1hτ≤ 1\,h) eliminates the spurious temperature significance while preserving all genuine causal links. With h=4h=4 and one confounder, the joint embedding space has dimensionality d=4+1+1=6d=4+1+1=6 for the source–target–conditioning triplet (total 77 dimensions including the target future), causing the Chebyshev k-nearest-neighbour balls to become sparse and the digamma-based CMI estimate to degrade. The ε -shrinkage correction (Configuration 2) and adaptive k (Configuration 3) partially compensate, but the most effective remedy is reducing the embedding dimension directly. With h≤2h≤ 2, the maximum conditioning dimension is d=2+1=3d=2+1=3 (total 5D joint space), within the regime where the KSG estimator retains adequate sensitivity. Justification of the production configuration Configuration 4 was selected as the production configuration for main Fig. 2 because it occupies the optimal point on the bias–variance frontier: it includes sufficient target history (h=2h=2 at fine scales, capturing 3030–120min120\,min of autoregressive memory) and one confounder to isolate unique information transfer, while keeping the total conditioning dimension low enough (d≤3d≤ 3) for reliable KSG estimation. Configuration 5 (h=1h=1) sacrifices temporal context at the 1515–60min60\,min scales where the autoregressive memory of H2S extends beyond a single lag, and Configuration 6 (bivariate) abandons confounding control entirely. Configurations 1–3 retain excessive embedding depth that degrades estimator performance at fine scales, as evidenced by the spurious significance of temperature at 15min15\,min in Supplementary Table 8. Dependence between tests The corrections applied to this family assume either independence or positive regression dependence, and the tests are dependent by construction: the scales are nested coarse-grainings of one series and the three tendency variables are deterministic transforms of drivers already in the family. The reported significance is Benjamini–Hochberg, as pre-specified. A Benjamini–Yekutieli correction, valid under arbitrary dependence, was computed alongside it as a sensitivity analysis and is encoded on Fig. 2 as marker fill: 20 of the 29 significant cells survive it, including all six wind-direction and all seven pressure cells. Of the nine that do not, seven are tendency variables which the main text does not interpret; the other two are wind speed at 12 h and temperature at 15 min (Methods, “Causal analysis”). Split-half stability of the recovered hierarchy The six-configuration ablation above varies the estimator settings while holding the sample fixed. The complementary test, holding the settings fixed and varying the sample, was run separately by split-half resampling, because a hierarchy robust to every estimator choice could still be a property of one particular twelve months. The analysis was repeated independently on January–June and July–December 2024 for the three directly measured drivers across the seven informative scales (21 driver–scale cells per half, 42 in total), inheriting the production specification without modification: B=2,000B=2,000 circular-shift surrogates under the add-one convention, block completeness 0.500.50, Theiler window W=4W=4, one confounder selected over driver variables, adaptive k, and causal derivatives. Benjamini–Hochberg correction was applied within each half over its own 21 cells. The only difference from the production run is the sample. Driver H1 (Jan–Jun) H2 (Jul–Dec) Full year WDsinWD_ 0.03706 0.02713 0.03244 Pressure 0.02012 0.02269 0.02444 WS 0.01516 0.00860 0.01753 Mean effective transfer entropy across the seven informative scales. The ordering WDsin>Pressure>WSWD_ >Pressure>WS is identical in both halves and in the full-year analysis. Supporting agreement: the sign of the effective transfer entropy matches in 19 of 21 cells; Spearman rank correlation between halves is 0.6230.623, and between each half and the full year 0.6710.671 (H1) and 0.8140.814 (H2). Benjamini–Hochberg significance is retained by 20 of 21 cells in H1 and 18 of 21 in H2, against 19 of 21 for the full year. That last comparison is the more informative one: each half carries half the data and therefore a wider surrogate null, so cells sitting close to the threshold should have dropped out on power alone. They largely did not. Two cells disagree in sign, both wind speed and both at coarse scales: +0.02387+0.02387 against −0.00231-0.00231 at 6 h, and +0.01692+0.01692 against −0.00370-0.00370 at 12 h, where the second half retains 714 and 358 coarse-grained blocks respectively (727 and 363 in the first). Both second-half estimates are consistent with zero. Wind speed ranks third in both halves regardless, so the ordering is unaffected, but the wind-speed signal at coarse scales is the least robust element of the hierarchy. Pressure and wind direction agree in sign at every scale in both halves. Two limitations attach. The halves differ seasonally rather than being randomly interleaved, so a disagreement would have been ambiguous between estimator instability and genuine seasonal variation in transport; the test could therefore only confirm the hierarchy, not refute it, and a pass under that confound is correspondingly stronger than a pass on interleaved data. More fundamentally, split-half resampling cannot detect a bias the estimator applies uniformly to both halves. The full-year estimate exceeds both half-year estimates for pressure and wind speed, which is the behaviour expected of a sample-size-dependent estimator bias and is a further reason to read the magnitudes as ordinal rather than absolute. Validation here is by stability of the recovered hierarchy, not by agreement with a second estimator. S6 Joint site analysis This Note expands the joint multi-site, multi-species analysis summarised in main Fig. 7. Four independently trained walk-forward nowcasters (H2S at MMF9, H2S at MMF2, CH4 at MMF9, CH4 at MMF2; each instantiated with the S4 dual-pathway architecture of Supplementary Methods, “S4D state space model” and the optimiser schedule of Supplementary Methods, “Optimiser schedule”) feed a Bayesian log-odds evidence accumulator that emits a continuous P(network event)P(network event) and an ordinal tier in 0,1,2,3\0,1,2,3\ at the native 15-minute cadence. The four-tier scheme is designed to be operationally interpretable in the public-health framework of Supplementary Note 1: each tier corresponds to a distinct, escalating layer of the source–pathway–receptor response already in regulatory use, from sub-annoyance background through short-term wellbeing exceedance to the ATSDR Intermediate Minimal Risk Level invoked in chemical incident management. Ordinal alert tier definitions The classifier produces four sequential ordinal states. The cut-points on the continuous posterior probability P(network event)P(network event) have been chosen to align the path τ=0→1→2→3τ=0→ 1→ 2→ 3 with the operational decision landscape of Note 1 along a single monotone variable; the ordinal property is mathematically guaranteed because the thresholds are sequential on the same continuous quantity. Tier 0 (normal background, P<θ1P< _1). No coordinated multi-channel evidence above prior. Routine permit monitoring under best available techniques; no public-health escalation. Tier 1 (θ1≤P<θ2 _1≤ P< _2). The fused posterior has crossed a sensitive activation threshold but the multi-channel evidence is not yet corroborated. Operationally this aligns with the FIDOL frequency–intensity–duration regime of the World Health Organization 30-minute odour-annoyance guideline value (7μgm−37\, \,m^-3, Supplementary Note 1), at which community wellbeing impacts (headache, irritation, sleep disturbance) become detectable but acute health risk remains low. Tier 1 is the appropriate level at which the operator’s Odour Management Plan should be checked and gas-extraction performance verified. Tier 2 (θ2≤P<θ3 _2≤ P< _3). Multiple independent channels (across both target species and both receptor sites) corroborate the elevated state. Operationally this is the regime in which the regulatory Environmental Health Officer would be expected to confirm a statutory-nuisance assessment under Part I of the Environmental Protection Act 1990, and at which short-term mitigation actions (extraction-rate increase, surface flux survey, public-information notice) are warranted. Tier 3 (P≥θ3P≥ _3). The fused posterior is at or near the maximum value attainable under all four channels concurrently active, indicating coordinated multi-site exceedance. The predicted tier is a strict function of the fused posterior: no raw target-species concentration enters the predicted pipeline at inference, so escalation remains meteorology-only. Raw concentrations appear solely in the independent rules-based ground-truth tier of Section S6, against which the predicted tier is scored; that comparator is never an input to the prediction and likewise carries no override. This tier is the information-theoretic counterpart of the regulatory threshold at which formal multi-agency coordination escalates under Local Resilience Forum structures, with convening of Strategic and Tactical Coordinating Groups and a Scientific and Technical Advisory Cell where impacts become severe. Hysteresis debouncing. Each channel’s binary state is passed through a two-parameter latch with asymmetric onset/clearance counts (Non,Noff)(N_on,N_off). The latch opens after Non=3N_on=3 consecutive ticks above threshold (45 min) and closes after Noff=2N_off=2 consecutive ticks below threshold (30 min). The asymmetry is deliberate: a longer onset suppresses single-spike voltage transients and short-lived turbulent fluctuations that do not represent sustained exposure, while a shorter clearance prevents an episode from being prematurely declared resolved during a brief lull in advection. The numerical values map directly onto the WHO 30-minute averaging convention: 30 min is the shortest interval over which odour annoyance is meaningfully assessed (the minimum clearance window), and 45 min is the shortest interval that places a multi-tick event securely outside this averaging window (the activation window). The full parameter set is given in Supplementary Table S1. Algorithm 1 Ordinal alert tier classifier (Tier-Recipe). Revised: the raw-sensor extreme override has been removed from the inputs, the debounced override computation and the Tier 3 condition. The predicted tier is a strict function of the fused posterior. 1: Per-channel walk-forward probabilities pc,tp_c,t for c∈c with =H2S–MMF9,H2S–MMF2,CH4–MMF9,CH4–MMF2C=\H$_2$S--MMF9,\,H$_2$S--MMF2,\,CH$_4$--MMF9,\,CH$_4$--MMF2\; prior π and cut-points 0<θ1<θ2<θ3<10< _1< _2< _3<1; per-channel active/quiet likelihood ratios LRcon,LRcoffLR^on_c,LR^off_c; hysteresis counts (Non,Noff)(N_on,N_off); activation threshold p⋆p 2: Ordinal tier τt∈0,1,2,3 _t∈\0,1,2,3\ at every grid step t 3: for each channel c in C do 4: ac,t←(pc,t>p⋆)a_c,t 1(p_c,t>p ) ⊳ raw per-step activation 5: ℓc,t←Hysteresis(ac,t,Non,Noff) _c,t← Hysteresis(a_c,t;\,N_on,N_off) ⊳ latched state 6: mc,t←(pc,tisNaN)m_c,t 1(p_c,t\ is\ NaN) ⊳ outage mask 7: end for 8: for each grid step t do 9: logOt←log(π/(1−π)) O_t← (π/(1-π)) ⊳ Eq. (S3) 10: for each channel c in C do 11: if mc,t=1m_c,t=1 then 12: LRc,t←1LR_c,t← 1 ⊳ NaN-neutrality 13: else if ℓc,t=1 _c,t=1 then 14: LRc,t←LRconLR_c,t ^on_c 15: else 16: LRc,t←LRcoffLR_c,t ^off_c ⊳ CH4 off =1=1; H2S off <1<1 17: end if 18: logOt←logOt+logLRc,t O_t← O_t+ _c,t 19: end for 20: Pt←Ot/(1+Ot)P_t← O_t/(1+O_t) ⊳ Eq. (S6) 21: end for 22: for each grid step t do 23: τt←0 _t← 0 24: if Pt≥θ1P_t≥ _1 then 25: τt←1 _t← 1 26: end if 27: if Pt≥θ2P_t≥ _2 then 28: τt←2 _t← 2 29: end if 30: if Pt≥θ3P_t≥ _3 then 31: τt←3 _t← 3 32: end if 33: end for 34: return τt\ _t\ Table S1: Ordinal alert tier classifier parameters and the operational interpretation attached to each. The concentration thresholds follow from the regulatory and physiological levels reviewed in Supplementary Note 1. The hysteresis counts, the prior π, the per-channel activation threshold p⋆p and the likelihood ratios were grid-search parameters, and the tier cut-points are operational choices that the documented search does not reproduce; the third column below states how each value is interpreted in operation, not how it was derived. Provenance for each is given in “Bayesian classifier parameter set” below, and the ±25%± 25\% likelihood-ratio sensitivity analysis is in Supplementary Table S4. The latched-state hysteresis counts (Non,Noff)=(3,2)(N_on,N_off)=(3,2) correspond to (45 min onset, 30 min clearance) at the 15-minute grid. No raw-sensor concentration enters the predicted tier; the classifier is meteorology-only at inference time. Symbol Value Operational interpretation CthrH2SC^H_2S_thr 7.0μgm−37.0\, \,m^-3 WHO 30-min odour-annoyance guideline value (Note 1); ground-truth activation threshold for each H2S channel and the training boundary that defined the High class for the H2S CAIRN nowcasters CthrCH4C^CH_4_thr 2.756ppm2.756\,ppm Linear High-class boundary from H2S=5.55CH4−8.29H_2S=5.55\,CH_4-8.29 (Supplementary Table 6); ground-truth activation threshold for each CH4 channel and the training boundary that defined the High class for the CH4 CAIRN nowcasters p⋆p 0.500.50 Grid-search parameter (range 0.300.30–0.500.50); binarises model probability pc,tp_c,t before the hysteresis latch NonN_on 33 ticks (45 min) Minimum sustained-event window: one full 30-min WHO averaging interval plus a 15-min safety margin against single-tick voltage transients NoffN_off 22 ticks (30 min) Matches the WHO 30-min averaging window so the latch will not declare an episode resolved within a single regulatory averaging interval π 0.100.10 Grid-search parameter (range 0.030.03–0.100.10); at the selected value the classifier reverts to the observed network-wide High-class rate when all channels are missing, i.e. is neutrally calibrated at zero information θ1 _1 0.150.15 Tier 1 trigger: single-channel sensitivity regime aligned with FIDOL nuisance assessment θ2 _2 0.500.50 Tier 2 trigger: multi-channel corroboration warranting statutory-nuisance confirmation; the fused posterior cannot exceed this value from any single channel alone under the LR set in Supplementary Table S3) θ3 _3 0.920.92 Tier 3 trigger: just below the analytical posterior ceiling Pmax=0.949P_ =0.949 under all four channels concurrently active and no missing data; Tier 3 is therefore reached only by fully coordinated multi-channel meteorological evidence) Physical relevance of the parameter set The classifier parameters have been chosen to map onto independently established physical, physiological or regulatory anchors wherever such anchors exist, with the soft posterior cut-points calibrated to align the four-tier decision path with the operational decision landscape of Note 1. Activation concentrations CthrH2SC^H_2S_thr and CthrCH4C^CH_4_thr. The H2S activation threshold is the WHO 30-minute odour-annoyance guideline value (103), the benchmark identified in Note 1 as the relevant level for community wellbeing. The CH4 threshold is its co-emission equivalent under the linear mapping H2S=5.55CH4−8.29H_2S=5.55\,CH_4-8.29 established for the case-study site (Supplementary Table 6), reflecting the strong sulphate-reducing-bacteria coupling identified by the chemical fingerprint analysis (Supplementary Note 2): activating both species at concentrations corresponding to the same underlying release process avoids over-counting jointly informative evidence. At the principal receptor the mapped CH4 classes hold 78.0%78.0\% (Low), 10.6%10.6\% (Medium) and 11.4%11.4\% (High) of observations, against 10.95%10.95\% High for H2S — the linear mapping reproduces class prevalence across species almost exactly, which the ordinary least-squares construction does not force; the receptor CH4 baseline (median 1.491.49 ppm) sits below the nominal global background, but the class boundaries derive from the co-fitted regression on the same instrument scale, so any absolute offset is common-mode and does not affect classification. Hysteresis counts (Non,Noff)(N_on,N_off). The 30-min clearance (Noff=2N_off=2) matches the WHO averaging convention so the latch will not declare an episode resolved within a single regulatory averaging interval. The 45-min activation (Non=3N_on=3) corresponds to one full 30-min WHO averaging window plus a single 15-min safety-margin step, ensuring any declared event is necessarily supported by evidence persisting beyond the WHO interval and is robust to the kind of isolated voltage transient or one-step turbulent fluctuation that would briefly raise a single 30-min average above threshold. The asymmetry Non>NoffN_on>N_off implements the operationally cautious behaviour of being slow to escalate but fast to acknowledge resolution risk. The same latch is applied identically to the predicted and the rules-based ground-truth tiers, so the two are compared on equal debouncing terms. Prior π. The prior π=0.10π=0.10 is the observed network-wide rate of coordinated High-class evidence across the four channels during the Jan–Mar 2025 validation window. Setting the prior to the data’s own base rate makes the classifier neutrally calibrated at zero information: under all four channels missing simultaneously, the posterior reverts to the empirical occurrence rate. Posterior cut-points θ1,θ2,θ3 _1, _2, _3. The tier cut-points (θ1,θ2,θ3)=(0.15,0.50,0.92)( _1, _2, _3)=(0.15,0.50,0.92) are operational choices; they are not reproduced by the documented grid search, whose lattice does not contain the deployed θ1 _1 and θ2 _2. The per-channel activation threshold p⋆p was itself a search parameter, the fifth alongside the hysteresis counts, likelihood ratios, prior and tier thresholds; the deployed π=0.10π=0.10 and p⋆=0.50p =0.50 both sit at the upper edge of their documented lattices, and the flatness quantified below bounds the consequence. The leave-one-week-out refit selects θ1=0.20 _1=0.20 in all thirteen folds, and substituting it into the deployed configuration changes κw _w by 0.0130.013, comparable to the ±25%± 25\% likelihood-ratio perturbation bound of 0.0090.009; the tier mapping’s performance is therefore only weakly sensitive to the cut-point choice within this range. The operational interpretations attached to each tier (Supplementary Note 1) describe how the cut-points are used, not how they were derived. θ1=0.15 _1=0.15 is the smallest posterior increment requiring at least one likelihood ratio above the prior, i.e. at least one channel actively contributing evidence. θ2=0.50 _2=0.50 is the regime in which the fused posterior cannot be sustained by any single channel alone, so that multi-channel corroboration is required–matching the Tier 2 operational interpretation. θ3=0.92 _3=0.92 lies just below the analytical maximum of the posterior under all four channels active and no missing data, Pmax=Λπ/(1−π)1+Λπ/(1−π)=0.949,Λ=∏cLRcon=2×3×7×4=168,P_ = \,π/(1-π)1+ \,π/(1-π)=0.949, = _cLR^on_c=2× 3× 7× 4=168, for the LR set in Supplementary Table S3 with π=0.10π=0.10, so this tier is reached only when the fused evidence is at or near the maximum attainable by the network. Selection protocol and held-out performance. The likelihood ratios and cut-points were chosen by grid search maximising tier-match κw _w over nine of the thirteen Jan–Mar 2025 weeks, the remaining weeks 1,3,8,10\1,3,8,10\ forming a stratified held-out subset on which the same configuration attains κw=0.528 _w=0.528 and strict accuracy 0.7160.716. That value lies below the lower bound of the block-bootstrap interval on the headline figure; the interval quantifies sampling variability conditional on the tuned parameters, whereas the held-out value additionally removes tuning optimism, which is a bias rather than a variance, so falling below the interval is expected rather than anomalous. The headline κw=0.709 _w=0.709 is scored over all thirteen weeks and is therefore a mixed quantity, in sample with respect to the fusion parameters on nine weeks and out of sample on four. Leave-one-week-out refit. To quantify the tuning optimism directly, the likelihood ratios and cut-points were refitted with each week in turn withheld and the held-out predictions concatenated, so that every timestep is scored out of sample with respect to them. The hysteresis counts and the prior are held fixed across folds: the same latch is applied to the predicted and the ground-truth tiers, so refitting it per fold would move the ground truth itself and make κw _w incomparable across folds. This gives κw=0.703 _w=0.703 and a complaint correlation of +0.721+0.721, against 0.7090.709 and +0.729+0.729 in sample. The ground-truth series’ own moving-block interval is wide ([+0.34,+0.88][+0.34,+0.88]), dominated by a small number of high-count blocks — its Spearman ρ=+0.580ρ=+0.580 is correspondingly more stable — so the paired difference, not the marginal intervals, is the appropriate comparison. All thirteen folds independently selected the deployed H2S likelihood ratios (2.0,3.0)(2.0,3.0); all thirteen selected θ1=0.20 _1=0.20 against the deployed 0.150.15, whose substitution changes κw _w by 0.0130.013 — the objective surface is weakly peaked rather than flat, and the consequence is bounded at ≤0.013≤ 0.013 (“Posterior cut-points” below). The CH4 likelihood ratios and the upper two cut-points were not unanimous: four distinct CH4 pairs were selected across the thirteen folds, θ2=0.50 _2=0.50 in eleven of thirteen and θ3 _3 took all three lattice values, so the flatness is specific to the parameters the deployed configuration fixes rather than general. The hysteresis counts (3,2)(3,2), the prior π and the per-channel activation threshold p⋆p were themselves grid outputs on the nine-week split and are held fixed across folds, so a small residual optimism passes through them; the ±25%± 25\% perturbation analysis (Supplementary Table S4) bounds it. Table S2: Decomposition of the alert-tier agreement by what each configuration holds out. All rows use the identical four constituent nowcasters, ground-truth construction and Jan–Mar 2025 window; they differ only in which fusion parameters were fitted on the data being scored. The 0.5280.528 subset figure and the 0.7030.703 leave-one-week-out figure differ because the former scores a single four-week subset under one fixed configuration, while the latter concatenates held-out predictions across all thirteen weeks; the subset figure’s deviation from the whole-window value reflects week-composition variance, not additional tuning bias — which the leave-one-week-out comparison isolates at 0.0060.006. Configuration κw _w Complaint r Isolates Published, whole window 0.709 +0.729+0.729 In-sample reference; mixed, nine weeks in sample and four held out Stratified held-out weeks 1,3,8,10\1,3,8,10\ 0.528 — Fixed-configuration generalisation to a four-week subset Leave-one-week-out refit 0.703 +0.721+0.721 Fusion-parameter tuning optimism Constrained-ordering leave-one-week-out 0.692 — Sensitivity to the likelihood-ratio ordering Likelihood ratios. Each channel-state likelihood ratio is a Bayes factor expressed in the standard log-odds-update form (59), LRcon=Pr(ℓc,t=1∣network event)Pr(ℓc,t=1∣¬network event),LR^on_c= ( _c,t=1 event) ( _c,t=1 event), (S1) i.e. the relative likelihood of a latched “active” state under the event versus non-event hypotheses. The likelihood-ratio values are grid outputs; the dependence-discounting rationale given below described a design intention, not the selection mechanism. A monotonicity-constrained refit imposing LR(H2S)≥LR(CH4)LR(H_2S) (CH_4) inside the same leave-one-week-out loop changes κw _w by 0.0110.011 on the concatenated held-out series, and the paired per-fold comparison does not distinguish the two orderings (p=0.69p=0.69, Wilcoxon signed-rank over thirteen folds); the paired per-fold differences reverse sign between aggregation conventions, as the pooled and unweighted conventions do elsewhere in this work (Supplementary Table S13). No interpretation of the ordering is therefore supported: the objective surface is flat across orderings, and no result in this work rests on the CH4-heavy configuration. (The deployed ordering coincides with the per-channel F1-High ranking, though the flatness means this coincidence carries no evidential weight.) The bare precision-prevalence ratio of each constituent nowcaster lies in the range 1010–1919 for the four channels; the integer values in Supplementary Table S3 are smaller than that ratio, which is consistent with three sources of inter-channel dependence that violate the conditional-independence assumption underlying a naive product of channel-specific Bayes factors: • the two H2S nowcasters share a sulphate-reducing chemical source term and a partially overlapping advection field (westerly principal bearing, Supplementary Note 2); • the two CH4 nowcasters share the same primary chemistry (sulphate-reducer/methanogen co-emission, Supplementary Note 2); • the H2S and CH4 channels at the same receptor share local meteorology, so a high-pressure inversion event registers simultaneously on both species at that receptor. The values in Supplementary Table S3 therefore function as calibrated multipliers chosen to keep the multi-channel posterior in a well-conditioned range under positive inter-channel correlation, rather than as a direct application of the precision-prevalence ratio. A copula or empirical-Bayes treatment of the joint channel-state distribution would refine the calibration but is left to future work; the present integer values are deliberately conservative. H2S quiet states (LRoff<1LR^off<1) provide mild negative evidence because silence at the canary receptor (MMF9, the most exposed site) is itself informative: a long-running plume that is not detected at MMF9 is unlikely to constitute a network event. CH4 quiet states carry LRoff=1LR^off=1 (no negative evidence) because CH4 exhibits a higher baseline rate of background activation than H2S; allowing CH4 silence to suppress the posterior risks falsely clearing an H2S-led event. Missing data (NaN) is treated as LR=1LR=1 at every channel in keeping with the principle that an absent sensor provides no evidence in either direction. Table S3: Channel-state likelihood ratios for the Bayesian fusion stage. Each ratio is the multiplicative update applied to the prior odds when the corresponding channel is in the indicated state (Eq. (S4)). Active-state values are calibrated multipliers, substantially smaller than the bare precision-prevalence ratio of each constituent nowcaster (column 2; the bare ratio is in the range 1010–1919 across the four channels at π=0.10π=0.10) to account for inter-channel dependence (shared chemistry between H2S receptors, between CH4 receptors, and between species at the same receptor). H2S-quiet states carry moderate negative evidence because the H2S channels are the most direct receptor measurements of the public-health-relevant species; CH4-quiet states carry LR=1LR=1 because the higher CH4 background would make suppressive quiet evidence operationally unsafe. Missing data is always neutral. Robustness of these values is established by the ±25%± 25\% sensitivity analysis of Supplementary Table S4. Channel Native precision LRconLR^on_c LRcoffLR^off_c LRcNaNLR^NaN_c H2S @ MMF9 (broad-coverage) 0.526 2.00 0.50 1.00 H2S @ MMF2 (tight-coverage confirmer) 0.636 3.00 0.50 1.00 CH4 @ MMF9 (chemical corroboration) 0.632 7.00 1.00 1.00 CH4 @ MMF2 (chemical corroboration) 0.673 4.00 1.00 1.00 Bayesian fusion of multi-site, multi-species evidence The fusion stage maps the four meteorology-trained walk-forward nowcaster outputs to a continuous network-event probability and an ordinal alert tier on the synchronous 15-minute grid. The classifier is purely meteorology-driven at inference time: the predicted tier τt _t never reads any raw target-species concentration; only the ground-truth label τtGTτ^GT_t uses raw concentrations. Per-channel state. For channel c∈c at timestep t, the latched binary state ℓc,t _c,t and the missing-data mask mc,tm_c,t are computed from the model probability pc,tp_c,t by ℓc,t=Hysteresis((pc,t>p⋆),Non,Noff),mc,t=(pc,tis missing.CLOSE _c,t= Hysteresis (1(p_c,t>p );\,N_on,N_off ), m_c,t=1(p_c,t\;is missing. (S2) with p⋆=0.5p =0.5, Non=3N_on=3 (45 min onset) and Noff=2N_off=2 (30 min clearance, matched to the WHO 30-min averaging convention). Log-odds initialisation. The prior log-odds is anchored at the empirical network-event base rate: logO0=logπ1−π,π=0.10. O_0= \! π1-π, π=0.10. (S3) Channel-state log-odds update. Each channel contributes additively in log-odds: logLRc,t=0,mc,t=1(sensor missing)logLRcon,mc,t=0∧ℓc,t=1logLRcoff,mc,t=0∧ℓc,t=0. _c,t= cases0,&m_c,t=1 (sensor missing)\\[1.0pt] ^on_c,&m_c,t=0 _c,t=1\\[1.0pt] ^off_c,&m_c,t=0 _c,t=0. cases (S4) CH4-quiet contributions are exactly neutral (LRCH4off=1LR^off_CH_4=1); only H2S-quiet states deliver logLR<0 <0 (mild suppression). Posterior probability. The fused log-odds and posterior are logOt O_t =logO0+∑c∈logLRc,t, = O_0+ _c _c,t, (S5) Pt P_t =σ(logOt)=Ot/(1+Ot),σ(x)=(1+e−x)−1. =σ( O_t)=O_t/(1+O_t), σ(x)=(1+e^-x)^-1. (S6) Tier mapping (pure meteorology). The continuous posterior is mapped to the ordinal tier by sequential thresholding only, with no raw-sensor override on the predicted side: τt=∑k=13(Pt≥θk). _t= _k=1^31 (P_t≥ _k ). (S7) with (θ1,θ2,θ3)=(0.15,0.50,0.92)( _1, _2, _3)=(0.15,0.50,0.92). Because θ1<θ2<θ3 _1< _2< _3 and the indicator sum is monotone in PtP_t, the trajectory 0→1→2→30→ 1→ 2→ 3 cannot skip a tier on the soft path. The Tier 3 cut-point sits just below the analytical posterior ceiling under all four channels concurrently active (Pmax≈0.949P_ ≈ 0.949 for the LR set in Supplementary Table S3), so Tier 3 is reached only by fully coordinated multi-channel meteorological evidence. Bayesian classifier parameter set The hysteresis counts, prior, likelihood ratios and tier cut-points in Supplementary Table S1 were selected to maximise tier-match κw _w against the deterministic ground-truth tier of Eq.(S9) on nine of the thirteen Jan–Mar 2025 validation weeks (“Selection protocol and held-out performance” above), so the window-wide figure is a mixed quantity. A ±25%± 25\% perturbation of each LRconLR^on_c in turn (eight perturbations, all other parameters held fixed) changes the window-wide κw _w by at most 0.009 (mean |Δκw|=0.002| _w|=0.002 across the eight perturbations, Supplementary Table S4). The framework’s agreement with the ground truth is therefore not contingent on the precise LR values within the operationally meaningful range; it reflects the structural choice to fuse four meteorology-trained channels under a log-odds accumulator with H2S-quiet suppression and CH4-quiet neutrality. Table S4: ±25%± 25\% perturbation sensitivity of the window-wide tier-match κw _w to each active-state likelihood ratio. Each row scales one LRconLR^on_c by the indicated factor with all other parameters held fixed (other LRs, hysteresis, prior, tier cut-points). Δκw _w is the deviation from the baseline κw=0.709 _w=0.709. The maximum absolute deviation is 0.009 (one row), indicating that the headline performance is robust to the precise LR choice. Perturbed parameter Factor κw _w Δκw _w H2S MMF9 LRonLR^on ×0.75× 0.75 0.710 +0.001+0.001 H2S MMF9 LRonLR^on ×1.25× 1.25 0.709 ++0.000 H2S MMF2 LRonLR^on ×0.75× 0.75 0.709 ++0.000 H2S MMF2 LRonLR^on ×1.25× 1.25 0.707 −0.002-0.002 CH4 MMF9 LRonLR^on ×0.75× 0.75 0.705 −0.004-0.004 CH4 MMF9 LRonLR^on ×1.25× 1.25 0.709 ++0.000 CH4 MMF2 LRonLR^on ×0.75× 0.75 0.700 −0.009-0.009 CH4 MMF2 LRonLR^on ×1.25× 1.25 0.709 ++0.000 Inter-channel dependence. The conditional-independence assumption underlying Eq.(S5) is violated in practice: the two H2S channels share a chemical source and a partially overlapping advection field, the two CH4 channels share the same primary chemistry, and the H2S/CH4 channels at a single receptor share local meteorology. A formally correct fusion under dependent channel states would require a copula model of the joint ℓc,tc∈\ _c,t\_c distribution or an empirical-Bayes treatment from joint state frequencies. We do not adopt either here; the LR values in Supplementary Table S3 absorb the positive inter-channel correlation by remaining substantially below the bare precision-prevalence ratio of each constituent classifier. The sensitivity analysis above demonstrates that this absorption is operationally adequate for the present agreement level; a principled dependence-aware fusion is left to follow-up work and is expected to deliver an incremental rather than transformative improvement to κw _w given the modest sensitivity margins reported above. Ground-truth tier definition and headline performance Ground-truth tier construction. The ground-truth tier τtGTτ^GT_t is constructed from the raw 15-min sensor record by a deterministic channel-vote rule on debounced threshold exceedances. For each channel c define the raw indicator and its latched form: bc,tGT=(rc,t≥Ccthr),ℓc,tGT=Hysteresis(bc,tGT,Non,Noff),b^GT_c,t=1(r_c,t≥ C^thr_c), ^GT_c,t= Hysteresis(b^GT_c,t;\,N_on,N_off), (S8) where CthrH2S=7μgm−3C^H_2S_thr=7\, \,m^-3 and CthrCH4=2.756ppmC^CH_4_thr=2.756\,ppm, the same class boundaries used to define each CAIRN nowcaster’s training target. Missing data is treated conservatively (bc,tGT=0b^GT_c,t=0) so that an absent sensor cannot raise the ground-truth tier. The number of latched channels is ntGT=∑c∈ℓc,tGTn^GT_t= _c ^GT_c,t and the ground-truth tier is τtGT=∑k=13(ntGT≥k),τ^GT_t= _k=1^31 (n^GT_t≥ k ), (S9) i.e. Tier 1 when at least one channel is latched, Tier 2 when at least two, Tier 3 when at least three. This is the standard information-theoretic ground truth for a multi-receptor network: a single isolated exceedance warrants Tier 1, multi-channel corroboration is required for Tier 2, and convergent evidence across both species or both receptors for Tier 3. The same hysteresis (Non,Noff)(N_on,N_off) is applied to both sides, so τt _t and τtGTτ^GT_t live on the same ordinal scale. The ground-truth label is deterministic and reproducible given the threshold and hysteresis declaration above, but not independent of those choices: a reviewer adopting a different threshold scheme would obtain a different labelling. Headline tier-match performance. On the synchronous Jan–Mar 2025 validation window (n=8,536n=8,536 shared-grid timesteps), the meteorology-only predicted tier agrees with the deterministic raw-sensor ground truth at quadratic-weighted Cohen’s kappa κw=0.709(95% CI: 0.557, 0.852),strict accuracy=79.4% _w=0.709 (95\% CI: 0.557,\;0.852), accuracy=79.4\% where the 95% confidence interval is obtained by a block bootstrap with B=1,000B=1,000 resamples and a 24-hour block length (96 ticks at 15-minute cadence) to accommodate within-episode autocorrelation. The matched-block bootstrap has a κw _w standard error of 0.075. The headline κw _w value sits in the “substantial agreement” band of the Landis–Koch interpretive scale (64). Per-tier metrics with bootstrap 95full confusion matrix is in Supplementary Table S5. Comparison to a vote-count baseline. A natural baseline applies the same channel-vote rule of Eq. (S9) directly to the four predicted latched activations ℓc,t\ _c,t\, replacing the Bayesian log-odds accumulator with a hard count of latched channels and holding all other pipeline stages fixed. On the same window this baseline returns κwvote=0.639 _w^vote=0.639 (95% CI: 0.481, 0.792). The Bayesian fusion improves on it by Δκw=+0.070 _w=+0.070; under a paired block bootstrap, in which the same resampled blocks are used for both estimators so that shared fold-to-fold variation cancels, the improvement is significant (95% CI 0.0250.025, 0.1090.109; one-sided p=0.003p=0.003). An unpaired bootstrap, which resamples the two κ values independently, approximately triples the interval width and is not the appropriate test for two estimators evaluated on identical timesteps. The Bayesian framework is retained on both statistical and operational grounds: it produces a calibrated continuous posterior PtP_t, exposes channel-specific evidence weights, and accommodates missing data and CH4-quiet neutrality natively. The vote-count baseline delivers none of these. Table S5: Confusion matrix of predicted versus ground-truth ordinal alert tiers across the Jan–Mar 2025 validation window (n=8,536n=8,536 synchronous 15-min timesteps). Rows are the ground-truth tier τtGTτ^GT_t; columns are the meteorology-only predicted tier τt _t. Diagonal cells are typeset in bold. Parenthesised values are row percentages. Predicted τt _t Ground truth τtGTτ^GT_t Tier 0 Tier 1 Tier 2 Tier 3 Support Tier 0 6,026 (88%) 620 (9%) 141 (2%) 39 (1%) 6,826 Tier 1 270 (38%) 338 (47%) 92 (13%) 13 (2%) 713 Tier 2 77 (13%) 259 (44%) 215 (36%) 42 (7%) 593 Tier 3 17 (4%) 55 (14%) 133 (33%) 199 (49%) 404 Table S6: Per-tier and macro classification metrics for the meteorology-only multi-channel Bayesian tier classifier. Block-bootstrap 95% CIs (B=1,000B=1,000, block = 24 h) in brackets. Window-wide κw=0.709 _w=0.709 (CI 0.557, 0.852); strict element-wise accuracy 0.794. Tier Precision Recall F1 Support Population share Tier 0 0.943 [0.925, 0.961] 0.883 [0.843, 0.920] 0.912 [0.888, 0.935] 6,826 79.9% Tier 1 0.266 [0.179, 0.348] 0.474 [0.363, 0.578] 0.341 [0.247, 0.420] 713 8.4% Tier 2 0.370 [0.247, 0.496] 0.363 [0.225, 0.519] 0.366 [0.245, 0.479] 593 6.9% Tier 3 0.679 [0.377, 0.832] 0.493 [0.228, 0.689] 0.571 [0.298, 0.717] 404 4.7% Macro avg 0.564 0.553 0.547 8,536 — Accuracy 0.794 — 8,536 — κw _w 0.709 [0.557, 0.852] — 8,536 — Effective sample size at the extremes. The bootstrap CI on Tier 3 F1 is wide ([0.30, 0.72]) because Tier 3 episodes have an effective sample size much smaller than the nominal count of 404 timesteps. Episodes are temporally autocorrelated on 1–3 hour scales (4–12 timesteps), giving an effective independent count of roughly 30–100 episodes. Per-tier point estimates at the extremes should be read against this CI rather than as deterministic operational guarantees, particularly when planning deployment thresholds. Per-week tier-match performance. Supplementary Table S7 reports the within-week κw _w breakdown with single-week bootstrap CIs. Single-week κw _w values are unstable when one or more per-tier support drops below ∼10 10: the statistic’s variance grows by an order of magnitude under these conditions, and weekly values in those regimes should be read as descriptive rather than as calibrated indicators of tier-match quality. Week 7 contains no ground-truth deviation from Tier 0 (all 672 timesteps GT =0=0) and yields a degenerate κw _w undefined from the absence of contingency-table variation; the strict accuracy of 0.970 in that week is informative as a specificity check on the classifier. The headline κw=0.709 _w=0.709 is the pooled estimate over the entire n=8,536n=8,536 synchronous timeseries, not a sample-size-weighted average of the per-week values. Table S7: Per-week tier-match performance with bootstrap 95% CIs. Each row reports the within-week quadratic-weighted Cohen’s kappa, strict element-wise accuracy, and the per-tier support counts in the ground-truth labels. Bold marks weekly blocks with point estimate κw≥0.70 _w≥ 0.70. Bootstrap CIs (B=1,000B=1,000, within-week resampling) capture sampling variation within the week and do not account for cross-week autocorrelation. Week 7 has degenerate (κw _w undefined) tier variation. Ground-truth tier support Wk κw _w 95% CI Acc. Tier 0 Tier 1 Tier 2 Tier 3 nwkn_wk 1 0.528 [0.451, 0.598] 0.571 423 82 132 27 664 2 0.708 [0.661, 0.752] 0.594 314 19 120 219 672 3 0.477 [0.415, 0.539] 0.542 538 33 26 75 672 4 0.500 [0.407, 0.589] 0.805 609 21 42 0 672 5 0.287 [0.200, 0.381] 0.750 604 50 18 0 672 6 0.936 [0.897, 0.964] 0.955 604 7 11 50 672 7 — — 0.970 672 0 0 0 672 8 0.148 [0.037, 0.261] 0.875 608 64 0 0 672 9 0.807 [0.776, 0.839] 0.757 418 109 128 17 672 10 0.420 [0.323, 0.512] 0.876 587 49 24 12 672 11 0.754 [0.692, 0.809] 0.871 545 91 36 0 672 12 0.698 [0.633, 0.759] 0.862 495 127 46 4 672 13 0.762 [0.654, 0.843] 0.931 409 61 10 0 480 Pooled 0.709 [0.557, 0.852] 0.794 6,826 713 593 404 8,536 External validation against community odour complaints The Jan–Mar 2025 validation window includes a co-registered daily record of community odour-complaint counts at the principal receptor (n = 89 days; mean 58 complaints/day, max 1,137 on 2025-01-13). This record was not used in training, in the hysteresis or LR selection, or in any tier-mapping decision. It provides a fully external validation of the classifier against the operational endpoint the manuscript actually targets (community exposure). Daily-mean alert tier versus daily complaints. Aggregating the 15-min predicted tier τt _t to daily means and correlating with daily complaint count gives Pearson r=0.729r=0.729 (95% CI: 0.549, 0.837; bootstrap with within-day resampling, B=1,000B=1,000), R2=0.53R^2=0.53, Spearman ρ=0.617ρ=0.617, p=5.4×10−16p=5.4× 10^-16. The same aggregation on the deterministic ground-truth tier τtGTτ^GT_t gives an upper-bound reference of r=0.793r=0.793 (95% CI: 0.626, 0.897), R2=0.63R^2=0.63. The CAIRN classifier therefore captures approximately 84% of the ground-truth’s explained variance in complaints (0.53/0.63=0.840.53/0.63=0.84), reaching within 0.064 Pearson units of the operational ceiling that a raw-sensor oracle could deliver. Lagged cross-correlation. The lagged Pearson cross-correlation between daily-mean predicted tier and daily complaint count has its unique global maximum at lag 0 (r=0.729r=0.729), with approximately symmetric roll-off: r=+0.575r=+0.575 at lag −1-1 day, r=+0.552r=+0.552 at lag +1+1 day, and r≤0.40r≤ 0.40 at |lag|≥2| lag|≥ 2 days (Supplementary Table S8). The dominant lag-0 peak confirms that the classifier is a nowcaster of community exposure, not a delayed reflector of past complaint-driving events (which would peak at negative lag) and not a forecaster of future events (which would peak at positive lag). Table S8: Lagged Pearson cross-correlation between daily-mean CAIRN predicted tier and daily community complaint count (n=89n=89 days, ±7± 7-day window). Lag is in days; lag >0>0 means the model leads complaints. The unique global maximum at lag 0 confirms nowcast alignment. Lag (days) −7-7 −6-6 −5-5 −4-4 −3-3 −2-2 −1-1 0 +1+1 +2+2 +3+3 +4+4 +5+5 +6+6 +7+7 Pearson r 0.33 0.22 0.21 0.15 0.16 0.40 0.58 0.73 0.55 0.23 0.15 0.17 0.31 0.32 0.39 Operational implication and limitations The headline tier-match performance supports a qualified operational claim. The classifier reliably identifies routine background (Tier 0 F1=0.91F_1=0.91, the dominant class accounting for 80% of the validation window) and recovers most of the ground-truth Tier 3 support at the operationally relevant cost (Tier 3 F1=0.57F_1=0.57, recall 0.490.49; roughly half of Tier 3 episodes are correctly flagged from meteorology alone). The intermediate Tier 1 and Tier 2 tiers carry the bulk of the off-by-one disagreement: each has F1≈0.35F_1≈ 0.35 with the classifier over-triggering Tier 1 (recall 0.47, precision 0.27) and under-triggering Tier 2 (recall 0.36, precision 0.37). These limitations are inherited from the marginal precision of the constituent walk-forward classifiers (Supplementary Tables S13, S14) and from the difficulty of recovering multi-channel agreement on a ∼15 15-min temporal scale. Three deployment caveats follow: (i) The headline κw=0.709 _w=0.709 — a mixed quantity, in sample with respect to the fusion parameters on nine of the thirteen weeks and out of sample on four — has a wide block-bootstrap confidence interval [0.557,0.852][0.557,0.852]. Operational decisions sensitive to the precise agreement level should use the lower end of this interval as a conservative planning floor. (i) The improvement over a same-architecture vote-count baseline is Δκw=+0.070 _w=+0.070 (paired block bootstrap, 95% CI 0.0250.025, 0.1090.109); the predicted tier nonetheless coincides with the vote count minus one on 94.2%94.2\% of timesteps, so the fusion is a calibrated refinement of channel counting rather than a categorically different rule. (i) External validation against community complaints (r=0.729r=0.729 at lag 0, R2=0.53R^2=0.53, p=5×10−16p=5× 10^-16, within 0.06 of the GT ceiling of r=0.793r=0.793) is the more operationally meaningful endpoint. The classifier’s strong performance on this independent signal, which was not used in any stage of training or parameter selection, is the principal external defence of the framework. Supplementary Methods This section holds the full mathematical and procedural detail for the methods summarised in main Methods. Six subsections correspond, in order, to the six explicit cross-references in main Methods: “Environmental characterisation”, “Multiscale transfer entropy”, “Accumulated Local Effects”, “S4D state space model”, “Optimiser schedule” and “Walk-forward validation”. All data-handling, software-environment, hyperparameter and class-imbalance choices are stated in main Methods and are not repeated here. Environmental characterisation This subsection holds the per-panel mathematical detail for Fig. 1 of the main text, summarised in main Methods, “Environmental characterisation”. The quantitative results are tabulated in Supplementary Table 1 and discussed panel-by-panel in Supplementary Note 2. Conditional probability function (Fig. 1a). For each station the full 360∘360 compass is partitioned into ns=36n_s=36 equal Δθ=10∘ θ=10 sectors. Each 15-minute observation is assigned to the sector θk _k containing its recorded wind direction. The CPF for sector k is CPF(θk)=Nexc(θk)N(θk),CPF( _k)= N_exc( _k)N( _k), (S10) where Nexc(θk)N_exc( _k) is the count of observations in sector k with H2S>Cthr=5μgm−3H_2S>C_thr=5\, \,m^-3, N(θk)N( _k) is the total count, and sectors with N(θk)=0N( _k)=0 are assigned CPF=0CPF=0 (5). Plots are rendered on a polar axis with North at 0∘0 proceeding clockwise. Site coordinates are converted from WGS84 to British National Grid (OSGB36, EPSG:27700) via standard transverse Mercator projection for the receptor map. Chemical fingerprint (Fig. 1b). Pearson’s product-moment correlation r=∑i=1n(xi−x¯)(yi−y¯)∑i(xi−x¯)2∑i(yi−y¯)2r= _i=1^n(x_i- x)(y_i- y) _i(x_i- x)^2 _i(y_i- y)^2 (S11) and Spearman’s rank correlation ρ are computed on all paired H2S–CH4 observations after listwise deletion of missing values; significance is via two-tailed t-test with n−2n-2 degrees of freedom. A 96-sample (24-h) sliding-window rolling Pearson r(t)r(t) quantifies temporal stability. Spike co-occurrence on the top-5% spike sets SH2S,SCH4S_H_2S,S_CH_4 is quantified by the Jaccard index J=|SH2S∩SCH4||SH2S∪SCH4|,J= |S_H_2S∩ S_CH_4||S_H_2S∪ S_CH_4|, (S12) with significance via χ2χ^2 test of independence on the ×22\!×\!2 contingency table. A power-law H2S=a⋅CH4bH_2S=a·CH_4^b is fitted by ordinary least-squares (OLS) on log-transformed concentrations. Co-emission evidence is classified as Strong if r>0.7r>0.7 and J>0.5J>0.5, Moderate if r>0.4r>0.4 and J>0.3J>0.3, Weak if r>0.2r>0.2. Diurnal hysteresis (Fig. 1c). Observations are grouped by hour-of-day h∈0,…,23h∈\0,…,23\ to form the diurnal mean trajectory in (TEMP,H2S)(TEMP,H_2S) phase space: T¯h=1Nh∑i:hi=hTi,H2S¯h=1Nh∑i:hi=hci, T_h= 1N_h _i:h_i=hT_i, H_2S_h= 1N_h _i:h_i=hc_i, (S13) with hours containing fewer than max(5,n/240) (5,n/240) observations excluded. The closed 24-point loop is characterised by its enclosed area, computed via the Shoelace surveyors’ formula A=12|∑h=023(T¯hH2S¯h+1−T¯h+1H2S¯h)|,A= 12 | _h=0^23 ( T_h H_2S_h+1- T_h+1 H_2S_h ) |, (S14) with indices modulo 24. The dimensionless normalised area is A∗=A(T¯max−T¯min)(H2S¯max−H2S¯min).A^*= A( T_ - T_ )\,( H_2S_ - H_2S_ ). (S15) A directional-asymmetry index from cooling- (18:00–06:00) and warming-limb (06:00–14:00) OLS slopes is Hindex=|βcool−βwarm||βcool|+|βwarm|+ϵ,ϵ=10−6.H_index= | _cool- _warm|| _cool|+| _warm|+ε, ε=10^-6. (S16) Normalised areas A∗<0.05A^*<0.05, 0.050.05–0.150.15, 0.150.15–0.300.30, >0.30>0.30 are classified as none / weak / moderate / strong hysteresis. Diurnal profile (Fig. 1d). Observations are grouped by hour-of-day; the hourly mean and 95th percentile are reported, with the shaded envelope spanning mean to P95. Hourly variation is tested by a 24-group Kruskal–Wallis H-test. The peak hour is the hour of maximum mean concentration. Sunrise alignment (Fig. 1e). Astronomical sunrise times tsr(di)t_sr(d_i) are computed for each calendar date did_i using the astral Python library (v2.2) at the site latitude/longitude. For each 15-minute observation the elapsed time since sunrise is Δti=ti−tsr(di), t_i=t_i-t_sr(d_i), (S17) restricted to Δti∈[−2,+6] t_i∈[-2,+6] h. The pre-sunrise baseline is the mean concentration over Δt∈[−2,0] t∈[-2,0] h. A one-sided Mann–Whitney U test U=∑i∈post∑j∈pre(ci>cj)U= _i _j 1(c_i>c_j) (S18) is applied with the alternative hypothesis that pre-sunrise concentrations exceed post-sunrise. Evidence is classified Strong if p<0.01p<0.01 and peak/baseline ratio >1.5>1.5, Moderate if p<0.05p<0.05 and ratio >1.2>1.2, Weak if p<0.10p<0.10. Bivariate mechanism scatter (Fig. 1f). For each predictor x∈WS,TEMP,dT/dtx∈\WS,TEMP,dT/dt\ and response c (H2S), Pearson’s r and an OLS regression c^i=β0+β1xi,(β^0,β^1)=argmin∑iβ(ci−β0−β1xi)2 c_i= _0+ _1x_i, ( β_0, β_1)= _β _i(c_i- _0- _1x_i)^2 (S19) are reported. Significance on β^1 β_1 is tested via two-tailed t-test with n−2n-2 degrees of freedom. All three associations are predicted a priori to be negative: higher wind speed dilutes the plume, warmer temperatures promote vertical mixing, and positive dT/dtdT/dt indicates convective destabilisation of the boundary layer. Multiscale transfer entropy This subsection holds the full procedural specification and pseudo-code for the MSTE algorithm summarised in main Methods, “Causal analysis”. The information-theoretic framework follows 91 and 40. Temporal coarse-graining. At each scale τ with aggregation factor m=τ/(15min)m=τ/(15\,min), the coarse-grained series is obtained by non-overlapping block averaging: x¯j(τ)=1m∑i=(j−1)m+1jmxi, x^(τ)_j= 1m _i=(j-1)m+1^jmx_i, (S20) with right-closed, right-labelled blocks to prevent information leakage across aggregation boundaries. Block boundaries at coarsest scales (τ≥3τ≥ 3 h) are aligned to 18:00 UTC to centre the nocturnal accumulation period within a single block. Gap handling. The 15-minute record for calendar year 2024 at the case-study site comprises 35,136 slots. Missing observations are unevenly distributed across channels: wind direction and wind speed 0.04% each, temperature 0.63%, pressure 1.75% and H2S 1.50%. Applying the listwise rule, 1,121 slots (3.19%) lack at least one of the five series entering the analysis, in 47 runs exceeding 30 minutes and with a longest contiguous run of 54.2 h. Gaps are forward-filled over at most 30 minutes; longer gaps are left missing and the affected 15-minute slots are excluded. That bound recovers 386 of the 1,121 slots (34.4%), leaving 735 (2.09% of the year) unusable. Describing the input as the “raw 15-minute series” is therefore accurate only under this bounded-fill policy, which is stated explicitly because it determines how many slots enter each estimate. Usable N after masking. Coarse-graining is applied to the masked series, so the number of blocks entering each estimate falls with scale. A block contributes if it contains at least one complete slot; the stricter requirement that every slot be complete is reported alongside, because it is the quantity that degrades: Scale Blocks Contributing % All slots complete 15 min 35,136 34,015 96.8 34,015 (96.8%) 30 min 17,568 17,218 98.0 16,797 (95.6%) 1 h 8,784 8,631 98.3 8,187 (93.2%) 2 h 4,392 4,331 98.6 3,879 (88.3%) 3 h 2,928 2,893 98.8 2,442 (83.4%) 6 h 1,464 1,448 98.9 1,010 (69.0%) 12 h 732 725 99.0 330 (45.1%) 24 h 366 364 99.5 0 (0%) 48 h 183 183 100 0 (0%) 96 h 91 91 100 0 (0%) These counts apply the block-mean rule; the multiscale transfer-entropy analysis additionally requires at least 50% of a block’s nominal sub-samples to survive, which is the stricter rule under which the 96 h scale retains n=89n=89 and is reported as not estimable (Methods, “Causal analysis”). At τ≥24τ≥ 24 h no block is free of at least one filled or missing slot, so every coarse-scale block mean is computed over a partially reconstructed window. At the 6 h scale that carries the architectural anchor, 31.0% of blocks are affected. This is a stated boundary on the coarse-scale estimates rather than a property of the system, and the estimates at τ≥24τ≥ 24 h should be read as descriptive: with 91–366 blocks they are additionally sample-limited. Conditional transfer entropy and confounder selection. At each scale, the causal influence of driver X on target Y (H2S) is quantified by the conditional transfer entropy (92), formulated as a conditional mutual information (CMI): TX→Y|=I(Yt+1;Xt|Yt(h),t),T_X→ Y =I\! (Y_t+1;\,X_t\; |\;Y_t^(h),\,Z_t ), (S21) where Yt(h)=(Yt,…,Yt−h+1)Y_t^(h)=(Y_t,…,Y_t-h+1) is the target history of order h and tZ_t is a parsimonious confounder set selected by greedy mutual-information ranking with a Pearson |r|<0.85|r|<0.85 redundancy filter (excluding the source itself), with ||=1|Z|=1 throughout. Bivariate transfer entropy (=∅Z= ) cannot distinguish genuine causal influence from confounded association in this multivariate system (91, Fig. 2a); the parsimonious single-confounder choice balances confounder control against the (d2/k)O(d^2/k) scaling of kkNN density-estimator bias with conditioning dimension d (62). History embedding depth. The embedding order h is set adaptively to approximate 3 h of physical memory: h=max(1,min(2,⌊3h/τ⌋)).h= \! (1,\; \! (2,\; 3\,h\,/\,τ ) ). (S22) This yields h=2h=2 at τ≤1τ≤ 1 h (capturing 30–120 min of target memory) and h=1h=1 at coarser scales. The cap h≤2h≤ 2 keeps the joint conditioning dimension at d≤3d≤ 3, within the regime where the KSG estimator retains adequate sensitivity (Note 5). Frenzel–Pompe KSG estimator. The CMI in Eq. (S21) is estimated using the k-nearest-neighbour algorithm of 37, an extension of the KSG estimator (62) to conditional mutual information. For N observations in the joint space (,,)=(Xt,Yt+1,[Yt(h),t])(x,y,z)=(X_t,\,Y_t+1,\,[Y_t^(h),Z_t]), the procedure constructs a k-d tree in the full joint space using the Chebyshev (L∞L^∞) norm, identifies for each point i the k-th nearest-neighbour distance εi _i (excluding all points j with |j−i|<W|j-i|<W, the Theiler window), applies the strict-inequality correction εi←εi(1−10−10) _i← _i(1-10^-10) (because k-d-tree queries return points within a closed ≤ε≤ ball whereas the KSG Algorithm requires the open ball <ε< ), counts marginal neighbours n(i)n_xz(i), n(i)n_yz(i), n(i)n_z(i) within distance εi _i in each subspace, and computes I^(;∣)=ψ(k)−1N∑i=1N[ψ(n(i)+1)+ψ(n(i)+1)−ψ(n(i)+1)], I(x;y )=ψ(k)- 1N _i=1^N\! [ψ(n_xz(i)+1)+ψ(n_yz(i)+1)-ψ(n_z(i)+1) ], (S23) where ψ(⋅)ψ(·) is the digamma function. Theiler window. A Theiler exclusion window (95) of W=4W=4 time steps (1 h at 15-min sampling) is applied in both the εi _i determination and the marginal neighbour counts. This exceeds the lag-1 autocorrelation timescale of the principal drivers, removing the upward bias in CMI caused by selecting nearest neighbours that are close in state space simply because the atmosphere evolves smoothly between consecutive readings (91). Rank transform and adaptive k. All variables are mapped to Uniform[0,1]Uniform[0,1] marginals via the rank transform x~i=(rank(xi)−0.5)/N x_i=(rank(x_i)-0.5)/N, with a small (0,10−8)N(0,10^-8) jitter added to break sensor-quantisation ties and prevent degenerate zero-radius kkNN balls (62). The number of neighbours is k=max(10,⌊Nτ0.4⌋, 2d+6),k= \! (10,\; N_τ^0.4 ,\;2d+6 ), (S24) with NτN_τ the valid sample count at scale τ. The N0.4N^0.4 term keeps the relative ball radius stable as effective sample size falls at coarser scales; the 2d+62d+6 floor grows with conditioning dimensionality, ensuring adequate neighbours for stable digamma estimates. Surrogate testing. Surrogate count and reproducibility of the null. Surrogate seeds are drawn as range(B) and the tie-breaking jitter from a fixed seed, so the surrogate mean is exactly recoverable for any given B; this makes the surrogate count of a stored result verifiable rather than assumed, and it is the check by which the count reported here was confirmed. The significance underpinning Fig. 2 is computed at B=2,000B=2,000; the estimator ablation of Supplementary Note 5 runs at B=10B=10, whose resolution of Δp=0.1 p=0.1 cannot support false-discovery correction across a family of this size and is reported as a coarse screen only. Significance is assessed by B=2,000B=2,000 circular-shift surrogates. For each surrogate b the source series is shifted by a random offset δb∼Uniform1,…,Nτ−1 _b \1,…,N_τ-1\: Xt(b)=X(t+δb)modNτ,X^(b)_t=X_(t+ _b) N_τ, (S25) while target, target history and confounder series are held fixed. Circular shifting (rather than random permutation) preserves the internal autocorrelation structure, periodicity and marginal distribution of the source exactly, breaking only its temporal alignment with the target; this produces a conservative null under which significance reflects driver–target timing rather than the smoothness of the driver series (91, cf.). The same KSG–Theiler estimator is used for both observed TE and surrogates. Effective transfer entropy removes the finite-sample positive bias of the estimator: ETEX→Y(τ)=T^X→Y(τ)−1B∑b=1BT^X(b)→Y(τ).ETE_X→ Y^(τ)= T_X→ Y^(τ)- 1B _b=1^B T_X^(b)→ Y^(τ). (S26) Negative ETE indicates no detectable coupling. Benjamini–Hochberg false-discovery-rate correction (7) at α=0.05α=0.05 is then applied across a single family of all 70 driver–scale tests. Pseudo-code. Algorithms 2 and 3 give the full algorithmic specification. Robustness to estimator hyperparameter choice is established by the six-configuration ablation reported in Note 5 and Supplementary Tables 7–9. Algorithm 2 Multiscale Transfer Entropy (MSTE) 1: Raw multivariate time series D at Δt=15 t=15 min; driver set X; target Y (H2S); scales =15m,30m,1h,2h,3h,6hT=\15m,30m,1h,2h,3h,6h\ with aggregation factors =1,2,4,8,12,24m=\1,2,4,8,12,24\; surrogate count B; Theiler window W; FDR level α 2: ETE and adjusted p-value for each (X,τ)(X,τ) 3: for each scale τ∈τ with aggregation factor m do 4: ¯(τ)←BlockMean(,m) D^(τ)← BlockMean(D,m) ⊳ Eq. (S20) 5: Nτ←N_τ← valid rows in ¯(τ) D^(τ) 6: h←max(1,min(2,⌊3h/τ⌋))h← (1, (2, 3h/τ )) ⊳ Eq. (S22) 7: for each driver X∈X do 8: Rank =∖XC=X \X\ by I^(Cj,Y) I(C_j;Y) descending 9: ←Z← top-1 with |r|<0.85|r|<0.85 vs. already selected 10: d←h+||d← h+|Z|; k←max(10,⌊Nτ0.4⌋,2d+6)k← (10, N_τ^0.4 ,2d+6) ⊳ Eq. (S24) 11: Build embeddings; rank-transform; jitter 12: T^X→Y(τ)←KSG-CMI(~,~,~,k,W) T_X→ Y^(τ)← KSG-CMI( x, y, z,k,W) ⊳ Algorithm 3 13: for b=1,…,Bb=1,…,B do 14: δb∼Uniform1,…,Nτ−1 _b \1,…,N_τ-1\ 15: ~(b)←CircularShift(~,δb) x^(b)← CircularShift( x, _b) ⊳ Eq. (S25) 16: T^(b)←KSG-CMI(~(b),~,~,k,W) T^(b)← KSG-CMI( x^(b), y, z,k,W) 17: end for 18: ETEX(τ)←T^X→Y(τ)−1B∑bT^(b)ETE_X^(τ)← T_X→ Y^(τ)- 1B _b T^(b) ⊳ Eq. (S26) 19: pX(τ)←1B∑b[T^(b)≥T^X→Y(τ)]p_X^(τ)← 1B _b1[ T^(b)≥ T_X→ Y^(τ)] 20: end for 21: end for 22: padj←Benjamini–Hochberg(pX(τ),α)\p_adj\← Benjamini--Hochberg(\p_X^(τ)\,α) 23: return ETEX(τ),padj,X(τ)\ETE_X^(τ),p_adj,X^(τ)\ Algorithm 3 KSG-CMI: Frenzel–Pompe CMI with Theiler window 1: Rank-transformed ~∈ℝN x ^N, ~∈ℝN y ^N, ~∈ℝN×d z ^N× d; neighbours k; Theiler window W 2: I^(~;~∣~) I( x; y z) in nats 3: Build k-d tree jointT_joint on [~,~,~][ x, y, z] using Chebyshev norm 4: for i=1,…,Ni=1,…,N do 5: Query k+2W′+1k+2W +1 neighbours, W′=max(1,W)W = (1,W) 6: Discard neighbours j with |j−i|<W′|j-i|<W 7: εi← _i← Chebyshev distance to k-th surviving neighbour 8: εi←εi(1−10−10) _i← _i(1-10^-10) ⊳ open-ball correction 9: end for 10: Build marginal trees ,,T_xz,T_yz,T_z 11: for i=1,…,Ni=1,…,N do 12: n(i),n(i),n(i)←|j:∥⋅∥∞<εi∧|j−i|≥W′|n_xz(i),n_yz(i),n_z(i)←|\j:\|·\|_∞< _i |j-i|≥ W \| 13: end for 14: return ψ(k)−1N∑i[ψ(n(i)+1)+ψ(n(i)+1)−ψ(n(i)+1)]ψ(k)- 1N _i[ψ(n_xz(i)+1)+ψ(n_yz(i)+1)-ψ(n_z(i)+1)] ⊳ Eq. (S23) Accumulated Local Effects This subsection holds the full mathematical derivation for the ALE interpretability analysis summarised in main Methods, “XGBoost seasonal nowcaster”. ALE (3) quantifies the marginal effect of each feature on the predicted class probabilities while remaining unbiased in the presence of correlated predictors, a critical advantage over partial-dependence plots (38) for meteorological feature sets where wind speed, temperature and stability indices are physically coupled. Definition. For a continuous feature xsx_s, the first-order ALE is the centred accumulated integral of the local partial derivative of the prediction function f f along the feature axis: ALE^(xs)=∫xminxs[∂f^()∂Xs|Xs=zs]dzs−const. ALE(x_s)= _x_ ^x_s\!E\! [ ∂ f(X)∂ X_s\; |\;X_s=z_s ]\!dz_s-const. (S27) In practice, the feature range is partitioned into K equal-count intervals [zk−1,zk)k=1K\[z_k-1,z_k)\_k=1^K, and the local effect within each interval is estimated by finite differences over samples in that bin: ALE^(xs)=∑k=1kxs1nk∑i:xi,s∈[zk−1,zk)[f^(zk,i,∖s)−f^(zk−1,i,∖s)]−const., ALE(x_s)= _k=1^k_x_s 1n_k _i:\,x_i,s∈[z_k-1,z_k)\! [ f(z_k,x_i, s)- f(z_k-1,x_i, s) ]-const., (S28) where nkn_k is the count of samples in bin k, kxsk_x_s the bin index containing xsx_s, and i,∖sx_i, s all features except xsx_s held at their observed values. The centring constant ensures [ALE^]=0E[ ALE]=0 over the data distribution. ALE perturbs only within observed intervals and conditions on the empirical neighbourhood, avoiding extrapolation into unlikely regions of feature space (3). Multiclass aggregation. For the three-class XGBoost classifier, ALE is computed separately for each class c∈Low,Medium,Highc∈\Low,Medium,High\ using the class-specific soft probability f^c()=P(Y=c∣) f_c(x)=P(Y=c ) from the multi:softprob XGBoost output, yielding three ALE curves per feature per seasonal classifier. Feature importance. Importance is the bin-size-weighted variance of the ALE effect: Is(c)=∑k=1KnkN(ALE^k(c)−ALE¯(c))2,ALE¯(c)=∑knkNALE^k(c).I_s^(c)= _k=1^K n_kN ( ALE_k^(c)- ALE^(c) )^2, ALE^(c)= _k n_kN ALE_k^(c). (S29) The total importance is Is=∑cIs(c)I_s= _cI_s^(c) and the class-mean importance reported in Fig. 4a is I¯s=Is/3 I_s=I_s/3. ALE plots are computed on held-out validation data using the PyALE library with K=20K=20 quantile-based grid points, independently per seasonal classifier. For the cross-seasonal comparison in Fig. 4(a,b), per-class ALE curves and importance values are averaged across the three classes at each grid point. S4D state space model This subsection holds the full mathematical foundations of the S4D state space architecture summarised in main Methods, “S4 dual-pathway nowcaster” (Eqs. 1–2 of the main text). The S4 family was introduced in 47 and 48; we use the diagonal (S4D) simplification of 49. Continuous-time SSM. The S4 architecture represents the mapping from a one-dimensional input signal u(t)∈ℝu(t) to an output signal y(t)∈ℝy(t) as a continuous-time linear time-invariant dynamical system (47): ˙(t) x(t) =(t)+u(t), =A\,x(t)+B\,u(t), (S30) y(t) y(t) =(t)+Du(t), =C\,x(t)+D\,u(t), (S31) with latent state (t)∈ℂNx(t) ^N, ∈ℂN×NA ^N× N, ∈ℂN×1B ^N× 1, ∈ℂ1×NC ^1× N and D∈ℝD . Zero-order-hold discretisation. For a discrete uniformly sampled input uk=u(kΔ)u_k=u(k ) at step Δ>0 >0, the ZOH discretisation (47) of Eqs. (S30)–(S31) yields the recurrence k _k =¯k−1+¯uk,yk=k+Duk, = Ax_k-1+ Bu_k, y_k=Cx_k+D\,u_k, (S32) ¯ A =exp(Δ),¯=(Δ)−1(exp(Δ)−)Δ. = ( ), B=( )^-1( ( )-I)\, \,B. (S33) Unrolling Eq. (S32) from a zero initial state gives the equivalent convolutional form yk=∑j=0k¯j¯uk−j+Duk=(¯∗u)k+Duk,y_k= _j=0^kC\, A^\,j\, B\,u_k-j+D\,u_k=( K u)_k+D\,u_k, (S34) with the SSM convolution kernel given in main text Eq. (1). The kernel decomposition has three properties exploited directly: (i) strict causality, since j∈0,…,kj∈\0,…,k\; (i) (LlogL)O(L L) inference cost via FFT-based convolution, enabling the L=384L=384-step (96-h) context without quadratic scaling; (i) a continuum of learned timescales governed by ¯=exp(Δ) A= ( ), with slow coordinates (|Re(λi)|≪1|Re( _i)| 1) holding information over many time steps and fast coordinates (|Re(λi)|≫1|Re( _i)| 1) decaying rapidly. Diagonal (S4D) parameterisation. The S4D simplification of 49 restricts A to be diagonal, retaining the HiPPO-LegS spectrum that underpins long-memory behaviour while eliminating the normal-plus-low-rank kernel construction. With =diag(λ1,…,λN)A=diag( _1,…, _N), the kernel reduces to a complex Vandermonde evaluation: K¯k=2Re[∑n=1N/2Cn⋅exp(kΔλn)−1λn⋅exp((k−1)Δλn)]. K_k=2\,Re\! [ _n=1^N/2C_n· (k _n)-1 _n· ((k-1) _n) ]. (S35) The eigenvalues are parametrised as λn=−exp(logAnreal)+iπ(n−1),n=1,…,N/2, _n=- ( A^real_n)+iπ(n-1), n=1,…,N/2, (S36) so that Re(λn)<0Re( _n)<0 is enforced by construction (stability) and gradients remain well-scaled across many orders of magnitude in |A||A|. The step size is similarly parametrised as Δ=exp(logΔ) = ( ); C is stored as real/imaginary parts with Xavier-scale initialisation Cstd=1/HNC_std=1/ HN (39) to prevent the readout from dominating at initialisation. Equation (S35) is assembled with a numerically stable real-arithmetic factorisation: expm1(a)=ea−1expm1(a)=e^a-1 is used to retain precision when Re(Δλ)Re( λ) is small, and the decay/phase accumulation exp(kΔλ) (k λ) is computed in forced FP32 to prevent TF32 tensor-core truncation from collapsing slow-lane values. S4D block. Each S4D block wraps the SSM kernel in a pre-normalised, gated residual structure following the template of contemporary sequence models (101; 105; 23). For block input ∈ℝB×L×HU ^B× L× H: ~ U =LayerNorm(), =LayerNorm(U), (S37) ssm[:,:,h] _ssm[:,:,h] =(¯(h)∗~[:,:,h])+Dh~[:,:,h],h=1,…,H, =( K^(h) U[:,:,h])+D_h\, U[:,:,h],\;\;h=1,…,H, (S38) drop _drop =Dropout(ssm), =Dropout(Y_ssm), (S39) [;] [V;G] =gludrop+glu, =W_glu\,Y_drop+b_glu, (S40) =⊙σ(), =V σ(G), (S41) ′ =+out+out, =U+W_outZ+b_out, (S42) ′ =′+FFN(LayerNorm(′)), =U +FFN(LayerNorm(U )), (S43) where σ is the logistic sigmoid, ⊙ is element-wise product, and the FFN is a two-layer GELU network (53) of inner width 2H2H. Pre-Norm rather than Post-Norm (105) is essential for stable gradient flow in deep residual stacks and protects the long-tailed H2S concentration distribution from saturating the kernel’s learned timescales. Gated Linear Units with sigmoid gate (23) preserve amplitude information better than a pointwise non-linearity, allowing slow-lane channels to remain alive through the non-linearity even when their instantaneous magnitude is small. Zero-initialised skip (Dh=0D_h=0) forces the block, at the start of training, to route information through the SSM memory pathway rather than through a direct feedthrough. To regularise the shallow (nlayers=3n_layers=3) stack we apply stochastic depth (56) with linearly increasing per-layer drop probability pℓ=ℓ⋅pmax/(nlayers−1)p_ = · p_ /(n_layers-1) for pmax=0.15p_ =0.15, scaling the residual update by an independent Bernoulli mask at each forward pass. End-to-end model. For an input window ∈ℝB×L×FX ^B× L× F of L=384L=384 steps and F standardised features, the full architecture is (0) ^(0) =in+in,in∈ℝF×H, =XW_in+b_in, _in ^F× H, (S44) (ℓ) ^( ) =(ℓ−1)+DropPathpℓ(S4DBlockℓ((ℓ−1))−(ℓ−1)),ℓ=1,…,nlayers, =H^( -1)+DropPath_p_ \! (S4DBlock_ (H^( -1))-H^( -1) ), =1,…,n_layers, (S45) L−1 _L-1 =(nlayers):,L−1,:∈ℝB×H, =H^(n_layers)_:,L-1,: ^B× H, (S46) y^reg y_reg =MLPreg(L−1),^=softmax(MLPcls(L−1)). =MLP_reg(h_L-1), p=softmax\! (MLP_cls(h_L-1) ). (S47) Both heads are LayerNorm→ (H,H/2)→(H,H/2) → → (H/2,⋅)(H/2,·) with zero-initialised final layers, so the model starts training with calibrated uniform class probabilities and zero-mean regression output. Causality guarantees. Strict causality is enforced at three points: (i) the SSM kernel in Eq. (S35) is evaluated at non-negative time indices k=0,…,L−1k=0,…,L-1 only, with kernel length asserted equal to input length so no padding can introduce an acausal shift; (i) the FFT convolution in Eq. (S38) is computed with an explicit 2L2L-length FFT and truncated to the first L outputs, giving the same result as a direct causal convolution yk=∑j=0kKjuk−jy_k= _j=0^kK_j\,u_k-j with no bidirectional, circular or future-padded variant; (i) the sequence representation passed to the heads is the final-timestep slice L−1h_L-1 (Eq. (S46)), not a mean or max over the full window, preserving a single well-defined “prediction-time” receptive field. Production hyperparameters. The configuration used for all results in the main text and in Extended Data Tables ED1–ED3 is: H=128H=128 (model width), N=256N=256 (state size), nlayers=3n_layers=3, L=384L=384 (15-min steps; 96-h context), f=0.5f=0.5 (fast/slow channel split), |Afastreal|=10.0|A^real_fast|=10.0, |Aslowreal|=0.5|A^real_slow|=0.5, log-space spread s=0.5s=0.5, dropout 0.290.29, DropPath pmax=0.15p_ =0.15, focal-loss γ=2.62γ=2.62, class weights (αLow,αMed,αHigh)=(1.00,11.26,14.07)( _Low, _Med, _High)=(1.00,11.26,14.07). Optimiser schedule This subsection holds the full optimiser configuration for the S4 dual-pathway nowcaster summarised in main Methods, “S4 dual-pathway nowcaster”. Joint loss. The training objective is a weighted sum of focal cross-entropy classification (68) and standardised MSE regression: ℒ=λclsℒFL(^,y)+λregℒMSE(y^reg,yregstd),L= _cls\,L_FL( p,y)+ _reg\,L_MSE( y_reg,y_reg^std), (S48) with λcls=1.0 _cls=1.0 and λreg=0.0 _reg=0.0 in the production configuration (the regression head therefore receives no gradient. Its final layer is zero-initialised and, being untrained, emits a constant; it is present in the architecture but contributes nothing to any reported result, and no interpretability analysis reported here uses its output). The focal cross-entropy term is ℒFL(^,y)=−1B∑i=1Bαyi(1−p^i,yi)γlogp^i,yi,L_FL( p,y)=- 1B _i=1^B _y_i(1- p_i,y_i)^γ\, p_i,y_i, (S49) with γ=2.62γ=2.62 and per-class weights (αLow,αMed,αHigh)=(1.00,11.26,14.07)( _Low, _Med, _High)=(1.00,11.26,14.07). Weights are imposed manually rather than derived from inverse frequencies and deliberately allocate more weight to the High class than its ∼ 9:1 inverse prior, with the focusing factor down-weighting easy high-confidence examples so that gradient updates concentrate on borderline samples between Medium and High. This design directly targets High-class precision and recall as the operationally relevant metrics. In parallel, an inverse-prior balanced minibatch sampler (WeightedRandomSampler) renders the class marginal approximately uniform within each batch despite the ∼ 13:2:1 Low:Medium:High imbalance, a two-pronged correction combining sampling and weighting (51; 19). Differential learning-rate groups. Optimisation uses AdamW (70) with five parameter groups designed to fix known pathologies of S4D optimisation. Non-SSM parameters (input projection, FFN layers, output heads, C readout) use base learning rate ηbase=2.31×10−2 _base=2.31× 10^-2 and weight decay 2.7×10−22.7× 10^-2. SSM log-step (logΔ ) and imaginary-A parameters use a much smaller ηs4=10−4 _s4=10^-4 to prevent rapid migration of the timescale prior away from the physics-anchored initialisation. SSM real-A parameters (logAreal A^real) use ηA=10−4 _A=10^-4 (manually configured here, overriding the typical 20×20× default boost). The skip parameter D uses ηD=0.1ηbase _D=0.1\, _base to prevent it from short-circuiting the SSM memory path. Gradients are clipped globally at norm 1.0. Warmup and freeze schedule. For the first W=6W=6 warmup epochs, logAreal A^real and AimagA^imag are frozen, allowing C to learn to discriminate the fixed physics-anchored timescales before the A spectrum is allowed to migrate; this prevents the optimiser’s early high-loss gradient from collapsing the slow-lane memory during the first epoch. When warmup ends at epoch W, the A matrix is unfrozen and ηbase _base is transiently multiplied by a warmup-LR factor of 3.0×3.0× to accelerate subsequent physics fine-tuning. A ReduceLROnPlateau scheduler then monitors validation High-class F1 and halves the learning rate after pplateau=2p_plateau=2 epochs without improvement, down to a floor of 10−510^-5. Early stopping monitors the same metric with patience 7 and restores the best-epoch weights. Maximum training is 60 epochs. Walk-forward validation This subsection holds the per-week pipeline, no-leakage guarantees, metric definitions and post-hoc attribution procedures (lane ablation and feature sensitivity) for the walk-forward evaluation summarised in main Methods, “Walk-forward validation and metrics”. Validation horizon and training cutoff. The validation horizon (1 October–31 December 2024) is partitioned into 13 non-overlapping weekly blocks ww=113\V_w\_w=1^13 with tw+1start=twend+1dayt_w+1^start=t_w^end+1\,day. For each block w, the training set is the expanding window w=t∈all:t≤twstart−LΔtbase,T_w=\t _all:t≤ t_w^start-L\, t_base\, (S50) with allD_all the full 15-min record from 2023-09-01. The training cutoff is retreated by L−1=383L-1=383 steps (≈96≈ 96 h) before the first validation timestamp; the intervening L−1L-1 steps form a context reserve wC_w used only as input to the sliding-window convolution that produces the first prediction at twstartt_w^start. No label in wC_w is back-propagated through. Per-week pipeline. For each of the 13 weeks the following procedure is executed independently: [leftmargin=*, itemsep=2pt] 1. Build features on allD_all once (hour sin/cos, WD sin/cos, four Fourier seasonality harmonics), then partition into wT_w and w∪wV_w _w via Eq. (S50). 2. Drop rows with missing labels or features. 3. Fit a StandardScaler on wT_w alone (zero-mean, unit-variance per feature) and apply to w∪wV_w _w. The training regression target is standardised using wT_w statistics only. 4. Instantiate a fresh model (no carry-over of weights, optimiser state or LR schedule from week w−1w-1) and train for up to 60 epochs with the objective and optimiser of Supplementary Methods, “Optimiser schedule”, monitoring validation F1-High for early stopping. 5. Compute predicted class probabilities ^i p_i at every timestamp ti∈wt_i _w and persist alongside true labels, hard predictions and timestamps. 6. Compute week-w metrics (definitions below). Per-week metrics are aggregated by sample-size weighted averaging across the 13 weeks; the aggregates appear in Extended Data Table ED2 and Supplementary Table 5. No-leakage guarantees. Five independent leakage pathways are eliminated by construction: [leftmargin=*, itemsep=2pt] • Label leakage: Eq. (S50) guarantees that no training sample has a label whose timestamp falls within w∪wV_w _w. • Feature leakage: every feature is either a pointwise function of the current timestamp (hour sin/cos, Fourier seasonality, WD sin/cos) or the raw meteorology itself; no rolling mean, lag, derivative or stagnation index is present to carry information across the train/validation boundary, and no scaling or target statistic is fitted on validation data. • Model leakage: the SSM kernel, FFT convolution and final-timestep pool are strictly causal; no bidirectional, attention-based or global-pool operation is used; the implementation asserts kernel length equal to input length so any acausal padding becomes a runtime error. • Hyperparameter leakage: the hyperparameter set used for the 13 weekly retrains is fixed in advance, carried from an independent hyperopt experiment on Sep 2023–Sep 2024 data whose validation period does not overlap with ⋃w _wV_w; a Week-0 sanity check compares the first weekly retrain against the hyperopt’s held-out test metrics and flags any relative macro-F1 discrepancy exceeding 20%. • Retrain leakage: each week’s model is instantiated from scratch (weights, optimiser state and LR schedule all reinitialised), so no information from week w’s validation period can re-enter week w′>w >w via carried-over parameters. Metric definitions. For each weekly block we report hard-classification metrics (per-class precision PcP_c, recall RcR_c, F(c)1_1^(c); macro F1; false-alarm rate FAR), probabilistic metrics (Brier-High, binary ranked probability score, log-loss), and conditional-rolling metrics aligned with the WHO 30-minute averaging convention. Define yi(c)=(yi=c)y_i^(c)=1(y_i=c) as the one-hot true label and p^i(c) p_i^(c) the model’s predicted probability of class c. The Brier score for the High class (binary one-vs-rest) is the mean-squared probabilistic error (12): Brier-High=1N∑i=1N(p^i(High)−yi(High))2.Brier -High= 1N _i=1^N ( p_i^(High)-y_i^(High) )^2. (S51) The binary ranked probability score (32) for the ordinal Low << Medium << High classes is RPSi=1K−1∑k=1K−1(F^i(k)−Fi(k))2,F^i(k)=∑c≤kp^i(c),Fi(k)=∑c≤kyi(c),RPS_i= 1K-1 _k=1^K-1 ( F_i(k)-F_i(k) )^2, F_i(k)= _c≤ k p_i^(c),\;\;F_i(k)= _c≤ ky_i^(c), (S52) with K=3K=3. The multiclass log-loss is L=−1N∑i=1N∑c=1Kyi(c)logp^i(c).L=- 1N _i=1^N _c=1^Ky_i^(c) p_i^(c). (S53) Conditional-rolling F1/precision/recall and FAR are computed in a 4-h moving window aligned with the WHO 30-minute averaging convention. With window half-width Tw=8T_w=8 samples (2 h on each side, 16 samples total at 15-min cadence) centred on timestep t, Mroll(t)=M((yj,y^j)j:|j−t|≤Tw),M^roll(t)=M\! (\(y_j, y_j)\_j:|j-t|≤ T_w ), (S54) where M∈F1,P,R,FARM∈\F_1,P,R,FAR\ and the function on the right-hand side is the standard hard-classification metric applied to the windowed predictions and labels. The rolling metrics are evaluated only at timesteps for which the window contains ≥Tw≥ T_w samples. False-alarm rate is FAR=FP/(FP+TN)FAR=FP/(FP+TN) for the High vs. non-High binary partition. Lane ablation. To attribute the trained model’s High-class skill to each architectural lane, an inference-time ablation is applied to every weekly checkpoint. For S4D layer ℓ , a forward hook on the SSM kernel module intercepts the output of ¯(ℓ)∈ℝH×L K^( ) ^H× L and zeros the rows corresponding to the targeted lane: ¯ablated(ℓ)[h,k]=0,h∈ℐlane¯(ℓ)[h,k],otherwise, K^( )_ablated[h,k]= cases0,&h _lane\\ K^( )[h,k],&otherwise, cases (S55) where ℐfast=0,…,fH−1I_fast=\0,…,fH-1\ and ℐslow=fH,…,H−1I_slow=\fH,…,H-1\ are the fixed channel indices of the two lanes. The mask is applied after kernel construction but before FFT convolution (Eq. (S38)), so the SSM contribution of the masked lane is suppressed while the D-skip path, GLU gate, output projection and downstream FFN cross-channel mixing remain intact. This isolates the SSM contribution of the lane rather than the entire signal-flow pathway. Method dependence of the lane attribution. Under the kernel-path mask of Eq. (S55) the lane-zeroing ablation apportions +0.083+0.083 to the fast lane, +0.074+0.074 to the slow lane and +0.121+0.121 to their interaction, sample-weighted across the 13 weeks. The two lanes are not separately distinguishable: the paired per-fold difference is −0.002-0.002 (fast leads on 8 of 13 folds; sign test p=0.58p=0.58, Wilcoxon p=1.00p=1.00; bootstrap 95% CI [−0.079,+0.072][-0.079,+0.072]), the interaction exceeds either marginal contribution, and the ordering reverses between aggregation rules. These values are therefore a decomposition that does not resolve, not an attribution of mechanism to either lane. The attribution reported below is sensitive to where in the signal path the mask is applied, and this sensitivity must be stated because it determines the sign of the fast/slow comparison. Equation (S55) specifies a kernel-path mask: the targeted rows of ¯(ℓ) K^( ) are zeroed after kernel construction and before FFT convolution, leaving the D-skip path and GLU gate intact. An alternative block-path formulation, in which the whole S4 block output for the targeted channels is suppressed, additionally removes the skip and gate contributions and therefore attributes to each lane every downstream pathway that lane feeds. Under the block-path formulation the ordering of the two lanes reverses relative to the kernel-path values, with the fast lane carrying the larger contribution. The two formulations answer different questions: the kernel-path mask isolates the SSM contribution, the block-path mask the total pathway contribution. Neither is uniquely correct. The values above should accordingly be read as a statement about the SSM contribution under the kernel-path definition, and not as evidence that either lane encodes a specific atmospheric mechanism. Four forward passes per checkpoint give the metric values Mfull,Mno_fast,Mno_slow,Mno_bothM_full,M_no\_fast,M_no\_slow,M_no\_both. The per-lane contribution is ΔMlane=Mfull−Mno_lane, M_lane=M_full-M_no\_lane, (S56) positive when the lane increases M, and the interaction term ℐM=(Mfull−Mno_both)−(ΔMfast+ΔMslow)I_M=(M_full-M_no\_both)-( M_fast+ M_slow) (S57) is positive when the two lanes are complementary, zero when they act independently, and negative when partially redundant. Reported aggregates are sample-weighted means across the 13 weeks for F1-High and macro F1. Feature sensitivity over time. Per-feature per-lane sensitivity is computed analytically from each weekly checkpoint. The S4D kernel ¯(ℓ)∈ℝH×L K^( ) ^H× L is reconstructed from the saved logΔ(ℓ),logAreal,(ℓ),Aimag,(ℓ),(ℓ)\ ^( ), A^real,( ),A^imag,( ),C^( )\ via Eq. (S35), and combined with the input projection inW_in to give the per-lane sensitivity of feature f at lag k: Slane(f,k)=∑h∈ℐlane|Win[h,f]|⋅|¯(ℓ)[h,k]|,k=0,…,L−1.S_lane(f,k)= _h _lane|W_in[h,f]|·| K^( )[h,k]|, k=0,…,L-1. (S58) Column-normalisation at each lag gives the fractional contribution of each feature to the lane’s response: S~lane(f,k)=Slane(f,k)∑f′Slane(f′,k)×100%, S_lane(f,k)= S_lane(f,k) _f S_lane(f ,k)× 100\%, (S59) suitable for visualisation as stacked-area charts over the 96-h input window (Fig. 5c). The formulation is causal (k=0k=0 is present, k=L−1k=L-1 is the earliest input), respects the lane partition, and is linear in the feature step but multiplicative through the kernel, highlighting only the (feature, lag, lane) combinations that the trained model genuinely uses. Per-week sensitivity surfaces are averaged across the 13 walk-forward folds before plotting to suppress single-checkpoint noise. Supplementary Tables Table S9: Quantitative summary of the environmental characterisation of receptor-level H2S exposure (main Fig. 1). All statistics are evaluated at the principal receptor (MMF9, n=34,609n=34,609 15-minute records with valid H2S across 2024) unless otherwise indicated. Panel Statistic Value Source attribution (Fig. 1a) Peak CPF bearing 265∘(W)265 \ (W) Peak CPF value ≈0.37≈ 0.37 Peak/opposing-sector ratio 4.09×4.09× Sector-mean (W, NW, E) 11.211.2, 10.610.6, 1.1μgm−31.1\, \,m^-3 Spike threshold (P95) 15.3μgm−315.3\, \,m^-3 Circular–linear correlation r=0.149r=0.149, p<10−3p<10^-3 Spike Fisher’s odds ratio 1.431.43, p=2.9×10−4p=2.9× 10^-4 Chemical fingerprint (Fig. 1b) Pearson H2S–CH4 r=0.830r=0.830, p<10−15p<10^-15 Spearman H2S–CH4 ρ=0.683ρ=0.683, p<10−15p<10^-15 Power-law exponent b=1.95b=1.95 Jaccard spike index J=0.654J=0.654 Spike co-occurrence 1,386 vs 87.7 expected, χ2=21,298χ^2=21,298 PM10 negative control r=−0.028r=-0.028, p<10−7p<10^-7 PM2.5 negative control r=−0.009r=-0.009, p=0.084p=0.084 NOx confounded correlation r=0.377r=0.377 (6-h diurnal phase offset) Pooled multi-site H2S–CH4 r=0.80r=0.80, 95% CI 0.740.74–0.840.84, n=4n=4, I2=99.8%I^2=99.8\% Temperature–H2S hysteresis (Fig. 1c) Normalised loop area A∗=0.369A^*=0.369 Cooling-limb slope βcool=−2.83 _cool=-2.83 Warming-limb slope βwarm=−2.45 _warm=-2.45 Hysteresis-asymmetry index Hindex=0.072H_index=0.072 Wind-speed–H2S loop area A∗=0.012A^*=0.012 (no hysteresis) Diurnal profile (Fig. 1d) Peak (02:00 LST) mean 9.69μgm−39.69\, \,m^-3 Peak (02:00 LST) P95 51.6μgm−351.6\, \,m^-3 Trough (12:00–15:00 LST) mean 1.31μgm−31.31\, \,m^-3 Trough (15:00 LST) P95 4.2μgm−34.2\, \,m^-3 Mean peak/trough ratio 7.4×7.4× P95 peak/trough ratio 12.3×12.3× Hourly variation Kruskal–Wallis H=338.5H=338.5, p=7.0×10−58p=7.0× 10^-58 Pasquill A–B fraction 40.0%40.0\% Pasquill E–F fraction 26.1%26.1\% Feb monthly peak 14.97μgm−314.97\, \,m^-3 Aug monthly minimum 0.99μgm−30.99\, \,m^-3 Winter/summer ratio 15.1×15.1×, p<10−15p<10^-15 Sunrise-aligned inversion break-up (Fig. 1e) Pre-sunrise baseline (−2-2 to 00 h) 9.49μgm−39.49\, \,m^-3 (n=2,920n=2,920) Maximum at Δt=−2 t=-2 h 10.42μgm−310.42\, \,m^-3 +1+1 h, +2+2 h, +4+4 h 2.652.65, 1.541.54, 1.30μgm−31.30\, \,m^-3 Pre-/post-sunrise reduction 7.3×7.3× One-sided Mann–Whitney U=1.15×107U=1.15× 10^7, p<10−15p<10^-15 Mechanism scatter (Fig. 1f) WS slope −1.35μgm−3-1.35\, \,m^-3 per m s-1 (r=−0.122r=-0.122) 1/WS1/WS slope 10.75μgm−310.75\, \,m^-3 per (m s-1)-1, 95% CI 9.49.4–12.112.1 TEMP slope −0.78μgm−3-0.78\, \,m^-3 per ∘C (r=−0.173r=-0.173) dT/dtdT/dt slope −6.33μgm−3-6.33\, \,m^-3 per ∘C hr-1 (r=−0.143r=-0.143) Mean conc. at u<1u<1 m s-1 11.8μgm−311.8\, \,m^-3 Mean conc. at u>5u>5 m s-1 1.8μgm−31.8\, \,m^-3 Spike WS vs non-spike WS 1.751.75 vs 3.673.67 m s-1, p<10−15p<10^-15 Peak cooling rate −0.6∘-0.6\, C hr-1 Per-mechanism R2R^2 range 0.0150.015–0.0300.030 Table S10: Optimal hyperparameters selected by Bayesian Tree-structured Parzen Estimator optimisation for each seasonal XGBoost classifier. Each season uses an independent one-week holdout for hyperparameter validation, temporally isolated from the evaluation quarter by a one-week gap. Parameter Winter Spring Summer Autumn max_depth 3 7 7 5 η (learning rate) 0.092 0.118 0.132 0.040 n_estimators 150 100 150 300 subsample 0.883 0.978 0.832 0.981 colsample_bytree 0.798 0.790 0.789 0.740 α (L1) 0.078 1.131 0.326 0.930 λ (L2) 1.792 1.867 1.978 1.195 min_child_weight 4.505 4.957 4.664 4.758 Hyperopt trials 12 15 19 14 Hyperopt val. F1 0.649 0.676 0.658 0.617 Table S11: XGBoost seasonal classification performance on the 2024 calendar-year evaluation periods. Precision (P), recall (R) and F1 score are reported per class; accuracy summarises overall performance on the evaluation quarter. Sample counts per class are given in the lower block; the High-class fraction varies from 5.2%5.2\% in Summer to 12.8%12.8\% in Winter, reflecting seasonal emission and dispersion patterns. Low Medium High Season Acc. P R F1 P R F1 P R F1 Winter 0.764 0.887 0.851 0.869 0.536 0.573 0.554 0.523 0.583 0.551 Spring 0.817 0.962 0.848 0.902 0.567 0.739 0.642 0.402 0.644 0.495 Summer 0.836 0.980 0.864 0.918 0.355 0.631 0.455 0.474 0.749 0.581 Autumn 0.778 0.963 0.826 0.889 0.386 0.591 0.467 0.392 0.600 0.474 Annual (sample-weighted) 0.799 see per-class rows above Season Low Medium High Total High % Winter 5,741 1,591 1,073 8,405 12.8% Spring 6,578 1,232 666 8,476 7.9% Summer 7,339 792 443 8,574 5.2% Autumn 6,438 845 879 8,162 10.8% Annual 26,096 4,460 3,061 33,617 9.1% Six Autumn samples immediately following a seven-hour meteorological outage on 23 October 2024 are absent under the causal derivative formulation, which requires a longer leading history than the centred filter it replaces; all six fall in the Low class, so the Medium and High counts are unchanged. Table S12: Comparison of XGBoost hyperopt validation F1 (one-week holdout) and seasonal evaluation F1 (full quarter). The close agreement (|ΔF1|≤0.009| _1|≤ 0.009) confirms that the temporal isolation strategy successfully prevents hyperparameter overfitting. Season Hyperopt F1 Eval F1 Δ 1 Winter 0.649 0.658 ++0.009 Spring 0.676 0.679 ++0.003 Summer 0.658 0.651 −-0.006 Autumn 0.617 0.610 −-0.007 Table S13: Weekly walk-forward classification performance: S4 SSM versus XGBoost over the Oct–Dec 2024 validation window. Recall, precision and FAR are computed by taking the argmax of the three-class posterior; F1 is the High-class harmonic mean. The argmax rule is not equivalent to thresholding P(High)P(High) at 0.50.5: the latter would give 531/565/341 rather than the reported 539/612/333 true positives, false positives and false negatives. Bold entries are the weekly winner per metric. nHn_H is the number of true High-class observations in the week. The column is computed from the evaluation record and sums to 872, with a mean of 67.1. Period-wide aggregate counts are in Extended Data Table ED2. Aggregation convention: the x¯ x row is the unweighted mean of the 13 weekly values. Extended Data Table ED2 instead pools the confusion counts over all 8,040 timesteps before computing each metric, which is why the two tables report 0.528 and 0.533 for the same quantity. Under the matched protocol the pooled value is the larger of the two, because the weeks with the smallest High-class support are also the weakest and an unweighted mean gives them disproportionate influence. F1 High ↑ Recall ↑ Precision ↑ FAR ↓ Brier H ↓ Wk S4 XGB S4 XGB S4 XGB S4 XGB S4 XGB nHn_H 0 0.807 0.617 0.902 0.569 0.730 0.674 0.028 0.023 0.026 0.037 51 1 0.527 0.538 0.744 0.641 0.408 0.463 0.068 0.047 0.048 0.046 39 2 0.588 0.276 0.625 0.500 0.556 0.190 0.007 0.029 0.011 0.021 8 3 0.506 0.585 0.568 0.838 0.457 0.449 0.059 0.089 0.051 0.047 37 4 0.381 0.485 0.679 0.755 0.265 0.357 0.165 0.119 0.088 0.076 53 5 0.491 0.464 0.897 0.897 0.338 0.313 0.123 0.138 0.067 0.066 29 6 0.639 0.680 0.943 0.934 0.483 0.535 0.196 0.158 0.094 0.078 106 7 0.474 0.353 0.419 0.349 0.545 0.357 0.025 0.045 0.053 0.057 43 8 0.791 0.733 0.968 0.947 0.669 0.597 0.080 0.106 0.048 0.052 94 9 0.417 0.296 0.652 0.522 0.306 0.207 0.054 0.073 0.034 0.051 23 10 0.470 0.390 0.349 0.356 0.718 0.430 0.039 0.135 0.140 0.150 146 11 0.260 0.083 0.180 0.047 0.466 0.368 0.062 0.024 0.187 0.202 150 12 0.514 0.545 0.796 0.774 0.379 0.421 0.213 0.174 0.111 0.104 93 x¯ x 0.528 0.465 0.671 0.625 0.486 0.413 0.086 0.089 0.074 0.076 67.1 Table S14: Weekly walk-forward performance of the two S4 models on their respective High-class targets. CH4 thresholds are linearly corrected from the H2S thresholds via H2S=5.55CH4−8.29H_2S=5.55\,CH_4-8.29. Bold highlights the higher per-week value for the species pair on F1-High. Both columns report the published selection protocol, not the matched inner-probe protocol of Extended Data Table ED2 and Supplementary Table 5. The cross-species comparison requires the two models to share a protocol, and the CH4 arm was not re-run under the inner-probe correction; regenerating only the H2S column would make the comparison unmatched, which is the defect the correction removes. The H2S values here therefore differ from Supplementary Table 5 – period-mean F1-High 0.5770.577 against 0.5280.528 – and the difference is the selection protocol, not the model. What this table supports is that the same architecture transfers to a co-emitted species under a common protocol; it does not add to the benchmark comparison. S4 H2S S4 CH4 Wk F1 Prec Rec Brier L F1 Prec Rec Brier L 0 0.807 0.706 0.941 0.026 0.079 0.870 0.803 0.950 0.023 0.076 1 0.563 0.453 0.744 0.045 0.145 0.622 0.472 0.911 0.065 0.192 2 0.588 0.556 0.625 0.011 0.037 0.576 0.436 0.850 0.026 0.079 3 0.505 0.414 0.649 0.053 0.151 0.769 0.625 1.000 0.034 0.111 4 0.411 0.288 0.717 0.091 0.260 0.621 0.465 0.935 0.084 0.246 5 0.677 0.636 0.724 0.037 0.110 0.625 0.500 0.833 0.055 0.166 6 0.711 0.581 0.915 0.078 0.243 0.705 0.587 0.882 0.118 0.360 7 0.487 0.543 0.442 0.055 0.224 0.477 0.337 0.817 0.135 0.397 8 0.841 0.750 0.957 0.040 0.134 0.818 0.701 0.981 0.047 0.154 9 0.575 0.420 0.913 0.031 0.092 0.609 0.483 0.824 0.035 0.112 10 0.617 0.586 0.651 0.121 0.405 0.701 0.571 0.910 0.111 0.329 11 0.186 0.308 0.133 0.210 0.844 0.455 0.327 0.750 0.083 0.260 12 0.533 0.403 0.785 0.098 0.291 0.447 0.303 0.852 0.101 0.292 x¯ x 0.577 0.511 0.707 0.069 0.232 0.638 0.508 0.884 0.074 0.236 σ 0.164 0.138 0.219 0.054 0.207 0.128 0.142 0.069 0.040 0.108 Table S15: Ablation configurations for the KSG–Theiler conditional mutual information estimator used in the multiscale transfer entropy analysis. Each configuration is additive: subsequent rows inherit all prior fixes and modify a single additional parameter. Configuration 4 (bold) is the production configuration used in main Fig. 2. ID Description hmaxh_ Confounders k ε -correction 1 Baseline (original) 4 1 10 (fixed) No 2 + Boundary fix (ε -shrinkage) 4 1 10 (fixed) Yes 3 + Adaptive k 4 1 adaptive (Methods) Yes 4 + Reduced history (h≤2h≤ 2) 2 1 adaptive (Methods) Yes 5 + Minimal conditioning (h=1h=1) 1 1 adaptive (Methods) Yes 6 Bivariate (no conditioning) 1 0 adaptive (Methods) Yes Table S16: Effective transfer entropy (ETE, nats ×103× 10^3) for each driver–scale combination under six estimator configurations. Columns labelled 1–6 refer to the configurations defined in Supplementary Table 7. Asterisks mark cells in which zero of ten circular-shift surrogates exceeded the observed value; at B=10B=10 this is a coarse screen and not a false-discovery-corrected test (see the section text). Bold entries in column 4 indicate the production configuration used in main Fig. 2. Dashes denote ETE ≤0≤ 0 (no detectable coupling). Driver Scale 1 (B) 2 (BF) 3 (AK) 4 (RH) 5 (MC) 6 (BV) Core drivers (significant across all configurations) WDsin 15 min 17.9* 20.0* 12.2* 18.1* 24.2* 42.6* 1 h 23.5* 27.3* 18.6* 25.6* 33.1* 51.4* 6 h 56.5* 61.6* 52.8* 52.8* 52.8* 64.0* WS 15 min 10.1* 11.6* 5.3* 8.2* 10.5* 9.3* 1 h 8.9* 10.1* 8.3* 12.8* 18.8* 22.4* 3 h 18.5 21.5* 14.9* 14.9* 14.9* 22.9* Pressure 15 min 5.9* 6.4* 2.6* 2.9* 3.7* 12.6* 1 h 7.0* 9.7* 9.2* 12.6* 12.5* 23.4* 6 h 25.2* 26.6* 29.2* 29.2* 29.2* 24.3* Rate-of-change drivers (scale-dependent significance) dP/dtdP/dt 15 min — 0.3 — — — 0.6 2 h — — 5.2* 5.2* 5.2* 14.2* 6 h 12.6 11.6 8.5* 8.5* 8.5* 19.5* d(WS)/dtd(WS)/dt 15 min 4.2* 4.3* — — 0.1 — 2 h 5.1 4.9 7.9* 7.9* 7.9* 16.2* 3 h 20.5* 18.2* 6.9* 6.9* 6.9* 13.7* dT/dtdT/dt 2 h 9.6* 8.8* 11.9* 11.9* 11.9* 5.8* 6 h — — — — — — Slow driver (significant only at coarsest scale) TEMP 15 min 2.5* 3.2* — — — — 1 h 1.2 — 1.7 0.9 0.3 — 6 h 12.0* 15.9* 17.5* 17.5* 17.5* 10.5* Table S17: Aggregate ablation statistics across all 42 driver–scale combinations. “Screened links” counts the cells in which zero of ten circular-shift surrogates exceeded the observed transfer entropy. At B=10B=10 this is a coarse screen, not a false-discovery-corrected test; it is reported to compare stability across estimator settings, and the significance underpinning main Fig. 2 comes from the production run at B=2,000B=2,000. ID Configuration Mean ETE (nats) Screened links (/42) Notes 1 Baseline +0.0098+0.0098 27 Spurious TEMP significance (fine scales) 2 + Boundary fix +0.0108+0.0108 27 Marginal improvement 3 + Adaptive k +0.0083+0.0083 24 More conservative; fewer false positives 4 + Reduced h +0.0088+0.0088 25 Optimal bias–variance trade-off 5 + Minimal cond. +0.0099+0.0099 26 Loses 3030 min target memory 6 Bivariate +0.0143+0.0143 29 Inflated; confounders uncontrolled Table S18: Three-arm decomposition isolating the contribution of memory. All three arms share the walk-forward protocol, the 7-day adjacent inner probe, the argmax decision rule and the same 8,040 evaluation timesteps (872 High). They differ in one respect each. XGB-16 receives CAIRN’s sixteen raw inputs, none of which carries a lag, rolling window or derivative, and is therefore memoryless by construction; hyperparameters were re-tuned on the reduced set so the comparison is symmetric. XGB-24 adds the engineered feature set, in which Butterworth derivatives and stagnation indices supply memory by hand. CAIRN returns to the raw sixteen and supplies memory through the state-space kernels. Bold marks the best value in each row. The ordering is monotone on every High-class metric and on false alarms, but not on overall accuracy or multiclass log-loss, where CAIRN is last: its focal loss and High-class weighting trade specificity for sensitivity (Results). XGB-16 XGB-24 CAIRN Memory none hand-coded learnt Inputs 16 raw 24 engineered 16 raw High-class detection F1-High 0.4625 0.5013 0.5329 Precision-High 0.4000 0.4445 0.4683 Recall-High 0.5482 0.5745 0.6181 Brier-High 0.0835 0.0773 0.0748 Confusion (pooled) True positives 478 501 539 False positives 717 626 612 False negatives 394 371 333 Whole-distribution (operational target of no arm) Accuracy 0.7692 0.7866 0.6944 Log-loss 0.5733 0.5372 0.8997 Increment over the memoryless floor pooled Δ 1 folds won Wilcoxon p bootstrap 95% CI Δengineering _engineering = XGB-24 −- XGB-16 +0.040+0.040 10/13 0.0270.027 [−0.001,+0.087][-0.001,+0.087] Δarchitecture _architecture = CAIRN −- XGB-16 +0.073+0.073 10/13 0.0220.022 [+0.019,+0.150][+0.019,+0.150] endpoint difference = CAIRN −- XGB-24 +0.032+0.032 8/13 0.1100.110 [−0.017,+0.081][-0.017,+0.081] The paired per-fold Wilcoxon test is the primary test used throughout this work, and both increments over the memoryless floor clear it where the endpoint difference does not; neither clears a two-sided sign test (p=0.092p=0.092 for both), so the evidence is the pair of increments rather than either alone. Confidence intervals are block bootstrap over the 13 weekly folds; a row-wise bootstrap ignores autocorrelation and returns spuriously narrow intervals. Δarchitecture _architecture falls to +0.042+0.042 when the two folds whose inner probe contains no High-class sample are excluded, while the ordering XGB-16 << XGB-24 << CAIRN survives every jackknife. The fully crossed design would add a state-space model on the engineered features; it is not reported, because CAIRN is specified for raw channels and the paper makes no claim about it consuming an engineered representation. Supplementary Figures Figure S1: Co-located 15-minute time series of H2S and CH4 at the principal receptor across the autumn 2024 walk-forward validation window. Top: H2S concentration with WHO Medium (2μgm−32\, \,m^-3) and High (7μgm−37\, \,m^-3) thresholds. Bottom: simultaneously measured CH4 concentration on the linearly corrected class-boundary scale H2S=5.55CH4−8.29H_2S=5.55\,CH_4-8.29 (Low <1.854<1.854, Medium 1.8541.854–2.7562.756, High ≥2.756≥ 2.756 ppm). The strong concurrent coupling between the two species is the basis for the cross-species transfer experiment reported in main Fig. 6, panels b (right reliability diagram), c (CH4 cross-correlogram with complaints) and d (right panel: cross-model agreement r=0.777r=0.777). Methods: “Cross-species CH4 nowcaster”.