Paper deep dive
Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)
AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 8/19/2026, 5:13:15 AM
Summary
This paper presents a validated protocol for label-free test-time adaptation (TTA) of drone-based crowd counting models for mass gathering events like the 2034 FIFA World Cup. Using a CSRNet backbone and the DroneCrowd corpus, the authors evaluate five adaptation methods, finding that entropy minimization (TENT) and batch normalization realignment (AdaBN) are effective. A physics-informed population conservation prior was tested but found to be less effective than normalisation-driven adaptation. The study establishes a 'severity law' for error recovery and a 'stability budget' for deployment safety, concluding that unconditional adaptation with tail monitoring is superior to shift-gated policies.
Entities (10)
Relation Signals (7)
2034 FIFA World Cup → locatedin → Saudi Arabia
confidence 99% · Saudi Arabia will host the 2034 FIFA World Cup
TENT → istypeof → Label-free test-time adaptation
confidence 97% · TENT [2] adds a single objective, minimising prediction entropy through the BN affine parameters
AdaBN → istypeof → Label-free test-time adaptation
confidence 97% · AdaBN [1] recomputes batch-normalisation statistics on the target data
DroneCrowd → usedforevaluation → Label-free test-time adaptation
confidence 96% · We answer each of these on DroneCrowd [9] with controlled corruptions
RAFT → usedfor → Population conservation prior
confidence 95% · estimated with a frozen pretrained RAFT network [8]... population-conservation loss
CSRNet → usedin → Label-free test-time adaptation
confidence 95% · We build on a CSRNet density regressor [7]... adapted at test time
Label-free test-time adaptation → improves → CSRNet
confidence 93% · adaptation repairs the dense-scene undercounting that would otherwise under-report a forming crush
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting must hold accuracy on footage unlike anything in its training corpus, without labels, and must warn of dangerous inflow before a crush forms. We deliver a validated answer built on 525 controlled runs, a full-resolution corpus study, five falsification ablations, and a five-condition safety-interlock evaluation. Label-free adaptation recovers 31-49% of shift-induced error across four corruptions and five severities, with the strongest method gaining 41.8 MAE over the frozen source (95% CI [34.1, 49.6], p=7.5x10^-10, d=2.52). We establish a severity law separating methods with a constant absolute margin from the one whose margin grows, and a stability budget identifying which configuration is safe to fly. On a full-resolution corpus carrying a genuine +48 MAE aerial gap (source retrained to 14.6 validation MAE, a 34% improvement), adaptation repairs the dense-scene undercounting that would otherwise under-report a forming crush, and the flux-based risk module fires on real congestion episodes in 2 of 6 full-length clips. We localise the recoverable error: in a regime built to favor a physics-informed conservation prior (300-frame clips at 200ms spacing, five times wider than standard), the adaptation signal is normalisation-driven, not flow-driven; the continuity residual is invariant to the proportional counting errors domain shift produces, confirmed by four on/off ablations correlated at r=0.999 and a 40% input corruption moving accuracy by only 0.05 MAE. A label-free shift gate shows shift magnitude and accuracy damage are rank-independent (Spearman rho=0.20; rho=-0.60 among genuine shifts), quantifying the 58% of headroom a magnitude gate forgoes. We establish unconditional adaptation with tail monitoring as policy, closing with a six-point protocol.
Tags
Links
- Source: https://arxiv.org/abs/2608.17625v1
- Canonical: https://arxiv.org/abs/2608.17625v1
Trouble viewing inline? Open PDF directly →
Full Text
43,465 characters extracted from source content.
Expand or collapse full text
Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia) AlAnoud AllGhayth* AlJawharh AlOtaibi* Jude AlSubaie* Affiliation: alanoud@daldata.ai, aljawharh@daldata.ai, jude@daldata.ai Affiliation: Riyadh, Saudi Arabia Abstract Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting for such venues must hold accuracy on footage unlike anything in its training corpus, without labels, and must warn of dangerous inflow before a crush forms. We deliver a validated answer built on 525525 controlled runs, a full-resolution corpus study, five falsification ablations, and a five-condition evaluation of a safety interlock, and we resolve three questions that a deployment decision depends on. We validate the adaptation stage. Label-free adaptation is decisive and holds up as conditions worsen: it recovers 3131–49%49\% of shift-induced error across four corruptions and five severities, with the strongest single method gaining 41.841.8 mae over the frozen source (95% CI [34.1,49.6][34.1,49.6], p=7.5×10−10p=7.5×10^-10, d=2.52d=2.52). We establish a severity law separating methods whose absolute protective margin is constant from the one whose margin grows, and a stability budget that identifies which configuration is safe to fly. On a full-resolution corpus carrying a genuine +48+48 mae aerial gap (reached after retraining the source model to 14.614.6 validation mae, a 34%34\% improvement), adaptation repairs the dense-scene undercounting that would otherwise cause a monitor to under-report a forming crush, and the flux-based risk module fires on real congestion episodes in 22 of 66 full-length target clips. We localise where the recoverable error lives. Building the regime a physics-informed conservation prior asks for (300300-frame clips at 200200 ms spacing, five times wider than standard, so genuine motion exists between frames), we determine that the adaptation signal in this task is normalisation-driven rather than flow-driven: the continuity residual is provably invariant to the proportional counting errors that domain shift actually produces, a result confirmed by four on/off ablations correlated at r=0.999r=0.999 and by a 40%40\% input corruption that moves accuracy by 0.050.05 mae. This tells practitioners where to spend adaptation capacity and where not to. We derive the optimal deployment policy. Evaluating a label-free shift gate as a decision policy, we show that shift magnitude and accuracy damage are rank-independent (Spearman ρ=0.20ρ=0.20; ρ=−0.60ρ=-0.60 among genuine shifts), quantify the 58%58\% of available headroom a magnitude-based gate forgoes, and establish unconditional adaptation with tail monitoring as the evidence-backed policy. We close with a six-point protocol and the acceptance criteria for the next build. 1 Introduction Crowd disasters are failures of monitoring before they are failures of crowd control. In nearly every modern stadium and pilgrimage tragedy, the dangerous build-up of density was under way for minutes before anyone acted on it. The 2034 FIFA World Cup in Saudi Arabia, and the Hajj gatherings the country manages each year, will place enormous crowds under exactly the conditions in which such build-ups form. A system that could watch these crowds from the air and raise a warning while there is still time to intervene would address a problem that existing, manual monitoring handles poorly. Drone-mounted cameras are the natural sensor, and crowd counting from aerial video is a mature enough technique to estimate density in principle. In practice it breaks at the first contact with a real event. A counting model trained on one corpus loses accuracy the moment the footage differs in altitude, illumination, optics, or transmission quality, and event footage always differs. Worse, no ground-truth counts exist during a live event to correct the model. It must adapt to the incoming stream using no labels at all. Label-free test-time adaptation (TTA), which updates the model from a self-supervised objective on the test stream itself, is the only practical response [2, 1]. This setting raises a specific and appealing idea. Between two consecutive frames the number of people in a region can change only through movement across its boundary: people are conserved. If a counting network’s density predictions are inconsistent with the motion measured by optical flow, that inconsistency is an error the network can correct, without labels. The same quantity, the flux of people across a line, is also a natural early-warning signal for congestion. A single physical law might therefore supply both the adaptation signal and the safety signal the deployment needs. This paper asks whether it does. We build the pipeline our idea implies: a CSRNet density regressor [7] adapted at test time under a population-conservation loss computed from RAFT optical flow [8], and evaluate it against the requirements a safety deployment actually imposes. Those requirements are stricter than a single benchmark average. An integrator must know which parts of the pipeline carry the accuracy, how that benefit changes as conditions worsen toward the tail where danger lives, how the system behaves on its worst runs rather than its average ones, and whether it can decide unaided when adaptation is warranted. We answer each of these on DroneCrowd [9] with controlled corruptions and a full-resolution transfer study, using paired statistics, effect sizes, and Holm correction, and two ablations built to expose a component that contributes nothing. Because the conservation prior is expected to be weakest when frames are close together, we also grant it the regime it favours: a full-resolution retrain on the complete corpus (validation mae 22.3→14.622.3→ 14.6) with frames sampled five times further apart than the default. Our findings are as follows. 1. Adaptation is effective and its benefit is predictable. It recovers 4040–46%46\% of shift-induced error at the reference severity and 3030–49%49\% across a five-level severity sweep (Section 5). 2. The benefit does not degrade as corruption worsens; we report a per-method severity law and identify a stability cost in the combined method, whose worst runs occur at low severity (Section 5). 3. On full-resolution transfer, adaptation removes the dense-scene undercounting that dominates source error, and the flux indicator fires on real congestion episodes (Sections 6, 9). 4. The conservation prior does not improve on entropy minimisation in any condition we tested, including the wide-spacing regime built to favour it. We explain this with an invariance argument and localise the recoverable error to normalisation statistics (Section 7). 5. A label-free shift score is a poor basis for gating adaptation, because its magnitude does not track the accuracy damage a shift causes; we therefore recommend unconditional adaptation with monitoring of the worst-run tail (Sections 8, 10). 2 Related Work Test-time adaptation. Adapting a model to the test stream without labels has converged on the normalisation layers as the point of intervention. AdaBN [1] recomputes batch-normalisation statistics on the target data and needs no gradient step; TENT [2] adds a single objective, minimising prediction entropy through the BN affine parameters while every convolutional weight stays frozen. The robustness-oriented successors, CoTTA [3] against error accumulation, EATA [4] through sample selection and anti-forgetting, and SAR [5] through sharpness-aware updates, as well as the gradient-free LAME [6], all inherit this frozen-backbone, normalisation-centred design. That shared design is what makes the family the right setting for our question: with capacity confined to the same small parameter space, any advantage a physics prior offers must show up there or nowhere. We benchmark against AdaBN and TENT and position the robust variants as the next comparison (Section 10). Crowd counting. Density-map regression with dilated convolutions, as in CSRNet [7], remains the standard treatment of congested scenes, and we adopt it unchanged so that our findings concern the adaptation objective rather than a new architecture. DroneCrowd [9] is the corpus throughout this study; its scale, altitude range, and dense aerial viewpoints are representative of the mass-gathering setting we target. VisDrone [10] defines the adjacent aerial benchmark, and we are explicit that we report no results on it: we name DroneCrowd→ the external-validity milestone this protocol is built to be carried into (Section 10). Physics-informed priors. Physics-informed learning [11] supervises a network with a law its outputs must satisfy, and succeeds where that law genuinely constrains the solution. Population conservation is the natural instance for counting: with a displacement field from RAFT [8], the change in count within a region must equal the flux across its boundary. Our contribution to this programme is a sharp negative characterisation: the precise conditions under which the conservation residual carries gradient for counting, and the invariance that empties it under the shifts that actually occur (Section 7). Because the argument is stated at the level of the residual rather than the architecture, it transfers to any density-regression task tempted by the same prior. Shift detection. Label-free detection of distribution shift [12] is well developed, but a safety interlock imposes a stronger requirement than the literature usually asks of it: the score must be monotone not in whether a shift occurred but in how much accuracy it costs. We show these are different quantities in this task, and that a magnitude score, however well it detects shift, is the wrong basis for gating adaptation (Section 8). 3 Method 3.1 Backbone and adaptation family We build on a CSRNet density regressor [7], mapping each frame to a density map Dt(x)D_t(x) whose integral over a region Ω is the predicted count Ct(Ω)=∫ΩDtxC_t( )= _ D_t\,dx. Following the fully test-time protocol of TENT [2], only batch-normalisation parameters are updated on the test stream; all convolutional weights stay frozen. Holding everything fixed except the objective ensures each comparison isolates the loss rather than a difference in model capacity. We compare five configurations: Source (frozen, no adaptation), AdaBN (test-stream BN statistics), TENT (entropy minimisation), Ours (conservation residual alone), and TENT+Ours (both objectives). 3.2 Population-conservation prior People are neither created nor destroyed between consecutive frames, so the count inside a region can change only through motion across its boundary. Absent sources or sinks in Ω , ∂t∫ΩDtx+∮∂ΩDtt⋅ℓ= 0, ∂ t _ D_t\,dx\;+\; _∂ D_t\,v_t·n\,d \;=\;0, (1) where tv_t is the pixel-wise displacement field between frames t and t+1t+1, estimated with a frozen pretrained RAFT network [8]. Written in divergence form and discretised on the pixel grid, this yields the per-pixel continuity residual rt=Dt+1−Dt+∇⋅(Dtt),r_t\;=\;D_t+1-D_t+∇\!·\! (D_tv_t ), (2) whose squared magnitude ℒphys=‖rt‖22L_phys=\|r_t\|_2^2 we minimise either alone or added to the entropy objective, ℒ=ℒent+λℒphysL=L_ent+ _phys. Where predictions obey the law the two terms cancel; any imbalance is a candidate label-free error signal, and Section 7 determines precisely which errors it can and cannot see. 3.3 Flux-based risk indicator The boundary integral in Eq. (1) yields as a by-product an inward-flux signal Φt(Ω)=−∮∂ΩDtt⋅dℓ _t( )=- _∂ D_tv_t·n\,d : a region taking in people faster than they leave registers sustained positive Φt _t before it becomes dangerously dense. We use Φt _t as a relative congestion-onset indicator, which is the form in which it is operationally useful today. Expressing it as an absolute crush threshold requires density in people/m2/m^2 and hence a meters-per-pixel scale to the ground plane; Section 9 specifies that calibration as the acceptance criterion for the absolute mode. 3.4 Shift-gated safeguard Adaptation modifies a model at inference time, so a mature system should be able to decide without labels whether to intervene. We instrument a gate that compares the batch-normalisation statistics of the incoming stream against those cached from clean data, producing a scalar shift score s, and adapts only when s>τs>τ, with τ=2scleanτ=2s_clean. Section 8 evaluates it as a decision policy, comparing what it delivered against what each alternative policy would have delivered, which is the form a deployment decision requires. 4 Experimental Protocol Table 1: The three experimental tracks. All adaptation is label-free and updates only BN parameters. Base models are stated explicitly, because Track B uses a stronger retrained source and the two error scales are reported separately throughout. Track A Track B Purpose controlled shift real domain gap Corpus subset, n=750n=750 full release, full-res Clip length short 300 frames Frame spacing 40 ms ∼200 200 ms Source val mae 26.1 14.6 Runs 525 ablations + risk Track C: shift-gated policy, 5 conditions Track A: controlled corruption benchmark. We evaluate CSRNet on a drone crowd-counting stream of n=750n=750 frames. To isolate robustness from scene variability, we hold scene content fixed and apply four synthetic corruptions that emulate documented failure modes of aerial capture: additive Gaussian noise (sensor noise), motion blur (platform and subject motion), low light (dusk and night operation), and JPEG compression (bandwidth-limited transmission), each measured against a clean reference. The five-method benchmark runs at the reference severity for 55 conditions × 55 methods × 55 seeded replicates =125=125 runs. The severity sweep extends this over five severity levels for four methods: 4×5×4×5=4004× 5× 4× 5=400 runs. Total: 525 runs. Track B: full-resolution corpus with real inter-frame motion. Track B is the engineering centrepiece of the study and was purpose-built to test the conservation prior in its strongest regime while simultaneously providing the realistic transfer setting the deployment case needs. We ingested the full 1111 GB release, converted trajectory annotations from the native .mat format, retrained the source model at full resolution (validation mae 14.614.6, improved from 22.322.3), and rebuilt pair sampling to draw 300300-frame clips at ∼200 200 ms spacing, five times wider than Track A, so genuine displacement exists between paired frames for Eq. (2) to constrain. Track B additionally carries a real aerial domain gap (+48+48 mae source degradation) rather than a synthetic corruption, and its full-length clips are what make the risk module measurable (Section 9). Track C: policy evaluation of the safeguard. The gate is evaluated on the five Track-A conditions against both the frozen source and the adapted model, and scored as one of four candidate policies rather than as a binary classifier. Metrics and analysis. We report mean absolute error (mae) and root-mean-square error (RMSE) of the predicted count. Because replicates share seeds across methods, comparisons are paired: we use paired t-tests with 95% confidence intervals and the paired effect size Cohen’s dzd_z, corroborated by Wilcoxon signed-rank tests, and we control families of per-shift tests with the Holm–Bonferroni procedure. The Source model is deterministic across seeds, so its comparisons are one-sample tests of each adaptive method’s replicates against the Source constant. Stability is reported as across-replicate coefficient of variation (CV) and worst-replicate error, because a safety application is governed by its tail. Reporting discipline on base models. Track A results use a source model early-stopped at validation mae 26.126.1; Track B uses the full-resolution retrain at 14.614.6. Every comparison in this paper is within-track on a single fixed base model, and no quantity is pooled across tracks. The two tracks are designed to converge on conclusions, not on absolute error levels, and they do. 5 Validated: Adaptation Efficacy and the Severity Law Figure 1: Adaptation across corruption severity. (a) Typical accuracy (median count mae, lower is better): the frozen Source degrades steeply while every adaptive method holds far below it and the protective gap widens. (b) Stability (worst replicate, min–max band): TENT+Ours carries the extreme tail, peaking near 113113 mae at motion-blur severity 1 (best typical accuracy, widest spread). Panel (b) is the basis for the stability budget in Section 10. Result 1: adaptation recovers most of the cost of domain shift. Averaged over the four corruptions at the reference severity, the unadapted Source model reaches 96.196.1 mae, while every adaptive method lands in the 5252–5858 range (Table 5), a 4040–45%45\% reduction of shift-induced error. Aggregating the strongest single method (AdaBN) over its 2020 shifted replicates, the gain over Source is 41.841.8 mae (95% CI [34.1,49.6][34.1,49.6], p=7.5×10−10p=7.5× 10^-10, Cohen’s d=2.52d=2.52), several times the conventional threshold for a large effect. The gain is concentrated in the simplest available mechanism, realigning batch-norm statistics to the incoming stream, which is an operationally welcome finding: the component doing the work is parameter-free, cheap, and stable. RMSE reproduces the same ordering. Result 2: the severity law. Table 6 reports the sweep. Source error climbs monotonically from 73.673.6 to 112.0112.0 mae across severities 11–55 and no adaptive method follows it: at severity 55, TENT holds 76.776.7 and TENT+Ours 66.166.1. The structure of the benefit separates the methods cleanly, and the distinction is the practically important one. Entropy-only adaptation maintains a near-constant absolute protective margin (35.735.7 mae recovered at severity 11; 35.335.3 at severity 55), which corresponds to a relative recovery falling from 48.5%48.5\% to 31.5%31.5\% as the corruption intensifies. The combined objective instead grows its absolute margin (30.4→45.830.4→ 45.8 mae), overtaking TENT from severity 22 onward. Stated for a deployment brief: the protective margin against severe corruption is at minimum preserved and at maximum increasing, the adapted and unadapted curves never reconverge, and one configuration converts additional severity into proportionally additional benefit. This is a law about the methods, not a single benchmark number, and it is what lets an integrator predict behaviour at severities not yet observed. Result 3: a stability budget, and a diagnosis of its source. Mean across-replicate CV over the sweep is 11.3%11.3\% for TENT and 10.9%10.9\% for Ours, against 20.6%20.6\% for TENT+Ours, peaking at 70.8%70.8\% in a single cell. The worst individual run in the sweep is 113.5113.5 mae for TENT+Ours (motion blur, severity 1) versus 91.891.8 for TENT, and the location matters as much as the magnitude: TENT’s worst run occurs where an operator would expect it, at maximum severity, whereas TENT+Ours’ worst run occurs at minimum severity. The paired seed design lets us go further and identify the source. In the five-method benchmark the same replicate destabilises all adaptive methods on clean data (AdaBN 65.265.2, TENT 64.764.7, Ours 60.560.5, TENT+Ours 156.1156.1 against a median of 34.634.6), which establishes that combining objectives amplifies a pre-existing adaptation instability rather than introducing one. That is a transferable diagnosis: the instability belongs to test-time adaptation under low-shift conditions, and any method stacked on top of it inherits and magnifies it. Outcome. Validated for deployment: BN realignment with entropy minimisation as a single-objective adaptation stage, operated within the stability budget of Section 10. The combined objective is held back from flight on tail behaviour despite its superior mean, a decision the paired design made possible to justify quantitatively. 6 Validated: Full-Corpus Transfer Track A establishes that adaptation repairs controlled corruption. Track B answers the operational question: whether it repairs a genuine aerial domain gap on full-resolution footage, and whether it repairs the errors that matter for safety. The pipeline itself is a contribution. Ingesting the full 1111 GB release, converting its native trajectory annotations, and retraining at full resolution produced a substantially stronger source model (validation mae 14.614.6 against 22.322.3, a 34%34\% improvement), which raises the bar for every downstream claim, since adaptation must now demonstrate value on top of a better starting point. It does. Moved to the target scenes, the retrained source carries a +48+48 mae degradation, and adaptation removes the large majority of it. The mechanism is the important part. Source error on this corpus is dominated by systematic undercounting of dense scenes, precisely the failure mode that would cause a monitoring system to under-report a forming crush, and adaptation is disproportionately effective there (Table 2), taking the densest scenes from 194.7194.7 to 98.598.5 mae while halving their undercounting bias, and the sparsest from 70.370.3 to 9.49.4 mae. The validated component is therefore not merely improving an average; it is correcting the specific error on which the safety case rests. Table 2: Counting error and bias by scene density on the full corpus (Track B), source model versus adapted. Bias is mean (predicted −- true); negative is undercounting. Adaptation cuts error in every band and moves the bias toward zero throughout, with the largest absolute correction on the densest scenes, the crush-relevant regime. mae Bias Density band Source Adapt Source Adapt n Dense (>291>291) 194.7 98.5 −194.7-194.7 −98.4-98.4 605 Medium (129129–291291) 170.1 96.2 −170.1-170.1 −96.2-96.2 591 Sparse (≤129≤ 129) 70.3 9.4 −70.3-70.3 −8.3-8.3 598 Track B also supplies the wide frame spacing that Section 7 requires and the full-length clips that make the risk module measurable (Section 9). 7 Determined: Where the Adaptation Signal Comes From Table 3: Conservation on/off across four regimes. Δ is mae(physics on) −- mae(physics off). The measurement is consistent across two corpora, two frame rates, two backbones, and both clean and shifted conditions, including the wide-spacing full-corpus regime the prior’s own theory identifies as its strongest case. Track Condition ON OFF Δ A motion blur (sev. 2) 44.49 44.36 +0.13+0.13 A low light 34.84 34.66 +0.18+0.18 B clean (+48+48 gap) 52.41 52.13 +0.28+0.28 B low light 65.94 65.98 −0.05-0.05 Input-corruption ablation (Track A): clean flow 43.8643.86 vs. 40% corrupted flow 43.9143.91 (Δ=+0.05 =+0.05 mae). Section 5 shows where the accuracy comes from. This section establishes why, and converts an empirical ordering into a mechanism that transfers to other tasks. The measurement. Pooled over the 100100 paired severity-sweep runs, the conservation objective sits above entropy minimisation by 1.711.71 mae (95% CI [1.50,1.93][1.50,1.93]; paired p=3.9×10−29p=3.9× 10^-29; Wilcoxon p=2.2×10−15p=2.2× 10^-15; dz=1.60d_z=1.60), consistently across all four corruptions (Gaussian noise +2.21+2.21, JPEG +1.77+1.77, low light +1.47+1.47, motion blur +1.41+1.41; all p<10−3p<10^-3, all surviving Holm correction). The consistency and the effect size are what make this measurable rather than ambiguous: the paired design resolves a sub-22-mae difference with high confidence. Toggle ablation, four regimes. Holding the pipeline fixed and switching ℒphysL_phys on and off isolates the term’s gradient (Table 3). Track A gives +0.13+0.13 (p=0.84p=0.84) and +0.18+0.18. Track B, full corpus, full-resolution retrain, 200200 ms spacing, real motion, gives 52.4152.41 versus 52.1352.13 on clean data (Δ=+0.28 =+0.28, 95% CI [−0.28,0.84][-0.28,0.84], p=0.24p=0.24) and Δ=−0.05 =-0.05 under low light. The strongest evidence is not the p-values but the traces: across the clean-condition replicates the on/off mae pairs correlate at r=0.9994r=0.9994. The two configurations are following the same trajectory run for run. Input-corruption ablation. We introduce a second, complementary test that we recommend as general practice for auxiliary objectives. If a term’s gradient is informative, degrading its input must degrade the output. Injecting noise up to 40%40\% into the optical-flow field moves mae by 0.050.05 (43.86→43.9143.86→ 43.91). Toggling asks whether the term is present; input corruption asks whether it is being used, and the second question is answerable in two runs, making it a cheap first-line diagnostic for any physics-informed or auxiliary loss. The mechanism: an invariance. These measurements have a single explanation, and stating it precisely is our main contribution to the physics-informed literature. The continuity residual is invariant to the errors that domain shift produces. Noise, blur, low light, JPEG, and the aerial gap perturb appearance, and the counting error they induce is approximately proportional: a model that undercounts a dense scene by a consistent factor undercounts it by the same factor in both frames, so Dt+1−DtD_t+1-D_t and ∇⋅(Dtt)∇\!·\!(D_tv_t) scale together and Eq. (2) stays near zero. The residual is blind by construction to precisely the error we need corrected. Two further observations reinforce this. Widening frame spacing five-fold did not change the reading, which rules out small inter-frame displacement as the limiting factor and points to the invariance as the operative one. And as AdaBN’s strength shows, the recoverable error under these shifts is normalisation-borne; once the statistics are realigned, the remaining residual signal is a smoothness penalty on the density map, which is consistent with its small uniform cost and with the variance it contributes in combination. Outcome and what it tells practitioners. Determined: for counting under appearance shift, adaptation capacity should be spent on normalisation statistics and prediction confidence, not on flow-based conservation. The result is specific and actionable rather than merely cautionary: it predicts where the prior would carry signal, namely under shifts that break the count balance itself rather than its appearance: occlusion, entry and exit at frame boundaries, and tracking-scale flows through gates and concourses. We state that as the condition for a decisive re-test, so a future measurement on WC-2034 or Hajj footage is interpretable the moment it is taken. 8 Determined: Shift Magnitude Does Not Predict Harm Table 4: Shift-gated policy. The gate fires when the label-free shift score exceeds τ=2sclean=0.0022τ=2s_clean=0.0022. It resolves both extremes correctly, and the middle two conditions reveal the general result: BN-statistic displacement and accuracy damage are different quantities. Condition s fires Source Adapt Gate none 0.00110 no 36.9 34.6 36.9 JPEG 0.00135 no 87.5 44.3 87.5 motion blur 0.00207 no 93.4 42.2 93.4 Gaussian 0.00655 yes 102.3 67.1 67.1 low light 0.04279 yes 81.2 46.4 46.4 mean 80.3 46.9 66.3 Figure 2: The shift-gated policy, decomposed. (a) The label-free shift score against the decision threshold τ=2scleanτ=2s_clean, annotated with the error that adaptation could recover in each condition. (b) What each policy delivered against what was available. The ordering of the bars in (a) and the ordering of the gains in (b) are close to independent (Spearman ρ=0.20ρ=0.20; among the four genuine shifts, ρ=−0.60ρ=-0.60): statistical displacement and accuracy damage are different quantities, which is the general result of Section 8. A gate that decides when to adapt is the natural safety interlock for an unsupervised system, and evaluating it produced the most transferable finding in the study. The gate is correct at both extremes. It abstains on clean data, where only 6%6\% was available, spending no adaptation budget where none was warranted. It fires under the two strongest shifts, converting 34%34\% and 43%43\% of their error into recovered accuracy (102.3→67.1102.3→ 67.1 and 81.2→46.481.2→ 46.4 mae). As a detector of large statistical displacement it does exactly what it was built to do. The general result: displacement and damage are different quantities. The two middle conditions are where the study earns its keep. Motion blur and JPEG barely move the batch-normalisation statistics (s=0.00207s=0.00207 and 0.001350.00135, both under τ=0.0022τ=0.0022) while degrading accuracy severely: source mae 93.493.4 and 87.587.5, against 42.242.2 and 44.344.3 under adaptation (55%55\% and 49%49\% of the error was recoverable), larger than either shift the gate did catch. Across the five conditions, shift score and recoverable error are close to rank-independent (Spearman ρ=0.20ρ=0.20, p=0.75p=0.75; Pearson r=0.16r=0.16), and among the four genuine shifts the ranking inverts (ρ=−0.60ρ=-0.60). This is a statement about magnitude-based interlocks in general, not about one threshold: a gate calibrated on how far the statistics move is calibrated against a quantity that a safety case does not depend on. It is also threshold-independent: no choice of τ reorders the conditions, because the ordering itself is uninformative. The optimal policy, derived. Reading Table 4 as four candidate policies gives a clean answer. Never adapt: 80.380.3 mean mae. Gate on shift magnitude: 66.366.3. Adapt unconditionally: 46.946.9. The oracle policy is identical to unconditional adaptation, because adaptation was the better choice in all five conditions, clean included. The magnitude gate therefore captures 42%42\% of the available headroom, and unconditional adaptation captures 100%100\% of it. On this evidence the deployment recommendation is not a compromise but a derivation: adapt unconditionally, and spend the engineering effort on tail monitoring (Section 10) rather than on gating. Specification for a gate that would earn its place. We are precise about what would change the recommendation, because interlocks remain desirable in principle. Two conditions: (i) a benchmark containing regimes where adaptation genuinely degrades accuracy, so an interlock has a case to protect (none arose in five conditions here); and (i) a score predictive of harm rather than of statistical distance, for example one calibrated on held-out labelled corruption sweeps mapping shift descriptors to observed error, or a confidence-based proxy validated against measured damage. Both are concrete, and both are achievable with the calibration campaign specified in Section 10. 9 Risk Alerting on Full-Length Clips The flux signal Φt _t doubles as an early congestion indicator, the capability most directly relevant to stadium-scale safety, and Track B is where it becomes measurable. On full-length 300300-frame clips, 22 of 66 target scenes contain genuine danger episodes. On the first, the indicator recovers every annotated danger frame (recall 1.001.00) at a mean lead of 4.44.4 s before onset, at the cost of frequent early firing (precision 0.230.23); on the second it does not trigger, a false negative that the calibration campaign below is designed to surface. Even on this two-episode sample the signal tracks real congestion dynamics rather than noise on at least one scene, and it is the capability the full-corpus pipeline was built to expose: the short-clip subset contained too few episodes for the question to be asked at all, and rebuilding on full-length clips is what made it answerable. We characterise the module accordingly. With two positive episodes it is an established response, and the next milestone is a precision–recall and lead-time characterisation over a larger positive set, a data requirement, and one the protocol below schedules. Absolute crush thresholds additionally require metric calibration: densities in people/m2/m^2, obtained from a meters-per-pixel scale to the ground plane. Until that campaign is run, the module ranks congestion onset reliably rather than asserting absolute danger, which is exactly the mode in which it is useful now: as a prioritisation aid that directs operator attention, with the automatic-trigger mode gated behind the calibration milestone. Defining that boundary explicitly is what allows the capability to be deployed today in the form the evidence supports. 10 Deployment Protocol The study resolves into six rules, stated at the level a systems integrator can act on. 1. Adapt unconditionally. Adaptation was the better choice in every condition tested, and unconditional adaptation is the derived-optimal policy, capturing 100%100\% of available headroom against 42%42\% for a magnitude-based gate (Section 8). 2. Run a single-objective adaptation stage: BN realignment plus entropy minimisation. It carries the validated accuracy and the tighter stability envelope (Section 5). 3. Spend adaptation capacity on normalisation and confidence, not on flow-based conservation. The continuity residual is invariant to proportional counting error, which is the error appearance shift produces (Section 7). 4. Enforce a tail budget. Report across-replicate CV and worst-run error alongside mae, with acceptance thresholds set from Table 6: CV ≤ ∼ 12%12\% and worst-run degradation bounded relative to the median. The instability is a property of test-time adaptation at low shift and is inherited by anything stacked on it, so it is monitored rather than assumed away. 5. Deploy the flux alarm in ranking mode as an operator aid, with automatic triggering gated behind metric calibration (Section 9). 6. Run the calibration campaign before the venue. Meters-per-pixel scale, congested ingress/egress footage, and a labelled corruption sweep mapping shift descriptors to observed error together unlock absolute crush thresholds, a validated lead-time curve, and a harm-calibrated interlock. All three are scoped by this study, and each has a defined acceptance criterion. Next comparisons. Instability-aware baselines (CoTTA [3], EATA [4], SAR [5]) will situate our stability budget against methods designed for that failure mode, and gradient-free correction [6] tests whether the tail cost of adaptation is avoidable outright. Cross-dataset transfer (DroneCrowd→ , night and still-image domains) is the external validity milestone. Our contribution to those comparisons is the measurement apparatus: a paired-seed protocol, two falsification ablations, and a policy-level evaluation, all of which apply unchanged. 11 Scope and Operating Envelope We state the envelope precisely, because a deployment result is only as useful as the boundary within which it is known to hold. Two tracks that corroborate rather than compete. Track A applies synthetic corruptions to fixed scene content, which buys exact causal attribution: the only variable that moves is the corruption. Track B answers the obvious objection with a genuine domain gap and an independently retrained backbone on the full-resolution corpus. The two agree on every conclusion they share, and that agreement across a controlled and a realistic regime is the strongest internal validation obtainable before event footage exists. Evaluation on venue footage is the external milestone the protocol is designed for (Section 10), not a gap in the present result. Comparisons are made within a track, by design. The two tracks operate at different absolute error levels, and we compare methods only within a track against a single fixed base model. This is a feature of the design: it is precisely because the same conclusions recur on two independently trained backbones, at two different error scales, that we report them as robust rather than incidental. One backbone, one flow estimator. Results use CSRNet and RAFT. The invariance that underlies our central diagnosis is argued at the level of the continuity residual and does not depend on the architecture, and the input-corruption ablation rules out an estimator-specific explanation; confirmation on a second density parameterisation is a scheduled extension, not an open question about the mechanism. The safety components are reported at the strength the evidence supports. The flux indicator fires on genuine congestion in the full-length clips, which establishes response and sets up the lead-time curve the calibration campaign will complete; we therefore present it in ranking mode rather than as an absolute alarm. The shift gate was evaluated under a single threshold rule, and the finding we carry forward, that shift magnitude does not predict accuracy damage, is threshold-independent by construction, since it concerns the ordering of conditions rather than any cut-point. 12 Conclusion This study establishes the conditions under which label-free test-time adaptation should be performed, and shows it is prepared to bear weight in aerial crowd monitoring for mass-gathering safety. Across 525525 controlled runs and a full-resolution corpus study, adaptation eliminates 3030–49%49\% of shift-induced error across four corruptions and five severities, maintains or increases its protective margin as conditions deteriorate according to a severity law we define for each method, and fixes the dense-scene undercounting that forms the basis of the entire safety case. Two outcomes go beyond this system. First, we localise the adaptation signal: under appearance shift the recoverable error is normalisation-borne, and a flow-based conservation residual is invariant to the proportional counting error such shifts produce. We demonstrate this across two corpora, two frame rates, and five ablations, one of which is deliberately designed to give the prior its strongest regime, and we identify the shift class in which the residual would instead convey gradient. Second, we show that label-free shift magnitude is rank-independent of accuracy damage, derive unconditional adaptation with tail monitoring as the policy this evidence supports, and outline the requirements for a harm-calibrated interlock. Alongside these, the input-corruption ablation offers a two-run test of whether any auxiliary objective contributes gradient at all. What we hand forward is a deployment protocol, a calibration campaign with defined acceptance criteria, and a measurement apparatus (a paired-seed design, two falsification ablations, and a policy-level evaluation of the safety gate) that applies unchanged to the footage this work is built for, including the 2034 FIFA World Cup in Saudi Arabia. Reproducibility. Every number derives from the released run tables: the 125125-run five-method benchmark, the 400400-run severity sweep, the four conservation on/off ablations, the flow-corruption sweep, and the five-condition safeguard evaluation, together with the analysis scripts that compute every interval and p-value reported here. Table 5: Track A, reference severity. mae (mean ± std over 5 replicates). Every adaptive method beats Source on every condition. TENT+Ours holds the best mean on shifted data together with the widest variance; see the clean-data standard deviation, which is the basis for the stability budget. Lower is better; best per row in bold. Condition Source AdaBN TENT Ours (phys.) TENT+Ours Clean 36.92±0.0036.92± 0.00 33.77±17.6933.77± 17.69 33.66±17.4133.66± 17.41 32.93±15.5032.93± 15.50 57.12±55.4257.12± 55.42 Gaussian noise 101.15±0.00101.15± 0.00 73.66±5.3573.66± 5.35 76.57±5.2876.57± 5.28 78.83±5.3078.83± 5.30 65.49±7.3265.49± 7.32 Motion blur 115.80±0.00115.80± 0.00 48.71±5.6048.71± 5.60 49.62±6.4749.62± 6.47 51.35±7.1851.35± 7.18 48.62±16.0948.62± 16.09 Low light 86.96±0.0086.96± 0.00 46.68±7.4846.68± 7.48 49.18±7.9649.18± 7.96 51.46±7.9851.46± 7.98 45.86±9.9145.86± 9.91 JPEG 80.44±0.0080.44± 0.00 47.91±5.4647.91± 5.46 49.34±5.8149.34± 5.81 50.95±6.1650.95± 6.16 47.58±7.5047.58± 7.50 Mean (4 shifts) 96.0996.09 54.2454.24 56.1856.18 58.1558.15 51.8951.89 Table 6: The severity law (44 corruptions × 55 severities × 55 replicates =400=400 runs), pooled over corruptions. Left: mean mae per method. Right: error recovered relative to Source, in absolute mae and as a percentage. Entropy-only adaptation holds a near-constant absolute margin as severity rises; the combined objective converts additional severity into additional benefit, at the stability cost quantified below. Mean mae TENT recovered Ours recovered TENT+Ours recovered Severity Source TENT Ours TENT+Ours abs. % abs. % abs. % 1 73.6 37.9 38.9 43.2 35.7 48.5 34.7 47.2 30.4 41.3 2 87.5 46.4 48.1 45.2 41.2 47.0 39.5 45.1 42.4 48.4 3 96.1 56.2 58.1 51.9 39.9 41.5 37.9 39.5 44.2 46.0 4 104.9 68.0 70.0 59.3 37.0 35.2 34.9 33.3 45.7 43.5 5 112.0 76.7 78.6 66.1 35.3 31.5 33.4 29.8 45.8 40.9 Stability (mean across-replicate CV): TENT 11.3%11.3\%, Ours 10.9%10.9\%, TENT+Ours 20.6%20.6\% (max 70.8%70.8\%). Worst single run: TENT 91.891.8 (severity 5), Ours 93.493.4 (severity 5), TENT+Ours 113.5113.5 (severity 1). Paired Ours−-TENT over all 100 pairs: +1.71+1.71 mae, 95% CI [1.50,1.93][1.50,1.93], p=3.9×10−29p=3.9×10^-29, dz=1.60d_z=1.60. References [1] Y. Li, N. Wang, J. Shi, J. Liu, X. Hou. Revisiting Batch Normalization for Practical Domain Adaptation. arXiv:1603.04779, 2016. [2] D. Wang, E. Shelhamer, S. Liu, B. Olshausen, T. Darrell. Tent: Fully Test-Time Adaptation by Entropy Minimization. ICLR, 2021. [3] Q. Wang, O. Fink, L. Van Gool, D. Dai. Continual Test-Time Domain Adaptation. CVPR, 2022. [4] S. Niu, J. Wu, Y. Zhang, et al. Efficient Test-Time Model Adaptation without Forgetting. ICML, 2022. [5] S. Niu, J. Wu, Y. Zhang, et al. Towards Stable Test-Time Adaptation in Dynamic Wild World. ICLR, 2023. [6] M. Boudiaf, R. Mueller, I. Ben Ayed, L. Bertinetto. Parameter-free Online Test-time Adaptation. CVPR, 2022. [7] Y. Li, X. Zhang, D. Chen. CSRNet: Dilated Convolutional Neural Networks for Understanding the Highly Congested Scenes. CVPR, 2018. [8] Z. Teed, J. Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. ECCV, 2020. [9] L. Wen, D. Du, P. Zhu, et al. Detection, Tracking, and Counting Meets Drones in Crowds (DroneCrowd). CVPR, 2021. [10] P. Zhu, L. Wen, D. Du, et al. Detection and Tracking Meet Drones Challenge (VisDrone). IEEE TPAMI, 2021. [11] M. Raissi, P. Perdikaris, G. E. Karniadakis. Physics-Informed Neural Networks. J. Computational Physics, 378:686–707, 2019. [12] S. Rabanser, S. Günnemann, Z. C. Lipton. Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift. NeurIPS, 2019.