Paper deep dive
Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference
Usef Faghihi, Amir Saki
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Structured potential outcomes such as microscopy images may be recorded after an unknown, unit-specific transformation. If that transformation can depend on treatment, covariates or the intrinsic outcome, raw-coordinate analyses may mix biological effects with acquisition geometry. We study the unrestricted observation model X = {\Gamma} . Y(A) and characterize its observable information: a target is uniformly recoverable exactly when it is constant on group orbits, while a Borel maximal invariant retains every measurable invariant target. We then distinguish observability from statistical losslessness. A quotient-faithful reconstruction theorem shows that quotient reduction is sufficient for the full transformed experiment exactly when the conditional law of the raw observation given treatment, covariates and the quotient has a parameter-free version. Conditional Haar contamination on a compact group yields Blackwell equivalence as a special case; it is not imposed in the main model. We also separate independent site-specific product actions from shared diagonal actions and show why componentwise canonicalization can discard relative cross-site information. Under explicit metric and kernel regularity, an approximate-contamination theorem bounds quotient-law Wasserstein error and the induced perturbation of population maximum mean discrepancy. For finite-support multichannel lattice images, we construct a maximal invariant under integer translations and quarter turns, combine its characteristic Gaussian kernel with a complete paired-swap test, and retain the original simulations and RxRx1 HUVEC study. Under the sharp null, the quotient test rejected in 0.052 of simulation replicates; at unit effect strength its power was 0.992. The primary RxRx1 contrast had an enumerated paired-swap p-value of 0.0078.
Tags
Links
- Source: https://arxiv.org/abs/2608.11954v1
- Canonical: https://arxiv.org/abs/2608.11954v1
Trouble viewing inline? Open PDF directly →
Full Text
59,394 characters extracted from source content.
Expand or collapse full text
Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference Usef Faghihi Amir Saki[0.5em] Département de mathématiques et d’informatiqueUniversité du Québec à Trois-RivièresTrois-Rivières, Québec, Canada [0.3em] These authors contributed equally to this work.Usef Faghihi: usef.faghihi@uqtr.caAmir Saki: amir.saki@uqtr.ca Abstract Structured potential outcomes such as microscopy images may be recorded after an unknown, unit-specific transformation. If that transformation can depend on treatment, covariates or the intrinsic outcome, raw-coordinate analyses may mix biological effects with acquisition geometry. We study the unrestricted observation model X=Γ⋅Y(A)X= · Y(A) and characterize its observable information: a target is uniformly recoverable exactly when it is constant on group orbits, while a Borel maximal invariant retains every measurable invariant target. We then distinguish observability from statistical losslessness. A quotient-faithful reconstruction theorem shows that quotient reduction is sufficient for the full transformed experiment exactly when the conditional law of the raw observation given treatment, covariates and the quotient has a parameter-free version. Conditional Haar contamination on a compact group yields Blackwell equivalence as a special case; it is not imposed in the main model. We also separate independent site-specific product actions from shared diagonal actions and show why componentwise canonicalization can discard relative cross-site information. Under explicit metric and kernel regularity, an approximate-contamination theorem bounds quotient-law Wasserstein error and the induced perturbation of population maximum mean discrepancy. For finite-support multichannel lattice images, we construct a maximal invariant under integer translations and quarter turns, combine its characteristic Gaussian kernel with a complete paired-swap test, and retain the original simulations and RxRx1 HUVEC study. Under the sharp null, the quotient test rejected in 0.052 of simulation replicates; at unit effect strength its power was 0.992. The primary RxRx1 contrast had an enumerated paired-swap p-value of 0.0078. Keywords: causal inference; group action; maximal invariant; Blackwell sufficiency; microscopy image; quotient space; randomization test; structured outcome. 1 Introduction Many experimental outcomes are structured objects rather than scalars. An image, shape, network or function may be represented in a coordinate system that varies from unit to unit. In microscopy, for example, translation or orientation can arise from acquisition and preprocessing rather than from the biological intervention. We study the observation model X=Γ⋅Y(A),X= · Y(A), (1.1) where A is treatment, Y(a)Y(a) is the intrinsic structured potential outcome and Γ is an unobserved group element. Crucially, no independence between Γ and treatment, covariates or potential outcomes is imposed. Averaging over a posited nuisance distribution is therefore neither required nor generally available. The first question is an information boundary: which targets of Y(a)Y(a) are determined by the transformed observation? Orbit invariance gives an exact answer. The second question is causal: under what design or exchangeability conditions does the observed quotient distribution identify its interventional counterpart? A third, logically stronger question is statistical: when does discarding the within-orbit coordinate lose no information about the full family of observed-data laws? Maximal invariance answers the first question, but not the third. Conflating those statements would overstate what quotienting achieves. This paper develops a theorem-to-algorithm-to-application account of these three layers. The principal additions to the original outcome-level framework are: 1. a sharp sufficiency boundary based on a common conditional law and a quotient-faithful reconstruction kernel; 2. a conditional-Haar Blackwell-equivalence corollary, stated only as a special case of the unrestricted contamination framework; 3. a product-versus-diagonal action result that identifies when sitewise canonicalization is appropriate and when it erases relative configuration; 4. an approximate-contamination stability result under explicit orbit-metric and reproducing-kernel assumptions. The individual ingredients—maximal invariants, measurable factorization, statistical sufficiency, Blackwell comparison, Haar measure, characteristic kernels and Fisher randomization—are classical [2, 5, 25, 6]. The contribution claimed here is their outcome-level causal synthesis under unrestricted latent group contamination, together with an explicit noncompact lattice construction and a finite-sample inferential workflow. We do not claim novelty for those classical ingredients separately. Existing causal methods for metric- and random-object outcomes generally take the structured outcome as directly observed [11, 20, 1]; registration-based functional methods address a related preprocessing problem [15]. Topological transforms can be injective on specified shape classes when sufficiently rich directional information is available [24, 4]. Our setting is complementary: the observation equation itself contains a latent, unit-specific action, and the target is the full observable orbit rather than a chosen low-dimensional summary. The application uses two-site, six-channel RxRx1 wells. Under separate site-specific acquisition motions the nuisance group is a product of lattice rigid-motion groups. Translation normalization and finite rotational minimization yield a measurable maximal invariant without placing a probability law on the noncompact translation group. A Gaussian kernel on canonical representatives is characteristic for quotient laws, and complete enumeration of the reported paired assignment mechanism yields finite-sample inference for the quotient Fisher sharp null. Section 2 develops observability and quotient-law identification. Section 3 gives the statistical-losslessness boundary and Haar special case. Section 4 distinguishes product and diagonal actions. Section 5 proves stability under approximate contamination. Sections 6 and 7 give the image construction and exact test. Sections 8 and 9 report the preserved empirical results. Proof details that would interrupt the main argument appear in Appendix A. 2 Observable quotient targets and causal identification 2.1 Measurable group contamination Let ,,Y,C,S and G be standard Borel spaces. Assume that G is a measurable group and acts on Y through a jointly Borel map (g,y)↦g⋅y(g,y) g· y. The orbit of y is [y]G=g⋅y:g∈G[y]_G=\g· y:g∈ G\. A Borel map τ:→τ:Y is invariant if τ(g⋅y)=τ(y)τ(g· y)=τ(y) for all (g,y)(g,y). Let C denote the observed pretreatment covariates, taking values in the standard Borel space C, and let A denote the realized treatment assignment, taking values in a finite set A. For each a∈a , let Y(a)Y(a) be the intrinsic potential outcome under treatment a in the potential-outcomes framework [19]. We observe O=(C,A,X)O=(C,A,X), where X satisfies (1.1). The conditional distribution of Γ given C, A and the potential outcomes is unrestricted. Theorem 2.1 (Sharp observability). For a Borel target τ:→τ:Y , the following are equivalent: 1. there is a Borel decoder D:→D:Y such that D(g⋅y)=τ(y)D(g· y)=τ(y) for every g∈Gg∈ G and y∈y ; 2. τ is G-invariant. Proof. If τ is invariant, take D=τD=τ. Conversely, the identity element gives D(z)=τ(z)D(z)=τ(z) for every z. Hence τ(g⋅y)=D(g⋅y)=τ(y)τ(g· y)=D(g· y)=τ(y). ∎ Corollary 2.2 (Nonidentifiability outside the quotient). If τ is not invariant, two data-generating processes satisfying (1.1) can induce the same observed law while giving different laws of τY(a)τ\Y(a)\ for any fixed a. Proof. Choose y and g with τ(g⋅y)≠τ(y)τ(g· y)≠τ(y). In one process set Y(a)=yY(a)=y and Γ=g =g; in the other set Y(a)=g⋅yY(a)=g· y and Γ=e =e. Both give X=g⋅yX=g· y, but their target values differ. ∎ 2.2 Maximal invariant and factorization Definition 2.3. A Borel map M:→M:Y is a maximal invariant when M(y)=M(z)⟺z=g⋅y for some g∈G.M(y)=M(z) z=g· y for some g∈ G. Theorem 2.4 (Borel factorization). Let ,,Y,S,T be standard Borel spaces and M:→M:Y a Borel maximal invariant. Every Borel invariant h:→h:Y admits a Borel factorization h=h¯∘Mh= h M. The restriction of h¯ h to M()M(Y) is unique. The proof, given in Appendix A, uses saturation of h−1(B)h^-1(B), Lusin separation of two disjoint analytic images and the measurable factorization lemma; these are standard tools of descriptive set theory [10]. Thus M retains every measurable invariant target; it does not follow that M(X)M(X) is sufficient for a statistical parameter. 2.3 Identification of the interventional quotient law For each a∈a , define the quotient potential outcome Q(a)=MY(a),Q(a)=M\Y(a)\, and define the observed quotient outcome by Q=M(X)Q=M(X). Invariance then gives Q=M(X)=MΓ⋅Y(A)=MY(A)=Q(A)Q=M(X)=M\ · Y(A)\=M\Y(A)\=Q(A) (2.1) without imposing any restriction on the conditional distribution of Γ . Theorem 2.5 (Identification of the full quotient law). Suppose, for every a∈a , that quotient consistency (2.1) holds, Q(a)⟂A|CQ(a) \!\!\! A C, and ℙ(A=a∣C)>0P(A=a C)>0 almost surely. Then, for every Borel B⊆B , ℙQ(a)∈B=∫ℙ(Q∈B∣A=a,C=c)dPC(c).P\Q(a)∈ B\= (Q∈ B A=a,C=c)\,dP_C(c). (2.2) Consequently, the entire interventional law of the maximal invariant is identified. Proof. Standard Borel spaces admit regular conditional probabilities [9]. Apply iterated expectation, conditional exchangeability, positivity and consistency in that order. Equality for every Borel B determines the probability measure on S. ∎ By Theorem 2.4, the interventional law of any Borel invariant target is the pushforward of (2.2) through h¯ h. This is the ordinary adjustment formula [17, 18] applied to a quotient-valued outcome, and it remains distinct from the sufficiency question considered next. 3 When quotient reduction is statistically lossless Let Θ index a family of observed-data laws. We use the standard comparison-of-experiments framework [12, 23]. Put Ofull=(C,A,X),Oquot=T(Ofull)=(C,A,M(X)),O_full=(C,A,X), O_quot=T(O_full)=(C,A,M(X)), and write ℰfull=Pθfull:θ∈ΘE_full=\P_θ^full:θ∈ \ and ℰquot=Pθquot:θ∈ΘE_quot=\P_θ^quot:θ∈ \. Since T is deterministic, the full experiment Blackwell-dominates the quotient experiment. The reverse comparison requires an additional condition. Definition 3.1 (Quotient-faithful kernel). A Markov kernel K from the quotient sample space to the full sample space is quotient-faithful for ℰquotE_quot if K(s,T−1(s))=1Pθquot-almost surely for every θ∈Θ.K\! (s,T^-1(\s\) )=1 P_θ^quot-almost surely for every θ∈ . Thus, up to quotient-null sets under each member of the experiment, the kernel reconstructs a raw representative without changing the supplied covariates, treatment, or quotient value. Theorem 3.2 (Boundary of lossless quotient reduction). Assume the full and quotient sample spaces are standard Borel. The following are equivalent: 1. there is a Markov kernel K, independent of θ, such that, for every θ∈Θθ∈ , K is a regular conditional distribution of OfullO_full given OquotO_quot under PθfullP_θ^full; equivalently, K(s,⋅)=Lawθ(Ofull∣Oquot=s)Pθquot-almost surely.K(s,·)= _θ(O_full O_quot=s) P_θ^quot-almost surely. 2. there is a parameter-free quotient-faithful kernel K such that PθquotK=Pθfullfor every θ∈Θ.P_θ^quotK=P_θ^full every θ∈ . (3.1) When these conditions hold, OquotO_quot is sufficient for ℰfullE_full and ℰfull≡BℰquotE_full _BE_quot. Hence every bounded decision problem has the same attainable risk set under the two experiments. If the family is dominated, the conditions are also equivalent to Fisher–Neyman factorization through OquotO_quot. Proof. Write Qθ:=PθquotQ_θ:=P_θ^quot. Suppose first that (1) holds. By the disintegration identity for the common regular conditional distribution, for every measurable set B in the full sample space, Pθfull(B)=∫K(s,B)Qθ(s)=(QθK)(B).P_θ^full(B)= K(s,B)\,Q_θ(ds)=(Q_θK)(B). Hence QθK=PθfullQ_θK=P_θ^full for every θ. Moreover, since Oquot=T(Ofull)O_quot=T(O_full) deterministically, a regular conditional law of OfullO_full given Oquot=sO_quot=s is supported on T−1(s)T^-1(\s\), QθQ_θ-almost surely. Thus K is quotient-faithful, proving (2). Conversely, suppose that (2) holds. Let B be measurable in the full sample space and let D be measurable in the quotient sample space. Quotient faithfulness implies that, QθQ_θ-almost surely, K(s,B∩T−1(D))=D(s)K(s,B).K\! (s,B∩ T^-1(D) )=1_D(s)K(s,B). Indeed, conditional on input s, the kernel is supported on T−1(s)T^-1(\s\); this fibre is contained in T−1(D)T^-1(D) when s∈Ds∈ D and is disjoint from T−1(D)T^-1(D) when s∉Ds∉ D. Using QθK=PθfullQ_θK=P_θ^full with the measurable set B∩T−1(D)B∩ T^-1(D) therefore gives Pθfull(B∩T−1(D)) P_θ^full\! (B∩ T^-1(D) ) =∫K(s,B∩T−1(D))Qθ(ds) = K\! (s,B∩ T^-1(D) )\,Q_θ(ds) =∫DK(s,B)Qθ(ds). = _DK(s,B)\,Q_θ(ds). Since T(Ofull)=OquotT(O_full)=O_quot, this is precisely Pθfull(Ofull∈B,Oquot∈D)=∫DK(s,B)Pθquot(s),P_θ^full\! (O_full∈ B,\ O_quot∈ D )= _DK(s,B)\,P_θ^quot(ds), the defining identity for K to be a regular conditional distribution of OfullO_full given OquotO_quot. The kernel is common to all θ, so (1) follows. The deterministic map T gives ℰfull⪰BℰquotE_full _BE_quot, while K gives ℰquot⪰BℰfullE_quot _BE_full. Hence the experiments are Blackwell equivalent. The decision-theoretic assertion follows from Blackwell equivalence, and the dominated factorization equivalence is the standard Fisher–Neyman sufficiency theorem. ∎ The quotient-faithfulness clause is essential for the stated equivalence. An arbitrary reverse kernel may reproduce each full marginal law while scrambling quotient values; such a kernel need not be a conditional distribution given the supplied quotient. Example 3.3 (Maximal invariance is not sufficiency). Let G=−1,1G=\-1,1\ act on −1,1\-1,1\ by multiplication. The maximal invariant is constant. Let Y=1Y=1 and let ℙθ(Γ=1)=θP_θ( =1)=θ, so X=ΓX= and ℙθ(X=1)=θP_θ(X=1)=θ. The raw observation identifies θ, whereas the quotient law is identical for every θ. Thus the conditional law of X given the quotient depends on θ and the quotient is not sufficient. Under unrestricted contamination, within-orbit coordinates may carry information about the nuisance mechanism even though they carry no uniformly observable information about the intrinsic representative. 3.1 Conditional-Haar contamination as a special case Let G be compact and second countable, and let μG _G denote its normalized Haar probability measure. Assume that G acts jointly Borel measurably on Y and that M:→M:Y is a Borel maximal invariant. Put 0=M(),S_0=M(Y), and assume that 0S_0 is a Borel subset of S. Let s:0→s:S_0 be a Borel section satisfying Ms(m)=mfor every m∈0.M\s(m)\=m every m _0. For m∈0m _0 and Borel B⊆B , define the orbit kernel Km(B)=∫GBg⋅s(m)dμG(g).K_m(B)= _G1_B\g· s(m)\\,d _G(g). (3.2) The kernel may be extended arbitrarily and measurably from 0S_0 to S; its values outside 0S_0 are irrelevant because M(X)∈0M(X) _0 almost surely. Right invariance of Haar measure implies that KmK_m does not depend on the particular representative selected from the orbit indexed by m. Corollary 3.4 (Conditional-Haar Blackwell equivalence). Suppose that, for every θ∈Θθ∈ , LawθΓ∣C,A,Y(A)=μG _θ\ C,A,Y(A)\= _G (3.3) almost surely. Then, for PθquotP_θ^quot-almost every (C,A,m)(C,A,m), LawθX∣C,A,M(X)=m=Km _θ\X C,A,M(X)=m\=K_m has a version that is independent of θ. Consequently, ℰfull≡Bℰquot.E_full _BE_quot. Proof. Fix a Borel set B⊆B . The joint measurability of the action and the measurability of the section ensure that m↦Km(B)m K_m(B) is Borel. Condition first on (C,A,Y(A))(C,A,Y(A)). Under (3.3), ℙθX∈B∣C,A,Y(A) _θ\X∈ B C,A,Y(A)\ =∫GB(g⋅Y(A))dμG(g). = _G1_B\! (g· Y(A) )\,d _G(g). Let y∈y and put m=M(y)m=M(y). Since s(m)s(m) and y lie in the same orbit, maximality of M yields an h∈Gh∈ G such that y=h⋅s(m).y=h· s(m). Consequently, ∫GB(g⋅y)dμG(g)=∫GB((gh)⋅s(m))dμG(g). _G1_B(g· y)\,d _G(g)= _G1_B\! ((gh)· s(m) )\,d _G(g). By right invariance of normalized Haar probability, ghgh has law μG _G whenever g has law μG _G. Hence ∫GB(g⋅y)dμG(g)=∫GB(g⋅s(m))dμG(g)=Km(B). _G1_B(g· y)\,d _G(g)= _G1_B\! (g· s(m) )\,d _G(g)=K_m(B). It follows that ℙθX∈B∣C,A,Y(A)=KMY(A)(B)almost surely.P_θ\X∈ B C,A,Y(A)\=K_M\Y(A)\(B) surely. Taking conditional expectations given σC,A,M(Y(A))σ\C,A,M(Y(A))\ therefore gives ℙθX∈B∣C,A,M(Y(A)) _θ\X∈ B C,A,M(Y(A))\ =θ[KMY(A)(B)∣C,A,M(Y(A))] =E_θ\! [K_M\Y(A)\(B) C,A,M(Y(A)) ] =KMY(A)(B)almost surely. =K_M\Y(A)\(B) surely. Since M is invariant and X=Γ⋅Y(A)X= · Y(A), M(X)=MΓ⋅Y(A)=MY(A)almost surely.M(X)=M\ · Y(A)\=M\Y(A)\ surely. Thus LawθX∣C,A,M(X)=m=Km _θ\X C,A,M(X)=m\=K_m for PθquotP_θ^quot-almost every (C,A,m)(C,A,m), and this version is independent of θ. For completeness, define a kernel from the quotient sample space to the full sample space by K~((c,a,m),B):=∫B(c,a,x)Km(x), K ((c,a,m),B ):= _Y1_B(c,a,x)\,K_m(dx), for every measurable B in the full sample space. This kernel is parameter-free. Moreover, KmK_m is supported on M−1(m)M^-1(\m\), since Mg⋅s(m)=Ms(m)=mfor every g∈G.M\g· s(m)\=M\s(m)\=m every g∈ G. Hence K~ K is quotient-faithful. It is a common regular conditional distribution of OfullO_full given OquotO_quot, so Theorem 3.2 yields ℰfull≡Bℰquot.E_full _BE_quot. ∎ The logical hierarchy is therefore: unrestricted Γ:MY(a) is the maximal uniformly observable target;parameter-free within-orbit law:M(X) is statistically sufficient;conditional Haar Γ:a concrete compact-group condition implying sufficiency. array[]lunrestricted :&M\Y(a)\ is the maximal uniformly observable target;\\ parameter-free within-orbit law:&M(X) is statistically sufficient;\\ conditional Haar :&a concrete compact-group condition implying sufficiency. array The main results and empirical construction use the first line. The Haar assumption is not needed for observability, quotient identification or exact randomization inference. 4 Independent product actions and shared diagonal actions Structured units often contain labeled components, such as two microscopy sites. The appropriate group depends on whether acquisition transformations are component-specific or shared. Proposition 4.1 (Product versus diagonal actions). Let M:→M:Y be a maximal invariant for a G-action and let J≥2J≥ 2. Define M×J(y1,…,yJ)=(M(y1),…,M(yJ)).M^× J(y_1,…,y_J)= (M(y_1),…,M(y_J) ). 1. Under the product action of GJG^J, (g1,…,gJ)⋅(y1,…,yJ)=(g1⋅y1,…,gJ⋅yJ),(g_1,…,g_J)·(y_1,…,y_J)=(g_1· y_1,…,g_J· y_J), M×JM^× J is a maximal invariant. 2. Under the diagonal action of G, g⋅(y1,…,yJ)=(g⋅y1,…,g⋅yJ),g·(y_1,…,y_J)=(g· y_1,…,g· y_J), M×JM^× J is invariant. Moreover, on any diagonally G-invariant model subset ⊆JD ^J, it is maximal for the restricted diagonal action if and only if, whenever y,z∈y,z and zj=gj⋅yjz_j=g_j· y_j for component-specific gjg_j, there is a single g∈Gg∈ G satisfying zj=g⋅yjz_j=g· y_j for every j. Thus componentwise canonicalization can discard relative cross-component information under a shared transformation. Proof. For (1), suppose first that M×J(y1,…,yJ)=M×J(z1,…,zJ).M^× J(y_1,…,y_J)=M^× J(z_1,…,z_J). Then M(yj)=M(zj)M(y_j)=M(z_j) for every j. By maximality of M, for each j there exists gj∈Gg_j∈ G such that zj=gj⋅yjz_j=g_j· y_j. Therefore (z1,…,zJ)=(g1,…,gJ)⋅(y1,…,yJ),(z_1,…,z_J)=(g_1,…,g_J)·(y_1,…,y_J), so the two points belong to the same GJG^J-orbit. The converse follows immediately from the invariance of M in each component. Hence M×JM^× J is maximal for the product action. For (2), invariance under the diagonal action follows because M(g⋅yj)=M(yj)for every j.M(g· y_j)=M(y_j) every j. Now let y,z∈y,z . Equality M×J(y)=M×J(z)M^× J(y)=M^× J(z) is equivalent, by maximality of M, to the existence of possibly different elements g1,…,gJ∈Gg_1,…,g_J∈ G satisfying zj=gj⋅yj,j=1,…,J.z_j=g_j· y_j, j=1,…,J. The points y and z belong to the same diagonal orbit precisely when these component-specific transformations can be replaced by a single g∈Gg∈ G satisfying zj=g⋅yjz_j=g· y_j for every j. This is exactly the stated condition. ∎ Example 4.2 (Relative displacement). Let G=(ℝ,+)G=(R,+) act on ℝR by translation. A one-component maximal invariant is constant because the action is transitive. For a pair (y1,y2)(y_1,y_2) under the diagonal action, the difference y2−y1y_2-y_1 is invariant and separates diagonal orbits. Componentwise quotienting is constant and therefore destroys this relative displacement. The image analogue is relative position or orientation between sites subjected to one common acquisition motion. In contrast, when the two sites undergo genuinely separate motions, relative pose is not uniformly observable and the product quotient is appropriate. For the RxRx1 stress test below, the imposed transformation generator assigns separate site-specific motions, so the product action is the declared model. Proposition 4.1(1) prevents that application-specific choice from being misread as a general prescription. 5 Stability under approximate contamination Exact invariance is action-specific. To obtain a correct perturbation result, additional metric regularity is needed. Let (,d)(Y,d) be a metric space on which G acts by isometries and define the orbit pseudometric dG(y,z)=infg∈Gd(y,g⋅z).d_G(y,z)= _g∈ Gd(y,g· z). (5.1) Let (,d)(S,d_S) contain a quotient representation M satisfying dM(y),M(z)≤LMdG(y,z).d_S\M(y),M(z)\≤ L_Md_G(y,z). (5.2) Let 1()P_1(S) denote the Borel probability measures on S having a finite first moment with respect to d_S. For a measurable positive-semidefinite kernel k with reproducing kernel Hilbert space ℋkH_k and Borel feature map Φ:→ℋk :S _k, define μk(P)=∫Φ(s)P(s) _k(P)= _S (s)\,P(ds) whenever this Bochner integral exists, and put MMDk(P,Q)=‖μk(P)−μk(Q)‖ℋk. _k(P,Q)=\| _k(P)- _k(Q)\|_H_k. Theorem 5.1 (Quotient-law and MMD stability). For arms a∈0,1a∈\0,1\, let (Xa,Ya)(X_a,Y_a) be a coupling of the approximately contaminated and intrinsic outcomes, and assume dG(Xa,Ya)≤εa.Ed_G(X_a,Y_a)≤ _a. Write P~a=LawM(Xa),Pa=LawM(Ya), P_a= \M(X_a)\, P_a= \M(Y_a)\, and assume that P~a,Pa∈1() P_a,P_a _1(S). Then W1,d(P~a,Pa)≤LMεa.W_1,d_S( P_a,P_a)≤ L_M _a. (5.3) If the mean embeddings of P~a P_a and PaP_a exist and the feature map satisfies ‖Φ(s)−Φ(t)‖ℋk≤Lkd(s,t),\| (s)- (t)\|_H_k≤ L_kd_S(s,t), (5.4) then |MMDk(P~1,P~0)−MMDk(P1,P0)|≤LkLM(ε1+ε0). | _k( P_1, P_0)- _k(P_1,P_0) |≤ L_kL_M( _1+ _0). (5.5) If the perturbation bounds hold almost surely, the same conclusions follow with those uniform bounds. Proof. The supplied coupling and (5.2) give W1,d(P~a,Pa)≤dM(Xa),M(Ya)≤LMεa.W_1,d_S( P_a,P_a) _S\M(X_a),M(Y_a)\≤ L_M _a. Let μ(R)=S∼RΦ(S)μ(R)=E_S R (S). By the same coupling and (5.4), ‖μ(P~a)−μ(Pa)‖ℋk≤LkLMεa.\|μ( P_a)-μ(P_a)\|_H_k≤ L_kL_M _a. Since MMDk(R,S)=‖μ(R)−μ(S)‖ℋk _k(R,S)=\|μ(R)-μ(S)\|_H_k, the reverse triangle inequality followed by the last two bounds proves (5.5). ∎ For a Gaussian kernel k(s,t)=exp−∥s−t∥2/(2σ2)k(s,t)= \-\|s-t\|^2/(2σ^2)\ on a Hilbert quotient embedding, ‖Φ(s)−Φ(t)‖2=21−k(s,t)≤‖s−t‖2/σ2,\| (s)- (t)\|^2=2\1-k(s,t)\≤\|s-t\|^2/σ^2, so Lk=1/σL_k=1/σ. The theorem is intentionally conditional: the lexicographic lattice canonicalizer in Section 6 need not be globally Lipschitz near support changes or canonicalization ties. The approximate-action experiment in Table 5 therefore remains an empirical sensitivity analysis; it is not retroactively certified by Theorem 5.1 unless (5.2) is verified on the relevant image class. 6 A maximal invariant for lattice microscopy images 6.1 Lattice rigid motions Fix L<∞L<∞. Let q,LX_q,L contain q-channel, 8-bit functions on ℤ2Z^2 with finite support whose nonempty bounding box has at most L rows and columns, together with the zero image. Let C4=0,1,2,3C_4=\0,1,2,3\, with addition understood modulo four, and let RrR_r denote the counterclockwise rotation of ℤ2Z^2 through rπ/2rπ/2. The action of C4C_4 on ℤ2Z^2 is u↦Rru R_ru. Thus the semidirect product Grig=ℤ2⋊C4G_ rig=Z^2 C_4 has multiplication and inverse (t,r)(u,s)=(t+Rru,r+s),(t,r)−1=(−Rr−1t,−r).(t,r)(u,s)= (t+R_ru,r+s ), (t,r)^-1= (-R_r^-1t,-r ). Its action on an image x is (t,r)⋅x(v)=xRr−1(v−t),v∈ℤ2.\(t,r)· x\(v)=x\R_r^-1(v-t)\, v ^2. (6.1) Here t,u,v∈ℤ2t,u,v ^2 are lattice coordinates or translations, whereas r,s∈C4r,s∈ C_4 index quarter turns. For nonzero x, let b(x)b(x) be the lower corner of its support bounding box and set N(x)=(−b(x),0)⋅xN(x)=(-b(x),0)· x. For each r∈C4r∈ C_4, form xr=N(0,r)⋅xx_r=N\(0,r)· x\. The key κ(xr)κ(x_r) is the finite vector containing the bounding-box height and width, followed by all channel bytes in a fixed site–channel–row–column order. These finite vectors are compared using the lexicographic order. Define c(x)=xr⋆,r⋆=minargminr∈C4κ(xr),c(x)=x_r , r = _r∈ C_4κ(x_r), (6.2) and set c(0)=0c(0)=0. Theorem 6.1 (Lattice rigid-motion canonicalization). The map c is Borel and is a maximal invariant for ℤ2⋊C4Z^2 C_4. Proof. Because q,LX_q,L is a countable union of finite-dimensional finite-valued image spaces, the support map, bounding-box map, translation normalization, finite serialization and lexicographic comparison operations are Borel. Since the minimization in (6.2) is over the finite set C4C_4, the selected index r⋆r and hence c are Borel. We next prove invariance. Translation normalization removes every integer translation: N(t,0)⋅z=N(z)for every t∈ℤ2.N\(t,0)· z\=N(z) every t ^2. Let y=(t,s)⋅xy=(t,s)· x. For any r∈C4r∈ C_4, the semidirect-product law gives (0,r)(t,s)=(Rrt,r+s).(0,r)(t,s)=(R_rt,r+s). Therefore N(0,r)⋅y N\(0,r)· y\ =N(0,r)(t,s)⋅x =N\(0,r)(t,s)· x\ =N(Rrt,r+s)⋅x =N\(R_rt,r+s)· x\ =N(0,r+s)⋅x. =N\(0,r+s)· x\. Thus the four normalized rotational candidates of y are exactly the four normalized rotational candidates of x, with their indices permuted by r↦r+sr r+s. Their lexicographically smallest elements are consequently equal, and hence c(y)=c(x).c(y)=c(x). It remains to prove maximality. Suppose that c(x)=c(y)=zc(x)=c(y)=z. By the construction of c, there exist translations u,v∈ℤ2u,v ^2 and rotations r,s∈C4r,s∈ C_4 such that z=(u,r)⋅xandz=(v,s)⋅y.z=(u,r)· x z=(v,s)· y. Applying (v,s)−1(v,s)^-1 to the second equality and substituting the first gives y=(v,s)−1⋅z=(v,s)−1(u,r)⋅x.y=(v,s)^-1· z=(v,s)^-1(u,r)· x. Because (v,s)−1(u,r)∈Grig(v,s)^-1(u,r)∈ G_ rig, the images x and y belong to the same GrigG_ rig-orbit. Together with invariance, this shows c(x)=c(y)⟺y=g⋅x for some g∈Grig,c(x)=c(y) y=g· x for some g∈ G_ rig, so c is a maximal invariant. ∎ An RxRx1 well consists of two labeled sites. Let =q,L2W=X_q,L^2 be the raw well space, with a typical well written as w=(x1,x2)w=(x_1,x_2). Under the empirically imposed separate site-specific motions, the well-level nuisance group is Gwell=Grig×Grig.G_ well=G_ rig× G_ rig. Define the well-level canonicalization map Mwell(x1,x2)=(c(x1),c(x2)),M_ well(x_1,x_2)= (c(x_1),c(x_2) ), (6.3) and put can=Mwell().W_ can=M_ well(W). Theorem 6.1 and Proposition 4.1(1) imply that MwellM_ well is a maximal invariant for the product action. No probability measure is introduced on the noncompact translation group. 6.2 Characteristic quotient kernel and Euler signature Let ψ:can⟶ℝ2qL2ψ:W_ can ^2qL^2 place the two canonical crops on fixed L×L× L zero canvases and concatenate their entries in a fixed site–channel–row–column order. This map is a Borel injection. For σ>0σ>0, define kσ(w,w′)=exp[−‖ψMwell(w)−ψMwell(w′)‖222σ2].k_σ(w,w )= [- \|ψ\M_ well(w)\-ψ\M_ well(w )\\|_2^22σ^2 ]. (6.4) Proposition 6.2. The kernel (6.4) is positive semidefinite and invariant in each argument under the action of GwellG_ well. Its kernel mean embedding distinguishes all Borel probability laws on canW_ can. Proof. Let kG(u,v)=exp−‖u−v‖222σ2,u,v∈ℝ2qL2.k_ G(u,v)= \- \|u-v\|_2^22σ^2 \, u,v ^2qL^2. The kernel kGk_ G is positive semidefinite and characteristic on Euclidean space [21]. Equation (6.4) is its pullback through the Borel map ψ∘Mwellψ M_ well, and is therefore positive semidefinite. For every g∈Gwellg∈ G_ well, Mwell(g⋅w)=Mwell(w).M_ well(g· w)=M_ well(w). Consequently, kσ(g⋅w,w′)=kσ(w,w′)andkσ(w,g⋅w′)=kσ(w,w′),k_σ(g· w,w )=k_σ(w,w ) k_σ(w,g· w )=k_σ(w,w ), which proves invariance in each argument. Finally, let P and Q be Borel probability measures on canW_ can. Equality of their kernel mean embeddings implies, by the characteristic property of the Gaussian kernel, ψ#P=ψ#Q. _\#P= _\#Q. Because ψ is a Borel injection between standard Borel spaces, the Lusin–Souslin theorem implies that ψ(can)ψ(W_ can) is Borel and that ψ−1ψ^-1 is Borel on this image. Applying ψ−1ψ^-1 to the two equal pushforward laws therefore yields P=QP=Q. ∎ The secondary endpoint is a finite directional Euler vector. After canonicalization, Otsu thresholding [14] of the nuclear channel and removal of components smaller than four pixels produce a planar cubical complex. Euler characteristic is evaluated along 25 shape-adaptive thresholds in eight lattice directions at both sites. Because the calculation is applied after MwellM_ well, the vector is exactly invariant under GwellG_ well, but it is neither claimed nor used as a maximal invariant. 7 Exact design-based testing Suppose the confirmation experiment contains B paired blocks. In block b, two eligible wells receive conditions 00 and 11, once each. Order the two wells deterministically before observing their assigned conditions, and let qb,0,qb,1∈canq_b,0,q_b,1 _ can denote the quotient outcomes of the first and second wells, respectively. Let Zb∈0,1Z_b∈\0,1\ be the label of the condition assigned to the first well. Thus, if Zb=1Z_b=1, the first well receives condition 11 and the second receives condition 00; if Zb=0Z_b=0, these assignments are reversed. Conditional on the unordered eligible wells, their complete potential outcomes and the pretreatment design information, assume that Z=(Z1,…,ZB)Z=(Z_1,…,Z_B) is uniform on Ω=0,1B. =\0,1\^B. For an assignment z∈Ωz∈ , the quotient samples assigned to conditions 11 and 00 are, respectively, 1(z)=qb,1−zb:b=1,…,B,0(z)=qb,zb:b=1,…,B.Q_1(z)=\q_b,1-z_b:b=1,…,B\, _0(z)=\q_b,z_b:b=1,…,B\. Let P1P_1 and P0P_0 denote the corresponding population quotient-outcome laws under conditions 11 and 00. Their population kernel discrepancy is MMDk2(P1,P0)=k(U,U′)+k(V,V′)−2k(U,V), _k^2(P_1,P_0)=Ek(U,U )+Ek(V,V )-2Ek(U,V), (7.1) where U,U′U,U are independent with law P1P_1, V,V′V,V are independent with law P0P_0, and the four variables are mutually independent. The empirical statistic is the conventional two-sample U-statistic MMD^u2(z)= _u^2(z)= 1B(B−1)∑b≠b′k(qb,1−zb,qb′,1−zb′)+1B(B−1)∑b≠b′k(qb,zb,qb′,zb′) 1B(B-1) _b≠ b k(q_b,1-z_b,q_b ,1-z_b )+ 1B(B-1) _b≠ b k(q_b,z_b,q_b ,z_b ) −2B2∑b,b′k(qb,zb,qb′,1−zb′). - 2B^2 _b,b k(q_b,z_b,q_b ,1-z_b ). (7.2) Matched-block dependence can remove the usual unbiasedness interpretation; here (7) is a prespecified discrepancy statistic. The kernel two-sample construction follows the standard MMD formulation [7], but the exactness below does not require unbiased estimation of a superpopulation parameter. Let pool=qb,j:b=1,…,B,j∈0,1Q_ pool=\q_b,j:b=1,…,B,\ j∈\0,1\\. The Gaussian bandwidth is fixed before examining the condition labels by setting σ2=12median∥ψ(q)−ψ(q′)∥22:q,q′∈pool,q≠q′,∥ψ(q)−ψ(q′)∥2>0.σ^2= 12\,median \\|ψ(q)-ψ(q )\|_2^2:q,q _ pool,\ q≠ q ,\ \|ψ(q)-ψ(q )\|_2>0 \. Because poolQ_ pool does not change under paired label swaps, the bandwidth is constant over the randomization distribution. Let S:Ω→ℝS: be a deterministic test statistic; for the primary analysis, S(z)=MMD^u2(z)S(z)= _u^2(z). Define the exact upper-tail randomization p-value by p(Z)=2−B∑z∈ΩS(z)≥S(Z).p(Z)=2^-B _z∈ 1\S(z)≥ S(Z)\. (7.3) Theorem 7.1 (Finite-sample randomization validity). Assume well-defined reagent versions, no cross-well interference, swap-invariant block retention and conditional uniformity of Z on Ω . Under the Fisher sharp null that every confirmation well’s invariant potential outcome is unchanged by the two conditions, the test rejecting when p(Z)≤αp(Z)≤α has conditional rejection probability at most α for every α∈[0,1]α∈[0,1]. Proof. Under the sharp null, all quotient outcomes and label-invariant tuning parameters are fixed over Ω . Let N=2BN=2^B and r(z)=#u∈Ω:S(u)≥S(z)r(z)=\#\u∈ :S(u)≥ S(z)\. Then p(z)=r(z)/Np(z)=r(z)/N. For each integer m, at most m assignments have r(z)≤mr(z)≤ m; ties can only increase the upper-tail rank. Taking m=⌊αN⌋m= α N and using uniformity proves the result. ∎ The theorem is exact conditional on the reported assignment mechanism and is an instance of finite-sample randomization inference [6, 13]. The metadata audit can establish that paired siRNAs occur within the same plate randomization set; it cannot itself prove that the original allocation algorithm was uniform. The primary contrast uses one quotient-pixel test. Nine additional contrasts form a separate secondary family adjusted by Holm’s method [8]. 8 Simulation study Each of 250 replicates at each effect strength contains eight randomized blocks, two units per block, two sites per unit, six channels and 28×2828× 28 intrinsic lattice arrays. Treatment adds a concentric annulus with strength η∈0,0.35,0.70,1.00η∈\0,0.35,0.70,1.00\. At η=0η=0, the two treatment potential outcomes of each simulated unit are bytewise identical; different units may nevertheless have different intrinsic morphologies. After assignment, each site receives an integer shift and quarter turn from a deterministic generator depending on treatment, realized morphology, unit identity and a frozen replicate seed. The analyzer receives neither the group elements nor their generator inputs. All 28=2562^8=256 paired assignments are enumerated. Table 1: Paired-swap rejection fractions under treatment- and outcome-dependent acquisition transformations. The interval at η=0η=0 is the exact 95% Clopper–Pearson interval [3] over 250 independent replicates. Method η=0η=0 0.350.35 0.700.70 1.001.00 Assignments Exact quotient 0.052 (0.028, 0.087) 0.140 0.848 0.992 256 Raw pixels 1.000 (0.985, 1.000) 1.000 1.000 1.000 256 Moment registration 0.056 (0.031, 0.092) 0.540 0.996 1.000 256 Finite Euler signature 0.040 (0.019, 0.072) 0.184 1.000 1.000 256 For the exact quotient and post-canonical Euler endpoints, positive-η columns estimate power against a known intrinsic effect. Raw-pixel and moment-registration values are diagnostic rejection fractions because those endpoints are not invariant under the informative acquisition mechanism. 000.10.10.20.20.30.30.40.40.50.50.60.60.70.70.80.80.90.911000.20.20.40.40.60.60.80.811Intrinsic morphological effect strengthPaired-swap rejection fractionExact quotientRaw pixelsMoment registrationEuler secondaryNominal 0.05 Figure 1: Simulation rejection fractions. Finite-sample causal validity under informative acquisition pertains to the quotient and post-canonical Euler endpoints; the other methods are diagnostics. The exact quotient begins near the nominal 0.05 level and rises to 0.992, raw pixels remain at 1.0, and moment registration and the Euler endpoint begin near 0.05 and approach 1.0 as effect strength increases. 9 RxRx1 confirmation experiment 9.1 Units, selection and imposed acquisition The public RxRx1 release [16, 22] contains 125,510 six-channel 512×512512× 512 images, with two nonoverlapping sites per well, 51 experimental batches, four cell types and 1,138 siRNAs, including 1,108 noncontrol perturbations. HUVEC contributes 24 batches. The well is the experimental unit; sites and pixels are components of one structured outcome, not independent replicates. The conditions are applied siRNA reagents, so the estimand is reagent-specific rather than automatically gene-specific. HUVEC-01 through HUVEC-12 are used for discovery, HUVEC-13 through HUVEC-16 are reserved, and HUVEC-17 through HUVEC-24 supply eight confirmation blocks. Selection uses public 128-dimensional embeddings from discovery rows only. A split-half score combines within-half signal-to-noise with agreement of contrast direction. Eligibility requires one two-site well per condition in every discovery batch and identical plate signatures. Greedy matching yields ten disjoint pairs; the highest score defines the primary pair and nine pairs form the secondary family. The row split does not prove that the public embedding generator was trained independently of every confirmation image, so complete algorithmic independence remains an assumption. The frozen confirmation subset has 160 wells, 320 sites and 1,920 single-channel PNG files. Each site is downsampled to 64×6464× 64, retained as 8-bit data and zero-padded. For each published nuisance seed, a BLAKE2b digest of the seed, well identifier and site determines a quarter turn and a signed translation. The transformation also depends on treatment and a coordinate-free morphology score. All channels at one site share a motion, while the two sites receive separate motions. Parameters are logged for audit but excluded from every representation and test routine; the software aborts if nonzero support is cropped. 9.2 Confirmatory results The primary contrast is siRNA 384 (s38230) versus siRNA 747 (s37149). The completeness and same-plate gates passed for all frozen pairs and blocks. The discovery-only manifest has SHA-256 digest b6f350ca96fb9e7a6271845d6e432ed9dca556144a1a332d8f5af168642b968. Table 2: Primary RxRx1 pair under imposed lattice rigid motions. The quotient-pixel endpoint is the sole primary test; Euler is secondary; moment registration and raw pixels are acquisition-sensitive diagnostics. Representation MMD^u2 _u^2 Paired-swap p Canonical quotient pixels 0.0036 0.0078 Finite directional Euler signature 0.7535 0.0078 Moment registration -0.0065 0.2109 Raw transformed pixels 0.1664 0.0078 No secondary contrast remained significant after Holm adjustment. Exact invariance held across all acquisition seeds: maximum clean-versus-contaminated feature discrepancies were zero for the canonical pixels and Euler signature, as were the largest changes in quotient MMD^u2 _u^2 and exact p-value. The algebraic and software audit found zero canonical-code failures among 6,400 generated transformations; 29 software checks passed. Table 3: Canonical-quotient tests for the nine prespecified secondary contrasts. These contrasts are not biological replications of the primary contrast. Pair Conditions MMD^u2 _u^2 Swap p Holm p 1 s21474 vs s21486 0.0040 0.0859 0.1719 2 s29342 vs s28343 0.0045 0.0156 0.0703 3 s18592 vs s18763 0.1266 0.0078 0.0703 4 s20782 vs s21523 0.0031 0.0078 0.0703 5 s20401 vs s19178 0.0228 0.0156 0.0703 6 s29371 vs s28116 0.0063 0.0078 0.0703 7 s17831 vs s20468 0.0486 0.0078 0.0703 8 s37483 vs s224918 -0.0107 0.1562 0.1719 9 s37712 vs s224940 0.0761 0.0078 0.0703 Table 4: Acquisition-only sharp-null stress test over frozen pairs and nuisance seeds. Intervals are descriptive Clopper–Pearson summaries; pair–seed tests share biological blocks. Method Tests Rejections Fraction (95% interval) Exact quotient 200 0 0.000 (0.000, 0.018) Finite Euler signature 200 0 0.000 (0.000, 0.018) Moment registration 200 0 0.000 (0.000, 0.018) Raw pixels 200 200 1.000 (0.982, 1.000) Table 5: Approximate-action stress test for the primary pair. Feature error is relative to the exactly transformed baseline. These are empirical sensitivity results; Theorem 5.1 applies only when its Lipschitz assumptions are verified. Perturbation Method Relative feature error |ΔMMD^u2|| _u^2| |Δp|| p| Rotation 5∘5 Exact quotient 0.8432 0.0012 0.0000 Euler signature 0.0779 0.0860 0.0000 Moment registration 0.7295 0.0014 0.0234 Raw pixels 0.6540 0.0013 0.0000 Rotation 10∘10 Exact quotient 0.8967 0.0022 0.0000 Euler signature 0.0827 0.0530 0.0000 Moment registration 0.8124 0.0021 0.1250 Raw pixels 0.8038 0.0012 0.0000 Rotation 20∘20 Exact quotient 0.9715 0.0005 0.0000 Euler signature 0.1076 0.0898 0.0000 Moment registration 0.8662 0.0009 0.0469 Raw pixels 0.8802 0.0014 0.0000 Two-pixel support crop Exact quotient 0.8253 0.0026 0.0000 Euler signature 0.1478 0.0164 0.0000 Moment registration 0.6857 0.0203 0.2031 Raw pixels 0.3523 0.0001 0.0000 Gaussian noise Exact quotient 0.5949 0.0007 0.0000 Euler signature 0.0618 0.0244 0.0000 Moment registration 0.3870 0.0067 0.1328 Raw pixels 0.1775 0.0006 0.0000 10 Discussion Under unrestricted contamination, invariance is an observability requirement rather than a modeling preference. A noninvariant target can change while the observed law remains fixed; a maximal invariant captures the entire recoverable orbit. Causal identification then requires a separate design or exchangeability argument. The losslessness theorem adds a third distinction: even a maximal invariant may discard information about how probability is distributed within an orbit. Only a parameter-free conditional law, or a condition such as conditional Haar randomization that implies one, makes quotienting sufficient for the full statistical experiment. The distinction between product and diagonal actions has practical consequences. Sitewise canonicalization is correct for the imposed RxRx1 mechanism because each site receives a separate motion. If a microscope applies one common transformation to all sites, a diagonal quotient should retain relative placement and orientation. Treating that problem as a product action would erase scientifically meaningful relational structure. The approximate-contamination theorem links orbit error to quotient-law and MMD error, but it also clarifies the limits of the current stress test. Lexicographic canonicalization can be discontinuous. Large feature changes accompanied by stable p-values in Table 5 demonstrate test-level robustness for the frozen perturbations, not global representation stability. Establishing a Lipschitz quotient embedding on a scientifically relevant image class is a separate research problem. The RxRx1 analysis is a controlled unknown-acquisition experiment on real biological images. Its content and batch structure are real; the transformation law is imposed and exactly audited. The causal conclusion is conditional on the reported within-plate assignment mechanism, well-defined reagent versions, no cross-well interference and swap-invariant retention. The primary result concerns a reagent-condition contrast, not automatically a gene-level mechanism. With eight pairs, exact enumeration is a strength, but the randomization distribution has coarse resolution. 11 Conclusion Unknown transformations need not be estimated to conduct causal inference on the information that survives them. Under a measurable group action, orbit invariance is necessary and sufficient for uniform observability, and a maximal invariant contains every measurable observable target. Quotient reduction is statistically lossless only under the additional, exact boundary of a parameter-free conditional law given the quotient; conditional Haar contamination is one useful special case. Correct specification of product or diagonal actions determines whether relative component information is preserved. The lattice construction, characteristic quotient kernel and complete paired randomization test turn these distinctions into an auditable workflow for structured outcomes under informative acquisition. Data availability The analysis uses the public RxRx1 dataset subject to its stated licence. The accompanying software is described by the reproducibility statement below. Repository or archival identifiers will be added when the public code archive is released. Reproducibility statement The analysis software downloads the frozen RxRx1 subset, verifies file and manifest hashes, constructs all representations, enumerates the complete assignment distribution, performs multiplicity adjustment and sensitivity analyses, executes unit tests and writes the numerical outputs consumed by the manuscript. The source fingerprint reported for the analysed version is source-e9f042c387cf. Author contributions Usef Faghihi and Amir Saki contributed equally to the conception, methodology, theoretical development, analysis, interpretation of results and preparation of the manuscript. Both authors reviewed and approved the final manuscript. Funding This research received no specific grant from any funding agency in the public, commercial or not-for-profit sectors. Acknowledgements and use of AI-assisted tools During manuscript preparation, the authors used OpenAI Codex for language editing, theorem-structure review and LaTeX reformatting. The authors are responsible for verifying all mathematical statements, empirical results, citations and final wording. Appendix A Proof and measurability details A.1 Proof of Theorem 2.4 Proof. Fix an arbitrary Borel set B⊆B , and write E:=h−1(B).E:=h^-1(B). Since h is Borel, both E and ∖EY E are Borel subsets of Y. We first verify explicitly that E is saturated with respect to the fibres of M, that is, M−1M(E)=E.M^-1\M(E)\=E. The inclusion E⊆M−1M(E)E M^-1\M(E)\ is immediate. Conversely, suppose that y∈M−1M(E)y∈ M^-1\M(E)\. Then there exists z∈Ez∈ E such that M(y)=M(z)M(y)=M(z). By maximality of M, there is some g∈Gg∈ G for which y=g⋅zy=g· z. Invariance of h therefore gives h(y)=h(g⋅z)=h(z)∈B,h(y)=h(g· z)=h(z)∈ B, and hence y∈Ey∈ E. This proves the asserted saturation. Notice that this statement concerns the particular set E=h−1(B)E=h^-1(B): it says that every fibre of M is either entirely contained in E or entirely contained in its complement; it does not assert that the fibres of M are singletons. We next claim that M(E)∩M(∖E)=∅.M(E)∩ M(Y E)= . Indeed, if s belonged to this intersection, there would exist y∈Ey∈ E and z∈∖Ez E such that M(y)=s=M(z).M(y)=s=M(z). Maximality of M would then place y and z in the same G-orbit, so invariance of h would imply h(y)=h(z)h(y)=h(z). This is impossible because h(y)∈Bh(y)∈ B, whereas h(z)∉Bh(z)∉ B. Because Y and S are standard Borel spaces, and M is Borel, the images M(E)M(E) and M(∖E)M(Y E) are analytic subsets of S. They need not themselves be Borel, which is why a separation argument is required. By Lusin’s separation theorem, there exists a Borel set DB⊆D_B satisfying M(E)⊆DBandM(∖E)⊆∖DB.M(E) D_B M(Y E) D_B. We now show that E=M−1(DB).E=M^-1(D_B). If y∈Ey∈ E, then M(y)∈M(E)⊆DBM(y)∈ M(E) D_B, and hence y∈M−1(DB)y∈ M^-1(D_B). If y∉Ey∉ E, then M(y)∈M(∖E)⊆∖DBM(y)∈ M(Y E) D_B, and hence y∉M−1(DB)y∉ M^-1(D_B). The two inclusions therefore follow, and h−1(B)=E=M−1(DB)∈σ(M).h^-1(B)=E=M^-1(D_B)∈σ(M). Since B⊆B was an arbitrary Borel set, h is σ(M)σ(M)-measurable. The measurable factorization lemma, applicable because T is a standard Borel space, therefore yields a Borel map h¯:⟶ h:S such that h=h¯∘M.h= h M. It remains to establish the stated uniqueness. Suppose that h¯1,h¯2:→ h_1, h_2:S are two Borel maps satisfying h=h¯1∘M=h¯2∘M.h= h_1 M= h_2 M. For any s∈M()s∈ M(Y), choose y∈y with M(y)=sM(y)=s. Then h¯1(s)=h¯1M(y)=h(y)=h¯2M(y)=h¯2(s). h_1(s)= h_1\M(y)\=h(y)= h_2\M(y)\= h_2(s). Thus h¯1=h¯2 h_1= h_2 on M()M(Y). No uniqueness is claimed outside M()M(Y), since values of a factor map there do not affect its composition with M. ∎ A.2 Existence constructions for maximal invariants If a compact metrizable group acts continuously on a Polish space, the orbit-valued map y↦G⋅y G· y is a Borel maximal invariant taking values in the hyperspace of nonempty compact subsets with its Vietoris Borel structure. For a finite group H=h1,…,hmH=\h_1,…,h_m\ acting by Borel maps, any Borel injection ν:→[0,1]ν:Y→[0,1] gives the canonical representative cH(y)=hj⋆⋅y,j⋆=minargminjν(hj⋅y),c_H(y)=h_j · y, j = _jν(h_j· y), which is Borel and maximal. The lattice construction combines an explicit cross-section for noncompact translations with this finite residual minimization. A.3 Measurability of the Haar orbit kernel For compact second-countable G, a jointly Borel action and measurable section make m↦Km(B)m K_m(B) in (3.2) measurable for every Borel B by integration of the jointly measurable indicator (g,m)↦g⋅s(m)∈B(g,m) 1\g· s(m)∈ B\. Stabilizers may change the map from G onto an orbit but do not change the pushforward orbit measure. References [1] Bhattacharjee, S., Li, B., Wu, X. and Xue, L. (2025) Doubly robust estimation of causal effects for random object outcomes with continuous treatments. arXiv:2506.22754. [2] Blackwell, D. (1953) Equivalent comparisons of experiments. Ann. Math. Statist., 24, 265–272. [3] Clopper, C. J. and Pearson, E. S. (1934) The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika, 26, 404–413. [4] Curry, J., Mukherjee, S. and Turner, K. (2022) How many directions determine a shape and other sufficiency results for two topological transforms. Trans. Amer. Math. Soc. Ser. B, 9, 1006–1043. [5] Eaton, M. L. (1989) Group Invariance Applications in Statistics. Institute of Mathematical Statistics, Hayward, CA. [6] Fisher, R. A. (1935) The Design of Experiments. Oliver and Boyd, Edinburgh. [7] Gretton, A., Borgwardt, K. M., Rasch, M. J., Schölkopf, B. and Smola, A. (2012) A kernel two-sample test. J. Mach. Learn. Res., 13, 723–773. [8] Holm, S. (1979) A simple sequentially rejective multiple test procedure. Scand. J. Stat., 6, 65–70. [9] Kallenberg, O. (2021) Foundations of Modern Probability, 3rd edn. Springer, Cham. [10] Kechris, A. S. (1995) Classical Descriptive Set Theory. Springer, New York. [11] Kurisu, D., Zhou, Y., Otsu, T. and Müller, H.-G. (2024) Geodesic causal inference. arXiv:2406.19604. [12] Le Cam, L. (1986) Asymptotic Methods in Statistical Decision Theory. Springer, New York. [13] Lehmann, E. L. and Romano, J. P. (2005) Testing Statistical Hypotheses, 3rd edn. Springer, New York. [14] Otsu, N. (1979) A threshold selection method from gray-level histograms. IEEE Trans. Syst. Man Cybern., 9, 62–66. [15] Raykov, Y. P., Luo, H., Strait, J. D. and KhudaBukhsh, W. R. (2025) Kernel-based estimators for functional causal effects. arXiv:2503.05024. [16] [dataset] Recursion (2023) RxRx1: an image set for cellular morphological variation across many biological perturbations. https://w.rxrx.ai/rxrx1. [17] Robins, J. M. (1986) A new approach to causal inference in mortality studies with a sustained exposure period. Math. Modelling, 7, 1393–1512. [18] Rosenbaum, P. R. and Rubin, D. B. (1983) The central role of the propensity score in observational studies for causal effects. Biometrika, 70, 41–55. [19] Rubin, D. B. (1974) Estimating causal effects of treatments in randomized and nonrandomized studies. J. Educ. Psychol., 66, 688–701. [20] Shin, H.-Y., Kim, K., Lee, K. and Oh, H.-S. (2024) Absolute average and median treatment effects as causal estimands on metric spaces. arXiv:2407.03726. [21] Sriperumbudur, B. K., Gretton, A., Fukumizu, K., Schölkopf, B. and Lanckriet, G. R. G. (2010) Hilbert space embeddings and metrics on probability measures. J. Mach. Learn. Res., 11, 1517–1561. [22] Sypetkowski, M. et al. (2023) RxRx1: a dataset for evaluating experimental batch correction methods. arXiv:2301.05768. [23] Torgersen, E. (1991) Comparison of Statistical Experiments. Cambridge University Press, Cambridge. [24] Turner, K., Mukherjee, S. and Boyer, D. M. (2014) Persistent homology transform for modeling shapes and surfaces. Information and Inference, 3, 310–344. [25] Wijsman, R. A. (1990) Invariant Measures on Groups and Their Use in Statistics. Institute of Mathematical Statistics, Hayward, CA. Figure caption list Figure 1. Simulation rejection fractions across intrinsic effect strengths; causal validity under informative acquisition pertains to invariant endpoints. Table caption list Table 1. Simulation rejection fractions. Table 2. Primary RxRx1 contrast. Table 3. Prespecified secondary contrasts. Table 4. Acquisition-only sharp-null stress test. Table 5. Approximate-action stress test.