Paper deep dive
A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models
Trung Phung, Ilya Shpitser
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader class of Markov models, defined using solely conditional independence (CI) restrictions. In order to estimate finite-dimensional target parameters in such models efficiently, semi-parametric theory provides a principled framework for constructing regular and asymptotically linear estimators via influence functions (IFs). These estimators are asymptotically normal and root-$n$ consistent. Characterizing the class of all influence functions for a target parameter is crucial for statistically efficient inference in these models. For models that are Markov relative to directed acyclic graphs (DAGs), the orthogonal complement of the tangent space is known, implying that for any target the class of all influence functions can be derived once an influence function is obtained. On the other hand, for Markov models not equivalent to a DAG model -- such as ordinary Markov models associated with undirected graphs, chain graphs, or acyclic directed mixed graphs -- the orthogonal complement has not been characterized, impeding semi-parametric inference in these models. We derive closed form expressions for the orthogonal complement of the tangent space for general Markov models and illustrate our results by characterizing the class of influence functions for the conditional mean parameter in several graphical models.
Tags
Links
- Source: https://arxiv.org/abs/2607.23439v1
- Canonical: https://arxiv.org/abs/2607.23439v1
Trouble viewing inline? Open PDF directly →
Full Text
143,095 characters extracted from source content.
Expand or collapse full text
A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models Trung Phung Computer Science Department Johns Hopkins University Baltimore, Maryland, USA Ilya Shpitser Computer Science Department Johns Hopkins University Baltimore, Maryland, USA Abstract Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader class of Markov models, defined using solely conditional independence (CI) restrictions. In order to estimate finite-dimensional target parameters in such models efficiently, semi-parametric theory provides a principled framework for constructing regular and asymptotically linear estimators via influence functions (IFs). These estimators are asymptotically normal and root-n consistent. Characterizing the class of all influence functions for a target parameter is crucial for statistically efficient inference in these models. For models that are Markov relative to directed acyclic graphs (DAGs), the orthogonal complement of the tangent space is known, implying that for any target the class of all influence functions can be derived once an influence function is obtained. On the other hand, for Markov models not equivalent to a DAG model – such as ordinary Markov models associated with undirected graphs, chain graphs, or acyclic directed mixed graphs – the orthogonal complement has not been characterized, impeding semi-parametric inference in these models. We derive closed form expressions for the orthogonal complement of the tangent space for general Markov models and illustrate our results by characterizing the class of influence functions for the conditional mean parameter in several graphical models. 1 Introduction A common task in statistical and causal inference is estimating a finite dimensional target parameter from data. In parametric models where elements are indexed by finite dimensional nuisance parameters, efficient estimation is achieved using the theory of maximum likelihood estimation [vanderVaart.2000.AsymptoticStatistics]. In nonparametric or semiparametric models where the parameterization is infinite dimensional, regular and asymptotically linear (RAL) estimators are desirable, since they have nice properties such as asymptotic normality and root-n rate of convergence. RAL estimators are typically obtained by deriving influence function of the target parameter [Newey.1990.SemiparametricEfficiency, Bickel.Klaassen.ea.1993.EfficientAdaptive, VanDerVaart.2002.SemiparametricStatistics, Tsiatis.2006.SemiparametricTheory, Kennedy.2024.SemiparametricDoubly]. The class of influence functions of a parameter of interest in a given semiparametric model are elements of the linear variety constructed using one influence function and the orthogonal complement of the tangent space of the model (Tsiatis.2006.SemiparametricTheory, Theorem 4.3). Characterizing the class of all possible RAL estimators in a given semiparametric model thus entails characterizing the orthogonal complement of the tangent space of the model. Statistical independence is an important tool for encoding irrelevance, and is thus crucial for obtaining interpretable models for empirical phenomena where relationships among variables are local. Thus, in this paper, we consider Markov models, or conditional independence models, which are statistical models defined using solely marginal and conditional independences, with a focus on the subclass of Markov models associated with graphs, referred to as graphical Markov models [Pearl.1988.ProbabilisticReasoning, Lauritzen.1996.GraphicalModels, Maathuis.Drton.ea.2018.HandbookGraphical]. Graphical models have been associated with directed acyclic graphs (DAGs), undirected graphs (UGs), as well as mixed graphs such as chain graphs (CGs) [Frydenberg.1990.ChainGraph] and acyclic directed mixed graphs (ADMGs) [Pearl.1988.ProbabilisticReasoning, Lauritzen.1996.GraphicalModels, Richardson.2003.MarkovProperties, Maathuis.Drton.ea.2018.HandbookGraphical]. Models associated with DAGs have seen wide use due to their interpretability, and connections to causal models [Pearl.2000.CausalityModels, Spirtes.Glymour.ea.2001.CausationPrediction, Richardson.Robins.2013.SingleWorld]. However, other types of graphical models have also seen uses for associational analysis of gene regulatory networks [segal2003mni], protein signaling networks [sachs05causal], network data [TchetgenTchetgen.Fulcher.ea.2021.AutoGComputationCausal, Ogburn.Shpitser.ea.2020.CausalInference], and spatial data [lopes08spatial], as well as modeling systems at equilibrium [Lauritzen.Richardson.2002.ChainGraph]. In addition, ADMG models have been connected to causal models with hidden variables, as well as theory of identification in such models [Richardson.2003.MarkovProperties, Richardson.Evans.ea.2023.NestedMarkov, Shpitser.Richardson.ea.2022.MultivariateCounterfactual, Shpitser.Pearl.2006.IdentificationJointa]. Explicit likelihoods in terms of sets of variationally independent components have been derived for DAG models, leading to a natural characterization of their model tangent spaces in terms of orthogonal subspaces [Tsiatis.2006.SemiparametricTheory, Rotnitzky.Smucler.2020.EfficientAdjustment, Bhattacharya.Nabi.ea.2022.SemiparametricInference]. While such likelihoods are also known for UG and CG models [Shpitser.2023.LauritzenChenLikelihood], they feature overlapping components, meaning their tangent space cannot easily be expressed in terms of orthogonal subspaces. Furthermore, no general likelihoods for ADMG models are known. Since these Markov models associated with UGs, CGs, and ADMGs are not, in general, observationally equivalent to models associated with DAGs, this greatly complicates the derivation of the tangent space and hence its orthogonal complement, necessitating the development of new techniques. In this paper, we derive a characterization of the orthogonal complement of the tangent space for semiparametric Markov models defined by marginal and conditional independences, including graphical models associated with undirected graphs, chain graphs, and acyclic directed mixed graphs. In addition, we illustrate how our characterization may be used to derive all influence functions for target parameters arising in statistical and causal inference applications, with the conditional mean parameter as a worked example. 2 Preliminaries Prior to discussing our contribution, we first review the necessary preliminaries: conditional independence (Markov) models and the important subsclass of graphical Markov models, as well as the theory of statistical inference in semi-parametric models. 2.1 Markov Models Suppose p()p(V) is a distribution over the set of all random variables of interest V. If ,,X,Y,Z are disjoint subsets of V, a conditional independence (CI) statement “X and Y are independent given Z in p”, denoted ⟂∣X \!\!\! , means p(,|)=p(|)p(|)p(x,y|z)=p(x|z)p(y|z) for all values of ,,X,Y,Z such that p()>0p(z)>0. Given a list ℒ L of CI statements, a statistical model associated with these constraints, written as p(): all CIs in ℒ hold in p()\p(V): all CIs in L hold in p( V)\ is called a Markov model, or a conditional independence model [Maathuis.Drton.ea.2018.HandbookGraphical]. Important subclasses of Markov models are graphical models associated with directed acyclic graphs (DAGs), undirected graphs (UGs), chain graph (CGs), and acyclic directed mixed graph (ADMGs) [Maathuis.Drton.ea.2018.HandbookGraphical, Lauritzen.1996.GraphicalModels, Richardson.2003.MarkovProperties]. Undirected graphs (UGs) are graphs with only undirected edges (−-). Directed acyclic graphs (DAGs) are graphs with only directed edges (→) lacking directed cycles. Bidirected graphs (BGs) are graphs with only bidirected edges (↔ ) Chain graphs (CGs) are mixed graphs containing both directed and undirected edges with no partially directed cycles. Acyclic directed mixed graphs (ADMGs) are mixed graphs with directed and bidirected edges lacking directed cycles (for the purposes of such cycles the presence of bidirected edges is ignored). Simple examples of these graphs classes are shown in Fig. 1. CCDDAABB(a) The two pathway DAGCCDDAABB(b) The undirected squareCCDDAABB (c) The dyadic minimal complex CG CCDDAABB(d) The Bell scenarioCCDDAABB(e) The bidirected squareCCDDAABB(f) A complete 44 vertex DAG Figure 1: Six different graphs used to define six different graphical Markov models. Models associated with graphs in (b), (c), (d), and (e) are not equivalent to any DAG model, so their tangent spaces cannot be characterized using the technique in Section 3. Graphical models use graphical separation criteria to encode CI constraints. The (finite) set of all CIs that hold in a graphical model is given by a global Markov property, while local Markov properties provide a small list of CIs that logically imply all others. Below, we define local Markov properties for models associated with UGs, DAGs, CGs, bidirected graphs (BGs), and ADMGs, and refer reader to Maathuis.Drton.ea.2018.HandbookGraphical, Richardson.2003.MarkovProperties for additional details. To describe the local Markov properties, we need the following basic graphical concepts. If G is an UG, denote ne(V)≡Z∈:Z−V in ne_G(V)≡\Z :Z-V in G\ as the set of neighbors of V in G. If G is a DAG, denote pa(V)≡Z∈:Z→V in pa_G(V)≡\Z :Z→ V in G\ as the set of parents of V in G, ch(V)≡Z∈:V→Z in ch_G(V)≡\Z :V→ Z in G\ as the set of children of V in G, de(V)≡Z∈:V→⋯→Z in de_G(V)≡\Z :V→·s→ Z in G\ as the set of descendants of V in G, an(V)≡Z∈:Z→⋯→V in an_G(V)≡\Z :Z→·s→ V in G\ as the set of ancestors of V in G, and nd(V)≡∖de(V) nd_G(V) de_G(V) as the set of non-descendants of V in G. Note that by convention, V∈de(V)∩an(V)V∈ de_ G(V)∩ an_ G(V). For sets of variables, the sets of parents, ancestors and descendants are defined disjunctively. That is pa()≡⋃A∈pa(A) pa_ G( A)≡ _A∈ A pa_ G(A), an()≡⋃A∈an(A) an_ G( A)≡ _A∈ A an_ G(A) and de()≡⋃A∈de(A) de_ G( A)≡ _A∈ A de_ G(A). For CGs, define pa(V) pa_G(V) as in DAGs and ne(V) ne_ G(V) as in UGs. The set of descendants of V, de(V) de_ G(V) is defined as the set of all variables Z with a partially directed path from V to Z – a sequence of distinct vertices such that there is an edge (−- or →) between any pair of consecutive vertices, with all directed edges pointing in the same direction. As before, V∈de(V)V∈ de_ G(V) and nd(V)=∖de(V) nd_ G(V)= V de_ G(V). We further define the boundary of V, bd(V) bd_ G(V) as ne(V)∪pa(V) ne_ G(V)∪ pa_ G(V). Finally, we define a chain component in a CG G to be a connected component in the edge subgraph of G retaining only undirected edges. If G contains bidirected edges, denote sib(V)≡Z∈:Z↔V in sib_G(V)≡\Z :Z V in G\ as the set of siblings of V in G. Furthermore, a subset ⊆C is called a bidirected connected subset if it is a connected component in the edge subgraph of G retaining only bidirected edges. A district is a maximal bidirected connected subset, and dis(V) dis_G(V) denotes the district in G containing V. For ADMGs, all mentioned graphical concepts for DAGs and BGs apply, and define mb(V)=pa(dis(V))∪(dis(V)∖V) mb_G(V)= pa_G( dis_G(V))∪ ( dis_G(V) \V\ ) as the Markov blanket of V in G. Furthermore, a subset A is called ancestral if whenever V∈an()V∈ an_ G( A) then V∈V∈ A. Given a graph G with a vertex set V, and ⊆ A V, define G_ A, the subgraph of G induced by A, to be the subgraph of G containing only the vertex set A and edges among A. The list of CIs corresponding to the local Markov property for UGs, DAGs, CGs, BGs and ADMGs with the vertex set V are as follows • UGs: (V⟂∖(V∪ne(V))|ne(V)) (V \!\!\! (\V\∪ ne_G(V) ) | ne_G(V) ), for all V∈V . • DAGs: (V⟂nd(V)∖pa(V)|pa(V)) (V \!\!\! nd_G(V) pa_G(V) | pa_G(V) ), for all V∈V . • CGs: (V⟂nd(V)∖bd(V)|bd(V)) (V \!\!\! nd_G(V) bd_G(V) | bd_G(V) ), for all V∈V . • BDs: (V⟂∖dis(V)|dis(V)) (V \!\!\! A dis_ G_ A(V) | dis_ G_ A(V) ), for all V and all subsets A of V such that V∈V∈ A. • ADMGs: (V⟂∖(mb(V)∪V)|mb(V)) (V \!\!\! ( mb_G_ A(V)∪\V\ ) | mb_G_ A(V) ), for all V and ancestral set A such that V∈⊆nd(V)V nd_G(V). Note that, as expected, the local properties for the UG and DAG model are special cases of the local property for the CG model, and the local properties for the BG and DAG model are special cases of the local property for the ADMG model. For example, the following are the lists of CIs for the models associated with the graphs in Figure 1. (a) The two pathways DAG model: C⟂B∣AC \!\!\! A and D⟂A∣B,CD \!\!\! B,C. (b) The undirected square UG model: C⟂B∣A,DC \!\!\! A,D and D⟂A∣B,CD \!\!\! B,C. (c) The dyadic minimal complex CG model: A⟂BA \!\!\! , C⟂B∣A,DC \!\!\! A,D and D⟂A∣B,CD \!\!\! B,C. (d) The Bell scenario ADMG model: A⟂B,DA \!\!\! ,D and B⟂A,CB \!\!\! ,C. (e) The bidirected square scenario ADMG model: A⟂DA \!\!\! and B⟂CB \!\!\! . (f) The complete DAG: the empty list, as the model is saturated. Note that there are graphical models associated with ADMGs defined not only in terms of CIs but also generalized CIs, sometimes called “Verma constraints” [Richardson.Evans.ea.2023.NestedMarkov]. As an example, the model associated with Fig. 2 is defined by a CI: C⟂A∣BC \!\!\! B and a generalized independence constraint stating that ∑Bp(D∣C,B,A)p(B∣A) _Bp(D C,B,A)p(B A) is a function of only C and D. In general, both CIs and Verma constraints are needed to describe marginal models for hidden variable DAGs. For the purposes of this work, we only consider ADMG models defined exclusively via ordinary CI constraints, called ordinary Markov models [Richardson.2003.MarkovProperties, evans18margins]. AABBCCDD Figure 2: A graph representing a model with a Verma constraint. 2.2 Semi-parametric Inference As we further describe in Sections 3 and 4, there is a close relationship between Markov model restrictions and the tangent space of the model. This relationship is used to construct estimators via theory of semi-parametric models which we now briefly outline. For further details, we refer the reader to excellent tutorials in Kennedy.2024.SemiparametricDoubly and Tsiatis.2006.SemiparametricTheory. We refer to a semi-parametric model as a set of probability densities =p(;θ):θ∈Θ P=\p(v;θ):θ∈ \ relative to a measure μ, parameterized by an infinite dimensional nuisance parameters θ, with the true data generating distribution p0∈p_0∈ P. As an example, the semi-parametric model associated with the DAG model in Fig. 1a is (a)=p(a,b,c,d):p(a,b,c,d)=p(a)p(b∣a)p(c∣a)p(d∣b,c) P^(a)=\p(a,b,c,d):p(a,b,c,d)=p(a)p(b a)p(c a)p(d b,c)\. The contaminated parametric submodel is a smooth parametric model ε=pε():ε∈ℰ⊆ℝ P_ =\p_ (v): \ defined around the true distribution p0p_0 such that pε=0=p0p_ =0=p_0, and is contained in the semi-parametric model ε⊆ P_ P. In the subsequent discussion, we will refer to ε P_ as the parametric submodel, following Tsiatis.2006.SemiparametricTheory. We also denote the saturated (unrestricted) model as all P^all. As an example, the model associated with the complete DAG in Fig. 1f is saturated. By definition, all semiparametric models are contained in all P^all. For a parametric submodel ε P_ , the score function is the derivative of the log likelihood at the truth s()=∂εlogpε()|ε=0s(v)= ∂ p_ (v)|_ =0. This function is an element of the Hilbert space ℋH of all mean-zero L2(p0)L^2(p_0) functions f()f(v) in which the inner product is relative to the truth distribution ⟨f,g⟩=∫f()g()p0(),∀f,g∈ℋ f,g = f(v)g(v)p_0(v)dv,∀ f,g . The tangent space of a semi-parametric model, denoted T, is the closure of all parametric submodels’ score functions, thus T is a subspace of ℋH. For the saturated model all P^all, =ℋT=H. The orthogonal complement of the tangent space, denoted ⟂T , is the subspace of all functions in ℋH orthogonal to T, so ⟨f,g⟩=0 f,g =0 for all f∈⟂f and g∈g . The target of inference in a semi-parametric model is a smooth function ψ:↦ℝdψ: P ^d where d is a fixed integer. For simplicity, we will only consider scalar target, so d=1d=1. For example, an adjustment formula functional arising in causal inference applications for the examples in Figure 1 yields a scalar parameter ψ(p)=[[D∣A=a0,B,C]]ψ(p)=E[E[D A=a_0,B,C]]. The main result can be extended easily to d>1d>1 but the notation is more involved. In semi-parametric theory we are interested in the class of regular and asymptotically linear (RAL) estimators for ψ. Given data 1,…,nV_1,…,V_n drawn from a distribution p∈p∈ P, an estimator ψ^n ψ_n for ψ is called asymptotically linear if there is a function ϕ∈ℋφ such that n(ψ^n−ψ)=1n∑i=1nϕ(i)+op(1). n( ψ_n-ψ)= 1 n _i=1^nφ(V_i)+o_p(1). (1) By the central limit theorem, the limiting distribution of such an estimator ψ^n ψ_n is normal with root-n rate of convergence, n(ψ^n−ψ)↝(0,[ϕ2]), n( ψ_n-ψ) (0,E[φ^2]), (2) where ↝ denotes convergence in distribution. Estimators exhibit asymptotically linear property locally uniformly around the truth are called regular and asymptotically linear (RAL) estimators, which excludes super efficient estimators like the Hodges estimator. The function ϕφ is called an influence function for the estimator ψ^n ψ_n. RAL estimators are attractive because they are root-n consistent and asymptotically normal, and the variance equal the variance of its influence function. A closely related notion is the influence functions for the target parameter. Specifically when the distribution changes from p to p¯ p, one can approximate the change in the target functional using the following von Mises expansion ψ(p¯)−ψ(p)=∫φ()(p¯−p)()+R[(p¯−p)2].ψ( p)-ψ(p)= (v)( p-p)(v)dv+R[( p-p)^2]. (3) The integral is the first-order change of the target parameter with respect to the change in the distribution, while R is the second-order term of the expansion. The function φ∈ℋ characterizes the first-order derivative of the target functional, and is referred to as the influence function of the target parameter ψ. The von Mises expansion implies the following for all parametric submodels pεp_ with score function s∈s ∂εψ(pε)|ε=0=∫φ()s()p0(), ∂ ψ(p_ ) |_ =0= (v)s(v)p_0(v)dv, (4) where, under regularity conditions, it suffices to consider only parametric submodels of the form pε()=p0()(1+εs())p_ (v)=p_0(v)(1+ s(v)). Equation 4 is called the pathwise differentiability condition, which is implied by the Neyman orthogonality condition [Chernozhukov.Chetverikov.ea.2018.DoubleDebiased], and is central to the derivation of target’s influence function. In practice, one first obtains a target parameter’s influence function φ from Equation 4, then from it constructs a RAL estimator, whose influence function coincides with the target’s influence function, ϕ=φ= . We focus on the target parameter’s influence functions and will refer to them as just influence functions (IFs). From Equation 4, it is evident that if φ is an influence function and f∈⟂f , then φ+f +f is also an influence function, since ⟨f,s⟩=∫f()s()p0()=0 f,s = f(v)s(v)p_0(v)dv=0. Therefore, the class of all influence functions is φ⊕⟂ , where φ is any solution of Equation 4 and ⊕ denotes the direct sum (Tsiatis.2006.SemiparametricTheory, Theorem 4.3). Thus, knowing the tangent space’s orthogonal complement ⟂T allows one to derive the class of influence functions for any target of inference. 3 The orthocomplement of the Tangent Space of a DAG Model The tangent space of a DAG model and its orthogonal complement have been derived in multiple contexts, see Theorem 4.5 in Tsiatis.2006.SemiparametricTheory, Lemma 24 in Rotnitzky.Smucler.2020.EfficientAdjustment, and Lemma 3 in Bhattacharya.Nabi.ea.2022.SemiparametricInference. In the interests of being self-contained, we repeat the derivation using examples in Figure 1a in order to highlight the main techniques. Despite using specific examples, the derivation is without loss of generality. Subsequently, we explain the reason why these techniques fail to apply to Markov models beyond DAG models. 3.1 A Worked Example Consider the DAG in Figure 1a, whose semi-parametric model is (a) P^(a). To find the tangent space T, we first derive the score function for an arbitrary parametric submodel pεp_ , then take the closure. It turns out that T is the range of a projection operator Π , hence the orthogonal complement of T is simply h−Π(h)h- (h) for all h∈ℋh . These steps are demonstrated in detail below. The score function of a parametric submodel. Consider a parametric submodel ε P_ . By the DAG factorization property, elements of ε P_ satisfy pε(a,b,c,d)=pε(a)pε(b∣a)pε(c∣a)pε(d∣c,b).p_ (a,b,c,d)=p_ (a)p_ (b a)p_ (c a)p_ (d c,b). (5) This implies that the total score function is the sum s(a,b,c,d)=s(a)+s(b∣a)+s(c∣a)+s(d∣c,b),s(a,b,c,d)=s(a)+s(b a)+s(c a)+s(d c,b), (6) where, for any pair of disjoint subsets ,⊆X,Y , a component score function s(∣)s(x ) is defined as ∂logpε(∣)∂ε|ε=0 ∂ p_ (x )∂ |_ =0. Since a component score function s(∣)s(x ) has the property that its conditional mean is zero ∫s(∣)p0(∣)=0 s(x )p_0(x )dx=0, the component scores are in the following subspaces of ℋH s(a)∈A s(a) _A ≡[h|a]:∀h∈ℋ, ≡\E[h|a]:∀ h \, (7) s(b∣a)∈B∣A s(b a) _B A ≡[h|b,a]−[h|a]:∀h∈ℋ, ≡\E[h|b,a]-E[h|a]:∀ h \, s(c∣a)∈C∣A s(c a) _C A ≡[h|c,a]−[h|a]:∀h∈ℋ, ≡\E[h|c,a]-E[h|a]:∀ h \, s(d∣c,b)∈D∣CB s(d c,b) _D CB ≡[h|d,c,b]−[h|c,b]:∀h∈ℋ. ≡\E[h|d,c,b]-E[h|c,b]:∀ h \. The expectations are relative to p0p_0. In the Appendix, we show that these subspaces are pairwise orthogonal by checking that the inner products of their elements vanish. The model’s tangent space. The semi-parametric tangent space T is the closure of all parametric submodels’s score functions, hence it must be a closed subspace of the direct sum A⊕B|A⊕C|A⊕D|BCT_A _B|A _C|A _D|BC. In fact, it is actually this direct sum = = A⊕B|A⊕C|A⊕D|BC. _A _B|A _C|A _D|BC. (8) We refer to the subspaces on the r.h.s as tangent subspaces. Any element of T can be explicitly constructed as the sum of four elements of the four tangent subspaces, whose explicit forms are given by Equation 7. A detailed proof of Equation 8 is provided in the Appendix. In essence, the equality is established once it is shown that each tangent subspace is contained in T. To illustrate, consider showing that B|A⊆T_B|A . Pick any f∈B|Af _B|A and define the parametric model pε=p0(1+εf)p_ =p_0(1+ f), where ε is small enough so that pε≥0p_ ≥ 0. The score function at p0p_0 is ∂εlogpε|ε=0=f ∂ p_ |_ =0=f. This parametric model contains the true distribution p0p_0 at ε=0 =0, and it satisfies all the DAG constraints: pε(c∣b,a)=pε(c∣a)p_ (c b,a)=p_ (c a) and pε(d∣c,b,a)=pε(d∣c,b)p_ (d c,b,a)=p_ (d c,b). Therefore, pεp_ defines a valid parametric submodel of P, confirming B|A⊆T_B|A . The orthogonal complement of the model tangent space. If S is a closed subspace of ℋH, the orthogonal complement of S in ℋH, denoted ⟂S , is defined as the set of all residues h−Π(h∣)h- (h ) for all h∈ℋh , where Π(⋅∣) (· ) denotes the projection operator onto the subspace S. Therefore, we first derive Π(⋅∣) (· ) before constructing ⟂T . The expressions of A,B|A,C|A,D|CBT_A,T_B|A,T_C|A,T_D|CB in Equation 7 show that these subspaces are ranges of corresponding linear operators ΛA,ΛB|A,ΛC|A,ΛD|CB _A, _B|A, _C|A, _D|CB appeared in these expressions, respectively. For example, the linear operator ΛC|A _C|A is defined by ΛC|A(h)(c,a)≡[h∣c,a]−[h∣a] _C|A(h)(c,a) [h c,a]-E[h a], and it maps the Hilbert space ℋH into C|AT_C|A. In fact, these operators are actually projection operators onto these subspaces, which can be verified directly by checking the inner product between the residue and the image. For example, in the case of C|AT_C|A, we have ⟨h−ΛC|A(h),ΛC|A(h)⟩=[(h−[h∣c,a]+[h∣a])([h∣c,a]−[h∣a])]=0 h- _C|A(h), _C|A(h) =E [(h-E[h c,a]+E[h a])(E[h c,a]-E[h a]) ]=0 for all h∈ℋh . Since the subspaces A,B|A,C|A,D|CBT_A,T_B|A,T_C|A,T_D|CB are pairwise orthogonal, the projection of a function h∈ℋh onto the tangent space T is given by the sum of the individual projections Π(h∣)= (h )= ΛA(h)+ΛB|A(h)+ΛC|A(h)+ΛD|CB(h) _A(h)+ _B|A(h)+ _C|A(h)+ _D|CB(h) (9) = = [h∣d,c,b]−[h∣c,b]+[h∣c,a] [h d,c,b]-E[h c,b]+E[h c,a] +[h∣b,a]−[h∣a]. +E[h b,a]-E[h a]. Indeed, one can directly verify that this is the expression of the projection operator onto T by checking the inner product ⟨h−Π(h∣),Π(h∣)⟩=0 h- (h ), (h ) =0 for all h∈ℋh . Hence, the orthogonal complement of the tangent space is ⟂= =\ h−[h∣d,c,b]+[h∣c,b]−[h∣c,a] h-E[h d,c,b]+E[h c,b]-E[h c,a] (10) −[h∣b,a]+[h∣a]:∀h∈ℋ. -E[h b,a]+E[h a]:∀ h \. Given any DAG model, these described steps can be used to derive the tangent space and its orthogonal complement, as done in Tsiatis.2006.SemiparametricTheory, Rotnitzky.Smucler.2020.EfficientAdjustment, Bhattacharya.Nabi.ea.2022.SemiparametricInference. In the special case when the DAG model has only one constraint, we reproduce the following well-known lemma. Lemma 1. Suppose P is a single independence DAG model over variables V, defined by a single constraint ⟂∣X \!\!\! . The model’s tangent space T and its orthogonal complement ⟂T are, respectively = = h−[h∣,,]+[h∣,]−[h∣] \h-E[h ,y,z]+E[h ,z]-E[h ] (11) +[h∣,]:∀h∈ℋ, +E[h ,z]:∀ h \, ⟂= = [h∣,,]−[h∣,] \E[h ,y,z]-E[h ,z] −[h∣,]+[h∣]:∀h∈ℋ. -E[h ,z]+E[h ]:∀ h \. The proof of this lemma, which is in the Appendix, is similar to the above worked example, except that the model has only one constraint instead of two constraints, and disjoint subsets of variables replace the univariate variables. 3.2 Methodological obstacles for obtaining general tangent spaces In the worked example in the previous section, and indeed in all DAG models, both the tangent space T and its orthogonal complement ⟂T are expressible in closed form. That is, one can construct arbitrary elements of T and ⟂T using Equation 8 and 10, respectively. This is particularly useful for inference, since if φ is a known influence function for some target parameter ψ(p)ψ(p), then φ+h−[h∣d,c,b]+[h∣c,b]−[h∣c,a] +h-E[h d,c,b]+E[h c,b]-E[h c,a] (12) −[h∣b,a]+[h∣a] -E[h b,a]+E[h a] is also an influence function of that target, where h is an arbitrary function in ℋH. The key reason for the existence of the closed form expressions for T and ⟂T is the fact that the DAG model likelihood in Equation 5 is also in closed form, as a single factorization representing all constraints in the local Markov property of the model. On the other hand, in Markov models associated with ADMGs, no single factorization representing all model constraints is known for general state spaces. For example, in the Bell scenario in Figure 1d, a closed form factorization of pε(a,b,c,d)p_ (a,b,c,d) representing both constraints A⟂B,DA \!\!\! ,D and B⟂A,CB \!\!\! ,C is unknown in general settings. Instead, distributions satisfying both constraints may be obtained iteratively by solving the following system of two equations pε(a,b,c,d) p_ (a,b,c,d) =pε(a)pε(b,d)pε(c|d,b,a), =p_ (a)p_ (b,d)p_ (c|d,b,a), (13) pε(a,b,c,d) p_ (a,b,c,d) =pε(b)pε(c,a)pε(d|c,b,a). =p_ (b)p_ (c,a)p_ (d|c,b,a). (14) The first equation forces the distribution to satisfy A⟂B,DA \!\!\! ,D, while the second equation forces the likelihood to satisfy B⟂A,CB \!\!\! ,C. We will refer to them as the likelihood constraint equations. Mathematical objects defined as solutions of some system of equations, such as pεp_ above, is said to have implicit expressions. Subsequently, the score function of this parametric submodel is also defined implicitly as solutions to the following two equations s(a,b,c,d) s(a,b,c,d) =s(a)+s(b,d)+s(c|a,b,d), =s(a)+s(b,d)+s(c|a,b,d), (15) s(a,b,c,d) s(a,b,c,d) =s(b)+s(a,c)+s(d|b,a,c), =s(b)+s(a,c)+s(d|b,a,c), (16) which are referred to as the score constraint equations. To see that these two equation solves for the same object, we can equivalently reexpress these component score functions in terms of the total score function s(a,b,c,d)s(a,b,c,d) as [s∣a,b,d] [s a,b,d] =[s∣a]+[s∣b,d], =E[s a]+E[s b,d], (17) [s∣b,a,c] [s b,a,c] =[s∣b]+[s∣a,c]. =E[s b]+E[s a,c]. (18) Similarly to Equations 13 and 14, solving this system of equations is an open problem in general settings. These problems are closely related to implicitization problems in algebraic geometry [Cox.Little.ea.2015.IdealsVarieties] and parameterization problems of graphical model likelihoods [Evans.Richardson.2019.SmoothIdentifiable, Shpitser.2023.LauritzenChenLikelihood], which may be viewed as the problem of transforming an implicit definition of a likelihood surface via constraints into an explicit definition of said likelihood indexed by variationally independent parameters. Absent such a likelihood definition, tangent spaces for graphical models such as the Bell scenario model will have an implicit expression of the form: =all common solutions of Equations 17 and 18T=\all common solutions of Equations 10000\ eq:bell_total_score_p1 and eq:bell_total_score_p2\. It is unknown how to construct the projection operator onto this subspace. Consequently, the orthogonal complement for the tangent space also has an implicit form: ⟂=all f orthogonal to all common solutions of 17 and 18T =\all f orthogonal to all common solutions of eq:bell_total_score_p1 and eq:bell_total_score_p2\. In other words, as opposed to the DAG case, elements of ⟂T cannot be written down in closed form. In the next section, we illustrate an alternative simple approach for obtaining the characterization of the orthogonal complement of the tangent space that bypasses the need for an explicit representation of the likelihood or the tangent space. 4 The orthocomplement of the Tangent Space of a CI Model A general Markov model may be viewed as an intersection of single independence models, each with one CI constraint. Thus, the tangent space of the general Markov model is the intersection of all the single independence models’s tangent spaces. Consequently, the orthogonal complement of this intersection is the span of all the orthogonal complements of the single independence models’s tangent spaces. This geometric characterization of the tangent space’s orthogonal complement for intersection models first appears in Lemma 1.7 of VanDerLaan.Robins.2003.UnifiedMethods; our argument is an application of this observation in the context of Markov models. 4.1 A worked example: the Bell scenario model We define the following two models with only one constraint 1 P_1 =all p(a,b,c,d) satisfying A⟂B,D, =\all p(a,b,c,d) satisfying A \!\!\! ,D\, (19) 2 P_2 =all p(a,b,c,d) satisfying B⟂A,C. =\all p(a,b,c,d) satisfying B \!\!\! ,C\. (20) The Bell scenario model satisfying both constraints is then the intersection of these two models, (d)=1∩2. P^(d)= P_1∩ P_2. (21) Graphical models are commonly defined in this way, that is as intersections of simpler models [Lauritzen.1996.GraphicalModels, Richardson.2003.MarkovProperties, Richardson.Evans.ea.2023.NestedMarkov]. Both models 1 P_1 and 2 P_2 are single independence DAG models – DAG models defined by a single constraint. Consequently, the technique described in Section 3 applies to these models. The parametric submodel likelihoods for 1 P_1 and 2 P_2 are Equations 13 and 14, respectively. The score functions for these parametric submodels are given by Equations 15 and 16 (equivalently, Equations 17 and 18), respectively. By Lemma 1, the tangent spaces for these models are 1 _1 =h−[h|a,b,d]+[h|a]+[h|b,d]:∀h∈ℋ, =\h-E[h|a,b,d]+E[h|a]+E[h|b,d]:∀ h \, (22) 2 _2 =h−[h|b,a,c]+[h|b]+[h|a,c]:∀h∈ℋ, =\h-E[h|b,a,c]+E[h|b]+E[h|a,c]:∀ h \, and their orthogonal complements are 1⟂ _1 =[h|a,b,d]−[h|a]−[h|b,d]:∀h∈ℋ, =\E[h|a,b,d]-E[h|a]-E[h|b,d]:∀ h \, (23) 2⟂ _2 =[h|b,a,c]−[h|b]−[h|a,c]:∀h∈ℋ. =\E[h|b,a,c]-E[h|b]-E[h|a,c]:∀ h \. The tangent space T of the Bell scenario model is the closure of all functions satisfying both Equations 17 and 18, =1∩2¯T= T_1 _2. Since both 1T_1 and 2T_2 are closed, =1∩2T=T_1 _2, the intersection model’s tangent space is the intersection of tangent spaces. For detail, see the proof of Theorem 3 in the Appendix. This implies ([VanDerLaan.Robins.2003.UnifiedMethods], Lemma 1.7) Lemma 2. The orthogonal complement of the tangent space of the Bell scenario model is the direct sum ⟂=1⟂⊕2⟂T =T_1 _2 , whose explicit expression is ⟂ =[h|a,d,b]−[h|a]−[h|d,b] =\E[h|a,d,b]-E[h|a]-E[h|d,b] (24) + + [g|b,c,a]−[g|b]−[g|c,a]:∀g,h∈ℋ. [g|b,c,a]-E[g|b]-E[g|c,a]:∀ g,h \. Proof. Since =1∩2T=T_1 _2, elements of T must be orthogonal to both 1⟂T_1 and 2⟂T_2 . This implies the span 1⟂⊕2⟂¯ T_1 _2 is in the orthogonal complement of T. Since both 1⟂T_1 and 2⟂T_2 are closed, 1⟂⊕2⟂¯=1⟂⊕2⟂⊆⟂ T_1 _2 =T_1 _2 . For other direction, pick any f∈(1⟂⊕2⟂)⟂f∈ (T_1 _2 ) . Then f must be orthogonal to just 1⟂T_1 , so f∈1f _1. Similarly, f∈2f _2, and therefore f∈f . We just show (1⟂⊕2⟂)⟂⊆ (T_1 _2 ) , and by taking the orthogonal complement we get ⟂⊆1⟂⊕2⟂T _1 _2 . Therefore, ⟂=1⟂⊕2⟂T =T_1 _2 . Finally, the explicit expression for ⟂T follows from the definition of direct sum and the explicit expressions for 1⟂T_1 and 2⟂T_2 . ∎ Consequently, if φ is a known influence function of a target parameter in this model, then φ +[h∣a,b,d]−[h∣a]−[h∣b,d] +E[h a,b,d]-E[h a]-E[h b,d] (25) +[g∣b,a,c]−[g∣b]−[g∣a,c] +E[g b,a,c]-E[g b]-E[g a,c] is also an influence function of that target parameter, where g,hg,h are arbitrary functions in ℋH. Equation 25 thus characterizes the class of influence functions for any target parameter in the Bell scenario model. 4.2 Main Result The result for the Bell scenario model described in the previous section can easily be extended to any Markov model, including all mentioned graphical models associated to UGs, CGs, and ADMGs. Consider a general Markov model defined by a list of K≥1K≥ 1 marginal and conditional independence constraints, as model without any constraint (K=0K=0) is the saturated model described in Tsiatis.2006.SemiparametricTheory. By definition, this Markov model is the intersection of K single independence DAG models defined by a single constraint =1∩…∩K P= P_1∩…∩ P_K. The tangent space is the intersection of the tangent spaces of these single independence models, =1∩…∩KT=T_1∩… _K, and its orthogonal complement is the direct sum ⟂=1⟂⊕…⊕K⟂T =T_1 … _K . Theorem 3. Suppose P is a general semi-parametric Markov model defined by K≥1K≥ 1 conditional independence constraints i⟂i∣iX_i \!\!\! _i _i, where i=1,…,Ki=1,…,K. For each i, let i P_i be the single independence DAG model defined by the i-th constraint. The orthogonal projection operator Π(⋅∣i⟂):ℋ↦ℋ (· _i ):H sends each function h∈ℋh to the orthogonal complement i⟂T_i of the tangent space iT_i of i P_i, Π(h∣i⟂)(i,i,i)=[h|i,i,i]−[h|i,i] (h _i )(x_i,y_i,z_i)=E[h|x_i,y_i,z_i]-E[h|x_i,z_i] (26) −[h|i,i]+[h|i] -E[h|y_i,z_i]+E[h|z_i] . The orthogonal complement of the tangent space of P is the direct sum of these orthogonal complements, ⟂=∑i=1KΠ(hi∣i⟂):∀h1,…,hK∈ℋ. = \ _i=1^K (h_i _i )\>:\>∀ h_1,…,h_K \. (27) Proof Sketch. Applying Lemma 1 to the single independence model iP_i, any element of the tangent space’s orthogonal complement i⟂T_i is Π(hi∣i⟂) (h_i _i ), with arbitrary hi∈ℋh_i . A general version of Lemma 2 shows that ⟂=1⟂⊕…⊕K⟂T =T_1 … _K . Hence, any element of ⟂T is a sum ∑i=1KΠ(hi∣i⟂) _i=1^K (h_i _i ) of elements of i⟂i=1K\T_i \_i=1^K, where hi∈ℋi=1K\h_i \_i=1^K are arbitrary (not necessarily equal) functions. ∎ This theorem applies to all Markov models – models defined by a list of CI constraints – which are not necessarily graphical models. Furthermore, for graphical Markov models such as those associated with UGs, DAGs, CGs, BGs, and ADMGs, the list of CI constraints can be obtained from global, local or any other Markov property, as long as this list fully describes the model. This means that it is possible to obtain several syntatically different representations of the same orthocomplement of the tangent space of the model by using different Markov properties. For example, the local Markov property for DAGs listed in Section 2.1 differs from the ordered local Markov property for DAGs, which is used in Bhattacharya.Nabi.ea.2022.SemiparametricInference, Tsiatis.2006.SemiparametricTheory, Rotnitzky.Smucler.2020.EfficientAdjustment. Despite the different expressions, the orthogonal complement of the tangent space is the same. Equation 81 suggests an operator Γ:ℋ↦⟂ :H sending functions in ℋH to functions in the orthogonal complement ⟂T , defined by Γ(f)=∑i=1KΠ(f∣i⟂) (f)= _i=1^K (f _i ). For general Markov model, it is unknown whether this operator is the projection operator onto ⟂T . The Appendix shows an example in the Bell scenario case. Therefore, finding the projection operator onto a general Markov model’s tangent space or its orthogonal complement is still an open problem. Since the efficient influence function (EIF) is φ−Π(φ∣⟂) - ( ), where φ is any influence function and Π(⋅∣⟂) (· ) is the projection onto the orthogonal complement ⟂T of the tangent space, the exact form of the efficient influence function for a target parameter in a general Markov model is also an open problem. A direct consequence of Theorem 3 is the explicit characterization of the class of all influence functions of any target parameter ψ(p)ψ(p) in a general Markov model. If φ is a known influence function of ψ(p)ψ(p), then φ+∑i=1KΠ(hi∣i⟂) + _i=1^K (h_i _i ) (28) is also an influence function of ψ(p)ψ(p), for any hi∈ℋi=1K\h_i \_i=1^K. An explicit example is shown in the next section. 5 Applications 5.1 The Class of Influence Functions In this subsection, we illustrate how our result streamline the derivations of the class of influence functions for the conditional mean parameter ψ(p)=[D∣A=a0]ψ(p)=E[D A=a_0] with a fixed value a0a_0, for all the Markov models corresponding to the graphs in Figure 1. Extensions to other target parameters are trivial. For example, if the target parameter is the adjustment formula [[D∣A=a0,C]]E[E[D A=a_0,C]], simply replace φ below by the augmented inverse probability weighting (AIPW) functional. The saturated model ABCDall P^all_ABCD in Figure 1f. The saturated model ABCDall P^all_ABCD over four variables A,B,C,D\A,B,C,D\ is a saturated extension over the saturated model ADall P^all_AD with only two variables A,D\A,D\, with B and C play the role of auxillary variables. Following Tsiatis.2006.SemiparametricTheory, Chapter 5.5, the set of influence functions in ABCDall P^all_ABCD is exactly the set of influence functions in ADall P^all_AD. This means, for the purpose of estimating [D∣A=a0]E[D A=a_0], we can safely ignore B and C as these variables do not provide efficiency gain. Moreover, the derivation of the influence functions in ADall P^all_AD involves only two variables and hence is much simpler. Any saturated model has 0\0\ as its orthogonal complement of the tangent space, so there is only one influence function, φ(a,b,c,d)=(a=a0)p(a)(d−[D|a]). (a,b,c,d)= I(a=a_0)p(a)(d-E[D|a]). (29) All other Markov models corresponding to the graphs in Figure 1a-e are submodels of ABCDall P^all_ABCD. Therefore, φ is also an influence function in any of these models. The other ingredient for the derivation the class of influence functions in these models is the orthogonal complement of the tangent space given by Theorem 3. The DAG model (a) P^(a) in Figure 1a. Let 1 P_1 and 2 P_2 be the simple models corresponding to the constraints C⟂B∣AC \!\!\! A and D⟂A∣C,BD \!\!\! C,B, respectively. Let 1⟂T_1 and 2⟂T_2 denote the orthocomplements of their respective tangent spaces. The projection operators onto these orthocomplements are Πa(g∣1⟂)= _a(g _1 )= [g|c,b,a]−[g|c,a]−[g|b,a]+[g|a], [g|c,b,a]-E[g|c,a]-E[g|b,a]+E[g|a], Πa(h∣2⟂)= _a(h _2 )= h−[h|d,c,b]−[h|a,c,b]+[h|c,b]. h-E[h|d,c,b]-E[h|a,c,b]+E[h|c,b]. Therefore, the orthogonal complement of the tangent space T for this DAG model is, according to Equation 81 ⟂=Πa(g∣1⟂)+Πa(h∣2⟂):∀g,h∈ℋ, =\ _a(g _1 )+ _a(h _2 ):∀ g,h \, (30) This expression differs from the one given in Equation 10. To reconcile the two, note that ⟂T is the direct sum of the orthogonal subspaces 1⟂T_1 and 2⟂T_2 . This geometric property holds for general DAGs: the orthocomplements of the tangent spaces corresponding to distinct CIs in the local Markov property are pairwise orthogonal. Since the projection operators onto these subspaces are known – namely, Πa(⋅∣1⟂) _a(· _1 ) and Πa(⋅∣2⟂) _a(· _2 ) – the projection of a function h∈ℋh onto ⟂T is given by Π(h∣⟂)=Πa(h∣1⟂)+Πa(h∣2⟂) (h )= _a(h _1 )+ _a(h _2 ). This shows that the two expressions for ⟂T in 10 and 30 describe the same subspace. The class of influence functions for ψ(p)ψ(p) in this model is φ+Πa(g∣1⟂)+Πa(h∣2⟂):∀g,h∈ℋ\ + _a(g _1 )+ _a(h _2 ):∀ g,h \. The UG model (b) P^(b) in Figure 1b. 1 P_1 and 2 P_2 correspond to C⟂B∣A,DC \!\!\! A,D and D⟂A∣C,BD \!\!\! C,B, respectively. Πb(g∣1⟂)= _b(g _1 )= g−[g∣c,a,d]−[g|b,a,d]+[g|a,d], g-E[g c,a,d]-E[g|b,a,d]+E[g|a,d], Πb(h∣2⟂)= _b(h _2 )= h−[h|d,c,b]−[h|a,c,b]+[h|c,b]. h-E[h|d,c,b]-E[h|a,c,b]+E[h|c,b]. The class of influence functions for ψ(p)ψ(p) in this model is φ+Πb(g∣1⟂)+Πb(h∣2⟂):∀g,h∈ℋ\ + _b(g _1 )+ _b(h _2 ):∀ g,h \. The CG model (c) P^(c) in Figure 1c. 1 P_1 and 2 P_2 are the same models in the UG case, while 3 P_3 corresponds to A⟂BA \!\!\! . The projection operator onto 3⟂T_3 is Πc(f∣3⟂)=[f∣a,b]−[f∣a]−[f∣b]. _c(f _3 )=E[f a,b]-E[f a]-E[f b]. The class of influence functions for ψ(p)ψ(p) in this model is φ+Πb(g|1⟂)+Πb(h|2⟂)+Πc(f|3⟂):∀f,g,h∈ℋ\ + _b(g|T_1 )+ _b(h|T_2 )+ _c(f|T_3 ):∀ f,g,h \. The bidirected square model (e) P^(e) in Figure 1e. 1 P_1 and 2 P_2 correspond to A⟂DA \!\!\! and B⟂CB \!\!\! , respectively. Πd(g∣1⟂) _d(g _1 ) =[g∣a,d]−[g∣a]−[g∣d], =E[g a,d]-E[g a]-E[g d], Πd(h∣2⟂) _d(h _2 ) =[h∣b,c]−[h∣b]−[h∣c]. =E[h b,c]-E[h b]-E[h c]. The class of influence functions for ψ(p)ψ(p) in this model is φ+Πd(g∣1⟂)+Πd(h∣2⟂):∀g,h∈ℋ\ + _d(g _1 )+ _d(h _2 ):∀ g,h \. 5.2 Improving Efficiency The variance of a RAL estimator ψ ψ for a target parameter ψ(p)ψ(p) is exactly the variance of the corresponding influence function φ∈ℋ . For influence functions have mean zero, the variance of φ is its squared norm [φ2]E[ ^2]. Thus, an estimator ψ^2 ψ_2 is at least as efficient as another estimator ψ^1 ψ_1 precisely when the variance of the influence function φ2 _2 of ψ^2 ψ_2 is equal or smaller than the variance of the influence functions φ1 _1 of ψ^1 ψ_1, [φ22]≤[φ12]E[ _2^2] [ _1^2] (Tsiatis.2006.SemiparametricTheory, Chapter 3). The efficient influence function is simply the influence function with the smallest variance in the class of all influence functions of the model. Therefore, in order to construct more efficient estimators, the first step is finding influence functions with smaller variances. Given an initial influence function φ0 _0 of a target parameter, our main result immediately implies an iterative method for constructing sequences of influence functions with non-increasing variances. As an example, suppose the Markov model has K≥2K≥ 2 CI constraints. Theorem 3 states that any ∑i=1KΠ(hi∣i⟂) _i=1^K (h_i _i ) is an element of ⟂T . Therefore, Π(φ0∣1⟂)∈⟂ ( _0 _1 ) . By Theorem 4.3 of Tsiatis.2006.SemiparametricTheory, adding an element in ⟂T to an influence function yields another influence function of the same target, hence φ1:=φ0−Π(φ0∣1⟂) _1:= _0- ( _0 _1 ) (31) is an influence function. Moreover, since Π(φ0∣1⟂) ( _0 _1 ) is the orthogonal projection of φ0 _0 onto 1⟂T_1 , the above φ1 _1 is the orthogonal projection of φ0 _0 onto 1T_1. The Pythagorean theorem for Hilbert spaces (Tsiatis.2006.SemiparametricTheory, Theorem 3.3) states that the projection of a vector cannot be longer than the vector itself, so E[φ12]≤E[φ02]E[ _1^2]≤ E[ _0^2]. This implies that φ1 _1 is an influence function with equal or smaller variance than φ0 _0. Next, consider φ2:=φ1−Π(φ1∣2⟂). _2:= _1- ( _1 _2 ). (32) By the same reasoning, φ2 _2 is an influence function with equal or smaller variance than φ1 _1. Repeating this process M times, we obtain a sequence of influence functions φ0,φ1,φ2,…,φM _0, _1, _2,…, _M, with the property that the variance of φm _m is equal or smaller than the variance of φm−1 _m-1, for all m=1,…,Mm=1,…,M. The influence function φm _m is referred to as an m-improvement of the initial φ0 _0. If we can construct an estimator ψ^M ψ_M based on the influence function φM _M, then this estimator is at least as efficient as the estimator ψ^0 ψ_0 based on the initial influence function φ0 _0. In the Appendix, we give a concrete example of such a construction with M=2M=2. Theorem 4. Consider a Markov model with K≥1K≥ 1 constraint(s), and let φ0 _0 be an influence function of a target parameter ψ(p)ψ(p). For any integer sequence i1,i2,…,iMi_1,i_2,…,i_M, where im∈1,…,Ki_m∈\1,…,K\ for all m∈1,…,Mm∈\1,…,M\, define φm:=φm−1−Π(φm−1∣im⟂),∀m∈1,…,M. _m= _m-1- ( _m-1 _i_m ), ∀ m∈\1,…,M\. (33) In words, φm _m is the projection of φm−1 _m-1 onto the tangent space imT_i_m of the DAG model corresponding to the imi_m-th constraint. Then φ0,…,φM _0,…, _M is a sequence of influence functions of ψ(p)ψ(p) with non-increasing variance, meaning E[φm2]≤E[φm−12]E[ _m^2]≤ E[ _m-1^2] for all m∈1,…,Mm∈\1,…,M\. If im+1=imi_m+1=i_m, then φm+1=φm _m+1= _m is the same influence function, because φm∈im _m _i_m and so Π(φm∣im+1⟂)=0 ( _m _i_m+1 )=0. Then, for single constraint models (K=1K=1), there is only one improvement, φ=φ0−Π(φ0∣⟂) = _0- ( _0 ), which is exactly the efficient influence function in these models. In models with K>1K>1, the connection between the efficient influence function and φM _M as M→∞M→∞ is an interesting question, yet outside the scope of our paper. 6 Conclusion We derived an explicit expression for the orthogonal complement of the tangent space for Markov models – statistical models defined using only marginal and conditional independence constraints. We demonstrated two applications of our result in these models. The first one is an explicit characterization of the class of influence functions for any target parameter. The second one is is an iterative method to improve efficiency of any initial RAL estimator of any target parameter. Our method does not yield an explicit expression for the likelihood or the tangent space of these models, so these questions remain open. Another important open problem is the projection of a function h∈ℋh onto the tangent space of these models. Derivation of such a projection is equivalent to the derivation of the efficient influence function in a semi-parametric Markov model. Our method provides an important step towards achieving this goal by providing the explicit form of the orthogonal complement of the tangent space. Acknowledgements.This work was supported by the Office of Naval Research (Grant No. N000142412701), the National Institutes of Health (Grant No. 1R56AI191526-01 and Grant no. 1R01LM014800-01A1), and the National Science Foundation CAREER Award (Grant No. 1942239). References A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models (Supplementary Material) Appendix A Additional Details about Semiparametric Theory The set of all influence function and the orthogonal complement of the tangent space. As mentioned in the main paper, an influence function of the target parameter ψ(p)ψ(p) is a solution of the pathwise differentiability condition, for all parametric submodel pεp_ (Tsiatis.2006.SemiparametricTheory, Theorem 3.2 and 4.2) ∂εψ(pε)|ε=0=∫φ()s()p0()d.(4) ∂ ψ(p_ ) |_ =0= (v)s(v)p_0(v)dv. eq:path_dev Suppose φ is an influence function of ψ(p)ψ(p) and ⟂T is the orthogonal complement of the tangent space of the model. Let f be any function in ⟂T . Since the score function s()∈s(v) , it is orthogonal to f, that is ⟨f,s⟩=∫f()s()p0()=0 f,s = f(v)s(v)p_0(v)dv=0. Then φ+f +f is also an influence function, because ∂εψ(pε)|ε=0=∫[φ()+f()]s()p0(). ∂ ψ(p_ ) |_ =0= [ (v)+f(v) ]s(v)p_0(v)dv. (34) On the other hand, if φ~ is any other influence function, meaning ∂εψ(pε)|ε=0=∫φ~()s()p0(). ∂ ψ(p_ ) |_ =0= (v)s(v)p_0(v)dv. (35) Then, for all parametric submodel pεp_ (equivalently, for all score function s()∈s(v) ) 0=∫[φ~()−φ()]s()p0().0= [ (v)- (v) ]s(v)p_0(v)dv. (36) This means φ~−φ∈⟂ - , i.e., φ~ is the sum of φ and an element of the orthogonal complement of the tangent space. This result is Theorem 4.3 in Tsiatis.2006.SemiparametricTheory, stating that: the set of all influence functions of a target parameter in a semiparametric model is the linear variety φ⊕⟂:=φ+f:f∈⟂ :=\ +f:f \, where φ is any influence function and ⟂T is the orthogonal complement of the tangent space. Hence, knowing the orthogonal complement of the tangent space ⟂T is vital to characterize the set of all influence functions. Variance of an influence function and efficiency of an estimator. As stated in the main paper, the limiting distribution of a RAL estimator ψ^n ψ_n of the target parameter ψ is n(ψ^n−ψ)↝(0,[φ2])(2) n( ψ_n-ψ) (0,E[ ^2]) eq:asymp_linear_clt where φ is the influence function corresponding to ψ ψ [Tsiatis.2006.SemiparametricTheory]. Since φ∈ℋ , it has zero mean, and therefore the variance of the influence function φ is its square norm in ℋH, Var(φ)=E[φ2]−E[φ]2=E[φ2]=⟨φ,φ⟩Var( )=E[ ^2]-E[ ]^2=E[ ^2]= , . An estimator ψ ψ is more efficient than another estimator ψ^′ ψ precisely when the variance of the limiting distribution of n(ψ^n−ψ) n( ψ_n-ψ) is equal or smaller than the variance of the limiting distribution of n(ψ^n′−ψ) n( ψ _n-ψ) [Tsiatis.2006.SemiparametricTheory]. This means the variance of the influence function φ of ψ ψ is equal or smaller than the variance of the influence function φ′ of ψ^′ ψ . Therefore, given an influence function φ and the orthogonal complement of the tangent space ⟂T , if one can construct φ−f -f, where f∈⟂f , such that [(φ−f)2]≤[φ2] *E[( -f)^2]≤ *E[ ^2], then the estimator obtained from φ−f -f will be equal or more efficient than the estimator obtained from φ . In particular, we derive the orthogonal complement of the tangent space for general Markov model in Theorem 3, and propose a method to improve the variance of an initial influence function in Theorem 4. The efficient influence function is the influence function φ(e) ^(e) in the class of all influence functions φ⊕⟂ with smallest variance, i.e., [(φ(e))2]≤[φ′2] *E[( ^(e))^2]≤ *E[ 2] for any influence function φ′∈φ⊕⟂ ∈ [Tsiatis.2006.SemiparametricTheory]. One way to find the efficient influence function is projecting any other influence function φ onto the tangent space T. Let Π(⋅∣) (· ) be the projection onto T, and Π(⋅∣⟂) (· ) is the projection onto the orthogonal complement ⟂T . We have Π(⋅∣⟂)=I−Π(⋅∣) (· )=I- (· ). For any influence function φ φ=φ−Π(φ∣⟂)+Π(φ∣⟂) = - ( )+ ( ) (37) Since Π(φ∣⟂) ( ) is an element in the orthogonal complement of the tangent space, φ−Π(φ∣⟂) - ( ) must be another influence function. Since φ−Π(φ∣⟂)=Π(φ∣) - ( )= ( ) is the projection of φ onto a subspace, its norm must be equal or less than the norm of φ , so [(φ−Π(φ∣⟂))2]≤[φ2] *E[( - ( ))^2]≤ *E[ ^2]. It can be shown (Tsiatis.2006.SemiparametricTheory, Theorem 3.5) that the efficient influence function is φ(e)=(φ−Π(φ∣⟂) ^(e)=( - ( ). Hence, to find the efficient influence function, one must first know the projection operator Π(⋅∣⟂) (· ) or Π(⋅∣) (· ). Appendix B Proofs in Section 3 The following lemma is the general version of Equation 7. Lemma 5 (Tsiatis.2006.SemiparametricTheory, Theorem 4.5). Let P be a semi-parametric model over variables V. Consider any disjoint subsets ,⊆X,Y . For any parametric submodel pεp_ , the score function s(∣)=∂εlogpε(∣)|ε=0s(x )= ∂ p_ (x )|_ =0 at p0p_0 is an element of the subspace ∣:=E[h∣,]−E[h∣]:∀h∈ℋ.T_X :=\E[h ,y]-E[h ]:∀ h \. (38) Proof. Let =∖(∪˙)Z=V (X ∪Y). The following is a really useful identify for the score function s(,,) s(x,y,z) :=∂εlogpε(,,)|ε=0=1p0(,,)∂εpε(,,)|ε=0. = ∂ p_ (x,y,z) |_ =0= 1p_0(x,y,z) ∂ p_ (x,y,z) |_ =0. (39) Then s(∣) s(x ) :=∂εlogpε(∣)|ε=0=∂εlogpε(,)|ε=0−∂εlogpε()|ε=0 = ∂ p_ (x ) |_ =0= ∂ p_ (x,y) |_ =0- ∂ p_ (y) |_ =0 (40) =1p0(,)∂ε∫pε(,,)|ε=0−1p0()∂ε∫pε(,,)|ε=0 = 1p_0(x,y) ∂ p_ (x,y,z)dz |_ =0- 1p_0(y) ∂ p_ (x,y,z)dxdz |_ =0 =1p0(,)∫(∂εpε(,,)|ε=0)−1p0()∫(∂εpε(,,)|ε=0) = 1p_0(x,y) ( ∂ p_ (x,y,z) |_ =0 )dz- 1p_0(y) ( ∂ p_ (x,y,z) |_ =0 )dxdz =1p0(,)∫s(,,)p0(,,)−1p0()∫s(,,)p0(,,) = 1p_0(x,y) s(x,y,z)p_0(x,y,z)dz- 1p_0(y) s(x,y,z)p_0(x,y,z)dxdz =∫s(,,)p0(∣,)−∫s(,,)p0(,∣) = s(x,y,z)p_0(z ,y)dz- s(x,y,z)p_0(x,z )dxdz =E[s()∣,]−E[s()∣]. =E[s(V) ,y]-E[s(V) ]. Since s()∈ℋs(v) , the claim is established. ∎ The following lemma shows that the tangent subspaces of the two pathway DAG model (Figure 1a) described in Equation 7 are orthogonal. Lemma 6 (Tsiatis.2006.SemiparametricTheory, Theorem 4.5). Let ,,,A,B,C,D be disjoint sets of random variables. The following 2 subspaces are orthogonal 1 _1 =[h∣,,,]−[h∣,,]:∀h∈ℋ, =\ *E[h ,b,c,d]- *E[h ,c,d]:∀ h \, (41) 2 _2 =[h∣,]−[h∣]:∀h∈ℋ. =\ *E[h ,d]- *E[h ]:∀ h \. Proof. Pick any [h1∣,,,]−[h1∣,,]∈1 *E[h_1 ,b,c,d]- *E[h_1 ,c,d] _1 and any [h2∣,]−[h∣]∈2 *E[h_2 ,d]- *E[h ] _2, then consider their inner product [([h1∣,,,]−[h1∣,,])([h2∣,]−[h2∣])] *E [ ( *E[h_1 ,B,C,D]- *E[h_1 ,C,D] ) ( *E[h_2 ,D]- *E[h_2 ] ) ] (42) =∫(∫([h1∣,,,]−[h1∣,,])([h2∣,]−[h2∣])p(,∣,))p(,) = ( ( *E[h_1 ,b,c,d]- *E[h_1 ,c,d] ) ( *E[h_2 ,d]- *E[h_2 ] )p(a,b ,d)dadb )p(c,d)dcdd =0 =0 by the law of iterated expectations. ∎ Lemma 7. The operator Π:ℋ↦ℋ :H defined by Π(h)=[h∣,]−[h∣] (h)= *E[h ,y]- *E[h ] (43) is the orthogonal projection onto the subspace ∣:=E[h∣,]−E[h∣]:∀h∈ℋT_X :=\E[h ,y]-E[h ]:∀ h \. Proof. By definition, ∣=Π(h):∀h∈ℋT_X =\ (h):∀ h \. For any f,g∈ℋf,g ⟨Π(f),(I−Π)(g)⟩ (f),(I- )(g) =[([f∣,]−[f∣])(g−[g∣,]+[g∣])] = *E [ ( *E[f ,Y]- *E[f ] ) (g- *E[g ,Y]+ *E[g ] ) ] (44) =[([f∣,]−[f∣])([g∣,]−[g∣,]+[g∣])] = *E [ ( *E[f ,Y]- *E[f ] ) ( *E[g ,Y]- *E[g ,Y]+ *E[g ] ) ] =[([f∣,]−[f∣])[g∣]] = *E [ ( *E[f ,Y]- *E[f ] ) *E[g ] ] =0, =0, by the law of iterated expectations. Any function h∈ℋh can be decomposed into the sum h=(h−Π(h))+Π(h)h= (h- (h) )+ (h) of two orthogonal terms, where Π(h)∈∣ (h) _X and the error term h−Π(h)h- (h) is orthogonal to ∣T_X . Therefore, Π is the orthogonal projection operator onto ∣T_X [Luenberger.1997.OptimizationVector]. ∎ Lemma 8. The tangent space of the two pathway DAG model in Figure 1a is the direct sum =A⊕B|A⊕C|A⊕D|BC =T_A _B|A _C|A _D|BC (8) Proof. Equation 7 shows that ⊆A⊕B|A⊕C|A⊕D|BCT _A _B|A _C|A _D|BC. To prove the equality, all we need to prove is that each subspace on the r.h.s. is contained in T. In each case, we pick arbitrary f in the subspace under consideration and define a parametric model ε P_ whose elements are pε=p0(1+εf)p_ =p_0(1+ f). We ensure pεp_ is a valid density using the conventional technique in Tsiatis.2006.SemiparametricTheory: (i) non-negativity pε(a,b,c,d)≥0p_ (a,b,c,d)≥ 0 is achieved by restricting the range of ε to be (−m,m)(-m,m) with m small enough, and (i) summation ∫pε(a,b,c,d)abcd=1+εE[f]=1 p_ (a,b,c,d)dadbdcdd=1+ E[f]=1 for all ε , because any f∈ℋf must have zero-mean, by definition of ℋH. This parametric model contain the truth p0p_0 at ε=0 =0. Moreover, the score function at p0p_0 is ∂εlogpε|ε=0=f ∂ p_ |_ =0=f. Finally, if we can show pεp_ satisfies the DAG’s constraints, then ε P_ is a parametric submodel of the DAG’s semiparametric model, and hence the chosen f must be in the tangent space T. We will directly check pε(c∣b,a)=pε(c∣a)p_ (c b,a)=p_ (c a) and pε(d∣c,b,a)=pε(d∣c,b)p_ (d c,b,a)=p_ (d c,b) to show that pεp_ satisfies the constraints of the DAG model. This will conclude that any function in that subspace is a valid score function of this DAG model. Note that at the truth p0p_0, the constraint D⟂A∣C,BD \!\!\! C,B implies p0(d∣c,b,a)=p0(d∣c,b)p_0(d c,b,a)=p_0(d c,b), and the constraint C⟂B∣AC \!\!\! A implies p0(c∣b,a)=p0(c∣a)p_0(c b,a)=p_0(c a) and p0(b∣c,a)=p0(b∣a)p_0(b c,a)=p_0(b a). We will use these equalities repeatedly in the proof. Finally, expectations are taken at the truth p0p_0. First case A⊆T_A : Pick any f=[h∣a]∈Af=E[h a] _A. Then pε(c,b,a) p_ (c,b,a) =∫p0(d,c,b,a)(1+ε[h∣a])d=p0(c,b,a)(1+ε[h∣a]) = p_0(d,c,b,a)(1+ [h a])d=p_0(c,b,a)(1+ [h a]) (45) Therefore, by definition of conditional distribution pε(c∣b,a) p_ (c b,a) =pε(c,b,a)pε(b,a)=p0(c∣b,a) = p_ (c,b,a)p_ (b,a)=p_0(c b,a) (46) pε(c∣a) p_ (c a) =pε(c,a)pε(a)=p0(c∣a) = p_ (c,a)p_ (a)=p_0(c a) pε(d∣c,b,a) p_ (d c,b,a) =pε(d,c,b,a)pε(c,b,a)=p0(d∣c,b,a) = p_ (d,c,b,a)p_ (c,b,a)=p_0(d c,b,a) This shows pε(c∣b,a)=p0(c∣b,a)=p0(c∣a)=pε(c∣a)p_ (c b,a)=p_0(c b,a)=p_0(c a)=p_ (c a). Next pε(d,c,b) p_ (d,c,b) =∫p0(d,c,b)p0(a∣c,b)(1+ε[h∣a])a = p_0(d,c,b)p_0(a c,b)(1+ [h a])da (A⟂D∣C,B) (A \!\!\! C,B) (47) =p0(d,c,b)(1+ε∫[h∣a]p0(a∣c,b)a) =p_0(d,c,b) (1+ [h a]p_0(a c,b)da ) pε(c,b) p_ (c,b) =p0(c,b)(1+ε∫[h∣a]p0(a∣c,b)a) =p_0(c,b) (1+ [h a]p_0(a c,b)da ) ⇒pε(d∣c,b) p_ (d c,b) =pε(d,c,b)pε(c,b)=p0(d∣c,b) = p_ (d,c,b)p_ (c,b)=p_0(d c,b) This shows pε(d∣c,b,a)=p0(d∣c,b,a)=p0(d∣c,b)=pε(d∣c,b)p_ (d c,b,a)=p_0(d c,b,a)=p_0(d c,b)=p_ (d c,b). Hence A⊆T_A . Second case B∣A⊆T_B A : Pick any f=[h∣b,a]−[h∣a]f=E[h b,a]-E[h a]. Then pε(c,b,a) p_ (c,b,a) =p0(c,b,a)(1+ε[h∣b,a]−ε[h∣a]) =p_0(c,b,a)(1+ [h b,a]- [h a]) (48) pε(b,a) p_ (b,a) =p0(b,a)(1+ε[h∣b,a]−ε[h∣a]) =p_0(b,a)(1+ [h b,a]- [h a]) pε(c,a) p_ (c,a) =p0(c,a)∫p0(b∣c,a)(1+ε[h∣b,a]−ε[h∣a])b =p_0(c,a) p_0(b c,a)(1+ [h b,a]- [h a])db =p0(c,a)∫p0(b∣a)(1+ε[h∣b,a]−ε[h∣a])b =p_0(c,a) p_0(b a)(1+ [h b,a]- [h a])db (C⟂B∣A) (C \!\!\! A) =p0(c,a) =p_0(c,a) Therefore, we get pε(c∣b,a)=p0(c∣b,a)=p0(c∣a)=pε(c∣a)p_ (c b,a)=p_0(c b,a)=p_0(c a)=p_ (c a). Next, pε(d,c,b) p_ (d,c,b) =p0(d,c,b)∫p0(a∣c,b)(1+ε[h∣b,a]−ε[h∣a])a =p_0(d,c,b) p_0(a c,b)(1+ [h b,a]- [h a])da (A⟂D∣C,B) (A \!\!\! C,B) (49) pε(c,b) p_ (c,b) =∫p0(c,b,a)(1+ε[h∣b,a]−ε[h∣a])a = p_0(c,b,a)(1+ [h b,a]- [h a])da =p0(c,b)∫p0(a∣c,b)(1+ε[h∣b,a]−ε[h∣a])a =p_0(c,b) p_0(a c,b)(1+ [h b,a]- [h a])da This shows pε(d∣c,b,a)=p0(d∣c,b,a)=p0(d∣c,b)=pε(d∣c,b)p_ (d c,b,a)=p_0(d c,b,a)=p_0(d c,b)=p_ (d c,b). Hence B|A⊆T_B|A . Third case C∣A⊆T_C A : Similar to B|A⊆T_B|A . Fourth case D∣CB⊆T_D CB : Pick f=[h∣d,c,b]−[h∣c,b]f=E[h d,c,b]-E[h c,b]. Then pε(c,b,a) p_ (c,b,a) =p0(c,b,a)∫p0(d∣c,b)(1+ε[h∣d,c,b]−ε[h∣c,b])d =p_0(c,b,a) p_0(d c,b)(1+ [h d,c,b]- [h c,b])d (D⟂A∣C,B) (D \!\!\! C,B) (50) =p0(c,b,a) =p_0(c,b,a) This show pε(c∣b,a)=p0(c∣b,a)=p0(c∣a)=pε(c∣a)p_ (c b,a)=p_0(c b,a)=p_0(c a)=p_ (c a). Moreover, pε(d∣c,b,a)=p0(d∣c,b,a)×∫fp0(d∣c,b)dp_ (d c,b,a)=p_0(d c,b,a)× fp_0(d c,b)d. Next pε(d,c,b) p_ (d,c,b) =p0(d,c,b)∫p0(a∣d,c,b)(1+ε[h∣d,c,b]−ε[h∣c,b])a =p_0(d,c,b) p_0(a d,c,b)(1+ [h d,c,b]- [h c,b])da (51) =p0(d,c,b)(1+ε[h∣d,c,b]−ε[h∣c,b]) =p_0(d,c,b)(1+ [h d,c,b]- [h c,b]) pε(c,b) p_ (c,b) =p0(c,b)∫p0(d∣c,b)(1+ε[h∣d,c,b]−ε[h∣c,b])d =p_0(c,b) p_0(d c,b)(1+ [h d,c,b]- [h c,b])d =p0(c,b) =p_0(c,b) This shows pε(d∣c,b)=p0(d∣c,b)×∫fp0(d∣c,b)dp_ (d c,b)=p_0(d c,b)× fp_0(d c,b)d. Therefore pε(d∣c,b,a)=pε(d∣c,b)p_ (d c,b,a)=p_ (d c,b). Hence D|CB⊆T_D|CB . The proof is established because A,B∣A,C∣A,D∣CB⊆T_A,T_B A,T_C A,T_D CB implies A⊕B∣A⊕C∣A⊕D∣CB⊆T_A _B A _C A _D CB , and the equality follows. ∎ We reproduce a special case of the well-known result about DAG model’s tangent space and its orthogonal complement [Tsiatis.2006.SemiparametricTheory, Bhattacharya.Nabi.ea.2022.SemiparametricInference, Rotnitzky.Smucler.2020.EfficientAdjustment] in the following lemma. Lemma 1. Suppose P is a single independence DAG model over variables V, defined by a single constraint ⟂∣X \!\!\! . The model’s tangent space T and its orthogonal complement ⟂T are, respectively = = h−[h∣,,]+[h∣,]−[h∣] \h-E[h ,y,z]+E[h ,z]-E[h ] +[h∣,]:∀h∈ℋ, +E[h ,z]:∀ h \, (11) ⟂= = [h∣,,]−[h∣,] \E[h ,y,z]-E[h ,z] −[h∣,]+[h∣]:∀h∈ℋ. -E[h ,z]+E[h ]:∀ h \. Proof. Let =∖(∪˙∪˙)W=V (X ∪Y ∪Z). For any parametric submodel pε()\p_ (v)\ of P, its elements factorize as pε()=pε(∣,,)pε(∣)pε(∣)pε()p_ (v)=p_ (w ,y,z)p_ (x )p_ (y )p_ (z) (52) The total score function is pε()=∂εpε()|ε=0p_ (v)= ∂ p_ (v)|_ =0. By Lemma 5, s(∣,,) s(w ,y,z) :=∂εpε(∣,,)|ε=0=s()−[s()∣,,] := ∂ p_ (w ,y,z)|_ =0=s(V)- *E[s(V) ,y,z] (53) s(∣) s(x ) :=∂εpε(∣)|ε=0=[s()∣,]−[s()∣] := ∂ p_ (x )|_ =0= *E[s(V) ,z]- *E[s(V) ] (54) s(∣) s(y ) :=∂εpε(∣)|ε=0=[s()∣,]−[s()∣] := ∂ p_ (y )|_ =0= *E[s(V) ,z]- *E[s(V) ] (55) s() s(z) :=∂εpε()|ε=0=[s()∣] := ∂ p_ (z)|_ =0= *E[s(V) ] (56) Define the following subspaces ∣ _W :=h−[h∣,,]:h∈ℋ := \h- *E[h ,y,z]:h \ (57) ∣ _X :=[h∣,]−[h∣]:h∈ℋ := \ *E[h ,z]- *E[h ]:h \ (58) ∣ _Y :=[h∣,]−[h∣]:h∈ℋ := \ *E[h ,z]- *E[h ]:h \ (59) _Z :=[h∣]:h∈ℋ := \ *E[h ]:h \ (60) It is evident that s(∣,,)∈∣s(w ,y,z) _W , s(∣)∈∣s(x ) _X , s(∣)∈∣s(y ) _Y and s()∈s(z) _Z. By Lemma 6, these subspaces are orthogonal. Moreover, due to Equation 52 s()=s(∣,,)+s(∣)+s(∣)+s()s(v)=s(w ,y,z)+s(x )+s(y )+s(z) (61) Then s()∈∣⊕∣⊕∣⊕s(v) _W _X _Y _Z. This means the tangent space of this model is ⊆∣⊕∣⊕∣⊕T _W _X _Y _Z. Next, we will show that each of the subspace is in T. To do this, we must show that any function in these subspaces can be used to construct parametric submodel of P. Pick any f=h−[h∣,,]∈∣f=h- *E[h ,y,z] _W and construct a parametric model by pε()=p0()(1+εf())p_ (v)=p_0(v)(1+ f(v)). Since [f]=0E[f]=0, pεp_ sums to 1. Moreover, the parameter ε∈(−δ,δ) ∈(-δ,δ) and δ is small enough so that pε()≥0p_ (v)≥ 0 at all v. Then pεp_ is a valid probability distribution. This parametric model contains p0p_0 at ε=0 =0, and the score function at p0p_0 is ∂εlogpε()|ε=0=f ∂ p_ (v)|_ =0=f. What left to show is that the constraint holds for all pεp_ . Indeed, for any pεp_ , pε(,,)=p0(,,)(1+ε∫f(,,,)p0(∣,,))=p0(,,),p_ (x,y,z)=p_0(x,y,z) (1+ f(x,y,z,w)p_0(w ,y,z)dw )=p_0(x,y,z), (62) since [f∣,,]=0 *E[f ,y,z]=0. Therefore the constraint ⟂∣X \!\!\! holds for pεp_ , pε(∣,)=pε(∣)p_ (x ,z)=p_ (x ), because it holds for p0p_0. This means pε():ε∈(−δ,δ)\p_ (v): ∈(-δ,δ)\ is also a parametric submodel of P, hence f∈f and therefore ∣⊆T_W . Pick any f=[h∣,]−[h∣]∈∣f= *E[h ,z]- *E[h ] _X and construct a parametric model by pε()=p0()(1+εf())p_ (v)=p_0(v)(1+ f(v)). By the same reasoning, all we need to show is the constraint holds for all pε()p_ (v). pε(,,) p_ (x,y,z) =p0(,,)(1+ε[f∣,,])=p0(,,)(1+ε[h∣,]−ε[h∣]) =p_0(x,y,z)(1+ *E[f ,y,z])=p_0(x,y,z)(1+ *E[h ,z]- *E[h ]) (63) pε(,) p_ (y,z) =p0(,)(1+ε[f∣,])=p0(,)(1+ε∫[h∣,]p0(∣,)−ε[h∣]) =p_0(y,z)(1+ *E[f ,z])=p_0(y,z) (1+ *E[h ,z]p_0(x ,z)dx- *E[h ] ) =p0(,)(1+ε∫[h∣,]p0(∣)−ε[h∣])(⟂∣ at p0) =p_0(y,z) (1+ *E[h ,z]p_0(x )dx- *E[h ] ) (X \!\!\! at p_0) =p0(,) =p_0(y,z) pε(,) p_ (x,z) =p0(,)(1+ε[f∣,])=p0(,)(1+ε[h∣,]−ε[h∣]) =p_0(x,z)(1+ *E[f ,z])=p_0(x,z)(1+ *E[h ,z]- *E[h ]) pε() p_ (z) =p0()(1+ε[f∣])=p0()(since [f∣]=0) =p_0(z)(1+ *E[f ])=p_0(z) (since *E[f ]=0) Therefore, pε(∣,) p_ (x ,z) :=pε(,,)pε(,)=p0(∣,)(1+ε[h∣,]−ε[h∣]) = p_ (x,y,z)p_ (y,z)=p_0(x ,z)(1+ *E[h ,z]- *E[h ]) (64) =p0(∣)(1+ε[h∣,]−ε[h∣])=pε(,)pε()=pε(∣) =p_0(x )(1+ *E[h ,z]- *E[h ])= p_ (x,z)p_ (z)=p_ (x ) Hence, the constraint holds for all pεp_ , so pε\p_ \ is a parametric submodel of P, and therefore f∈∣⊆f _X . Pick any f=[h∣,]−[h∣]∈∣f= *E[h ,z]- *E[h ] _X and construct a parametric model by pε()=p0()(1+εf())p_ (v)=p_0(v)(1+ f(v)). By the same reasoning, all we need to show is the constraint holds for all pε()p_ (v). pε(,,) p_ (x,y,z) =p0(,,)(1+ε[f∣,,])=p0(,,)(1+ε[h∣,]−ε[h∣]) =p_0(x,y,z)(1+ *E[f ,y,z])=p_0(x,y,z)(1+ *E[h ,z]- *E[h ]) (65) pε(,) p_ (y,z) =p0(,)(1+ε[f∣,])=p0(,)(1+ε[h∣,]−ε[h∣]) =p_0(y,z)(1+ *E[f ,z])=p_0(y,z)(1+ *E[h ,z]- *E[h ]) pε(,) p_ (x,z) =p0(,)(1+ε[f∣,])=p0(,)(1+ε∫[h∣,]p0(∣,)−ε[h∣]) =p_0(x,z)(1+ *E[f ,z])=p_0(x,z) (1+ *E[h ,z]p_0(y ,z)dy- *E[h ] ) =p0(,)(1+ε∫[h∣,]p0(∣)−ε[h∣])(⟂∣ at p0) =p_0(x,z) (1+ *E[h ,z]p_0(y )dx- *E[h ] ) (Y \!\!\! at p_0) =p0(,) =p_0(x,z) pε() p_ (z) =p0()(1+ε[f∣])=p0()(since [f∣]=0) =p_0(z)(1+ *E[f ])=p_0(z) (since *E[f ]=0) Therefore, pε(∣,) p_ (x ,z) :=pε(,,)pε(,)=p0(∣,)=p0(∣)=pε(∣) = p_ (x,y,z)p_ (y,z)=p_0(x ,z)=p_0(x )=p_ (x ) (66) Hence, the constraint holds for all pεp_ , so pε\p_ \ is a parametric submodel of P, and therefore f∈∣⊆f _Y . Pick any f=[h∣]∈f= *E[h ] _Z and construct a parametric model by pε()=p0()(1+εf())p_ (v)=p_0(v)(1+ f(v)). Then pε(,,) p_ (x,y,z) =p0(,,)(1+ε[f∣,,])=p0(,,)(1+ε[h∣]) =p_0(x,y,z)(1+ *E[f ,y,z])=p_0(x,y,z)(1+ *E[h ]) (67) pε(,) p_ (y,z) =p0(,)(1+ε[f∣,])=p0(,)(1+ε[h∣]) =p_0(y,z)(1+ *E[f ,z])=p_0(y,z)(1+ *E[h ]) pε(,) p_ (x,z) =p0(,)(1+ε[f∣,])=p0(,)(1+ε[h∣]) =p_0(x,z)(1+ *E[f ,z])=p_0(x,z)(1+ *E[h ]) pε() p_ (z) =p0()(1+ε[f∣])==p0()(1+ε[h∣]) =p_0(z)(1+ *E[f ])==p_0(z)(1+ *E[h ]) Therefore, pε(∣,) p_ (x ,z) :=pε(,,)pε(,)=p0(∣,)=p0(∣)=pε(∣) = p_ (x,y,z)p_ (y,z)=p_0(x ,z)=p_0(x )=p_ (x ) (68) Hence, the constraint holds for all pεp_ , so pε\p_ \ is a parametric submodel of P, and therefore f∈⊆f _Z . These results imply ∣⊕∣⊕∣⊕⊆T_W _X _Y _Z , and with previous result, we get =∣⊕∣⊕∣⊕T=T_W _X _Y _Z. Define the following operators Π(h∣) (h _W ) =h−[h∣,,] =h- *E[h ,y,z] (69) Π(h∣) (h _X ) =[h∣,]−[h∣] = *E[h ,z]- *E[h ] (70) Π(h∣) (h _Y ) =[h∣,]−[h∣] = *E[h ,z]- *E[h ] (71) Π(h∣) (h _Z) =[h∣] = *E[h ] (72) By Lemma 7, Π(⋅∣),Π(⋅∣),Π(⋅∣),Π(⋅∣) (· _W ), (· _X ), (· _Y ), (· _Z) are the orthogonal projection operators onto ∣,∣,∣,T_W ,T_X ,T_Y ,T_Z, respectively. Define the following operator Π(h) (h) =Π(h∣)+Π(h∣)+Π(h∣)+Π(h∣) = (h _W )+ (h _X )+ (h _Y )+ (h _Z) (73) =h−[h∣,,]+[h∣,]−[h∣]+[h∣,] =h- *E[h ,y,z]+ *E[h ,z]- *E[h ]+ *E[h ,z] We will show that Π is the projection operator onto the tangent space T. First, we will show =Π(h):∀h∈ℋT=\ (h):∀ h \. For the first direction, Π(h)∈ (h) for all h∈ℋh , obviously. For the other direction, since =∣⊕∣⊕∣⊕T=T_W _X _Y _Z, any f∈f can be written as, for some h1,h2,h3,h4∈ℋh_1,h_2,h_3,h_4 , f=h1−[h1∣,,]+[h2∣,]−[h2∣]+[h3∣,]−[h3∣]+[h4∣].f=h_1- *E[h_1 ,y,z]+ *E[h_2 ,z]- *E[h_2 ]+ *E[h_3 ,z]- *E[h_3 ]+ *E[h_4 ]. (74) Because the subspaces ∣,∣,∣,T_W ,T_X ,T_Y ,T_Z are orthogonal, Π(f∣) (f _W ) =h1−[h1∣,,] =h_1- *E[h_1 ,y,z] (75) Π(f∣) (f _X ) =[h2∣,]−[h2∣] = *E[h_2 ,z]- *E[h_2 ] Π(f∣) (f _Y ) =[h3∣,]−[h3∣] = *E[h_3 ,z]- *E[h_3 ] Π(f∣) (f _Z) =[h4∣]. = *E[h_4 ]. Therefore, f=Π(f)f= (f), hence ⊆Π(h):h∈ℋT \ (h):h \. This implies =Π(h):h∈ℋT=\ (h):h \. Second, we will show that f−Π(f)f- (f) and Π(g) (g) are orthogonal, for all f,g∈ℋf,g . [f−Π(f),Π(g)]=[fΠ(g)]−[Π(f)Π(g)] *E [f- (f), (g) ]= *E [f (g) ]- *E [ (f) (g) ] (76) =[([f∣,,]−[f∣,]+[f∣]−[f∣,])×(g−[g∣,,]+[g∣,]−[g∣]+[g∣,])] = *E [ ( *E[f ,Y,Z]- *E[f ,Z]+ *E[f ]- *E[f ,Z] )× (g- *E[g ,Y,Z]+ *E[g ,Z]- *E[g ]+ *E[g ,Z] ) ] =[([f∣,,]−[f∣,]+[f∣]−[f∣,])×([g∣,]−[g∣]+[g∣,])] = *E [ ( *E[f ,Y,Z]- *E[f ,Z]+ *E[f ]- *E[f ,Z] )× ( *E[g ,Z]- *E[g ]+ *E[g ,Z] ) ] =[[f∣,,][g∣,]]−[[f∣,][g∣,]]+[[f∣][g∣,]]−[[f∣,][g∣,]] = *E [ *E[f ,Y,Z] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ]+ *E [ *E[f ] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ] −[[f∣,,][g∣]]+[[f∣,][g∣]]−[[f∣][g∣]]+[[f∣,][g∣]] - *E [ *E[f ,Y,Z] *E[g ] ]+ *E [ *E[f ,Z] *E[g ] ]- *E [ *E[f ] *E[g ] ]+ *E [ *E[f ,Z] *E[g ] ] +[[f∣,,][g∣,]]−[[f∣,][g∣,]]+[[f∣][g∣,]]−[[f∣,][g∣,]] + *E [ *E[f ,Y,Z] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ]+ *E [ *E[f ] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ] =[[f∣,][g∣,]]−[[f∣,][g∣,]]+[[f∣][g∣]]−[[f∣,][g∣,]] = *E [ *E[f ,Z] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ]+ *E [ *E[f ] *E[g ] ]- *E [ *E[f ,Z] *E[g ,Z] ] −[[f∣][g∣]]+[[f∣][g∣]]−[[f∣][g∣]]+[[f∣][g∣]] - *E [ *E[f ] *E[g ] ]+ *E [ *E[f ] *E[g ] ]- *E [ *E[f ] *E[g ] ]+ *E [ *E[f ] *E[g ] ] +[[f∣,][g∣,]]−[[f∣,][g∣,]]+[[f∣][g∣]]−[[f∣,][g∣,]] + *E [ *E[f ,Z] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ]+ *E [ *E[f ] *E[g ] ]- *E [ *E[f ,Z] *E[g ,Z] ] =−[[f∣,][g∣,]]−[[f∣,][g∣,]]+2[[f∣][g∣]] =- *E [ *E[f ,Z] *E[g ,Z] ]- *E [ *E[f ,Z] *E[g ,Z] ]+2 *E [ *E[f ] *E[g ] ] Due to the constraint [[f∣,][g∣,]] *E [ *E[f ,Z] *E[g ,Z] ] =∫[f∣,][g∣,]p0(,,) = *E[f ,z] *E[g ,z]p_0(x,y,z)dxdydz (77) =∫[f∣,][g∣,]p0(∣)p0(∣)p0() = *E[f ,z] *E[g ,z]p_0(x )p_0(y )p_0(z)dxdydz =∫[f∣][g∣]p0() = *E[f ] *E[g ]p_0(z)dz =[[f∣][g∣]] = *E [ *E[f ] *E[g ] ] Similarly, [[f∣,][g∣,]]=[[f∣][g∣]] *E [ *E[f ,Z] *E[g ,Z] ]= *E [ *E[f ] *E[g ] ]. Therefore, the inner product vanishes, [f−Π(f),Π(g)]=0 *E [f- (f), (g) ]=0. In conclusion, the operator defined by Π(h∣)=h−[h∣,,]+[h∣,]−[h∣]+[h∣,] (h )=h- *E[h ,y,z]+ *E[h ,z]- *E[h ]+ *E[h ,z] (78) is the projection operator onto the tangent space of this model, hence =Π(h∣):∀h∈ℋT=\ (h ):∀ h \. The operator onto its orthogonal complement is Π(h∣⟂) (h ) =h−Π(h∣) =h- (h ) (79) =[h∣,,]−[h∣,]+[h∣]−[h∣,] = *E[h ,y,z]- *E[h ,z]+ *E[h ]- *E[h ,z] and ⟂=Π(h∣⟂):∀h∈ℋT =\ (h ):∀ h \. ∎ Appendix C Proofs in Section 4 Theorem 3. Suppose P is a general semi-parametric Markov model defined by K≥1K≥ 1 conditional independence constraints i⟂i∣iX_i \!\!\! _i _i, where i=1,…,Ki=1,…,K. For each i, let i P_i be the single independence DAG model defined by the i-th constraint. The orthogonal projection operator Π(⋅∣i⟂):ℋ↦ℋ (· _i ):H sends each function h∈ℋh to the orthogonal complement i⟂T_i of the tangent space iT_i of i P_i, Π(h∣i⟂)(i,i,i)=[h|i,i,i]−[h|i,i]−[h|i,i]+[h|i]. (h _i )(x_i,y_i,z_i)=E[h|x_i,y_i,z_i]-E[h|x_i,z_i]-E[h|y_i,z_i]+E[h|z_i]. (80) The orthogonal complement of the tangent space of P is the direct sum of these orthogonal complements, ⟂=∑i=1KΠ(hi∣i⟂):∀h1,…,hK∈ℋ. = \ _i=1^K (h_i _i )\>:\>∀ h_1,…,h_K \. (81) Proof. First, we will show =1∩…∩KT=T_1∩… _K. Let pε():ε∈(−δ,δ)\p_ (v): ∈(-δ,δ)\ be any parametric submodel of P, with total score function s()s(v) at p0p_0. For each i, pε()p_ (v) satisfies the constraint i⟂i∣iX_i \!\!\! _i _i. Therefore, pε():ε∈(−δ,δ)\p_ (v): ∈(-δ,δ)\ is also a parametric submodel of i P_i. Following the proof of Lemma 1, we have s∈is _i, which is the tangent space of i P_i. This is true for all i=1,…,Ki=1,…,K, so s∈1∩…∩K¯s∈ T_1∩… _K and therefore ⊆1∩…∩K¯T T_1∩… _K. For the other direction, pick any f∈1∩…∩Kf _1∩… _K. Define a parametric model by pε()=p0()(1+εf())p_ (v)=p_0(v)(1+ f(v)). Since [f]=0E[f]=0, pεp_ sums to 1. Moreover, the parameter ε∈(−δ,δ) ∈(-δ,δ) and δ is small enough so that pε()≥0p_ (v)≥ 0 at all v. Then pεp_ is a valid probability distribution. This parametric model contains p0p_0 at ε=0 =0, and the score function at p0p_0 is ∂εlogpε()|ε=0=f ∂ p_ (v)|_ =0=f. What left to show is that all K constraints hold for all pε()p_ (v). Indeed, for each i, f∈if _i. Following the proof of Lemma 1, the i-th constraint i⟂i∣iX_i \!\!\! _i _i must hold in pε()p_ (v). Therefore, 1∩…∩K⊆T_1∩… _K . By definition, T is closed, so 1∩…∩K¯⊆ T_1∩… _K . We have shown that =1∩…∩K¯T= T_1∩… _K. Since 1,…,KT_1,…,T_K are closed subspaces, =1∩…∩KT=T_1∩… _K. Next, we will show ⟂=1⟂⊕…⊕K⟂T =T_1 … _K . We have =1∩…∩KT=T_1∩… _K is orthogonal to all i⟂i=1K\T_i \_i=1^K, so T is orthogonal to the span 1⟂⊕…⊕K⟂¯ T_1 … _K . Since 1,…,KT_1,…,T_K are closed, this span is exactly the direct sum 1⟂⊕…⊕K⟂T_1 … _K . Therefore, 1⟂⊕…⊕K⟂⊆⟂T_1 … _K . For the other direction, every f∈(1⟂⊕…⊕K⟂)⟂f∈ (T_1 … _K ) must be orthogonal to all i⟂i=1K\T_i \_i=1^K. For each i, the fact that f is orthogonal to i⟂T_i implies f∈if _i. Therefore, f∈1∩…∩K=f _1∩… _K=T. Hence, (1⟂⊕…⊕K⟂)⟂⊆ (T_1 … _K ) , and 1⟂⊕…⊕K⟂⊇⟂T_1 … _K . Finally, by Lemma 1, the i-th subspace i⟂=Π(hi∣i⟂):∀h∈ℋT_i =\ (h_i _i ):∀ h \. Then by definition of direct sum, the tangent space T is the subspace of all functions of the form ∑i=1KΠ(hi∣i⟂) _i=1^K (h_i _i ) for arbitrary h1,…,hK∈ℋh_1,…,h_K . ∎ Example 1. The Bell scenario model associated with Figure 1d has two constraints: A⟂B,DA \!\!\! ,D and B⟂A,CB \!\!\! ,C. The operators correspond to these constraints are Π(⋅∣1⟂) (· _1 ) =[⋅|a,b,d]−[⋅|a]−[⋅|b,d] =E[·|a,b,d]-E[·|a]-E[·|b,d] (82) Π(⋅∣2⟂) (· _2 ) =[⋅|b,a,c]−[⋅|b]−[⋅|a,c]. =E[·|b,a,c]-E[·|b]-E[·|a,c]. Let Γ(⋅)=Π(⋅∣1⟂)+Π(⋅∣2⟂) (·)= (· _1 )+ (· _2 ). If Γ is a projection operator, then h−Γ(h)h- (h) and Γ(g) (g) must be orthogonal, for any h,g∈ℋh,g . We will compute their inner product. It is not evident that their inner product vanishes. [(h−Γ(h))Γ(g)]=[hΓ(g)]−[Γ(h)Γ(g)] *E [(h- (h)) (g) ]= *E [h (g) ]- *E [ (h) (g) ] For the first term term1 1 =[h×([g|A,B,D]−[g|A]−[g|B,D]+[g|B,A,C]−[g|B]−[g|A,C])] = *E [h× (E[g|A,B,D]-E[g|A]-E[g|B,D]+E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] =[[h|A,B,D]⋅[g|A,B,D]]+[[h|A]⋅[g|A]]−[[h|B,D]⋅[g|B,D]] = *E [E[h|A,B,D]·E[g|A,B,D] ]+ *E [E[h|A]·E[g|A] ]- *E [E[h|B,D]·E[g|B,D] ] +[[h|B,A,C]⋅[g|B,A,C]]−[[h|B]⋅[g|B]]−[[h|A,C]⋅[g|A,C]] + *E [E[h|B,A,C]·E[g|B,A,C] ]- *E [E[h|B]·E[g|B] ]- *E [E[h|A,C]·E[g|A,C] ] For the second term term2=[([h|A,B,D]−[h|A]−[h|B,D]+[h|B,A,C]−[h|B]−[h|A,C]) 2= *E [ (E[h|A,B,D]-E[h|A]-E[h|B,D]+E[h|B,A,C]-E[h|B]-E[h|A,C] ) ×([g|A,B,D]−[g|A]−[g|B,D]+[g|B,A,C]−[g|B]−[g|A,C])] × (E[g|A,B,D]-E[g|A]-E[g|B,D]+E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] =[[h|A,B,D]⋅[g|A,B,D]]+[[h|A]⋅[g|A]]−[[h|B,D]⋅[g|B,D]] = *E [E[h|A,B,D]·E[g|A,B,D] ]+ *E [E[h|A]·E[g|A] ]- *E [E[h|B,D]·E[g|B,D] ] (term2a) (term2a) +[[h|A,B,D]([g|B,A,C]−[g|B]−[g|A,C])] + *E [E[h|A,B,D] (E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] (term2b) (term2b) −[[h|A]([g|A,B,D]−[g|A]−[g|B,D])]−[[h|A]([g|B,A,C]−[g|B]−[g|A,C])] - *E [E[h|A] (E[g|A,B,D]-E[g|A]-E[g|B,D] ) ]- *E [E[h|A] (E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] (term2c) (term2c) −[[h|B,D]([g|A,B,D]−[g|A]−[g|B,D])]−[[h|B,D]([g|B,A,C]−[g|B]−[g|A,C])] - *E [E[h|B,D] (E[g|A,B,D]-E[g|A]-E[g|B,D] ) ]- *E [E[h|B,D] (E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] (term2d) (term2d) +[[h|B,A,C]⋅[g|B,A,C]]−[[h|B]⋅[g|B]]−[[h|A,C]⋅[g|A,C]] + *E [E[h|B,A,C]·E[g|B,A,C] ]- *E [E[h|B]·E[g|B] ]- *E [E[h|A,C]·E[g|A,C] ] (term2e) (term2e) +[[h|B,A,C]([g|A,B,D]−[g|A]−[g|B,D])] + *E [E[h|B,A,C] (E[g|A,B,D]-E[g|A]-E[g|B,D] ) ] (term2f) (term2f) −[[h|B]([g|A,B,D]−[g|A]−[g|B,D])]−[[h|B]([g|B,A,C]−[g|B]−[g|A,C])] - *E [E[h|B] (E[g|A,B,D]-E[g|A]-E[g|B,D] ) ]- *E [E[h|B] (E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] (term2g) (term2g) −[[h|A,C]([g|A,B,D]−[g|A]−[g|B,D])]−[[h|A,C]([g|B,A,C]−[g|B]−[g|A,C])] - *E [E[h|A,C] (E[g|A,B,D]-E[g|A]-E[g|B,D] ) ]- *E [E[h|A,C] (E[g|B,A,C]-E[g|B]-E[g|A,C] ) ] (term2h) (term2h) Note that term1=term2a+term2eterm1=term2a+term2e. For the others, term2b 2b =[[h|A,B,D][g|B,A,C]]−[[h|B][g|B]]−[[h|A,B,D][g|A,C]] = *E [E[h|A,B,D]E[g|B,A,C] ]- *E [E[h|B]E[g|B] ]- *E [E[h|A,B,D]E[g|A,C] ] term2c 2c =−[[h|A][g|A]]+[[h|A][g|A]]+[[h|A][g|B,D]] =- *E [E[h|A]E[g|A] ]+ *E [E[h|A]E[g|A] ]+ *E [E[h|A]E[g|B,D] ] −[[h|A][g|A]]+[[h|A][g|B]]+[[h|A][g|A]] - *E [E[h|A]E[g|A] ]+ *E [E[h|A]E[g|B] ]+ *E [E[h|A]E[g|A] ] term2d 2d =−[[h|B,D][g|B,D]]+[[h|B,D][g|A]]+[[h|B,D][g|B,D]] =- *E [E[h|B,D]E[g|B,D] ]+ *E [E[h|B,D]E[g|A] ]+ *E [E[h|B,D]E[g|B,D] ] −[[h|B,D][g|B,A,C]]+[[h|B][g|B]]+[[h|B,D][g|A,C]] - *E [E[h|B,D]E[g|B,A,C] ]+ *E [E[h|B]E[g|B] ]+ *E [E[h|B,D]E[g|A,C] ] and term2f 2f =[[h|B,A,C][g|A,B,D]]−[[h|A][g|A]]−[[h|B,A,C][g|B,D]] = *E [E[h|B,A,C]E[g|A,B,D] ]- *E [E[h|A]E[g|A] ]- *E [E[h|B,A,C]E[g|B,D] ] term2g 2g =−[[h|B][g|B]]+[[h|B][g|A]]+[[h|B][g|B]] =- *E [E[h|B]E[g|B] ]+ *E [E[h|B]E[g|A] ]+ *E [E[h|B]E[g|B] ] −[[h|B][g|B]]+[[h|B][g|B]]+[[h|B][g|A,C]] - *E [E[h|B]E[g|B] ]+ *E [E[h|B]E[g|B] ]+ *E [E[h|B]E[g|A,C] ] term2h 2h =−[[h|A,C][g|A,B,D]]+[[h|A][g|A]]+[[h|A,C][g|B,D]] =- *E [E[h|A,C]E[g|A,B,D] ]+ *E [E[h|A]E[g|A] ]+ *E [E[h|A,C]E[g|B,D] ] −[[h|A,C][g|A,C]]+[[h|A,C][g|B]]+[[h|A,C][g|A,C]] - *E [E[h|A,C]E[g|A,C] ]+ *E [E[h|A,C]E[g|B] ]+ *E [E[h|A,C]E[g|A,C] ] Then term2 2 =term1 =term1 +[[h|A,B,D][g|B,A,C]]−[[h|B][g|B]]−[[h|A,B,D][g|A,C]] + *E [E[h|A,B,D]E[g|B,A,C] ]- *E [E[h|B]E[g|B] ]- *E [E[h|A,B,D]E[g|A,C] ] +[[h|A][g|B,D]]+[[h|A][g|B]] + *E [E[h|A]E[g|B,D] ]+ *E [E[h|A]E[g|B] ] +[[h|B,D][g|A]]−[[h|B,D][g|B,A,C]]+[[h|B][g|B]]+[[h|B,D][g|A,C]] + *E [E[h|B,D]E[g|A] ]- *E [E[h|B,D]E[g|B,A,C] ]+ *E [E[h|B]E[g|B] ]+ *E [E[h|B,D]E[g|A,C] ] +[[h|B,A,C][g|A,B,D]]−[[h|A][g|A]]−[[h|B,A,C][g|B,D]] + *E [E[h|B,A,C]E[g|A,B,D] ]- *E [E[h|A]E[g|A] ]- *E [E[h|B,A,C]E[g|B,D] ] +[[h|B][g|A]]+[[h|B][g|A,C]] + *E [E[h|B]E[g|A] ]+ *E [E[h|B]E[g|A,C] ] −[[h|A,C][g|A,B,D]]+[[h|A][g|A]]+[[h|A,C][g|B,D]]+[[h|A,C][g|B]] - *E [E[h|A,C]E[g|A,B,D] ]+ *E [E[h|A]E[g|A] ]+ *E [E[h|A,C]E[g|B,D] ]+ *E [E[h|A,C]E[g|B] ] It is not evident that term2=term1term2=term1. Appendix D Proofs in Section 5 Theorem 4. Consider the Markov model with K≥1K≥ 1 constraint in Theorem 3, and let φ0 _0 be an influence function of the target parameter ψ(p)ψ(p). For any sequence of integers i1,i2,…,iMi_1,i_2,…,i_M where each im∈1,…,Ki_m∈\1,…,K\ for all m∈1,…,Mm∈\1,…,M\, define φm:=φm−1−Π(φm−1∣im⟂),∀m∈1,…,M. _m= _m-1- ( _m-1 _i_m ), ∀ m∈\1,…,M\. (83) In words, φm _m is the projection of φm−1 _m-1 onto the tangent space imT_i_m of the DAG model corresponding to the imi_m-th constraint. Then φ0,…,φM _0,…, _M is a sequence of influence functions of ψ(p)ψ(p) with non-increasing variance, meaning E[φm2]≤E[φm−12]E[ _m^2]≤ E[ _m-1^2] for all m∈1,…,Mm∈\1,…,M\. Proof. The proof of this theorem follows from the discussions in Section 5.2 of the main paper and Section A of the Appendix. ∎ The iterative process described in Theorem 4 along with the explicit form of the projections Π(⋅∣im⟂) (· _i_m ) in Theorem 3 suggests a natural procedure for constructing estimators. For example, given an IF of the following form (obtained from one step of the iteration): φm _m :=φm−1−Π(φm−1∣im⟂) = _m-1- ( _m-1 _i_m ) (84) =φm−1−E[φm−1∣im,im,im]+E[φm−1∣im,im]+E[φm−1∣im,im]−E[φm−1∣im], = _m-1-E[ _m-1 _i_m,y_i_m,z_i_m]+E[ _m-1 _i_m,z_i_m]+E[ _m-1 _i_m,z_i_m]-E[ _m-1 _i_m], an estimator may be constructed by adopting additional nuisance models for the four additional expectation terms. For some specific IFs, we maybe able to simplify φm _m because some terms from φm−1 _m-1 and the expectations may cancel. Note that this is only a single step in the iterative process, where multiple operators are applied sequentially. If we want to use an M-th improvement of some initial IF φ0 _0, the expression will involve nested expectations, and may become challenging to evaluate in practice. However, in some cases, we may derive explicit estimators that demonstrate efficiency improvement, as we demonstrate via the following example. Example 2. Consider the Bell scenario model corresponding to Figure 1d, whose constraints are A⟂B,DA \!\!\! ,D and B⟂A,CB \!\!\! ,C. Suppose that we want to estimate the target parameter ψ(p)=[C∣a0]ψ(p)= *E[C a_0] instead of [D∣a0] *E[D a_0], as D and A are independent in this model. The influence function for this parameter in the saturated model is: φ0(a,c)=(a=a0)p(a)(c−E[C∣a]) _0(a,c)= I(a=a_0)p(a)(c-E[C a]) (85) Let 1,2 P_1, P_2 be the single independence DAG models corresponding to the constraints A⟂B,DA \!\!\! B,D and B⟂A,CB \!\!\! A,C, respectively. According to Theorem 3 Π(⋅∣1⟂) (· _1 ) =[⋅|a,b,d]−[⋅|a]−[⋅|b,d] =E[·|a,b,d]-E[·|a]-E[·|b,d] Π(⋅∣2⟂) (· _2 ) =[⋅|b,a,c]−[⋅|b]−[⋅|a,c]. =E[·|b,a,c]-E[·|b]-E[·|a,c]. We will derive the first and second improvements of φ0 _0, namely φ1 _1 and φ2 _2, respectively, using the iterative method in Theorem 4, and explicitly derive the corresponding estimators. First improvement: φ1=φ0−[φ0|a,b,d]+[φ0|a]+[φ0|b,d], _1= _0-E[ _0|a,b,d]+E[ _0|a]+E[ _0|b,d], (86) where [φ0∣a,b,d] *E[ _0 a,b,d] =(a=a0)p(a)([C∣a,b,d]−[C∣a]) = I(a=a_0)p(a)( *E[C a,b,d]- *E[C a]) (87) [φ0∣a] *E[ _0 a] =(a=a0)p(a)([C∣a]−[C∣a])=0 = I(a=a_0)p(a)( *E[C a]- *E[C a])=0 [φ0∣b,d] *E[ _0 b,d] =∫[φ0∣a′,b,d]p(a′∣b,d)a′ = *E[ _0 a ,b,d]p(a b,d)da =∫[φ0∣a′,b,d]p(a′)a′ = *E[ _0 a ,b,d]p(a )da (A⟂B,D) (A \!\!\! B,D) =∫((a′=a0)p(a′)([C∣a′,b,d]−[C∣a′]))p(a′)a′ = ( I(a =a_0)p(a )( *E[C a ,b,d]- *E[C a ]) )p(a )da =[C∣a0,b,d]−[C∣a0] = *E[C a_0,b,d]- *E[C a_0] Substituting these expressions, cancelling terms, and noting that ψ(p)=[C∣a0]ψ(p)= *E[C a_0], we get φ1(a,b,c,d) _1(a,b,c,d) =(a=a0)p(a)(c−[C∣a,b,d])+[C∣a0,b,d]−ψ(p) = I(a=a_0)p(a)(c- *E[C a,b,d])+ *E[C a_0,b,d]-ψ(p) (88) Let p^(a) p(a) be an estimate of the distribution p(a)p(a), and ^[C∣a,b,d] *E[C a,b,d] be an estimate of the conditional mean. The estimation equation method gives us an estimate ψ^1 ψ_1 of the target ψ(p)ψ(p) via ψ^1=1n∑i=1n((Ai=a0)p^(Ai)(Ci−^[C∣Ai,Bi,Di])+^[C∣a0,Bi,Di]). ψ_1= 1n _i=1^n ( I(A_i=a_0) p(A_i)(C_i- *E[C A_i,B_i,D_i])+ *E[C a_0,B_i,D_i] ). (89) Note that this estimator is the augmented inverse probability weighted (AIPW) estimator arising in causal inference and missing data problems. For example, A is the treatment, C is the outcome, and B,D\B,D\ is the observed confounder. Then the Bell scenario is the observed confounder scenario, with randomization on A, explaining this “coincidence”. Second improvement: φ2=φ1−E[φ1∣b,a,c]+E[φ1∣b]+E[φ1∣a,c], _2= _1-E[ _1 b,a,c]+E[ _1 b]+E[ _1 a,c], (90) where [φ1∣b] *E[ _1 b] =∫(a′=a0)p(a′)c′⋅p(a′,c′∣b)a′c′−∫(a′=a0)p(a′)(∫[C∣a′,b,d′]p(d′∣b,a′)d′)p(a′∣b)a′ = I(a =a_0)p(a )c · p(a ,c b)da dc - I(a =a_0)p(a ) ( *E[C a ,b,d ]p(d b,a )d )p(a b)da (91) +∫[C∣a0,b,d′]p(d′∣b)d′−ψ(p) + *E[C a_0,b,d ]p(d b)d -ψ(p) =[C∣a0]−∫[C∣a0,b,d′]p(d′∣b,a0)d′(A,C⟂B) = *E[C a_0]- *E[C a_0,b,d ]p(d b,a_0)d (A,C \!\!\! B) +∫[C∣a0,b,d′]p(d′∣b)d′−ψ(p) + *E[C a_0,b,d ]p(d b)d -ψ(p) =∫[C∣a0,b,d′]p(d′∣b)d′−∫[C∣a0,b,d′]p(d′∣b,a0)d′([C∣a0]=ψ(p)) = *E[C a_0,b,d ]p(d b)d - *E[C a_0,b,d ]p(d b,a_0)d ( *E[C a_0]=ψ(p)) =∫[C∣a0,b,d′]p(d′∣b,a0)d′−∫[C∣a0,b,d′]p(d′∣b,a0)d′(D⟂A∣B) = *E[C a_0,b,d ]p(d b,a_0)d - *E[C a_0,b,d ]p(d b,a_0)d (D \!\!\! B) =0. =0. [φ1∣b,a,c] *E[ _1 b,a,c] =(a=a0)p(a)(c−∫[C∣a,b,d′]p(d′∣b,a,c)d′) = I(a=a_0)p(a) (c- *E[C a,b,d ]p(d b,a,c)d ) (92) +∫[C∣a0,b,d′]p(d′∣b,a,c)d′−ψ(p) + *E[C a_0,b,d ]p(d b,a,c)d -ψ(p) [φ1∣a,c] *E[ _1 a,c] =(a=a0)p(a)(c−∫[C∣a,b′,d′]p(d′∣b′,a,c)p(b′)d′db′)(B⟂A,C) = I(a=a_0)p(a) (c- *E[C a,b ,d ]p(d b ,a,c)p(b )d db ) (B \!\!\! A,C) +∫[C∣a0,b′,d′]p(d′∣b′,a,c)p(b′)d′db′−ψ(p)(B⟂A,C). + *E[C a_0,b ,d ]p(d b ,a,c)p(b )d db -ψ(p) (B \!\!\! A,C). Then φ2(a,b,c,d) _2(a,b,c,d) =(a=a0)p(a)(c−[C∣a,b,d])+[C∣a0,b,d]−ψ(p) = I(a=a_0)p(a)(c- *E[C a,b,d])+ *E[C a_0,b,d]-ψ(p) (93) +(a=a0)p(a)(∫[C∣a,b,d′]p(d′∣b,a,c)d′−∫[C∣a,b,d′]p(d′∣b′,a,c)p(b′)d′b′) + I(a=a_0)p(a) ( *E[C a,b,d ]p(d b,a,c)d - *E[C a,b,d ]p(d b ,a,c)p(b )d db ) −∫[C∣a0,b,d′]p(d′∣b,a,c)d′+∫[C∣a0,b,d′]p(d′∣b′,a,c)p(b′)d′b′ - *E[C a_0,b,d ]p(d b,a,c)d + *E[C a_0,b,d ]p(d b ,a,c)p(b )d db We posit models for p^(a),^[C∣a,b,d],p^(d∣a,c,b) p(a), *E[C a,b,d], p(d a,c,b) and p^(b) p(b). Importantly, these models must be consistent, i.e., there is a valid set of probability distributions p(a,b,c,d)p(a,b,c,d) satisfying the constraints and yield these models. Enforcing consistency is not always obvious. If the models are consistent, then the estimation equation method gives us the estimator ψ^2 ψ_2 for the target ψ(p)ψ(p) ψ^2 ψ_2 =1n∑i=1n((Ai=a0)p^(Ai)(Ci−^[C∣Ai,Bi,Di])+^[C∣a0,Bi,Di]) = 1n _i=1^n ( I(A_i=a_0) p(A_i)(C_i- *E[C A_i,B_i,D_i])+ *E[C a_0,B_i,D_i] ) (94) +1n∑i=1n(Ai=a)p^(Ai)(∫^[C∣Ai,Bi,d′]p^(d′∣Bi,Ai,Ci)d′−∫^[C∣Ai,b′,d′]p^(d′∣b′,Ai,Ci)p^(b′)d′b′) + 1n _i=1^n I(A_i=a) p(A_i) ( *E[C A_i,B_i,d ] p(d B_i,A_i,C_i)d - *E[C A_i,b ,d ] p(d b ,A_i,C_i) p(b )d db ) +1n∑i=1n(−∫^[C∣a0,Bi,d′]p^(d′∣Bi,Ai,Ci)d′+∫^[C∣a0,b′,d′]p^(d′∣b′,Ai,Ci)p^(b′)d′b′). + 1n _i=1^n (- *E[C a_0,B_i,d ] p(d B_i,A_i,C_i)d + *E[C a_0,b ,d ] p(d b ,A_i,C_i) p(b )d db ).