Paper deep dive
The Origins of Representation Manifolds in Large Language Models
Alexander Modell, Patrick Rubin-Delanchy, Nick Whiteley
Models: GPT-2 Small, Mistral 7B, OpenAI text-embedding-large-3
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/12/2026, 6:48:00 PM
Summary
The paper proposes a 'continuous correspondence hypothesis' to explain how large language models represent complex, multidimensional features as manifolds rather than just atomic vectors. It argues that neural representations are homeomorphic to the intrinsic geometry of the concepts they encode, and demonstrates that cosine similarity in representation space effectively captures this intrinsic geometry through shortest paths on these manifolds.
Entities (5)
Relation Signals (3)
Cosine Similarity â encodes â Intrinsic Geometry
confidence 95% ¡ cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths
Representation Manifold â ishomeomorphicto â Metric Space
confidence 95% ¡ Proposition 1. Under Hypothesis 1, the map Ďf:ZfâMf is a homeomorphism.
Sparse Autoencoders â recovers â Features
confidence 90% ¡ This model underpins the use of sparse autoencoders to recover features from representations.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key element of this effort is the linear representation hypothesis, which posits that neural representations are sparse linear combinations of `almost-orthogonal' direction vectors, reflecting the presence or absence of different features. This model underpins the use of sparse autoencoders to recover features from representations. Moving towards a fuller model of features, in which neural representations could encode not just the presence but also a potentially continuous and multidimensional value for a feature, has been a subject of intense recent discourse. We describe why and how a feature might be represented as a manifold, demonstrating in particular that cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths, potentially answering the question of how distance in representation space and relatedness in concept space could be connected. The critical assumptions and predictions of the theory are validated on text embeddings and token activations of large language models.
Tags
Links
Trouble viewing inline? Open PDF directly â
Full Text
101,516 characters extracted from source content.
Expand or collapse full text
The Origins of Representation Manifolds in Large Language Models Alexander Modell Department of Mathematics Imperial College London a.modell@imperial.ac.uk &Patrick Rubin-Delanchy School of Mathematics University of Edinburgh prd@ed.ac.uk &Nick Whiteley School of Mathematics University of Bristol nick.whiteley@bristol.ac.uk Abstract There is a large ongoing scientific effort in mechanistic interpretability to map embeddings and internal representations of AI systems into human-understandable concepts. A key element of this effort is the linear representation hypothesis, which posits that neural representations are sparse linear combinations of âalmost-orthogonalâ direction vectors, reflecting the presence or absence of different features. This model underpins the use of sparse autoencoders to recover features from representations. Moving towards a fuller model of features, in which neural representations could encode not just the presence but also a potentially continuous and multidimensional value for a feature, has been a subject of intense recent discourse. We describe why and how a feature might be represented as a manifold, demonstrating in particular that cosine similarity in representation space may encode the intrinsic geometry of a feature through shortest, on-manifold paths, potentially answering the question of how distance in representation space and relatedness in concept space could be connected. The critical assumptions and predictions of the theory are validated on text embeddings and token activations of large language models. 1 Introduction There is a large, ongoing, scientific effort in mechanistic interpretability to map internal representations used by AI systems into human-understandable concepts (Lin, 2024; Templeton et al., 2024), with broad implications for humanity including safety, alignment, and the future role of AI in science (Bostrom, 2014; Soares and Fallenstein, 2017; Wang et al., 2023). A key element of this effort is the linear representation hypothesis (LRH), which posits that language models represent human-interpretable features as directions in representation space, and that model representations are (literally) a sparse linear combination of these directions (Smolensky, 1990; Arora et al., 2018; Elhage et al., 2022). The methodology of sparse autoencoders (SAEs) (Elhage et al., 2022; Bricken et al., 2023) employs ideas from sparse coding (Elad, 2010) to estimate a dictionary of these directions from representations. This model and methodology reflect a radical goal of breaking representations down into basic, irreducible, atomic concepts which are meaningfully only described as present or absent (Cunningham et al., 2023; Bricken et al., 2023; Templeton et al., 2024). Commonly cited examples are features such as floppy_ears, Eiffel_Tower, or is_Arabic, the presence of which it would presumably be useful for an algorithm to infer (corresponding e.g. to cat/dog classification, the topic of a question, the language of a query). It is generally accepted that this breakdown of representation space into purely atomic features does not tell the whole story (Smith, 2024; Mendel, 2024; Bussmann et al., 2024; Olah, 2024; Engels et al., 2025). There is overwhelming empirical evidence that neural networks represent complex features in structures which unfold across multiple directions in potentially continuous, nonlinear ways: examples of curves (Hanna et al., 2023; Chang et al., 2022), swiss-roll-like manifolds (Cai et al., 2021), loops (Engels et al., 2025; Gorton, 2024), tori (Chang et al., 2022), hierarchical trees (Park et al., 2024) in real language models; topologically circular representations of numbers in toy models trained to perform modular arithmetic (Liu et al., 2022; Nanda et al., 2023a; Zhong et al., 2023; He et al., 2024) or simulated angular data (Olah and Batson, 2024), fractal geometry in simulated hidden Markov models (Shai et al., 2024); and broader phenomenology from local finite-state-automata (Bricken et al., 2023), to spatial âbrain-likeâ modularity (Li et al., 2025), to behaviour, such as deception (Templeton et al., 2024). SAEs are not made defunct by these discoveries, and in fact have often facilitated them through recombination of SAE directions (Bussmann et al., 2025; Engels et al., 2025). The LRH has been extended to allow this more flexible interpretation of the output of SAEs: Definition 1 (Multidimensional linear representation hypothesis). There exists a collection of features labeled â fâ Ftypewriter_f â typewriter_F and associated subspaces VââDsubscriptsuperscriptâV_ f ^DVtypewriter_f â blackboard_RD such that the functional relationship between an input xâx â X and its representation Ψâ˘(x)Ψ (x)Ψ ( x ) is Ψâ˘(x)=âââ˘(x)Ďâ˘(x)â˘vâ˘(x),vâ˘(x)âV⢠and â˘âvâ˘(x)â2=1,formulae-sequenceΨsubscriptsubscriptsubscriptsubscriptsubscript and subscriptnormsubscript21 (x)= _ fâ F(x) _ f(x)v_ f(x)% , v_ f(x)â V_ f and \|v_ f(x)\|_% 2=1,Ψ ( x ) = âtypewriter_f â typewriter_F ( x ) Ďtypewriter_f ( x ) vtypewriter_f ( x ) , vtypewriter_f ( x ) â Vtypewriter_f and ⼠vtypewriter_f ( x ) âĽ2 = 1 , (1) where Ďâ˘(x)subscript _ f(x)Ďtypewriter_f ( x ) is a non-negative scaling denoting the presence of the feature ftypewriter_f in x, and â˘(x)=:Ďâ˘(x)>0conditional-setsubscript0 F(x)=\ f\>:\> _ f(x)>0\typewriter_F ( x ) = typewriter_f : Ďtypewriter_f ( x ) > 0 is the set of features which are present in x. The standard LRH corresponds to the case where vâ˘(x)subscriptv_ f(x)vtypewriter_f ( x ) is constant in x (and VsubscriptV_ fVtypewriter_f one-dimensional), and the extension above is a slightly relaxed and reparametrised version of that which appears in Engels et al. (2025). Our paper concerns the representation of a feature ftypewriter_f as a manifold in VsubscriptV_ fVtypewriter_f, a phenomenon which is widely observed and intensely deliberated in the mechanistic interpretability community (Olah and Batson, 2024; Olah, 2024; Gorton, 2024; Engels et al., 2025). Despite numerous accounts (cited above) of a manifold clearly corresponding to some underlying ground truth feature (which may even be known exactly, e.g. in simulated data), there is no general description of this correspondence. We provide what we believe is a minimum viable mathematical theory to do this. Our most substantial, novel result establishes that under plausible hypotheses, cosine similarity in representation space encodes the intrinsic geometry of a feature through shortest, on-manifold paths. We develop this insight using concepts from metric geometry â the theory of length and shape in metric spaces (Burago et al., 2001). The widespread use of cosine similarity across data science could suggest many other applications for this result. More generally, our work provides a (hopefully) accessible explanation of why manifold structure might arise in representation space, how its topology and geometry might reflect a human conceptualisation of the feature, and suggests some simple diagnostic plots and statistical checks to explore critical assumptions and predictions of the theory. Related to the problem of mechanistic interpretability, there is enormous interest in the use of text embeddings (Li et al., 2020; Gao et al., 2021), which several entities now provide as a service, to support, for example, receiver augmented generation, search, recommendation, visualisation and classification. Again, these are representations which are usually unit vectors, and about which the cosine similarity is said to provide an effective measure of semantic similarity. The bulk of our results are also applicable to this area; in what follows, Ψâ˘(x)Ψ (x)Ψ ( x ) can be viewed as a generic representation of some input x, and vâ˘(x)subscriptv_ f(x)vtypewriter_f ( x ) some unit-norm representation of a feature in x. 2 The continuous correspondence hypothesis 2.1 What is a feature anyway? Before we begin our discussion on how features are geometrically represented in language models, we really ought to pin down what exactly we mean by âa featureâ. Definition 2. A feature, labeled ftypewriter_f, is a metric space (,)subscriptsubscript(Z_ f, d_ f)( Ztypewriter_f , sansserif_dtypewriter_f ). A metric space is simply a set equipped with a distance, and we find that it provides a simple yet highly expressive formal mathematical framework for discussing the abstract notion of a feature or concept. In particular, it allows us to readily talk about: 1. Atomic features: subscriptZ_ fZtypewriter_f a singleton set. 2. Hierarchical features: subscriptZ_ fZtypewriter_f a discrete set, subscript d_ fsansserif_dtypewriter_f a tree distance. 3. Continuous features: subscriptZ_ fZtypewriter_f an interval (equipped with e.g. (x,y)=|xây|) d_ f(x,y)=|x-y|)sansserif_dtypewriter_f ( x , y ) = | x - y | ), a circle (equipped with e.g. arc-length distance), multi-dimensional (equipped with e.g. the Euclidean distance), etc. Figure 1: Representation manifolds in large language models: colours, years and dates. The first and third example show text embeddings obtained from OpenAIâs text-embedding-large-3 model from prompts relating to English names for colours and dates of the year, respectivly. The second example shows token activations from layer 7 of GPT2-small, which were studied in Engels et al. (2025). The token activations were processed via an SAE to extract a feature corresponding to years of the twentieth century as in Engels et al. (2025), and normalized to have norm one. For each example, we perform principal component analysis (PCA) to reduce the dimension to three and display the resulting point clouds from two perspectives. The embeddings of English names for colours are displayed in their respective colour value. Years are coloured from blue (1900) through green to yellow (1999), and dates are coloured from white (1st Janurary) through blue to black (1st July) through red and back to white. We find that this formalism strikes a balance between the less expressive Euclidean and hyperspherical models often assumed in the learning theory literature (e.g. Zimmermann et al., 2021; Hyvärinen et al., 2024; Reizinger et al., 2025), and the more complicated and less accessible models which are often assumed in the disentanglement literature, such as Riemannian manifolds equipped with group structure (e.g. Higgins et al., 2018; Pfau et al., 2020). For any input x on which the feature ftypewriter_f is present (i.e. for which Ďnâ˘(x)>0subscript0 _n(x)>0Ďitalic_n ( x ) > 0), we assume the existence of a value zâ˘(x)subscriptz_ f(x)ztypewriter_f ( x ) which the input takes in subscriptZ_ fZtypewriter_f. For example, if the feature colour is present in an input x, then Ďâ˘(x)>0subscript0 _ colour(x)>0Ďtypewriter_colour ( x ) > 0 and zâ˘(x)subscriptz_ f(x)ztypewriter_f ( x ) might take a value describing the precise hue, saturation and lightness of that colour. As a final note, we will assume throughout this paper that each subscriptZ_ fZtypewriter_f is a compact set or, loosely speaking, âclosed and boundedâ: a standard assumption in manifold learning which avoids considerable and possibly distracting theoretical complications. 2.2 The continuous correspondence hypothesis Given the multidimensional linear representation hypothesis, and our definition of a feature, perhaps the most basic hypothesis that one can make is that there is some way of matching the representation directions vâ˘(x)subscriptv_ f(x)vtypewriter_f ( x ) to the abstract features zâ˘(x)subscriptz_ f(x)ztypewriter_f ( x ). Hypothesis 1 (continuous correspondence). The features zâ˘(x)subscriptz_ f(x)ztypewriter_f ( x ) and representation directions vâ˘(x)subscriptv_ f(x)vtypewriter_f ( x ) are in a continuous, one-to-one correspondence. Formally, there is a continuous invertible map from the metric space into the hypersphere, Ď:âDâ1:subscriptitalic-Ďâsubscriptsuperscript1 _ f:Z_ f ^D-1Ďtypewriter_f : Ztypewriter_f â blackboard_SD - 1, with image âłâĎâ˘()âsubscriptâłitalic-ĎM_ f Ď(Z)Mtypewriter_f â Ď ( Z ), such that vâ˘(x)=Ďâ˘(zâ˘(x))subscriptsubscriptitalic-Ďsubscriptv_ f(x)= _ f(z_ f(x))vtypewriter_f ( x ) = Ďtypewriter_f ( ztypewriter_f ( x ) ) for all xâx â X. Our correspondence hypothesis, combined with the prior assumption that subscriptZ_ fZtypewriter_f is compact, has an immediate implication for the topological relationship between subscriptZ_ fZtypewriter_f and âłsubscriptâłM_ fMtypewriter_f. Proposition 1. Under Hypothesis 1, the map Ď:ââł:subscriptitalic-Ďâsubscriptsubscriptâł _ f:Z_ f _ fĎtypewriter_f : Ztypewriter_f â Mtypewriter_f is a homeomorphism.111This is simply a restatement of the well-established fact that a continuous invertible map over a compact domain has a continuous inverse (Sutherland, 2009, Proposition 13.26). Proposition 1 tells us that under the continuous correspondence hypothesis, we should expect the representations directions vâ˘(x)subscriptv_ f(x)vtypewriter_f ( x ) to live on a manifold âłsubscriptâłM_ fMtypewriter_f that is topologically identical to subscriptZ_ fZtypewriter_f. So, if subscriptZ_ fZtypewriter_f is an interval, âłsubscriptâłM_ fMtypewriter_f is a one-dimensional curve in âDsuperscriptâR^Dblackboard_RD. If subscriptZ_ fZtypewriter_f is a circle, then âłsubscriptâłM_ fMtypewriter_f is a loop. If subscriptZ_ fZtypewriter_f is a discrete set comprising m values, âłsubscriptâłM_ fMtypewriter_f is a discrete set comprising m points. More generally, a homeomorphism preserves connected components, holes, branching points, and more. Figure 2: Representation manifolds in token activations from layer 8 of Mistral 7B, processed via an SAE to extract representations of âmonths of the yearâ and âdays of the weekâ, as in Engels et al. (2025). We normalise the representations to have norm one, and perform PCA into three dimensions. The top-down view of the first two principal components, which was shown in Engels et al. (2025), obscures manifold structure which weaves through the third principal component. 2.3 Representations reflect the topology of features in LLMs Figure 1 gives an indiction of the plausibility of Hypothesis 1 in some examples. The first and third subfigures show text embeddings obtained from OpenAIâs text-embedding-large-3 model, with inputs corresponding to colours222These inputs are of are of the form âThe color of the object is ÂĄcolorÂż. What color is the object?â. Color names and hex-codes were obtained from the XKCD color survey (Munroe, 2010), from which we removed entries with low saturation (<0.4absent0.4<0.4< 0.4), high brightness (>0.8absent0.8>0.8> 0.8), and whose names did not obviously refer to a color, such as fruits and gemstones. Some additional outliers were removed., and dates of the year333These inputs are of the form â1st Januaryâ, â2nd Januaryâ, ⌠â31st Decemberâ. respectively. This model returns 3,072 dimensional unit-norm embeddings which we reduce to three dimensions using PCA. We show two perspectives of each plot. In both cases, we see that the embeddings are roughly arranged around a loop which, perhaps after some stretching and bending, could seem consistent with the abstract circular model we might have for such concepts, such as the âcolour wheelâ or the âyearly cycleâ. In particular, observe that the colours are arranged in the same order as the standard colour wheel of hue: red, purple, blue, green, yellow, orange, and back to red. The second subfigure shows token activations of years of the twentieth century in layer 7 of GPT2-small. This example is taken from Engels et al. (2025) who use a sparse autoencoder to attempt to disentangle the feature representations from the full superposed representation (see Section 5 of their paper for addition details of this procedure). We subselect only tokens corresponding to the years in question, normalize each activation vector to have norm one and perform PCA into three dimensions on the resulting vectors. One observes a clear one-dimensional curve which weaves and bends through the dimensions of the space, again reminiscent up to some geometric distortion of the standard human concept of a âtime lineâ. Given what we see, our innate understanding of these concepts, and Proposition 1, we might conjecture that, allowing for error of different kinds, the shapes are homeomorphic to the following metric spaces: colour: colour=[0,2â˘Ď)subscriptcolour02Z_ colour=[0,2Ď)Zcolour = [ 0 , 2 Ď ), colourâ˘(x,y)=minâĄ(|xây|,2â˘Ďâ|xây|)subscriptcolour2 d_ colour(x,y)= (|x-y|,2Ď-|x-y|)sansserif_dcolour ( x , y ) = min ( | x - y | , 2 Ď - | x - y | ), these angles corresponding to hues 0: red, âŚ, Ď/33Ď/3Ď / 3: blue, âŚ, 2â˘Ď/3232Ď/32 Ď / 3: yellow. years: years=[1900,1999]subscriptyears19001999Z_ years=[1900,1999]Zyears = [ 1900 , 1999 ], yearâ˘(x,y)=|xây|subscriptyear d_ year(x,y)=|x-y|sansserif_dyear ( x , y ) = | x - y | dates: dates=[0,365)subscriptdates0365Z_ dates=[0,365)Zdates = [ 0 , 365 ), datesâ˘(x,y)=minâĄ(|xây|,365â|xây|)subscriptdates365 d_ dates(x,y)= (|x-y|,365-|x-y|)sansserif_ddates ( x , y ) = min ( | x - y | , 365 - | x - y | ) In the case of years, a simple statistic to assess the conjecture of homeomorphism presents itself: the rank correlation between the years and their corresponding position along the manifold. We approximate position along the manifold using a K-nearest neighbour graph with K=1010K=10K = 10 (picked as small as possible subject to the graph being connected), and rank the points according to weighted graph distance from the (mean) representation of 1900. The Kendall and Spearman rank correlations are 0.97 and over 0.99, respectively, telling us that the representations occur in very close to true temporal order along the manifold. In these examples it would clearly not be reasonable to say the shapes resembled circles or straight lines without any sort of geometric distortion, and in the coming section we provide a mechanistic argument for the presence of this geometric distortion in neural networks. This effect could be missed in some earlier papers due to 2D projection. The left and middle of panels of Figure 1 of Engels et al. (2025) show circular arrangements of day-of-the-week and month representations, but these seem subject to significant geometric distortion once we view the data in 3D, as in Figure 2. 2.4 Manifold geometry and computation Our investigations (see Figure 1) and those of many others (see e.g. Ansuini et al., 2019; Cai et al., 2021; Chang et al., 2022; Hanna et al., 2023), have found not only that representations tend to live on low-dimensional manifolds, but that these manifolds curve and bend to occupy higher dimensional spaces. Why might it be advantageous for a language model to embed a concept in a larger dimension than its intrinsic topology seems to require? To answer this question, we shall briefly illustrate how the space of functions which can be computed as a linear projection of Ďâ˘(z)subscriptitalic-Ď _ f(z)Ďtypewriter_f ( z ) relates to its geometry. Since linear operations are crucial component in how one layer of a neural network maps to the next, it seems a sensible working hypothesis that they would arrange their representations as to maximize the expressivity of these linear computations. For the purpose of this discussion, consider the case that subscriptZ_ fZtypewriter_f is a unit interval =[0,1]subscript01Z_ f=[0,1]Ztypewriter_f = [ 0 , 1 ]. If one simply wanted to be able to read z from Ďâ˘(z)subscriptitalic-Ď _ f(z)Ďtypewriter_f ( z ) using a linear projection, then it is sufficient to represent subscriptZ_ fZtypewriter_f as an arc on Dâ1superscript1S^D-1blackboard_SD - 1. For example, to set Ďâ˘(z)=b0â˘(z)â˘v0+b1â˘(z)â˘v1subscriptitalic-Ďsubscript0subscript0subscript1subscript1 _ f(z)=b_0(z)v_0+b_1(z)v_1Ďtypewriter_f ( z ) = b0 ( z ) v0 + b1 ( z ) v1 where v0,v1âDâ1subscript0subscript1superscript1v_0,v_1 ^D-1v0 , v1 â blackboard_SD - 1 are orthogonal unit-vectors, b1â˘(z)âzproportional-tosubscript1b_1(z) zb1 ( z ) â z, and b0â˘(z)subscript0b_0(z)b0 ( z ) is a function which ensures that âĎâ˘(z)â2=1subscriptnormsubscriptitalic-Ď21\| _ f(z)\|_2=1⼠Ďtypewriter_f ( z ) âĽ2 = 1. In this way, the identity operation idâĄ(z):=zassignidid(z):=zid ( z ) := z can be computed via a linear projection idâĄ(z)âv1â Ďâ˘(z)proportional-toidâ subscript1subscriptitalic-Ďid(z) v_1¡ _ f(z)id ( z ) â v1 â Ďtypewriter_f ( z ). If instead, one wanted to be able to represent a richer class of functions of z by linear projections of Ďâ˘(z)subscriptitalic-Ď _ f(z)Ďtypewriter_f ( z ), say, polynomials of order p, then one could do this by setting Ďâ˘(z)=b0â˘(z)â˘v0+âŻâ˘bp+1â˘(z)â˘vp+1subscriptitalic-Ďsubscript0subscript0âŻsubscript1subscript1 _ f(z)=b_0(z)v_0+¡s b_p+1(z)v_p+1Ďtypewriter_f ( z ) = b0 ( z ) v0 + ⯠bitalic_p + 1 ( z ) vitalic_p + 1 with b1â˘(z)â1,b2â˘(z)âz,b3âz2,formulae-sequenceproportional-tosubscript11formulae-sequenceproportional-tosubscript2proportional-tosubscript3superscript2b_1(z) 1,b_2(z) z,b_3 z^2,b1 ( z ) â 1 , b2 ( z ) â z , b3 â z2 , etc⌠Such a map represents the interval [0,1]01[0,1][ 0 , 1 ] as a continuous path which weaves through a p+22p+2p + 2-dimensional subspace of Dâ1superscript1S^D-1blackboard_SD - 1. Superposition. Under the Linear Representation Hypothesis (1), the language model cannot access Ďâ˘(zâ˘(x))subscriptitalic-Ďsubscript _ f(z_ f(x))Ďtypewriter_f ( ztypewriter_f ( x ) ) directly, but must do so via Ψâ˘(x)Ψ (x)Ψ ( x ). There is a generally agreed upon explanation, known as the superposition hypothesis, for how an algorithm might nonetheless be granted approximate access to Ďâ˘(zâ˘(x))subscriptitalic-Ďsubscript _ f(z_ f(x))Ďtypewriter_f ( ztypewriter_f ( x ) ), with only limited interference from other Ďâ˛â˘(zâ˛â˘(x))subscriptitalic-Ďsuperscriptâ˛subscriptsuperscriptⲠ_ f (z_ f (x))Ďtypewriter_fⲠ( ztypewriter_fⲠ( x ) ): features occur only sparsely (i.e. Ďâ˘(x)=0subscript0 _ f(x)=0Ďtypewriter_f ( x ) = 0 for most â fâ Ftypewriter_f â typewriter_F), and are represented in almost-orthogonal subspaces (Elhage et al., 2021, 2022), a hypothesis which, in particular, would be consistent with the total number of features being substantially greater than the available representation dimensions444see, for example, Theorem 1 in the appendix of Engels et al. (2025).. If we assume (for a real-valued feature) that the identity idâĄ(z)idid(z)id ( z ) is among the collection of functions linearly readable from Ďsubscriptitalic-Ď _ fĎtypewriter_f, the superposition hypothesis also explains the efficacy of linear probes (Alain and Bengio, 2017; Gurnee and Tegmark, 2023; Nanda et al., 2023b; Leask et al., 2024): low interference allows the feature of interest to be approximately recovered from the representation using linear regression. A similar story holds for discrete features accessed via linear classifiers. 3 The interpretation of distance on representation manifolds There is an open question in the mechanistic interpretability community about the meaning of distance in representation space, perhaps well-summarised in the commentary of Olah and Batson (2024): âWe suspect this idea that feature manifolds many be embedded in more complex ways than their topology suggests, in order to achieve a given distance metric, may actually be quite deep and important.â Figure 3: Evidence for Hypothesis 2 and its implications in Theorem 1. For each pair of representations, we plot their cosine similarities (first row) and estimated manifold distances (second row) against their (squared) distance in a putative metric space. We report the Chatterjee (Ξ) and Pearson (Ď) correlation coefficients, respectively. Colours correspond to the colourmaps described in Figure 1. If we accept there is a correspondence between features and their representations (Hypothesis 1), arguably the next most basic question we can ask is whether cosine similarity in representation space somehow tells us about distance between the corresponding feature values. Hypothesis 2 (cosine similarity reflects distance). Locally, the cosine similarity between feature representations and the distance their corresponding feature values are inversely related. Formally, there is some function gsubscriptg_ fgtypewriter_f with continuous second derivatives and with gâ˛â˘(0)<0superscriptsubscriptâ˛00g_ f (0)<0gtypewriter_fⲠ( 0 ) < 0, and some Ďľ>0italic-Ďľ0Îľ>0Ďľ > 0, such that âĄ(Ďâ˘(z),Ďâ˘(zâ˛))=gâ˘(â˘(z,zâ˛)2),subscriptitalic-Ďsubscriptitalic-Ďsuperscriptâ˛subscriptsubscriptsuperscriptsuperscriptâ˛2 CosSim ( _ f(z), _ % f(z ) )=g_ f( d_ f(z,z )^2% ),sansserif_CosSim ( Ďtypewriter_f ( z ) , Ďtypewriter_f ( zⲠ) ) = gtypewriter_f ( sansserif_dtypewriter_f ( z , zⲠ)2 ) , for all z,zâ˛âsuperscriptâ˛subscriptz,z _ fz , zⲠâ Ztypewriter_f such that â˘(z,zâ˛)â¤Ďľsubscriptsuperscriptâ˛italic-Ďľ d_ f(z,z )⤠_dtypewriter_f ( z , zⲠ) ⤠Ͼ. Strengthening just Hypothesis 1 to both Hypotheses 1 and 2 has formidable consequences: there is an intrinsic sense in which a feature and its representation are geometrically indistinguishable. This statement is made precise in Theorem 1. Metric spaces allow for a natural definition of a path which, loosely speaking, captures the idea of a continuous route from one point to another and there is an associated definition of the length of a path, denoted L, which generalises the usual Euclidean notion of length (Burago et al., 2001). Formally, a path in subscriptZ_ fZtypewriter_f is a continuous mapping Ρ from some interval [a,b][a,b][ a , b ] to subscriptZ_ fZtypewriter_f, and the length of such a path is Lâ˘(Ρ)âsupâi=1nâ˘(Ρti,Ρtiâ1),âsubscriptsupremumsuperscriptsubscript1subscriptsubscriptsubscriptsubscriptsubscript1L(Ρ) _T _i=1^n d_ f( _% t_i, _t_i-1),L ( Ρ ) â supcaligraphic_T âi = 1n sansserif_dtypewriter_f ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) , where the supremum is over all nâĽ11n⼠1n ⼠1 and =(t0,t1,âŚ,tn)subscript0subscript1âŚsubscriptT=(t_0,t_1,âŚ,t_n)T = ( t0 , t1 , ⌠, titalic_n ) such that t0=aâ¤t1â¤âŻâ¤tn=bsubscript0subscript1âŻsubscriptt_0=a⤠t_1â¤Âˇs⤠t_n=bt0 = a ⤠t1 ⤠⯠⤠titalic_n = b. Given a path in subscriptZ_ fZtypewriter_f, we can think of the image of this path when mapped through Ďsubscriptitalic-Ď _ fĎtypewriter_f, which we call the corresponding path on âłsubscriptâłM_ fMtypewriter_f. Its length is defined similarly, with Euclidean distance in place of subscript d_ fsansserif_dtypewriter_f. Theorem 1. Let Ρ be a path on subscriptZ_ fZtypewriter_f of finite length and, assuming Hypothesis 1, let Îł be the corresponding path on âłsubscriptâłM_ fMtypewriter_f. Then, under Hypothesis 2, Lâ˘(Îł)=â2â˘gâ˛â˘(0)â˘Lâ˘(Ρ).2superscriptsubscriptâ˛0L(Îł)= -2g_ f (0)L(Ρ).L ( Îł ) = square-root start_ARG - 2 gtypewriter_fⲠ( 0 ) end_ARG L ( Ρ ) . A proof of Theorem 1 is given in the appendix. Theorem 1 tells us that we can recover the intrinsic geometry of subscriptZ_ fZtypewriter_f, even though we know (almost) nothing about gsubscriptg_ fgtypewriter_f: shortest paths on âłsubscriptâłM_ fMtypewriter_f correspond to shortest paths on subscriptZ_ fZtypewriter_f, and their lengths are equal, up to a choice of unit (reflected by â2â˘gâ˛â˘(0)2superscriptsubscriptâ˛0 -2g_ f (0)square-root start_ARG - 2 gtypewriter_fⲠ( 0 ) end_ARG) Figure 4: Evidence against isometry with respect to the metric space years=[1900,1999]subscriptyears19001999Z_ years=[1900,1999]Zyears = [ 1900 , 1999 ], yearâ˘(x,y)=|xây|subscriptyear d_ year(x,y)=|x-y|sansserif_dyear ( x , y ) = | x - y |. There is no clear regular linear relationship between distances in this metric space and estimated distances on the representation manifold. The colours indicate that distances between more recent years are expanded on the manifold. 3.1 Geodesic distances on representation manifolds of LLMs are meaningful We now explore the plausibility of Hypothesis 2 in the same colour, year, date examples considered in Section 2.3 and Figure 1. In all cases, we find indications of isometry, with some important caveats. Given a putative metric space, Hypothesis 2 suggests two diagnostic tests. The first (direct) approach is to plot cosine similarity against squared distance, to check if the first appears to be a decreasing function of the second, around zero, up to noise. We quantify the global strength of functional dependence using Chatterjeeâs correlation coefficient Ξ (Chatterjee, 2021), which would be 1 if the cosine similarity was a deterministic function of distance. These experiments are shown in the first row of Figure 3. The second (indirect) approach is to test Theorem 1: geodesic distance on âłsubscriptâłM_ fMtypewriter_f (shortest path length) should be linear in the geodesic distance on subscriptZ_ fZtypewriter_f, up to noise (the slope being â2â˘gâ˛â˘(0)2superscriptâ˛0 -2g (0)square-root start_ARG - 2 gⲠ( 0 ) end_ARG). We estimate geodesics on âłsubscriptâłM_ fMtypewriter_f by constructing the K-nearest-neighbours graph over the representations, and reporting weighted graph distance, k chosen as small as possible subject to the graph being connected. We quantify the strength of isometry using Pearsonâs correlation Ď, which would be 1 if the distances were in a perfect proportional relationship. These experiments are shown in the second row of Figure 3. Across our experiments, we have found that a low-dimensional projection tends to be necessary for the representations to plausibly show isometry with a simple metric space. For our text embeddings, we find that projecting onto the first few (uncentered) principal components works well. The routine âlow-rankâ explanation that the remaining components are mostly noise seems disputable; these components often show clear structure. Our best explanation is that semantic similarity is much richer than the rudimentary metric spaces to which they are being compared. It is likely that we could achieve a deeper understanding of semantic similarity through improved metric space design. In the example of years, the process of extracting feature representation via an SAE automatically yields low-dimensional representations, so no PCA is applied in this case. Recall that we conjectured the following metric space for the years example: years=[1900,1999]subscriptyears19001999Z_ years=[1900,1999]Zyears = [ 1900 , 1999 ], yearâ˘(x,y)=|xây|subscriptyear d_ year(x,y)=|x-y|sansserif_dyear ( x , y ) = | x - y |. Although we found a rank correlation near 1, indicating homeomorphism, the evidence of the tests above is against isometry. The clearest indication in this direction is possibly provided by the right panel of Figure 4, which does not show a regular linear relationship, the colours suggesting that distances between more recent years are expanded on the manifold. In light of this, we consider a modified representation years=logâĄ(2019âyear):yearâ[1900,1999]subscriptyearsconditional-set2019yearyear19001999Z_ years=\ (2019-year):yearâ[1900,1999]\Zyears = log ( 2019 - year ) : year â [ 1900 , 1999 ] , yearâ˘(x,y)=|xây|subscriptyear d_ year(x,y)=|x-y|sansserif_dyear ( x , y ) = | x - y |, 2019 being the year GPT-2 was released (Radford et al., 2019). Observe that the rank correlation as computed in Section 2.3 remains unchanged: the two representations are homeomorphic, and cannot be distinguished on purely topological criteria. The tests are now in much stronger support of isometry. The top-middle plot of Figure 3 shows a trend which is clearly decreasing at zero, and a globally high functional dependence, 0.84. The bottom-middle panel of Figure 3 shows a clear linear fit, achieving a Pearson correlation of 0.99. Isometry is also found to be plausible for dates, and weakly plausible for colours, both with the original circular metric spaces conjectured in Section 2.3. The bottom-right panel shows a clear linear fit, achieving a Pearson correlation of 0.97. Observe in this case that the cosine similarity appears to be inconsistent with the metric for large distances, illustrating the point that Hypothesis 2 only requires an inverse functional relationship to hold locally. 4 Discussion, limitations and future work This work provides a formal mathematical framework to explain and interpret representation manifolds in large language models. By modeling features as metric spaces, we are able to accurately characterise the topological and geometric properties of their representations under some basic hypotheses. We perform some preliminary investigations on internal representations from GPT2-small and text embeddings from OpenAIâs text-embedding-large-3 model, which validate our theory and provide a nuanced, quantitative view of the correspondence between human concepts and their representation at different levels of geometric fidelity, namely homeomorphism and isometry. We find hints that these models encode distances in ways which are sometimes unexpected: years of the twentieth century in GPT2-small appear to be encoded on a logarithmic scale, with larger distances between more recent years, and colours in text-embedding-large-3 appear to be encoded in a cycle of hues, rather than representations that other systems might have chosen, such as RGB or wavelength. 4.1 Limitations In the spirit of scientific investigation, we have opted for a hypothesis-driven approach to structure discovery, in which we put down a possible metric space as a hypothesis and then assess evidence in favour or against. This âmanualâ approach is clearly not scalable, and moreover relies on there being some reasonable starting hypothesis for the metric space, which could be difficult for many features (say, emotions), as is evident from prior research (Li et al., 2023; Nanda et al., 2023b). There is an unexplored alternative approach of learning the metric, but we do not know exactly how this would proceed given that an interpretable solution would presumably remain a requirement. In our experiments on text-embedding-large-3, we use PCA with some success to isolate simple human-understandable distances. However, in reality, we expect that the true notions of distance used by the language model are more complex: we find additional structure in further principal components and think it is possible that a language model could encode distances in a way that is mechanistically useful, but does not correspond to any existing human understanding of the feature. Another limitation of our approach relates to the fundamental statistical difficulty of estimating manifolds in the presence of noise (Genovese et al., 2012). Here, we have opted for a simple and interpretable approach of using the K-nearest-neighbour graph to approximate the manifold, but this is prone to short-circuits causing enormous errors in the estimated manifold distances. It is often the case that one has to manually prune the graph in order to achieve reasonable manifold distance estimates, and we believe that more robust methodology for manifold estimation would be required to scale up our approach. 4.2 Implications for mechanistic interpretability research In mechanistic interpretability, one of the underlying motivations for understanding representation geometry is to be able to steer model outputs by making interventions on their internal representations. For features represented on manifolds, our insights suggest a path forward for doing this: learn the map Ďsubscriptitalic-Ď _ fĎtypewriter_f which maps the feature subscriptZ_ fZtypewriter_f onto its representation manifold âłsubscriptâłM_ fMtypewriter_f. Sparse autoencoders provide a potentially promising avenue for this. We conjecture that the sparsity penalty of a sparse autoencoder, trained on representation manifold, will encourage it to learn a collection of dictionary vectors which trace the manifold (see Section 4 of Engels et al. (2025) for an argument for this phenomenon on representation subspaces). We also conjecture that much observed feature splitting in SAEs is a result of this. We hope our work will encourage development of âmanifold-awareâ SAEs. Finally, mechanistic interpretability is a nascent field of research which is still developing a common language, and we hope that researchers will find the formalism of a feature as a metric space to be a useful possibility in future scientific discourse. References Alain and Bengio (2017) Guillaume Alain and Yoshua Bengio. Understanding intermediate layers using linear classifier probes. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings. OpenReview.net, 2017. URL https://openreview.net/forum?id=HJ4-rAVtl. Ansuini et al. (2019) Alessio Ansuini, Alessandro Laio, Jakob H Macke, and Davide Zoccolan. Intrinsic dimension of data representations in deep neural networks. Advances in Neural Information Processing Systems, 32, 2019. Arora et al. (2018) Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. Linear Algebraic Structure of Word Senses, with Applications to Polysemy. Transactions of the Association for Computational Linguistics, 6:483â495, December 2018. ISSN 2307-387X. doi: 10.1162/tacl_a_00034. URL https://direct.mit.edu/tacl/article/43451. Bostrom (2014) Nick Bostrom. Superintelligence: Paths, Dangers, Strategies. Oxford University Press, Oxford, UK, 2014. ISBN 978-0-19-967811-2. Bricken et al. (2023) Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread, 2023. URL https://transformer-circuits.pub/2023/monosemantic-features/index.html. Burago et al. (2001) Dmitri Burago, Yuri Burago, Sergei Ivanov, and others. A course in metric geometry. American Mathematical Society, 2001. Bussmann et al. (2024) Bart Bussmann, Michael Pearce, Patrick Leask, Joseph Bloom, Lee Sharkey, and Neel Nanda. Showing SAE Latents Are Not Atomic Using Meta-SAEs. Less Wrong, 2024. URL https://w.lesswrong.com/posts/TMAmHh4DdMr4nCSr5/showing-sae-latents-are-not-atomic-using-meta-saes. Bussmann et al. (2025) Bart Bussmann, Noa Nabeshima, Adam Karvonen, and Neel Nanda. Learning Multi-Level Features with Matryoshka Sparse Autoencoders. arXiv preprint arXiv:2503.17547, 2025. Cai et al. (2021) Xingyu Cai, Jiaji Huang, Yuchen Bian, and Kenneth Church. Isotropy in the contextual embedding space: Clusters and manifolds. In International conference on learning representations, 2021. Chang et al. (2022) Tyler A. Chang, Zhuowen Tu, and Benjamin K. Bergen. The Geometry of Multilingual Language Model Representations. In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, EMNLP 2022, Abu Dhabi, United Arab Emirates, December 7-11, 2022, pages 119â136. Association for Computational Linguistics, 2022. doi: 10.18653/V1/2022.EMNLP-MAIN.9. URL https://doi.org/10.18653/v1/2022.emnlp-main.9. Chatterjee (2021) Sourav Chatterjee. A new coefficient of correlation. Journal of the American Statistical Association, 116(536):2009â2022, 2021. Publisher: Taylor & Francis. Cunningham et al. (2023) Hoagy Cunningham, Aidan Ewart, Logan Riggs, Robert Huben, and Lee Sharkey. Sparse autoencoders find highly interpretable features in language models. arXiv preprint arXiv:2309.08600, 2023. Elad (2010) Michael Elad. Sparse and redundant representations: from theory to applications in signal and image processing. Springer Science & Business Media, 2010. Elhage et al. (2021) Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread, 2021. URL https://transformer-circuits.pub/2021/framework/index.html. Elhage et al. (2022) Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, Roger Grosse, Sam McCandlish, Jared Kaplan, Dario Amodei, Martin Wattenberg, and Christopher Olah. Toy Models of Superposition. Transformer Circuits Thread, 2022. URL https://transformer-circuits.pub/2022/toy_model/index.html. Engels et al. (2025) Joshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark. Not All Language Model Features Are One-Dimensionally Linear. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=d63a4AM4hb. Gao et al. (2021) Tianyu Gao, Xingcheng Yao, and Danqi Chen. SimCSE: Simple Contrastive Learning of Sentence Embeddings. In Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 6894â6910. Association for Computational Linguistics, 2021. doi: 10.18653/V1/2021.EMNLP-MAIN.552. URL https://doi.org/10.18653/v1/2021.emnlp-main.552. Genovese et al. (2012) Christopher R Genovese, Marco Perone-Pacifico, Isabella Verdinelli, and Larry Wasserman. Minimax manifold estimation. The Journal of Machine Learning Research, 13(1):1263â1291, 2012. Gorton (2024) Liv Gorton. Curve Detector Manifolds in InceptionV1, August 2024. URL https://livgorton.com/curve-detector-manifolds/. Gurnee and Tegmark (2023) Wes Gurnee and Max Tegmark. Language models represent space and time. arXiv preprint arXiv:2310.02207, 2023. Hanna et al. (2023) Michael Hanna, Roberto Zamparelli, David MareÄek, and others. The functional relevance of probed information: A case study. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 835â848. Association for Computational Linguistics, 2023. He et al. (2024) Tianyu He, Darshil Doshi, Aritra Das, and Andrey Gromov. Learning to grok: Emergence of in-context learning and skill composition in modular arithmetic tasks. Advances in Neural Information Processing Systems, 37:13244â13273, 2024. Higgins et al. (2018) Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner. Towards a definition of disentangled representations. arXiv preprint arXiv:1812.02230, 2018. Hyvärinen et al. (2024) Aapo Hyvärinen, Ilyes Khemakhem, and Ricardo Monti. Identifiability of latent-variable and structural-equation models: from linear to nonlinear. Annals of the Institute of Statistical Mathematics, 76(1):1â33, 2024. Publisher: Springer. Leask et al. (2024) Patrick Leask, Bart Bussmann, and Neel Nanda. Calendar feature geometry in GPT-2 layer 8 residual stream SAEs. Less Wrong, 2024. URL https://w.lesswrong.com/posts/WsPyunwpXYCM2iN6t/calendar-feature-geometry-in-gpt-2-layer-8-residual-stream. Li et al. (2020) Bohan Li, Hao Zhou, Junxian He, Mingxuan Wang, Yiming Yang, and Lei Li. On the Sentence Embeddings from Pre-trained Language Models. In Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020, pages 9119â9130. Association for Computational Linguistics, 2020. doi: 10.18653/V1/2020.EMNLP-MAIN.733. URL https://doi.org/10.18653/v1/2020.emnlp-main.733. Li et al. (2023) Kenneth Li, Aspen K Hopkins, David Bau, Fernanda ViĂŠgas, Hanspeter Pfister, and Martin Wattenberg. Emergent world representations: Exploring a sequence model trained on a synthetic task. ICLR, 2023. Publisher: ICLR. Li et al. (2025) Yuxiao Li, Eric J Michaud, David D Baek, Joshua Engels, Xiaoqing Sun, and Max Tegmark. The geometry of concepts: Sparse autoencoder feature structure. Entropy, 27(4):344, 2025. Publisher: MDPI. Lin (2024) Johnny Lin. Neuronpedia: Interactive reference and tooling for analyzing neural networks, 2024. URL https://w.neuronpedia.org. Liu et al. (2022) Ziming Liu, Ouail Kitouni, Niklas S Nolte, Eric Michaud, Max Tegmark, and Mike Williams. Towards understanding grokking: An effective theory of representation learning. Advances in Neural Information Processing Systems, 35:34651â34663, 2022. Mendel (2024) Jake Mendel. SAE feature geometry is outside the superposition hypothesis. AI Alignment Forum, 2024. URL https://w.alignmentforum.org/posts/MFBTjb2qf3ziWmzz6/sae-feature-geometry-is-outside-the-superposition-hypothesis. Munroe (2010) Randall Munroe. XKCD Color Name Survey Results, 2010. URL https://xkcd.com/color/rgb/. Nanda et al. (2023a) Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. Progress measures for grokking via mechanistic interpretability. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023a. URL https://openreview.net/forum?id=9XFSbDPmdW. Nanda et al. (2023b) Neel Nanda, Andrew Lee, and Martin Wattenberg. Emergent linear representations in world models of self-supervised sequence models. arXiv preprint arXiv:2309.00941, 2023b. Olah (2024) Chris Olah. What is a Linear Representation? What is a Multidimensional Feature? Transformer Circuits Thread, 2024. URL https://transformer-circuits.pub/2024/july-update/index.html#linear-representations. Olah and Batson (2024) Chris Olah and Josh Batson. Feature Manifold Toy Model. Transformer Circuits Thread, 2024. URL https://transformer-circuits.pub/2023/may-update/index.html#feature-manifolds. Park et al. (2024) Kiho Park, Yo Joong Choe, Yibo Jiang, and Victor Veitch. The geometry of categorical and hierarchical concepts in large language models. arXiv preprint arXiv:2406.01506, 2024. Pfau et al. (2020) David Pfau, Irina Higgins, Alex Botev, and SĂŠbastien Racanière. Disentangling by subspace diffusion. Advances in Neural Information Processing Systems, 33:17403â17415, 2020. Radford et al. (2019) Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language Models are Unsupervised Multitask Learners. 2019. Reizinger et al. (2025) Patrik Reizinger, Alice Bizeul, Attila Juhos, Julia E. Vogt, Randall Balestriero, Wieland Brendel, and David A. Klindt. Cross-Entropy Is All You Need To Invert the Data Generating Process. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URL https://openreview.net/forum?id=hrqNOxpItr. Shai et al. (2024) Adam S. Shai, Lucas Teixeira, Alexander Gietelink Oldenziel, Sarah Marzen, and Paul M. Riechers. Transformers Represent Belief State Geometry in their Residual Stream. In Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang, editors, Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. URL http://papers.nips.c/paper_files/paper/2024/hash/8936fa1691764912d9519e1b5673ea66-Abstract-Conference.html. Smith (2024) Lewis Smith. The âstrongâ feature hypothesis could be wrong. AI Alignment Forum, 2024. URL https://w.alignmentforum.org/posts/tojtPCCRpKLSHBdpn/the-strong-feature-hypothesis-could-be-wrong. Smolensky (1990) Paul Smolensky. Tensor product variable binding and the representation of symbolic structures in connectionist systems. Artificial intelligence, 46(1-2):159â216, 1990. Publisher: Elsevier. Soares and Fallenstein (2017) Nate Soares and Benya Fallenstein. Agent foundations for aligning machine intelligence with human interests: a technical research agenda. The technological singularity: Managing the journey, pages 103â125, 2017. Publisher: Springer. Sutherland (2009) Wilson A. Sutherland. Introduction to Metric and Topological Spaces. Oxford University Press, 2009. Templeton et al. (2024) Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L. Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Christopher Olah, and Tom Henighan. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread, 2024. URL https://transformer-circuits.pub/2024/scaling-monosemanticity/. Wang et al. (2023) Hanchen Wang, Tianfan Fu, Yuanqi Du, Wenhao Gao, Kexin Huang, Ziming Liu, Payal Chandak, Shengchao Liu, Peter Van Katwyk, Andreea Deac, Anima Anandkumar, Karianne Bergen, Carla P. Gomes, Shirley Ho, Pushmeet Kohli, Joan Lasenby, Jure Leskovec, Tie\-Yan Liu, Arjun Manrai, Debora Marks, Bharath Ramsundar, Le Song, Jimeng Sun, Jian Tang, Petar VeliÄkoviÄ, Max Welling, Linfeng Zhang, Connor W. Coley, Yoshua Bengio, and Marinka Zitnik. Scientific discovery in the age of artificial intelligence. Nature, 620(7972):47â60, August 2023. doi: 10.1038/s41586-023-06221-2. URL https://w.nature.com/articles/s41586-023-06221-2. Zhong et al. (2023) Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas. The clock and the pizza: Two stories in mechanistic explanation of neural networks. Advances in neural information processing systems, 36:27223â27250, 2023. Zimmermann et al. (2021) Roland S Zimmermann, Yash Sharma, Steffen Schneider, Matthias Bethge, and Wieland Brendel. Contrastive learning inverts the data generating process. In International conference on machine learning, pages 12979â12990. PMLR, 2021. Appendix Appendix A Code to reproduce the experiments in this paper Code to reproduce the experiments in this paper is made available at the GitHub repository https://github.com/alexandermodell/Representation-Manifolds. Appendix B Supporting definitions and proof of Theorem 1 Hypotheses 1 and 2, and the statement of the Theorem 1 concern the metric space (,)subscriptsubscript(Z_ f, d_ f)( Ztypewriter_f , sansserif_dtypewriter_f ), and functions Ďsubscriptitalic-Ď _ fĎtypewriter_f and gsubscriptg_ fgtypewriter_f associated with some particular feature ftypewriter_f. To de-clutter the proof we remove dependence on ftypewriter_f from the notation, and write simply (,)(Z, d)( Z , sansserif_d ), Ďitalic-ĎĎĎ and g. We shall need the following definitions, informed by (Burago et al., 2001). A path in ZZ is a continuous mapping Ρ from some interval [a,b][a,b][ a , b ] to ZZ. The length of such a path is: Lâ˘(Ρ)âsupâi=1nâ˘(Ρti,Ρtiâ1)âsubscriptsupremumsuperscriptsubscript1subscriptsubscriptsubscriptsubscript1L(Ρ) _T _i=1^n d( _t_i, _% t_i-1)L ( Ρ ) â supcaligraphic_T âi = 1n sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) where the supremum is over all =(t0,t1,âŚ,tn)subscript0subscript1âŚsubscriptT=(t_0,t_1,âŚ,t_n)T = ( t0 , t1 , ⌠, titalic_n ) such that nâĽ11n⼠1n ⼠1 and t0=aâ¤t1â¤âŻâ¤tn=bsubscript0subscript1âŻsubscriptt_0=a⤠t_1â¤Âˇs⤠t_n=bt0 = a ⤠t1 ⤠⯠⤠titalic_n = b. A path is said to be of finite length, or equivalently called rectifiable, if Lâ˘(Ρ)<âL(Ρ)<âL ( Ρ ) < â. For [aâ˛,bâ˛]â[a,b]superscriptâ˛[a ,b ] [a,b][ aⲠ, bⲠ] â [ a , b ] we write Lâ˘(Ρ,aâ˛,bâ˛)superscriptâ˛L(Ρ,a ,b )L ( Ρ , aⲠ, bⲠ) for the length of the restriction of Ρ to [aâ˛,bâ˛]superscriptâ˛[a ,b ][ aⲠ, bⲠ]. Any rectifiable path Ρ:[a,b]â:âΡ:[a,b] Ρ : [ a , b ] â Z admits a unit-speed parameterisation, meaning Ρ has the representation Ρ=Ρ~âĎ~Ρ= Ρ Ρ = over~ start_ARG Ρ end_ARG â Ď where Ρ~:[0,Lâ˘(Ρ)]â:~â0 Ρ:[0,L(Ρ)] ~ start_ARG Ρ end_ARG : [ 0 , L ( Ρ ) ] â Z is a path, Ď Ď is a continuous, nondecreasing map from [a,b][a,b][ a , b ] to [0,Lâ˘(Ρ)]0[0,L(Ρ)][ 0 , L ( Ρ ) ], and Lâ˘(Ρ~,s,t)=tâs~L( Ρ,s,t)=t-sL ( over~ start_ARG Ρ end_ARG , s , t ) = t - s (Burago et al., 2001, Prop. 2.5.9). Adopting this parameterisation does not change overall length of the path, in the sense that the stated properties of Ρ~~ Ρover~ start_ARG Ρ end_ARG imply Lâ˘(Ρ~)=Lâ˘(Ρ~,0,Lâ˘(Ρ))=Lâ˘(Ρ)~~0L( Ρ)=L( Ρ,0,L(Ρ))=L(Ρ)L ( over~ start_ARG Ρ end_ARG ) = L ( over~ start_ARG Ρ end_ARG , 0 , L ( Ρ ) ) = L ( Ρ ). We shall also consider paths on the unit hyper-sphere Dâ1âxââDâ1:âxâ2=1âsuperscript1conditional-setsuperscriptâ1subscriptnorm21S^D-1 \x ^D-1:\|x\|_2=1\blackboard_SD - 1 â x â blackboard_RD - 1 : ⼠x âĽ2 = 1 . The length of such a path, i.e., a continuous mapping Îł:[a,b]âDâ1:âsuperscript1Îł:[a,b] ^D-1Îł : [ a , b ] â blackboard_SD - 1 is: Lâ˘(Îł)âsupâi=1nâÎłtiâÎłtiâ1â2,âsubscriptsupremumsuperscriptsubscript1subscriptnormsubscriptsubscriptsubscriptsubscript12L(Îł) _T _i=1^n\| _t_i- _t_% i-1\|_2,L ( Îł ) â supcaligraphic_T âi = 1n ⼠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT - Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT âĽ2 , where again the supremum is over all =(t0,t1,âŚ,tn)subscript0subscript1âŚsubscriptT=(t_0,t_1,âŚ,t_n)T = ( t0 , t1 , ⌠, titalic_n ) such that nâĽ11n⼠1n ⼠1 and t0=aâ¤t1â¤âŻâ¤tn=bsubscript0subscript1âŻsubscriptt_0=a⤠t_1â¤Âˇs⤠t_n=bt0 = a ⤠t1 ⤠⯠⤠titalic_n = b. Proof of Theorem 1. Since the claim of the theorem depends on Ρ only through its length, we can assume w.l.o.g. that we are considering the unit-speed parameterisation of Ρ. That is [a,b]=[0,Lâ˘(Ρ)]0[a,b]=[0,L(Ρ)][ a , b ] = [ 0 , L ( Ρ ) ] and Ρ:[0,Lâ˘(Ρ)]â:â0Ρ:[0,L(Ρ)] Ρ : [ 0 , L ( Ρ ) ] â Z with Lâ˘(Ρ,s,t)=tâsL(Ρ,s,t)=t-sL ( Ρ , s , t ) = t - s. Under Hypothesis 1, Ďitalic-ĎĎĎ is continuous, hence Îł:[0,Lâ˘(Ρ)]âDâ1:â0superscript1Îł:[0,L(Ρ)] ^D-1Îł : [ 0 , L ( Ρ ) ] â blackboard_SD - 1 defined by Îłt=Ďâ˘(Ρt)subscriptitalic-Ďsubscript _t=Ď( _t)Îłitalic_t = Ď ( Ρitalic_t ) is a path on Dâ1superscript1S^D-1blackboard_SD - 1. For any =(t0,t1,,âŚ,tn)T=(t_0,t_1,,âŚ,t_n)T = ( t0 , t1 , , ⌠, titalic_n ) such that nâĽ11n⼠1n ⼠1 and t0=0â¤t1â¤âŻâ¤tn=Lâ˘(Ρ)subscript00subscript1âŻsubscriptt_0=0⤠t_1â¤Âˇs⤠t_n=L(Ρ)t0 = 0 ⤠t1 ⤠⯠⤠titalic_n = L ( Ρ ), introduce the notation: Sâ˘(Ρ,)ââi=1nâ˘(Ρti,Ρtiâ1),Sâ˘(Îł,)ââi=1nâÎłtiâÎłtiâ1â2.formulae-sequenceâsuperscriptsubscript1subscriptsubscriptsubscriptsubscript1âsuperscriptsubscript1subscriptnormsubscriptsubscriptsubscriptsubscript12S(Ρ,T) _i=1^n d( _t_i, _t_i-1% ), S(Îł,T) _i=1^n\| _t_i-Îł% _t_i-1\|_2.S ( Ρ , T ) â âi = 1n sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) , S ( Îł , T ) â âi = 1n ⼠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT - Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT âĽ2 . Fix any δ>00δ>0δ > 0. We shall prove that there exists δsubscriptT_δTitalic_δ such that : |Lâ˘(Îł)â2â˘gâ˛â˘(0)â˘Lâ˘(Ρ)|2superscriptâ˛0 |L(Îł)- -2g (0)L(Ρ) || L ( Îł ) - square-root start_ARG - 2 gⲠ( 0 ) end_ARG L ( Ρ ) | â¤|Lâ˘(Îł)âSâ˘(Îł,δ)|absentsubscript ⤠|L(Îł)-S(Îł,T_δ) |⤠| L ( Îł ) - S ( Îł , Titalic_δ ) | (2) +|Sâ˘(Îł,δ)â2â˘gâ˛â˘(0)â˘Sâ˘(Ρ,δ)|subscript2superscriptâ˛0subscript + |S(Îł,T_δ)- -2g (0)% S(Ρ,T_δ) |+ | S ( Îł , Titalic_δ ) - square-root start_ARG - 2 gⲠ( 0 ) end_ARG S ( Ρ , Titalic_δ ) | (3) +â2â˘gâ˛â˘(0)â˘|Sâ˘(Ρ,δ)âLâ˘(Ρ)|2superscriptâ˛0subscript + -2g (0) |S(Ρ,T_δ)-L(% Ρ) |+ square-root start_ARG - 2 gⲠ( 0 ) end_ARG | S ( Ρ , Titalic_δ ) - L ( Ρ ) | (4) â¤Î´3+δ3+δ3,absent333 ⤠δ3+ δ3+ δ3,⤠divide start_ARG δ end_ARG start_ARG 3 end_ARG + divide start_ARG δ end_ARG start_ARG 3 end_ARG + divide start_ARG δ end_ARG start_ARG 3 end_ARG , (5) which implies the claim of the theorem. We shall construct δsubscriptT_δTitalic_δ in the form δâδ(1)âŞÎ´(2)âŞÎ´(3)âsubscriptsuperscriptsubscript1superscriptsubscript2superscriptsubscript3T_δ _δ^(1) _δ% ^(2) _δ^(3)Titalic_δ â Titalic_δ( 1 ) ⪠Titalic_δ( 2 ) ⪠Titalic_δ( 3 ), i.e., δ(i)âδsuperscriptsubscriptsubscriptT_δ^(i) _δTitalic_δ( i ) â Titalic_δ for i=1,2,3123i=1,2,3i = 1 , 2 , 3, where δ(i)superscriptsubscriptT_δ^(i)Titalic_δ( i ) are defined in the remainder of the proof. We first consider a difference of the form |Sâ˘(Îł,â )â2â˘gâ˛â˘(0)â˘Sâ˘(Ρ,â )|â 2superscriptâ˛0â |S(Îł,¡)- -2g (0)S(Ρ,¡) || S ( Îł , â ) - square-root start_ARG - 2 gⲠ( 0 ) end_ARG S ( Ρ , â ) | as appears in (3). Noting that the mapping Ďitalic-ĎĎĎ by definition satisfies âĎâ˘(z)â2=1subscriptnormitalic-Ď21\|Ď(z)\|_2=1âĽ Ď ( z ) âĽ2 = 1 for all z, under Hypothesis 2, there exists Ďľ>0italic-Ďľ0Îľ>0Ďľ > 0 such that if â˘(z,zâ˛)<Ďľsuperscriptâ˛italic-Ďľ d(z,z )< _d ( z , zⲠ) < Ďľ, then â¨Ďâ˘(z),Ďâ˘(zâ˛)âŠ2=gâ˘(â˘(z,zâ˛)2)subscriptitalic-Ďitalic-Ďsuperscriptâ˛2superscriptsuperscriptâ˛2 Ď(z),Ď(z ) _2=g( d(z,z % )^2)â¨ Ď ( z ) , Ď ( zⲠ) âŠ2 = g ( sansserif_d ( z , zⲠ)2 ). Let C>00C>0C > 0 be any finite constant such that suprâ¤Ďľ|gâ˛â˘(r)|â¤Csubscriptsupremumitalic-ĎľsuperscriptⲠ_râ¤Îľ|g (r)|⤠Csupitalic_r ⤠Ͼ | gⲠⲠ( r ) | ⤠C. Such a constant exists because g is C2superscript2C^2C2 by assumption. Let δ(2)=(t0(2)=0,t1(2),âŚ,tn(2)(2)=Lâ˘(Ρ))superscriptsubscript2formulae-sequencesuperscriptsubscript020superscriptsubscript12âŚsuperscriptsubscriptsuperscript22T_δ^(2)=(t_0^(2)=0,t_1^(2),âŚ,t_n^(2)^(2)% =L(Ρ))Titalic_δ( 2 ) = ( t0( 2 ) = 0 , t1( 2 ) , ⌠, titalic_n( 2 )( 2 ) = L ( Ρ ) ) be defined by: n(2)ââ3â˘Câ˘|Lâ˘(Ρ)|2δâ¨Lâ˘(Ρ)Ďľâ,ti(2)âin(2)â˘Lâ˘(Ρ),i=0,âŚ,n(2).formulae-sequenceâsuperscript23superscript2italic-Ďľformulae-sequenceâsuperscriptsubscript2superscript20âŚsuperscript2n^(2) 3C|L(Ρ)|^2δ L(Ρ)% Îľ , t_i^(2) in^(2)L(Ρ),% i=0,âŚ,n^(2).n( 2 ) â â divide start_ARG 3 C | L ( Ρ ) |2 end_ARG start_ARG δ end_ARG ⨠divide start_ARG L ( Ρ ) end_ARG start_ARG Ďľ end_ARG â , titalic_i( 2 ) â divide start_ARG i end_ARG start_ARG n( 2 ) end_ARG L ( Ρ ) , i = 0 , ⌠, n( 2 ) . Using the fact that Ρ is unit-speed parameterised, it follows that, for 1â¤iâ¤n(2)1superscript21⤠i⤠n^(2)1 ⤠i ⤠n( 2 ), Lâ˘(Ρ,ti(2),tiâ1(2))=ti(2)âtiâ1(2)=Lâ˘(Ρ)n(2)â¤Î´3â˘Câ˘Lâ˘(Ρ)â§Ďľ.superscriptsubscript2superscriptsubscript12superscriptsubscript2superscriptsubscript12superscript23italic-ĎľL(Ρ,t_i^(2),t_i-1^(2))=t_i^(2)-t_i-1^(2)= L(Ρ)n^% (2)⤠δ3CL(Ρ) Îľ.L ( Ρ , titalic_i( 2 ) , titalic_i - 1( 2 ) ) = titalic_i( 2 ) - titalic_i - 1( 2 ) = divide start_ARG L ( Ρ ) end_ARG start_ARG n( 2 ) end_ARG ⤠divide start_ARG δ end_ARG start_ARG 3 C L ( Ρ ) end_ARG â§ Ďľ . (6) Now consider any =(t0,t1,,âŚ,tn)T=(t_0,t_1,,âŚ,t_n)T = ( t0 , t1 , , ⌠, titalic_n ) with nâĽn(2)superscript2n⼠n^(2)n ⼠n( 2 ), t0=0subscript00t_0=0t0 = 0 , tn=Lâ˘(Ρ)subscriptt_n=L(Ρ)titalic_n = L ( Ρ ) such that δ(2)âsuperscriptsubscript2T_δ^(2) _δ( 2 ) â T. Unit-speed parameterisation of Ρ combined with δ(2)âsuperscriptsubscript2T_δ^(2) _δ( 2 ) â T implies: max1â¤iâ¤nâĄLâ˘(Ρ,ti,tiâ1)=max1â¤iâ¤nâĄtiâtiâ1â¤max1â¤iâ¤n(2)âĄti(2)âtiâ1(2)=max1â¤iâ¤n(2)âĄLâ˘(Ρ,ti(2),tiâ1(2)),subscript1subscriptsubscript1subscript1subscriptsubscript1subscript1superscript2superscriptsubscript2superscriptsubscript12subscript1superscript2superscriptsubscript2superscriptsubscript12 _1⤠i⤠nL(Ρ,t_i,t_i-1)= _1⤠i⤠nt_i-t_i-1% ⤠_1⤠i⤠n^(2)t_i^(2)-t_i-1^(2)= _1⤠i⤠n^(% 2)L(Ρ,t_i^(2),t_i-1^(2)),max1 ⤠i ⤠n L ( Ρ , titalic_i , titalic_i - 1 ) = max1 ⤠i ⤠n titalic_i - titalic_i - 1 ⤠max1 ⤠i ⤠n( 2 ) titalic_i( 2 ) - titalic_i - 1( 2 ) = max1 ⤠i ⤠n( 2 ) L ( Ρ , titalic_i( 2 ) , titalic_i - 1( 2 ) ) , and it follows from the definition of length and the triangle inequality that â˘(Ρti,Ρtiâ1)â¤Lâ˘(Ρ,ti,tiâ1)subscriptsubscriptsubscriptsubscript1subscriptsubscript1 d( _t_i, _t_i-1)⤠L(Ρ,t_i,t_i-1)sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) ⤠L ( Ρ , titalic_i , titalic_i - 1 ) for all i=1,âŚ,n1âŚi=1,âŚ,ni = 1 , ⌠, n. Therefore using (6), we have: max1â¤iâ¤nâĄâ˘(Ρti,Ρtiâ1)â¤Î´3â˘Câ˘Lâ˘(Ρ)â§Ďľ.subscript1subscriptsubscriptsubscriptsubscript13italic-Ďľ _1⤠i⤠n d( _t_i, _t_i-1)⤠δ3% CL(Ρ) Îľ.max1 ⤠i ⤠n sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) ⤠divide start_ARG δ end_ARG start_ARG 3 C L ( Ρ ) end_ARG â§ Ďľ . (7) Using âĎâ˘(z)â2=1subscriptnormitalic-Ď21\|Ď(z)\|_2=1âĽ Ď ( z ) âĽ2 = 1 for all z, Îłt=Ďâ˘(Ρt)subscriptitalic-Ďsubscript _t=Ď( _t)Îłitalic_t = Ď ( Ρitalic_t ), the upper bound by Ďľitalic-ϾξϾ in (7) to enable the substitution â¨Ďâ˘(Ρti),Ďâ˘(Ρtiâ1)âŠ2=gâ˘(â˘(Ρti,Ρtiâ1)2)subscriptitalic-Ďsubscriptsubscriptitalic-Ďsubscriptsubscript12superscriptsubscriptsubscriptsubscriptsubscript12 Ď( _t_i),Ď( _t_i-1) _2=g (% d( _t_i, _t_i-1)^2 )â¨ Ď ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT ) , Ď ( Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) âŠ2 = g ( sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )2 ), and taking a Taylor expansion of g about zero, we have: 12â˘âÎłtiâÎłtiâ1â2212superscriptsubscriptnormsubscriptsubscriptsubscriptsubscript122 12\| _t_i- _t_i-1\|_2^2divide start_ARG 1 end_ARG start_ARG 2 end_ARG ⼠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT - Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT âĽ22 =1ââ¨Îłti,Îłtiâ1âŠ2absent1subscriptsubscriptsubscriptsubscriptsubscript12 =1- _t_i, _t_i-1 _2= 1 - ⨠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT âŠ2 =1ââ¨Ďâ˘(Ρti),Ďâ˘(Ρtiâ1)âŠ2absent1subscriptitalic-Ďsubscriptsubscriptitalic-Ďsubscriptsubscript12 =1- Ď( _t_i),Ď( _t_i-1) % _2= 1 - â¨ Ď ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT ) , Ď ( Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) âŠ2 (8) =gâ˘(0)âgâ˘(dâ˘(Ρti,Ρtiâ1)2)absent0superscriptsubscriptsubscriptsubscriptsubscript12 =g(0)-g (d( _t_i, _t_i-1)^2 )= g ( 0 ) - g ( d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )2 ) =âgâ˛â˘(0)â˘(Ρti,Ρtiâ1)2âgâ˛â˘(ci)2â˘(Ρti,Ρtiâ1)4,absentsuperscriptâ˛0superscriptsubscriptsubscriptsubscriptsubscript12superscriptâ˛subscript2superscriptsubscriptsubscriptsubscriptsubscript14 =-g (0) d( _t_i, _t_i-1)^2- % g (c_i)2 d( _t_i, _t_i-1)^4,= - gⲠ( 0 ) sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )2 - divide start_ARG gⲠⲠ( citalic_i ) end_ARG start_ARG 2 end_ARG sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )4 , (9) where cisubscriptc_icitalic_i is some point in the interval [0,â˘(Ρti,Ρtiâ1)2]0superscriptsubscriptsubscriptsubscriptsubscript12[0, d( _t_i, _t_i-1)^2][ 0 , sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )2 ]. Under Hypothesis 2, we have gâ˛â˘(0)<0superscriptâ˛00g (0)<0gⲠ( 0 ) < 0. Then using (9) and lemma 1 with Îą=âÎłtiâÎłtiâ1â2subscriptnormsubscriptsubscriptsubscriptsubscript12Îą=\| _t_i- _t_i-1\|_2Îą = ⼠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT - Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT âĽ2 and β=â2â˘gâ˛â˘(0)â˘(Ρti,Ρtiâ1)2superscriptâ˛0subscriptsubscriptsubscriptsubscript1β= -2g (0) d( _t_i, _t_i-1)β = square-root start_ARG - 2 gⲠ( 0 ) end_ARG sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ), |âÎłtiâÎłtiâ1ââ2â˘gâ˛â˘(0)â˘(Ρti,Ρtiâ1)|â¤|gâ˛â˘(ci)|1/2â˘(Ρti,Ρtiâ1)2,normsubscriptsubscriptsubscriptsubscript12superscriptâ˛0subscriptsubscriptsubscriptsubscript1superscriptsuperscriptâ˛subscript12superscriptsubscriptsubscriptsubscriptsubscript12 |\| _t_i- _t_i-1\|- -2g (0) d(% _t_i, _t_i-1) |â¤|g (c_i)|^1/2 % d( _t_i, _t_i-1)^2,| ⼠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT - Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ⼠- square-root start_ARG - 2 gⲠ( 0 ) end_ARG sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) | ⤠| gⲠⲠ( citalic_i ) |1 / 2 sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )2 , so that |Sâ˘(Îł,)â2â˘gâ˛â˘(0)â˘Sâ˘(Ρ,)|2superscriptâ˛0 |S(Îł,T)- -2g (0)S(Ρ, % T) || S ( Îł , T ) - square-root start_ARG - 2 gⲠ( 0 ) end_ARG S ( Ρ , T ) | â¤âi=1n|âÎłtiâÎłtiâ1â2â2â˘gâ˛â˘(0)â˘(Ρti,Ρtiâ1)|absentsuperscriptsubscript1subscriptnormsubscriptsubscriptsubscriptsubscript122superscriptâ˛0subscriptsubscriptsubscriptsubscript1 ⤠_i=1^n |\| _t_i- _t_i-1\|_2-% -2g (0) d( _t_i, _t_i-1) |⤠âi = 1n | ⼠γitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT - Îłitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT âĽ2 - square-root start_ARG - 2 gⲠ( 0 ) end_ARG sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) | â¤âi=1n|gâ˛â˘(ci)|1/2â˘(Ρti,Ρtiâ1)2absentsuperscriptsubscript1superscriptsuperscriptâ˛subscript12superscriptsubscriptsubscriptsubscriptsubscript12 ⤠_i=1^n|g (c_i)|^1/2 d( _% t_i, _t_i-1)^2⤠âi = 1n | gⲠⲠ( citalic_i ) |1 / 2 sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT )2 â¤Câ˘(max1â¤iâ¤nâĄâ˘(Ρti,Ρtiâ1))â˘âi=1nâ˘(Ρti,Ρtiâ1).absentsubscript1subscriptsubscriptsubscriptsubscript1superscriptsubscript1subscriptsubscriptsubscriptsubscript1 ⤠C ( _1⤠i⤠n d( _t_i, _t_% i-1) ) _i=1^n d( _t_i, _t_i-1).⤠C ( max1 ⤠i ⤠n sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) ) âi = 1n sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) . â¤Câ˘Lâ˘(Ρ)â˘max1â¤iâ¤nâĄâ˘(Ρti,Ρtiâ1)â¤Î´3,absentsubscript1subscriptsubscriptsubscriptsubscript13 ⤠CL(Ρ) _1⤠i⤠n d( _t_i, _t_% i-1)⤠δ3,⤠C L ( Ρ ) max1 ⤠i ⤠n sansserif_d ( Ρitalic_t start_POSTSUBSCRIPT i end_POSTSUBSCRIPT , Ρitalic_t start_POSTSUBSCRIPT i - 1 end_POSTSUBSCRIPT ) ⤠divide start_ARG δ end_ARG start_ARG 3 end_ARG , where the final inequality uses (7). In summary, we have shown that δ(2)ââ|Sâ˘(Îł,)â2â˘gâ˛â˘(0)â˘Sâ˘(Ρ,)|â¤Î´3.formulae-sequencesuperscriptsubscript2â2superscriptâ˛03T_δ^(2) |S(% Îł,T)- -2g (0)S(Ρ,T) |⤠% δ3.Titalic_δ( 2 ) â T â | S ( Îł , T ) - square-root start_ARG - 2 gⲠ( 0 ) end_ARG S ( Ρ , T ) | ⤠divide start_ARG δ end_ARG start_ARG 3 end_ARG . (10) Now consider |Lâ˘(Îł)âSâ˘(Îł,â )|â |L(Îł)-S(Îł,¡) || L ( Îł ) - S ( Îł , â ) | as appears in (2). By the definition of Lâ˘(Îł)L(Îł)L ( Îł ), there exists δ(1)=(t0(1)=0,t1(1),âŚ,tn(1)(1)=Lâ˘(Ρ))superscriptsubscript1formulae-sequencesuperscriptsubscript010superscriptsubscript11âŚsuperscriptsubscriptsuperscript11T_δ^(1)=(t_0^(1)=0,t_1^(1),âŚ,t_n^(1)^(1)% =L(Ρ))Titalic_δ( 1 ) = ( t0( 1 ) = 0 , t1( 1 ) , ⌠, titalic_n( 1 )( 1 ) = L ( Ρ ) ) such that: Lâ˘(Îł)âδ3â¤Sâ˘(Îł,δ(1))â¤Lâ˘(Îł).3superscriptsubscript1L(Îł)- δ3⤠S(Îł,T_δ^(1))⤠L(% Îł).L ( Îł ) - divide start_ARG δ end_ARG start_ARG 3 end_ARG ⤠S ( Îł , Titalic_δ( 1 ) ) ⤠L ( Îł ) . By applying the triangle inequality to the summands in Sâ˘(Îł,δ(1))superscriptsubscript1S(Îł,T_δ^(1))S ( Îł , Titalic_δ( 1 ) ) we have for any TT with δ(1)âsuperscriptsubscript1T_δ^(1) _δ( 1 ) â T, Sâ˘(Îł,δ(1))â¤Sâ˘(Îł,)â¤Lâ˘(Îł)superscriptsubscript1S(Îł,T_δ^(1))⤠S(Îł,T)⤠L(Îł)S ( Îł , Titalic_δ( 1 ) ) ⤠S ( Îł , T ) ⤠L ( Îł ), hence δ(1)ââ|Lâ˘(Îł)âSâ˘(Îł,)|â¤Î´3.formulae-sequencesuperscriptsubscript1â3T_δ^(1) |L(% Îł)-S(Îł,T) |⤠δ3.Titalic_δ( 1 ) â T â | L ( Îł ) - S ( Îł , T ) | ⤠divide start_ARG δ end_ARG start_ARG 3 end_ARG . (11) Now consider â2â˘gâ˛â˘(0)â˘|Sâ˘(Ρ,â )âLâ˘(Ρ)|2superscriptâ˛0â -2g (0) |S(Ρ,¡)-L(Ρ) |square-root start_ARG - 2 gⲠ( 0 ) end_ARG | S ( Ρ , â ) - L ( Ρ ) | as appears in (4). By similar arguments to those used above do establish (11), there exists δ(3)superscriptsubscript3T_δ^(3)Titalic_δ( 3 ) such that δ(3)âââ2â˘gâ˛â˘(0)â˘|Lâ˘(Ρ)âSâ˘(Ρ,)|â¤Î´3.formulae-sequencesuperscriptsubscript3â2superscriptâ˛03T_δ^(3) -2g^% (0) |L(Ρ)-S(Ρ,T) |⤠δ3.Titalic_δ( 3 ) â T â square-root start_ARG - 2 gⲠ( 0 ) end_ARG | L ( Ρ ) - S ( Ρ , T ) | ⤠divide start_ARG δ end_ARG start_ARG 3 end_ARG . (12) With δâδ(1)âŞÎ´(2)âŞÎ´(3)âsubscriptsuperscriptsubscript1superscriptsubscript2superscriptsubscript3T_δ _δ^(1) _δ% ^(2) _δ^(3)Titalic_δ â Titalic_δ( 1 ) ⪠Titalic_δ( 2 ) ⪠Titalic_δ( 3 ), the implications (10), (11), (12) together tell us that the inequality (5) holds, and this completes the proof of the theorem. â Lemma 1. For any Îą,βâĽ00Îą,β⼠0Îą , β ⼠0, |Îąâβ|â¤|Îą2âβ2|1/2superscriptsuperscript2superscript212|Îą-β|â¤|Îą^2-β^2|^1/2| Îą - β | ⤠| Îą2 - β2 |1 / 2. Proof. W.l.o.g., assume ÎąâĽÎ˛ÎąâĽÎ˛Îą ⼠β. Using the triangle inequality for the Euclidean norm in â2superscriptâ2R^2blackboard_R2, Îą=(β2+Îą2âβ2)1/2â¤Î˛+(Îą2âβ2)1/2superscriptsuperscript2superscript2superscript212superscriptsuperscript2superscript212Îą=(β^2+Îą^2-β^2)^1/2â¤Î˛+(Îą^2-β^2)^% 1/2Îą = ( β2 + Îą2 - β2 )1 / 2 ⤠β + ( Îą2 - β2 )1 / 2, i.e., Îąâβâ¤(Îą2âβ2)1/2superscriptsuperscript2superscript212Îą-βâ¤(Îą^2-β^2)^1/2Îą - β ⤠( Îą2 - β2 )1 / 2 . â