Paper deep dive
Composed Historical Image Retrieval by Modeling Temporal Representations
Adrià Molina Rodríguez, Oriol Ramos Terrades, Josep Lladós Canet
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/20/2026, 4:56:27 AM
Summary
This paper introduces Temporally Decomposable Image Representations (TDIR), a representation learning framework that separates historical image embeddings into orthogonal categorical and temporal subspaces. By leveraging proxy sets and a transitive temporal loss, TDIR enables composed image retrieval where users can query for specific objects within target time periods, either via labels or by injecting temporal information from reference images without explicit date supervision.
Entities (7)
Relation Signals (5)
TDIR → enables → Composed Image Retrieval
confidence 95% · TDIR enables a class of transitive operations on embedding spaces... All theoretical properties are grounded and validated in the real-world problem of Composed Image Retrieval
TDIR → enforces → Temporal Transitivity
confidence 93% · To enforce temporal transitivity (Definition 2), the shifted residuals are pulled toward their target year proxies
TDIR → uses → Proxy Loss
confidence 92% · We define three proxy sets... Proxy loss. All objectives share a single form... Category loss... Direct temporal loss... Transitive temporal loss.
TDIR → decomposes → Historical Photographs
confidence 90% · TDIR... decomposes historical photographs into separate date and content components through orthogonal subspaces.
DEW → usedby → Müller et al.
confidence 85% · Müller et al. introduce the DEW benchmark and treat date estimation as a regression problem
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While time evolves linearly, the geometry of neural embedding spaces is inherently multi-dimensional, often chaotic, and difficult to interpret. In principle, one could constrain an embedding space to a single temporal dimension; however, such a reduction would sacrifice performance on downstream tasks, as one-dimensional embeddings cannot retain sufficient expressive capacity. This paper asks whether it is possible to learn representations that preserve temporal structure while remaining effective for image and object retrieval, and answers this question by building the mathematical foundations of such a system. We propose Temporally Decomposable Image Representations (TDIR), a representation learning algorithm that decomposes historical photographs into separate date and content components through orthogonal subspaces. We define and prove the conditions under which such a decomposition is achievable, characterize the error incurred when those conditions are only partially met, and show that orthogonality between temporal and categorical subspaces emerges naturally from the joint optimization, without requiring it to be imposed explicitly. Beyond its geometric properties, TDIR enables a class of transitive operations on embedding spaces: the temporal information of one image can be extracted and injected into the representation of another, with no label supervision required. All theoretical properties are grounded and validated in the real-world problem of Composed Image Retrieval on historical photographs, where a query simultaneously specifies object content and a target time period, either through labels or through example images. This in-the-wild setting serves as a concrete backing for the propositions we derive, offering an intuitive and interpretable way to navigate photographic archives while maintaining competitive performance in both date estimation and object retrieval.
Tags
Links
- Source: https://arxiv.org/abs/2608.18694v1
- Canonical: https://arxiv.org/abs/2608.18694v1
Trouble viewing inline? Open PDF directly →
Full Text
88,555 characters extracted from source content.
Expand or collapse full text
Composed Historical Image Retrieval Composed Historical Image Retrieval by Modeling Temporal Representations Adrià Molina Oriol Ramos Terrades Josep Lladós Abstract While time evolves linearly, the geometry of neural embedding spaces is inherently multi-dimensional, often chaotic, and difficult to interpret. In principle, one could constrain an embedding space to a single temporal dimension; however, such a reduction would sacrifice performance on downstream tasks, as one-dimensional embeddings cannot retain sufficient expressive capacity. This paper asks whether it is possible to learn representations that preserve temporal structure while remaining effective for image and object retrieval, and answers this question by building the mathematical foundations of such a system. We propose Temporally Decomposable Image Representations (TDIR), a representation learning algorithm that decomposes historical photographs into separate date and content components through orthogonal subspaces. We define and prove the conditions under which such a decomposition is achievable, characterize the error incurred when those conditions are only partially met, and show that orthogonality between temporal and categorical subspaces emerges naturally from the joint optimization, without requiring it to be imposed explicitly. Beyond its geometric properties, TDIR enables a class of transitive operations on embedding spaces: the temporal information of one image can be extracted and injected into the representation of another, with no label supervision required. All theoretical properties are grounded and validated in the real-world problem of Composed Image Retrieval on historical photographs, where a query simultaneously specifies object content and a target time period, either through labels or through example images. This in-the-wild setting serves as a concrete backing for the propositions we derive, offering an intuitive and interpretable way to navigate photographic archives while maintaining competitive performance in both date estimation and object retrieval. †email: amolina@cvc.uab.cat†email: oriolrt@cvc.uab.cat†email: josep@cvc.uab.cat†affiliation: Centre de Visió per Computador Universitat Autònoma de Barcelona Bellaterra, Catalonia †affiliation: Computer Science Department Universitat Autònoma de Barcelona, Bellaterra, Catalonia 1 Introduction Despite photography emerging late during the 19th century, it still comprises an important portion of archival data. More precisely, around 5-10% of archival material is estimated to be non-textual (see Appendix B), despite covering a significantly shorter historical span than printed or handwritten documents. Unlike traditional historical records, however, photographs encode most of their information visually, making their description through discrete metadata inherently restrictive. In this context, many Digital Humanities projects aim to improve access to non-textual collections through semantic search [Yuan et al.(2025)Yuan, Li, Wang, Liu, and Zhang] or automatic metadata completion [Garcia et al.(2020)Garcia, Renoust, and Nakashima]. Among such metadata descriptors, the date of a photograph is one of the most prominent task in historical photography analysis [Palermo et al.(2012)Palermo, Hays, and Efros, Müller et al.(2017)Müller, Springstein, and Ewerth, Stacchio et al.(2020)Stacchio, Angeli, Lisanti, Calanca, and Marfia, Paplhám and Franc(2026), Barancová et al.(2023)Barancová, Wevers, and van Noord], as it enables historians and social scientists to contextualise much of the information contained in the image. Temporal information alone, however, is insufficient for a complete archival description. Archivists must also contextualise the depicted objects within their historical period. In archival science terms [Peterson(2014)], the preservation interest of a document is partially determined by its probative value: the capacity of the image to uniquely illustrate an event, person, technology, or phenomenon. This requires understanding not only when a photograph was taken, but also how the objects it contains relate to their historical moment. This archival reasoning has direct implications for image analysis. Methods for automatic date estimation have evolved from low-level visual predictors [Palermo et al.(2012)Palermo, Hays, and Efros] toward semantic and object-centric approaches [Net et al.(2024)Net, Hernández, Molina, and Gómez]. We argue that the latter is closer to the reasoning process of archivists, who infer dates through the historical signatures of the objects appearing in the scene. Objects such as cars, clothes, or posters evolve at different temporal rhythms, allowing trained observers to identify inconsistencies or confirming regularities at a glance. This object-sensitive reasoning, however, is difficult to replicate in standard neural representations, where colour, texture, objects, and date tend to be entangled in ways that obscure the individual temporal evolution of each category. More critically for archival practice, current retrieval systems do not allow a user to query an archive by composing an object of interest with a target time period, as an archivist naturally would. As illustrated in Figure 1, the ideal system would support queries such as: “find images of jackets from the 1940s”, or more powerfully, “find images of cars that look contemporaneous to this photograph of a politician” — without requiring any date label for the reference image. how did cars look during 1956?Query-by-label1956cars✓ how did look during 1989?Query-by-example1989 how did look during ?Query-by-examples ✓ (a) (b) (c) Figure 1: Visual abstract of the three proposed representation objectives, showing composed image retrieval for query-by-label (a), query-by-example (b) and query-by-examples (c), in the latter, the idea is to inject the temporal information of an image into the object representation of the other query. To enable this, we propose Temporally Decomposable Image Representations (TDIR), a formal framework for representation learning in which the temporal and categorical components of an image embedding are separated into orthogonal subspaces. We formalise the conditions under which such a decomposition holds, prove that the required orthogonality emerges naturally from joint optimisation, and characterise the error incurred when those conditions are only partially met. These theoretical contributions are, to our knowledge, the first formal treatment of temporal decomposability in visual embeddings and stand independently of any specific application. Concretely, we establish two formal properties: • Temporally Decomposable Image Representation: an embedding space is temporally decomposable when any representation can be expressed as a linear combination of independent category and year vectors, such that each component exclusively encodes its respective factor of variation. • Joint Proxy Optimization: We propose a joint optimization which promotes TDIR in real case scenarios. Thus, the year centroid, interpreted as a displacement vector, can transport any representation to a different date without altering its categorical content. Beyond their theoretical interest, both properties admit a practically relevant instantiation: Composed Image Retrieval on in-the-wild historical photographs, where a query simultaneously specifies object content and a target time period. This provides empirical evidence that the derived properties yield interpretable behavior in applied settings beyond the theoretical formulation of the framework. 2 Related Work Early approaches to photographic date estimation relied on low-level visual cues such as colour histograms and film grain [Palermo et al.(2012)Palermo, Hays, and Efros], while more recent methods leverage semantic features. Crucially, however, none of these works formalise the desired properties of an ideal temporal representation. Müller et al. introduce the DEW benchmark and treat date estimation as a regression problem [Müller et al.(2017)Müller, Springstein, and Ewerth], without characterising the geometric structure that makes a representation well-suited for this task. Post-hoc analyses have revealed that temporal structure does emerge in general-purpose embeddings: In [Molina et al.(2022)Molina, Gomez, Ramos Terrades, and Lladós], the authors propose a loss function that can re-organise vision embeddings according to a fixed temporal criteria, this temporal geometry is observed to naturally emerge through CLIP pretraining [Sonthalia et al.(2026)Sonthalia, Uselis, and Oh], where embeddings carry rankable temporal information. Neither work describes how to induce such structure by design, nor what formal properties an ideally disentangled temporal representation should satisfy. TDIR addresses precisely this gap. A complementary line of work argues that date estimation should be grounded in the objects present in a scene, mirroring the reasoning of trained archivists. Ashida et al. [Ashida et al.(2021)Ashida, Jatowt, Doucet, and Yoshikawa] propose an object-centred ensemble restricted to human subjects, and Net et al. [Net et al.(2024)Net, Hernández, Molina, and Gómez] extend this idea to a broader set of object categories via a transformer-based architecture trained on DEW, introducing the DEW-B benchmark. However, both works treat object-centricity as a means to improve date prediction, not as a representational goal in its own right: object identity is discarded once the date is estimated. Furthermore, since [Net et al.(2024)Net, Hernández, Molina, and Gómez] requires training a dedicated image encoder per object category, the method is evaluated on only six categories, which raises scalability concerns to larger object sets. By contrast, TDIR jointly preserves both categorical and temporal factors within a single unified backbone, enabling retrieval along either dimension independently and scaling naturally to a large number of object categories. Composed Image Retrieval (CIR), which involves querying by combining a reference image with a modification attribute, has been studied in general-domain settings [Wu et al.(2021)Wu, Gao, Guo, Al-Halah, Rennie, Grauman, and Feris, Ray et al.(2023)Ray, Radenovic, Dubey, Plummer, Krishna, and Saenko], but existing historical benchmarks do not support this paradigm. DEW provides date annotations without object labels [Müller et al.(2017)Müller, Springstein, and Ewerth]; IMAGO and yearbook-based datasets [Salem et al.(2016)Salem, Workman, Zhai, and Jacobs, Stacchio et al.(2020)Stacchio, Angeli, Lisanti, Calanca, and Marfia, Paplhám and Franc(2026)] are restricted to faces and thus offer only a single semantic category. EUFCC-CIR proposed a composed retrieval dataset for cultural heritage collections [Net and Gomez(2024)], but it contains no temporal metadata, making category-date composition impossible. No existing benchmark simultaneously provides object-category and year-level annotations over a diverse photographic archive. 3 Methodology 3.1 Decomposable and Transitive Temporal Representations Training operates on tuples T=(xA,xB)T=(x_A,x_B) where each image x carries two attributes: a category c∈1,…,Cc∈\1,…,C\ and a year y∈1,…,Yy∈\1,…,Y\. In the tuples, cA=cB=c_A=c_B=c and yA≠yBy_A≠ y_B. The geometric construction of our method rests on a key structural requirement for the year subspace: that temporal displacements between dates behave consistently across all categories. Definition 1 (Temporal Independence of Classes). Given class and date centroids (μcμ^c, μyμ^y) and their corresponding labels (c and y), a given representation is said to be temporally independent from the object classes when p(μc,c,μy,y)=p(μc,c)⋅p(μy,y)p(μ^c,c,μ^y,y)=p(μ^c,c)· p(μ^y,y) (1) From this definition, the pair (μc,c)(μ^c,c) is independent of (μy,y)(μ^y,y): category centroids carry no information about dates, and date centroids carry no information about categories. Definition 2 (Temporal Transitivity). For any ordered pair of dates (A,B)(A,B), define the displacement vector ΔA→By=μyB−μyA∈ℝd ^y_A→ B=μ^y_B-μ^y_A ^d. The year subspace is temporally transitive if, for any triple (A,B,C)(A,B,C), ΔA→By+ΔB→Cy=ΔA→Cy. ^y_A→ B+ ^y_B→ C= ^y_A→ C. (2) This property is necessary for zero-label inference: it guarantees that a year residual extracted from any image acts as a consistent temporal operator when transplanted onto another image. We now introduce the proxy and residual structures that make this property learnable. 3.2 Proxy Sets Let f:→ℝdf:X ^d be a CNN encoder. Each image x carries two attributes: a category c∈1,…,Cc∈\1,…,C\ and a year y∈1,…,Yy∈\1,…,Y\. We seek a representation where these attributes occupy orthogonal subspaces, admitting the decomposition: v=f(x)=μc+μy,v=f(x)=μ^c+μ^y, (3) where μc,μy∈ℝdμ^c,μ^y ^d are random vectors acting as class centroids for object and temporal information. The ideal decomposition satisfies I(μc,μy)=0I(μ^c;\,μ^y)=0: the two components are mutually uninformative, and their sum captures the full label-relevant content of v. A proxy p∈p is a learnable vector of the same dimensionality as the embedding space, representing the centroid of a contrastive class [Movshovitz-Attias et al.(2017)Movshovitz-Attias, Toshev, Leung, Ioffe, and Singh]. Unlike dataset-derived centroids, proxies are optimized end-to-end: they are simultaneously pulled toward the embeddings of their assigned class and pushed away from all others, converging to dynamic but class-specific attractors in ℝdR^d. We define three proxy sets: Category proxies c=μ1c,…,μCcP^c=\μ^c_1,…,μ^c_C\. One proxy per object category. Each μcμ^c integrates over all years of category c, converging to a year-agnostic centroid. These operate on unnormalized embeddings, so both magnitude and direction contribute to proxy assignment. Year proxies y=μ1y,…,μYyP^y=\μ^y_1,…,μ^y_Y\. One proxy per year, shared across all categories. These operate on ℓ2 _2-normalized residuals u^=u/‖u‖ u=u/\|u\|, reducing the dot product to cosine similarity and encoding temporal information purely in the angle of the representation — a design choice validated by our ablation study (Table 3). Auxiliary displacement proxies y=k1y,…,kYyK^y=\k^y_1,…,k^y_Y\. One learnable displacement vector per year, shared across categories. These provide learned temporal offsets used during training to enforce transitivity (Definition 2). The set yK^y is discarded at inference. Under the model assumptions of Definition 1, the following decomposability result holds: Proposition 1 (Temporally Decomposable Image Representation). Under the model assumptions of Definition 1: I(μc,μy,c,y)=I(μc,c)+I(μy,y)I(μ^c,μ^y;\,c,y)=I(μ^c;\,c)+I(μ^y;\,y) (4) See the proof in Appendix A.2. This implies that, under the class independence assumption, a visual embedding v=f(x)v=f(x) can be temporally decomposable with respect to the proxy sets cP^c and yP^y. Because v=μy+μcv=μ^y+μ^c, the mutual information between the joint labels and the full representation decomposes as: I(v,c,y)=I(μc,c)+I(μy,y)I(v;\,c,y)=I(μ^c;\,c)+I(μ^y;\,y) (5) Ideally, the total information about both factors contained in v is fully decomposable into independent contributions: categorical information I(μc,c)I(μ^c;\,c) and temporal information I(μy,y)I(μ^y;\,y). In the ideal case, μcμ^c and μyμ^y are full descriptors of v with respect to (c,y)(c,y): knowing both proxies captures everything v encodes about category and year, with no label-relevant information remaining. In practice, however, there is some statistical entanglement between the appearance of certain objects and their respective dates. We therefore account for a residual signal ε∈ℝd ^d present in v: v=μc+μy+ϵv=μ^c+μ^y+ε (6) Proposition 2 (Practical Learnability of TDIRs). Under the TDIR framework, the discrepancy term δ≜I(ε,c,y)δ I( ;\,c,y) depends exclusively on the statistical structure of the training data and not on the choice of proxy vectors μcμ^c or μyμ^y: I(v,c,y)≤I(μc,c)+I(μy,y)+δwhereδ≜I(ε,c,y)I(v;\,c,y)≤ I(μ^c;\,c)+I(μ^y;\,y)+δ δ I( ;\,c,y) (7) Consequently, δ constitutes a data-dependent constant with respect to the optimization, and the category and year centroids can be learned freely to satisfy the orthogonality constraint ⟨μc,μy⟩=0 μ^c,\,μ^y =0 without affecting this error bound (see proof in Appendix A.3). A representation v becomes more temporally decomposable as δ decreases. This leaves open a potential circularity: Proposition 2 assumes orthogonality to guarantee free learning of μcμ^c and μyμ^y, yet this orthogonality has not been enforced. Theorem 1 will resolve this by showing that orthogonality emerges implicitly from the proposed training strategy. 3.3 Residual Vectors Given the proxy sets above, we build four residual vectors from each training tuple T=(xA,xB)T=(x_A,x_B) with shared category c and distinct years yA≠yBy_A≠ y_B. The four-step geometric construction is illustrated in Figure 2. Direct residuals. Subtracting the category centroid μcμ^c re-centers each cluster at the origin, revealing the temporal component (Figure 2b): rA=vA−μc,rB=vB−μc.r_A=v_A-μ^c, r_B=v_B-μ^c. (8) In rAr_A and rBr_B, the information necessary to infer the category has been removed. Consequently, even if temporal information were present in the category subspace, it is stripped from the representation and will not propagate to subsequent steps. Shifted residuals (swapping trick). To enforce temporal transitivity, we additionally form two cross-year residuals using the auxiliary displacement proxies yK^y: rA→B=(vA−μc)+kBy,rB→A=(vB−μc)+kAy.r_A→ B=(v_A-μ^c)+k^y_B, r_B→ A=(v_B-μ^c)+k^y_A. (9) The vector rA→Br_A→ B displaces image A’s temporal residual toward year yBy_B; if the temporal manifold is truly transitive, this shifted residual should be indistinguishable from rBr_B — and therefore close to μyBμ^y_B. The symmetric argument holds for rB→Ar_B→ A. All four residuals are ℓ2 _2-normalized before being passed to the year loss: r^∗←r∗/‖r∗‖. r_*← r_*/\|r_*\|. (10) This normalization encodes temporal information purely as angular structure in the shared subspace (Figure 2c–d). Category Clustering (a) Centering- ProxyAProxy_A- ProxyBProxy_B- ProxyCProxy_C (b) Date Proxy Optimization (c) Category Translation+PA+P_A+PB+P_B+PC+P_C (d) Figure 2: Training pipeline. (a) Proxy loss forms year-agnostic category clusters. (b) Subtracting μcμ^c centers each cluster into a shared temporal subspace. (c) Year proxies are optimized on normalized residuals via cosine distance. (d) The swapping trick translates residuals across years, enforcing temporal transitivity. 3.4 Proxy Losses and Joint Objective Proxy loss. All objectives share a single form [Movshovitz-Attias et al.(2017)Movshovitz-Attias, Toshev, Leung, Ioffe, and Singh]. For an embedding u with ground-truth proxy μ∗∈μ^* , the proxy loss is: ℒ(u,μ∗,)=−logexp(u⊤μ∗)∑p∈exp(u⊤p).L(u,\,μ^*;\,P)=- (u μ^*) _p (u p). (11) This is a softmax cross-entropy over proxy assignments: it maximizes the score of the correct proxy μ∗μ^* relative to all others in P, pulling u into the neighborhood of μ∗μ^* while repelling it from every competing proxy. Category loss. The category subspace is optimized directly on the full (unnormalized) embeddings: ℒc(T)=ℒ(vA,μc,c)+ℒ(vB,μc,c).L^c(T)=L(v_A,\,μ^c;\,P^c)\;+\;L(v_B,\,μ^c;\,P^c). (12) Each proxy μcμ^c integrates over all years of category c (Figure 2a). At this stage, date information may still be present in the embedding through spurious correlations such as textures or color; the subsequent residual construction removes it. Direct temporal loss. The year subspace is optimized on the normalized direct residuals: ℒdirecty(T)=ℒ(r^A,μyA,y)+ℒ(r^B,μyB,y).L^y_direct(T)=L( r_A,\,μ^y_A;\,P^y)\;+\;L( r_B,\,μ^y_B;\,P^y). (13) This makes the year subspace estimable, but not yet geometrically consistent across categories. Transitive temporal loss. To enforce temporal transitivity (Definition 2), the shifted residuals are pulled toward their target year proxies — the year each residual has been displaced toward, rather than the image’s own year: ℒtransy(T)=ℒ(r^A→B,μyB,y)+ℒ(r^B→A,μyA,y).L^y_trans(T)=L( r_A→ B,\,μ^y_B;\,P^y)\;+\;L( r_B→ A,\,μ^y_A;\,P^y). (14) This cannot be satisfied unless the displacement kyk^y is consistent across all categories — that is, unless the year residuals of any two images from any two categories differ by the same offset in the temporal subspace. This swapping trick directly enforces Definition 2. Total year loss and joint objective. Combining direct and transitive terms (Figure 2d): ℒy(T)=ℒ(r^A,μyA,y)+ℒ(r^B,μyB,y)⏟direct: r≈μy+ℒ(r^A→B,μyB,y)+ℒ(r^B→A,μyA,y)⏟transitive: r+ky≈μy.L^y(T)= L( r_A,\,μ^y_A;\,P^y)+L( r_B,\,μ^y_B;\,P^y)_direct: r≈μ^y+ L( r_A→ B,\,μ^y_B;\,P^y)+L( r_B→ A,\,μ^y_A;\,P^y)_transitive: r+k^y≈μ^y. (15) The total minimized loss is: ℒ=ℒc(T)+ℒy(T).L=L^c(T)+L^y(T). (16) Theorem 1 (Emergent TDIR). Let ℒcL^c and ℒyL^y be the proxy loss functions defined over an embedding space ℝdR^d. Then: Part I (exact case). If I(ϵ,c,y)=0I(ε\,;\,c,y)=0, joint minimization of ℒ=ℒc+ℒyL=L^c+L^y implies TDIR: minθℒc+ℒy⟹p(μc,μy,c,y)=p(μc,c)⋅p(μy,y) _θ\,L^c+L^y p(μ^c,μ^y,c,y)=p(μ^c,c)· p(μ^y,y) (17) Part I (asymptotic case). If I(ϵ,c,y)≠0I(ε\,;\,c,y)≠ 0, joint minimization implies asymptotic TDIR: minθℒc+ℒy⟹⟨μc,μy⟩→O(I(ϵ,c,y)‖μy‖2)∀c,y _θ\,L^c+L^y μ^c,\,μ^y → O\! ( I(ε\,;\,c,y)\|μ^y\|^2 ) ∀\,c,\,y (18) See proof in Appendix A.4. Note that Proposition 2 states that μyμ^y and μcμ^c can be learned under the assumption that temporal and categorical centroids are orthogonal. Although this might seem a strong constraint, Theorem 1 shows that the joint optimization of Eq. (16) imposes this orthogonality as an emergent property of the centering and swapping tricks. It is therefore formally guaranteed that Algorithm 1 can yield TDIR in both the ideal and practical case, through an implicit orthogonality regularization arising from the swapping of temporal variables across common object categories. Algorithm 1 TDIR Training Loop 1: Initialize: c∈ℝC×dP^c ^C× d, y∈ℝY×dP^y ^Y× d, ∈ℝY×dK ^Y× d 2: for epoch e=1e=1 …E do 3: for (xA,xB,c)∈train(x_A,x_B,c) _train do 4: vA←f(xA)v_A← f(x_A); vB←f(xB)v_B← f(x_B) forward pass 5: ℒc←ℒ(vA,c,c)+ℒ(vB,c,c)L^c (v_A,c;P^c)+L(v_B,c;P^c) category subspace 6: rA=vA−μccr_A=v_A-μ^c_c; rB=vB−μccr_B=v_B-μ^c_c center by category (decomposable) 7: rA→B=vA+kB−μccr_A→ B=v_A+k_B-μ^c_c; rB→A=vB+kA−μccr_B→ A=v_B+k_A-μ^c_c swapping trick (transitive) 8: r^∗←r∗/‖r∗‖ r_*← r_*/\|r_*\| normalize for angular encoding 9: ℒy=ℒ(r^A,μyA,y)+ℒ(r^B,μyB,y)+L^y=L( r_A,μ^y_A;P^y)+L( r_B,μ^y_B;P^y) 1.99997pt+ estimable temporal embedding +ℒ(r^A→B,μyB,y)+ℒ(r^B→A,μyA,y)+ 1.84995ptL( r_A→ B,μ^y_B;P^y)+L( r_B→ A,μ^y_A;P^y) transitive temporal embedding 10: f,c,y,←f,c,y,−α∇(ℒc+ℒy)\f,P^c,P^y,K\←\f,P^c,P^y,K\-α∇(L^c+L^y) backward pass 11: end for 12: end for 3.5 Inference and Composed Historical Image Retrieval At inference yK^y is discarded. The three retrieval modes (Figure 3) follow directly from the decomposition f(x)=v≈μc+μyf(x)=v≈μ^c+μ^y. Label-based (Figure 3a). At inference, we construct a database of test embeddings f(xi)i=1N\f(x_i)\_i=1^N. As in any common embedding-based retrieval, we construct a query vector q∈ℝdq ^d and returning its nearest neighbors: x(1),x(2),…,x(k)=top-kxif(xi)⊤q‖f(xi)‖‖q‖.\x_(1),x_(2),…,x_(k)\= x_itop-k\; f(x_i) q\|f(x_i)\|\|q\|. (19) By the decomposability property, a query targeting category c at year y is simply the sum of their respective proxies: q=μc+μy.q=μ^c+μ^y. (20) Since ⟨μc,μy⟩=0 μ^c,μ^y =0, the query lies precisely at the intersection of both subspaces, and the retrieved images x(i)\x_(i)\ are those whose embeddings are simultaneously close to μcμ^c and μyμ^y — that is, images of category c from year y. Image + label (Figure 3b). When the target year y is known but no category label is provided, we replace μcμ^c with the actual image embedding: q=f(xc)+μy.q=f(x_c)+μ^y. (21) This is strictly more expressive than the label-based mode: rather than retrieving images close to the category centroid μcμ^c, the query anchors to the instance-specific content of xcx_c — including the ε residual — while the year direction is fully determined by μyμ^y. Retrieved images thus share the particular visual characteristics of xcx_c, translated to year y. Image + image, zero-label (Figure 3c) The most powerful inference mode requires no label information whatsoever. The user provides two images: a category image xcx_c — whose year is entirely unknown and irrelevant — and a temporal reference xyx_y, from which we wish to borrow the year. Step 1: Infer the category of xyx_y. Since no category label is provided, we assign xyx_y to its nearest category proxy: c^=argminc′‖f(xy)−μc′‖2. c= c \;\|f(x_y)-μ^c \|_2. (22) Step 2: Extract the temporal information Subtracting the inferred category proxy strips the category information from f(xy)f(x_y), leaving a category-agnostic vector: ry=f(xy)−μc^≈μy+ε,r_y=f(x_y)-μ c\;≈\;μ^y+ , (23) where μyμ^y is the (unknown to the user) year proxy and ε is the instance-specific residual. Crucially, neither μyμ^y nor the year of xyx_y need to be known — the vector ryr_y is the temporal information. Step 3: Temporal transplant. We add ryr_y to the category embedding of xcx_c: q=f(xc)+ry.q=f(x_c)+r_y. (24) By temporal transitivity (Definition 2), this displaces f(xc)f(x_c) along the temporal manifold toward the year of xyx_y, while preserving its category direction. The k-nearest neighbors of q are thus images of the same category as xcx_c, at the same period as xyx_y. ptraincp^c_train19201920f()f\! (\, [height]images/inference/subimages/imageA.png\, )19301930+p1970y+p^y_1970 (a) f()f\! (\, [height]images/inference/subimages/train_6.png\, )19701970f()f\! (\, [height]images/inference/subimages/imageA.png\, )19301930+p1930y+p^y_1930 (b) f()f\! (\, [height]images/inference/subimages/train_6.png\, )19701970f()f\! (\, [height]images/inference/subimages/imageA.png\, )19301930+f()−pfacec+\,f\! (\, [height]images/inference/subimages/head_0.png\, )- p c_ face (c) Figure 3: Inference modes. (a) Label-based retrieval via proxy addition. (b) Image-guided retrieval with a target year label. (c) Zero-label retrieval: the year residual of xyx_y is extracted and transplanted onto f(xc)f(x_c). 4 Experimental Set-Up 4.1 Dataset For the correct application and benchmarking of our proposed problem set-up, it is for us required to utilize a dataset where both the date and object annotations are available. For doing so, we take advantage of the Date Estimation in The Wild dataset (DEW) [Müller et al.(2017)Müller, Springstein, and Ewerth] with object-specific detections by using the same criteria as in [Net et al.(2024)Net, Hernández, Molina, and Gómez]. In this Section, we detail the dataset and how it differs from the original DEW images. DEW Dataset The Date Estimation in The Wild dataset contains 1M natural-scene photographs from 1930 to 1999, the date information has an excellent granularity of a year per photo. However, the images are natural scenes and, therefore, contain a variety of different objects which hinders the applicability of object-specific representations. Object-Centric DEW Dataset In [Net et al.(2024)Net, Hernández, Molina, and Gómez], the authors propose a detection pipeline to crop objects detected with a minimum 10000px resolution by using the DETR [Carion et al.(2020)Carion, Massa, Synnaeve, Usunier, Kirillov, and Zagoruyko] model. Although the object-level detection (bounding boxes) are publicly available, we note that the authors use a constrained set of 6 categories, which in our case could trivialize the object-sensitive subspace. Proposed Dataset To solve the aforementioned issue, we use the same resolution criteria, but utilize OWLv2 [Minderer et al.(2023)Minderer, Gritsenko, and Houlsby] to detect objects from all possible categories in the DETR HuggingFace implementation [Hugging Face(2026a)] with a 30% detection threshold11 1 This is a standard threshold when using OWL [Hugging Face(2026b)].. This leads us to 7,239,083 object detections of medium-to-high resolution of 586 objects and date-specific annotations from DEW (see Table 1). In this article we will focus on the sub-set of top-50 most predominant objects and 4,478,887 detections (see Appendix C). But the complete set of detections, which can be utilized to expand the presented method or for many other applications, is available to download22 2 Contact amolina@cvc.uab.cat with the string “[TDIR] - Access to the database” as subject.. The test partition follows the same image selection as the original DEW dataset, from which we separate the detections according to the image source. 1930s 1940s 1950s 1960s 1970s 1980s 1990s Jacket Poster Car Building Table 1: Example images and categories from the dataset. 4.2 Implementation Details The practical implementation of the Temporally Decomposable Image Representations has been evaluated using several backbone architectures in order to assess the robustness of the proposed methodology across both convolutional and transformer-based visual representations. In particular, experiments have been conducted with ConvNeXt-Base [Liu et al.(2022)Liu, Mao, Wu, Feichtenhofer, Darrell, and Xie], Vision Transformer B/32 (ViT-B/32) [Dosovitskiy et al.(2020)Dosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, et al.], ResNet50 [He et al.(2016)He, Zhang, Ren, and Sun], and VGG19-BN [Simonyan and Zisserman(2014)], all initialized with ImageNet pre-trained weights [Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei] using the official PyTorch implementations [PyTorch Contributors(2026)]. For all architectures, the original classification head was removed and replaced with a projection layer producing a fixed embedding dimension of 1024. This unified embedding size ensures a fair comparison between architectures and allows all models to operate under the same metric-learning framework. Unlike previous approaches proposed for solving the DEW dataset [Müller et al.(2017)Müller, Springstein, and Ewerth] in an object-centric manner [Net et al.(2024)Net, Hernández, Molina, and Gómez], which requires ensembles of specialized models for each semantic category, the proposed methodology employs a single unified backbone capable of handling all categories jointly. This significantly scales well with the number of object categories. All backbone architectures were trained using the same optimization setup to ensure experimental consistency. The AdamW optimizer jointly optimizes the backbone parameters, the learnable category proxies, and the auxiliary temporal embeddings (f, μ∗μ^*, and kyk^y) with a learning rate of 10−410^-4 during 10 epochs, using the ProxyNCALoss proposed by Movshovitz-Attias et al. [Movshovitz-Attias et al.(2017)Movshovitz-Attias, Toshev, Leung, Ioffe, and Singh] via the Pytorch-Metric-Learning library [Musgrave et al.(2020)Musgrave, Belongie, and Lim], with no explicit triplet mining. Training was conducted on a single NVIDIA A40 GPU. Due to the large number of object detections (approximately 4M cropped sub-images across 50 object categories), each epoch requires approximately 15 hours of computation, with GPU utilization analysis confirming that the available resources are fully saturated throughout training. Table 2: Comparison across backbone architectures and the CLIP baseline (K=10). Label+Label Image+Label Image+Image Backbone Obj. P@K Date P@K Obj. P@K Date P@K Obj. P@K Date P@K CLIP (arithmetic, see Appendix D.1) .152 .115 .387 .076 .389 .102 CLIP (prompting, see Appendix D.2) .415 .105 .372 .079 .520 .081 VGG + TDIR .397 .712 .388 .579 .669 .234 ResNet + TDIR .489 .803 .444 .665 .721 .306 ConvNeXt + TDIR .443 .824 .416 .664 .725 .393 ViT + TDIR .413 .863 .355 .708 .646 .429 4.3 Evaluation Protocol We evaluate each of the three inference modes (Section 3.5) using a unified set of metrics built around two axes: categorical fidelity (does retrieval respect the object class?) and temporal fidelity (does retrieval respect the target date?). As a standard practice in date estimation [Palermo et al.(2012)Palermo, Hays, and Efros, Müller et al.(2017)Müller, Springstein, and Ewerth], we consider a correct date classification whenever the retrieved date is within the same lustrum (5 years period). For each inference mode, we retrieve the top-K images for every valid query and report two complementary precision metrics: Object Precision@K Precision@K :#xi:ci=cK, : \#\x_i:c_i=c\K, (25) Date Precision@K Precision@K :#xi:yi≈yK, : \#\x_i:y_i≈ y\K, (26) where yi≈y_i≈ y denotes temporal correctness within the same lustrum (5-year window). Given that the DEW dataset contains images from 1930-1999, the random baseline is settled at 100×170/5=7.14%100× 170/5=7.14\% chance of returning a correctly classified image in terms of the date estimation. In the label-based mode, c and y are taken directly from the query labels; in the image-based modes, c is determined by c(xc)c(x_c) and y by the target year or the transferred residual ryr_y (see Section 3.5). 5 Results In this section, we present an exhaustive evaluation of our framework in a highly applied setup. First, in Table 2, we compare several widely used Computer Vision backbones against two CLIP baselines for reference (see Appendix D for the baselines implementation). We observe that our proposed approach follows the general trends observed in computer vision, with ConvNeXt and ViT emerging as the top-performing and overall comparable architectures. The CLIP baselines prove to be competitive on retrieving the correct category, but are not sufficiently sensitive to temporal information despite its demonstrated sensitivity to image dates in [Sonthalia et al.(2026)Sonthalia, Uselis, and Oh] and [Barancová et al.(2023)Barancová, Wevers, and van Noord]. Label+Label Image+Label Image+Image Swap Aux Norm Obj. P@K Date P@K Obj. P@K Date P@K Obj. P@K Date P@K ✗ ✓ ✓ .362 .712 .342 .568 .683 .358 ✓ ✗ ✓ .355 .822 .343 .676 .691 .354 ✓ ✓ ✗ .358 .816 .349 .675 .708 .362 ✓ ✓ ✓ .443 .824 .416 .664 .725 .393 Table 3: Ablation study (K=10, ConvNext) for our method including the usage of the swapping trick, auxiliary proxies and normalization of the temporal component. Figure 4: Calibration error bars when using query-by-examples year injection. Table 3 presents an ablation study. A key observation from the experiment is that the proposed swapping trick not only improves transitivity, as expected, but also leads to better date and object embedding subspaces. This is reflected in the improved performance observed not only for image-based queries but also for purely label-based inference, which does not rely on any transitive property induced by injecting a date into a temporal embedding. This observation partially supports the claim in Corollary 4 presented in Appendix A.4, heavily relies on the inclusion of a swapping term and states that this combination of losses imposes an implicit orthogonality regularization at the optimum. Category Date Top-1 Top-2 Top-3 Top-4 Top-5 Top-6 Poster 1975 1980 Table 4: Qualitative results with category error and temporal error. The inclusion of an auxiliary term (k∗k_*), instead of directly using the temporal proxy itself (μ∗yμ^y_*) in the swapping trick step, plays an important role in Image+Image inference, yielding significant gains in properly structuring the embedding space. However, when performing Image+Label inference, only a marginal performance gap is observed, likely due to the train–test mismatch at inference time caused by relying exclusively on the learned centroids (μ∗yμ^y_*). Lastly, the normalization of the temporal component appears to be the least impactful, although it provides a slight performance boost in Image+Image date estimation. Curiously, performance in the Label+Label object retrieval setup is maximized if and only if all three components are present in the training regime. Figure 4 reports object and date precision as a function of source year, target year, and their absolute difference. Object precision remains stable across all three axes, confirming that temporal transplantation does not corrupt the categorical content of the query: the category subspace is effectively insulated from the temporal injection. Date precision, by contrast, reveals a clear dependency on the target year and on the source-to-target gap, but notably not on the source year itself. This asymmetry rules out a straightforward directional explanation: if the error were caused by the displacement vector μyμ^y acting unidirectionally, pushing representations forward or backward in time indiscriminately, one would expect a monotonic pattern with respect to direction. Instead, the error follows a U-shaped curve over the target year axis, with decades at the extremes of the temporal range being harder to reach regardless of where the source lies. Furthermore, precision degrades consistently as the temporal gap widens, suggesting that the optimisation does not perfectly generalise large temporal displacements. Taken together, these findings indicate that the current additive translation mechanism, while effective at short and moderate temporal distances, is an approximation whose fidelity diminishes with displacement magnitude. A natural direction for future work is to replace the additive operator with a more expressive, potentially non-linear mechanism capable of modelling large temporal jumps more faithfully. Although the numerical results may appear modest, a qualitative inspection in Table 4 reveals that the method generally behaves as expected. While errors in the predicted years are present, they mostly correspond to inaccuracies within the same lustrum rather than severe temporal artifacts (see Appendix C for a detailed analysis). This behavior may stem from two limitations. First, the temporal subspace is constrained to be object-agnostic, thus enforcing the same temporal displacement across all object categories. As introduced in Section 1, not every object evolves at the same pace, which may inherently limit the achievable performance under this design choice. In many cases, such as the fourth row, the temporal transplant is qualitatively correct, but the retrieved samples exhibit category ambiguity due to images lying near the boundary between two classes. In that example, the target category is “poster”, yet images belonging to the “men” category are retrieved while still exhibiting the correct temporal attribution. As discussed, the error tends to increase with temporal distance, which can be observed in the fifth row, where the direction of the temporal shift is correct but the magnitude of the displacement is insufficient. Lastly, in some cases, such as the final row, the target category may simply contain too few old samples in the database because the object itself is inherently modern. As a result, the retrieved samples belong to the correct category but remain more recent than desired. 6 Conclusions In this paper, we have presented TDIR, a representation learning framework that formalises the decomposition of historical image embeddings into orthogonal temporal and categorical subspaces. We have established the theoretical conditions under which such a decomposition holds, proved that the required orthogonality emerges naturally from joint optimisation, and characterised the error incurred when those conditions are only partially met. To our knowledge, this constitutes the first formal treatment of temporal decomposability in visual embeddings, addressing a limitation of prior work that either observes temporal structure post-hoc or exploits it implicitly without a principled formulation. Beyond the theoretical contributions, we have grounded the framework in the novel problem of Composed Historical Image Retrieval, where a query simultaneously specifies object content and a target time period. This setting reflects a natural and practically relevant archival task, yet no existing benchmark supported its evaluation. We address this by extending the DEW dataset with object-level detections across 50 categories, enabling the first evaluation of compositional retrieval over a diverse photographic archive covering seven decades. Limitations include the assumption of a shared temporal displacement across all object categories, which may not reflect the uneven pace at which different objects evolve historically, and the use of an additive transplantation operator whose fidelity diminishes at large temporal distances. Future work will explore non-linear temporal operators and category-specific temporal subspaces, as well as the extension of the benchmark to support zero-shot evaluation over unseen object categories. Acknowledgments This work has been partially supported by the Spanish project PID2024-157778OB-I00, Ministerio de Ciencia e Innovación, the Departament de Cultura of the Generalitat de Catalunya, and the CERCA Program. Adrià Molina is funded with the PRE2022-101575 grant provided by MCIN / AEI / 10.13039 / 501100011033 and by ERDF/EU. References [Archives Nationales, France()] Archives Nationales, France. Corpus de documents numérisés des Archives Nationales. https://w.data.gouv.fr/datasets/corpus-de-documents-numerises-des-archives-nationales. Last visited: May 18, 2026. [Ashida et al.(2021)Ashida, Jatowt, Doucet, and Yoshikawa] Shota Ashida, Adam Jatowt, Antoine Doucet, and Masatoshi Yoshikawa. Determining image age with rank-consistent ordinal classification and object-centered ensemble. In Proceedings of the 2nd ACM International Conference on Multimedia in Asia, pages 1–8, 2021. [Barancová et al.(2023)Barancová, Wevers, and van Noord] Alexandra Barancová, Melvin Wevers, and Nanne van Noord. Blind dates: examining the expression of temporality in historical photographs. arXiv preprint arXiv:2310.06633, 2023. [Carion et al.(2020)Carion, Massa, Synnaeve, Usunier, Kirillov, and Zagoruyko] Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. End-to-end object detection with transformers. In European conference on computer vision, pages 213–229. Springer, 2020. [Deng et al.(2009)Deng, Dong, Socher, Li, Li, and Fei-Fei] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. [Dosovitskiy et al.(2020)Dosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, et al.] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [Garcia et al.(2020)Garcia, Renoust, and Nakashima] Noa Garcia, Benjamin Renoust, and Yuta Nakashima. Contextnet: representation and exploration for painting classification and retrieval in context. International Journal of Multimedia Information Retrieval, 9(1):17–30, 2020. [He et al.(2016)He, Zhang, Ren, and Sun] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 770–778, 2016. [Hugging Face(2026a)] Hugging Face. Detr. https://huggingface.co/docs/transformers/en/model_doc/detr, 2026a. Transformers documentation. Last visited: 2026-05-18. [Hugging Face(2026b)] Hugging Face. Owl-vit. https://huggingface.co/docs/transformers/model_doc/owlvit, 2026b. Transformers documentation. Last visited: 2026-05-18. [Ilharco et al.(2021)Ilharco, Wortsman, Wightman, Gordon, Carlini, Taori, Dave, Shankar, Namkoong, Miller, Hajishirzi, Farhadi, and Schmidt] Gabriel Ilharco, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, Vaishaal Shankar, Hongseok Namkoong, John Miller, Hannaneh Hajishirzi, Ali Farhadi, and Ludwig Schmidt. Openclip, July 2021. URL https://doi.org/10.5281/zenodo.5143773. If you use this software, please cite it as below. [Library of Congress()] Library of Congress. Library of Congress – Global Search. https://w.loc.gov/search/?all=true. Last visited: May 18, 2026. [Liu et al.(2022)Liu, Mao, Wu, Feichtenhofer, Darrell, and Xie] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11976–11986, 2022. [Minderer et al.(2023)Minderer, Gritsenko, and Houlsby] Matthias Minderer, Alexey Gritsenko, and Neil Houlsby. Scaling open-vocabulary object detection. Advances in Neural Information Processing Systems, 36:72983–73007, 2023. [Ministerio de Cultura, España()] Ministerio de Cultura, España. PARES: Portal de Archivos Españoles – Estadísticas. https://pares.cultura.gob.es/estadisticas.html. Last visited: May 18, 2026. [Ministerio de Cultura, España(2025)] Ministerio de Cultura, España. Anuario de Estadísticas Culturales 2025. https://w.cultura.gob.es/dam/jcr:daa6c0f8-9abb-48a8-8af3-7d54746d6e4a/anuario-de-estadisticas-culturales-2025.pdf, 2025. p. 319. Last visited: May 18, 2026. [Molina et al.(2022)Molina, Gomez, Ramos Terrades, and Lladós] Adrià Molina, Lluis Gomez, Oriol Ramos Terrades, and Josep Lladós. A generic image retrieval method for date estimation of historical document collections. In International Workshop on Document Analysis Systems, pages 583–597. Springer, 2022. [Movshovitz-Attias et al.(2017)Movshovitz-Attias, Toshev, Leung, Ioffe, and Singh] Yair Movshovitz-Attias, Alexander Toshev, Thomas K Leung, Sergey Ioffe, and Saurabh Singh. No fuss distance metric learning using proxies. In Proceedings of the IEEE international conference on computer vision, pages 360–368, 2017. [Müller et al.(2017)Müller, Springstein, and Ewerth] Eric Müller, Matthias Springstein, and Ralph Ewerth. “when was this picture taken?”–image date estimation in the wild. In European Conference on Information Retrieval, pages 619–625. Springer, 2017. [Musgrave et al.(2020)Musgrave, Belongie, and Lim] Kevin Musgrave, Serge Belongie, and Ser-Nam Lim. A metric learning reality check. In European Conference on Computer Vision, pages 681–699. Springer, 2020. [Net and Gomez(2024)] Francesc Net and Lluis Gomez. Eufcc-cir: A composed image retrieval dataset for glam collections. In European Conference on Computer Vision, pages 196–211. Springer, 2024. [Net et al.(2024)Net, Hernández, Molina, and Gómez] Francesc Net, Núria Hernández, Adriá Molina, and Lluis Gómez. A transformer-based object-centric approach for date estimation of historical photographs. In European Conference on Information Retrieval, pages 137–150. Springer, 2024. [Palermo et al.(2012)Palermo, Hays, and Efros] Frank Palermo, James Hays, and Alexei A Efros. Dating historical color images. In European Conference on Computer Vision, pages 499–512. Springer, 2012. [Paplhám and Franc(2026)] Jakub Paplhám and Vojtěch Franc. Photo dating by facial age aggregation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 8103–8112, 2026. [Peterson(2014)] Trudy Huskamp Peterson. The Probative Value of Archival Documents. Swisspeace Bern, 2014. [PyTorch Contributors(2026)] PyTorch Contributors. Torchvision models documentation. https://docs.pytorch.org/vision/main/models.html, 2026. Accessed: 2026-05-11. [Radford et al.(2021)Radford, Kim, Hallacy, Ramesh, Goh, Agarwal, Sastry, Askell, Mishkin, Clark, et al.] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021. [Ray et al.(2023)Ray, Radenovic, Dubey, Plummer, Krishna, and Saenko] Arijit Ray, Filip Radenovic, Abhimanyu Dubey, Bryan Plummer, Ranjay Krishna, and Kate Saenko. Cola: A benchmark for compositional text-to-image retrieval. Advances in Neural Information Processing Systems, 36:46433–46445, 2023. [Salem et al.(2016)Salem, Workman, Zhai, and Jacobs] Tawfiq Salem, Scott Workman, Menghua Zhai, and Nathan Jacobs. Analyzing human appearance as a cue for dating images. In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1–8. IEEE, 2016. [Schuhmann(2022)] C. et al. Schuhmann. Laion-5b: An open large-scale dataset for training next generation image-text models. In Advances in Neural Information Processing Systems, 2022. [Simonyan and Zisserman(2014)] Karen Simonyan and Andrew Zisserman. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [Sonthalia et al.(2026)Sonthalia, Uselis, and Oh] Ankit Sonthalia, Arnas Uselis, and Seong Joon Oh. On the rankability of visual embeddings. Advances in Neural Information Processing Systems, 38:66169–66203, 2026. [Stacchio et al.(2020)Stacchio, Angeli, Lisanti, Calanca, and Marfia] Lorenzo Stacchio, Alessia Angeli, Giuseppe Lisanti, Daniela Calanca, and Gustavo Marfia. Imago: A family photo album dataset for a socio-historical analysis of the twentieth century. arXiv preprint arXiv:2012.01955, 2020. [Wu et al.(2021)Wu, Gao, Guo, Al-Halah, Rennie, Grauman, and Feris] Hui Wu, Yupeng Gao, Xiaoxiao Guo, Ziad Al-Halah, Steven Rennie, Kristen Grauman, and Rogerio Feris. Fashion iq: A new dataset towards retrieving images by natural language feedback. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pages 11307–11317, 2021. [Yuan et al.(2025)Yuan, Li, Wang, Liu, and Zhang] Hua Yuan, Yuhan Li, Baohui Wang, Kaixuan Liu, and Junjie Zhang. Knowledge graph-based intelligent question answering system for ancient chinese costume heritage. npj Heritage Science, 13(1):198, 2025. Supplementary Material This document contains the supplementary material for the BMVC submission: Composed Historical Image Retrieval by Modeling Temporal Representations The supplementary material is organized as follows: Section Content A Derivations and Demonstrations B Photographic Material Calculation C Object-Centric Analysis D CLIP Baselines E Frequently Asked Questions (FAQs) The following sections provide theoretical derivations, empirical analyses, and additional experimental details supporting the main manuscript. Appendix A Derivations for Decomposable Representations A.1 Chain rule of mutual information Because the following derivations will heavily rely on rearranging terms in equalities using the chain rule, let us consider the definition on a general setting, using three random variables: A,B,and CA,B,and C. The chain rule of mutual information states that: I(A,C,B)=I(A,B)+I(C;B∣A)I(A,C;B)=I(A;B)+I(C;B A) (27) Intuitively, this means that the information that the joint representation (A,C)(A,C) carries about B can be decomposed into the information provided by A, plus the additional information provided by C, after conditioning on A. A.2 Temporally Decomposable Image Representations (TDIR) As expressed in the main corpus of the manuscript, we note that category and year classes are decomposable as a design choice, which means that the expression in Definition 1 p(μc,c,μy,y)=p(μc,c)p(μy,y)p(μ^c,c,μ^y,y)=p(μ^c,c)\,p(μ^y,y) (28) is satisfied. In short, the class centroids are distributed with no influence from the date information and vice-versa. Proposition 3 (Temporally Decomposable Image Representation). Under the model assumptions of Definition 1, we note that the following equality for the image representation holds: I(μc,μy,c,y)=I(μc,c)+I(μy,y)I(μ^c,μ^y;\,c,y)=I(μ^c;\,c)+I(μ^y;\,y) (29) Proof. We apply the chain rule for mutual information to decompose the left-hand side: I(μc,μy,c,y)=I(μc,c,y)⏟first term+I(μy;c,y∣μc)⏟second termI(μ^c,μ^y;\,c,y)= I(μ^c;\,c,y)_first term+ I(μ^y;\,c,y μ^c)_second term (30) We apply again the chain rule to I(μc,c,y)I(μ^c;\,c,y): I(μc,c,y)=I(μc,c)+I(μc;y∣c)I(μ^c;\,c,y)=I(μ^c;\,c)+I(μ^c;\,y c) (31) From Definition 1, the category centroid does not provide information about the date of the image, therefore it does not provide information when conditioned on the category itself: I(μc,c,y)=I(μc,c)+I(μc;y∣c)=I(μc,c)I(μ^c;\,c,y)=I(μ^c;\,c)+ I(μ^c;\,y c)=I(μ^c;\,c) (32) Hence the first term simplifies to I(μc,c)I(μ^c;\,c) if and only if the model assumption (μc⟂y|cμ^c y c) holds. Similarly, the term I(μy;c,y∣μc)I(μ^y;\,c,y μ^c) is expanded again by applying the chain rule of conditional mutual information: I(μy;c,y∣μc)=I(μy;c∣μc)+I(μy;y∣μc,c)I(μ^y;\,c,y μ^c)=I(μ^y;\,c μ^c)+I(μ^y;\,y μ^c,c) (33) From the Markov structure that emerges from Eq. (28), μc→candμy→y,μ^c→ c μ^y→ y, (34) together with the independence assumption μy⟂(μc,c)μ^y (μ^c,c), it follows that I(μy;c∣μc)=0,I(μ^y;\,c μ^c)=0, (35) and conditioning on (μc,c)(μ^c,c) does not affect the dependence between μyμ^y and y, yielding I(μy;y∣μc,c)=I(μy;y).I(μ^y;\,y μ^c,c)=I(μ^y;\,y). (36) Therefore, I(μy;c,y∣μc)=I(μy,y).I(μ^y;\,c,y μ^c)=I(μ^y;\,y). (37) Combining the results in Equations (32) and (33): I(μc,μy,c,y)=I(μc,c,y)⏟I(μc,c)+0+I(μy;c,y∣μc)⏟I(μy,y)=I(μc,c)+I(μy,y)I(μ^c,μ^y;\,c,y)= I(μ^c;\,c,y)_I(μ^c;\,c)+0+ I(μ^y;\,c,y μ^c)_I(μ^y;\,y)=I(μ^c;\,c)+I(μ^y;\,y) (38) which holds if and only if the initial assumption in Eq. (28) is satisfied. ∎ Lemma 2 (Orthogonality of Centroid Subspaces). Under Definition 1, the category and temporal centroids are orthogonal: Cov(μc,μy)=,Cov(μ^c,μ^y)=0, (39) and the total representational variance decomposes as Var(μc+μy)=Var(μc)+Var(μy).Var(μ^c+μ^y)=Var(μ^c)+Var(μ^y). (40) Proof. By Definition 1, marginalising over (c,y)(c,y) and applying Fubini’s theorem: p(μc,μy)=∫p(μc,c)⋅p(μy,y)cy=∫p(μc,c)dc⏟p(μc)⋅∫p(μy,y)dy⏟p(μy)=p(μc)⋅p(μy),p(μ^c,μ^y)= \!\! p(μ^c,c)· p(μ^y,y)\,dc\,dy= p(μ^c,c)\,dc_p(μ^c)· p(μ^y,y)\,dy_p(μ^y)=p(μ^c)· p(μ^y), (41) so μc⟂μyμ^c μ^y. Independence implies Cov(μc,μy)=Cov(μ^c,μ^y)=0, and the variance decomposition follows immediately from the bilinearity of covariance. ∎ A.3 Estimating the error in real scenarios Because the total independence μc⟂μyμ^c μ^y (Lemma 2) cannot be guaranteed in practical terms, where data might contain severe spurious correlations, one must account for the incorporation of a noise vector ε∈ℝd ^d which entangles both year and category information: v=f(x)=μc+μy+ϵandI(ϵ,c,y)≠0v=f(x)=μ^c+μ^y+ε I(ε;\,c,y)≠ 0 (42) Our model assumption is that epsilon can emerge due to correlation on the training data (hence, the labels y and c) but that an independent proxy vector can be learnt despite this entanglement33 3 With this, we do not negate the presence of noise and spurious correlations on the training; ϵε can heavily depend on the target labels, but one can impose a model restriction where μcμ^c and μyμ^y are learned to be orthogonal vectors.: ϵ⟂(μc,μy)ε (μ^c,μ^y) (43) Proposition 4 (Practical Learnability of TDIRs). Under the TDIR framework, the discrepancy term δ≜I(ε,c,y)δ I( ;\,c,y) depends only on the statistical structure of the training data and not on the choice of proxy vectors μcμ^c or μyμ^y. I(v,c,y)≤I(μc,c)+I(μy,y)+δwhereδ≜I(ε,c,y)I(v;\,c,y)≤ I(μ^c;\,c)+I(μ^y;\,y)+δ δ I( ;\,c,y) (44) Consequently, δ constitutes a data-dependent constant with respect to the optimization, and the category and year centroids can be learned freely to satisfy the orthogonality constraint ⟨μc,μy⟩=0 μ^c,\,μ^y =0 without affecting this error bound. Proof. The real mutual information expression entangles ϵε with the label variables: I(μc,μy,ϵ;c,y)=I(μc,μy;c,y)+I(ϵ;μc,μy∣c,y)I(μ^c,μ^y,ε;\,c,y)=I(μ^c,μ^y;\,c,y)+I(ε;\,μ^c,μ^y c,y) (45) Because the model is assumed to be trained under an independence condition ϵ⟂(μc,μy)ε (μ^c,μ^y), I(ϵ;μc,μy∣c,y)=I(ϵ;μc,μy)I(ε;\,μ^c,μ^y c,y)=I(ε;\,μ^c,μ^y) (46) Therefore, I(μc,μy,ϵ,c,y)=I(μc,μy,c,y)⏟I(μc,c)+I(μy,y)+I(ϵ,μc,μy)I(μ^c,μ^y,ε;\,c,y)= I(μ^c,μ^y;\,c,y)_I(μ^c;\,c)+I(μ^y;\,y)+I(ε;\,μ^c,μ^y) (47) To obtain the total expression for I(v,c,y)I(v;\,c,y), we apply the chain rule of mutual information: I(v⏟A,μc,μy,ϵ⏟C,c,y⏟B)=I(v,c,y)+I(μc,μy,ϵ;c,y∣v)I( v_A,\, μ^c,μ^y,ε_C;\, c,y_B)=I(v;\,c,y)+I(μ^c,μ^y,ε;\,c,y v) (48) Since v=μc+μy+ϵv=μ^c+μ^y+ε is a function of (μc,μy,ϵ)(μ^c,μ^y,ε), the joint tuple (v,μc,μy,ϵ)(v,μ^c,μ^y,ε) carries no more information about (c,y)(c,y) than (μc,μy,ϵ)(μ^c,μ^y,ε) alone: I(v,μc,μy,ϵ,c,y)=I(μc,μy,ϵ,c,y)I(v,μ^c,μ^y,ε;\,c,y)=I(μ^c,μ^y,ε;\,c,y) (49) Isolating I(v,c,y)I(v;\,c,y): I(v,c,y)=I(μc,μy,ϵ,c,y)−I(μc,μy,ϵ;c,y∣v)I(v;\,c,y)=I(μ^c,μ^y,ε;\,c,y)-I(μ^c,μ^y,ε;\,c,y v) (50) Substituting the expression in Eq. (50): I(v,c,y)=I(μc,c)+I(μy,y)+I(ϵ,c,y)−I(μc,μy,ϵ;c,y∣v) I(v;\,c,y)=I(μ^c;\,c)+I(μ^y;\,y)+I(ε;\,c,y)-I(μ^c,μ^y,ε;\,c,y v) (51) where the last term I(μc,μy,ϵ;c,y∣v)I(μ^c,μ^y,ε;\,c,y v) represents the information lost by collapsing the three components into a single vector v. Identifying δ: δ=I(ϵ,c,y)−I(μc,μy,ϵ;c,y∣v)⏟≥ 0δ=I(ε;\,c,y)- I(μ^c,μ^y,ε;\,c,y v)_≥\,0 (52) we can rewrite the exact relation compactly as: I(v,c,y)=I(μc,c)+I(μy,y)+δ⏟≤I(ϵ,c,y)I(v;\,c,y)=I(μ^c;\,c)+I(μ^y;\,y)+ δ_≤\,I(ε;\,c,y) (53) Since mutual information is always non-negative, δ≤I(ε,c,y)δ≤ I( ;\,c,y) provides an upper bound. To obtain a more precise approximation of δ, we expand I(ε,c,y)I( ;\,c,y) via the chain rule: I(ε,c,y)=I(ε,c)+I(ε;y∣c)I( ;\,c,y)=I( ;\,c)+I( ;\,y c) (54) Under the assumption that category c and year y are independent, I(ε;y∣c)≈I(ε,y)I( ;\,y c)≈ I( ;\,y), giving: δ≤I(ε,c)+I(ε,y)δ≤ I( ;\,c)+I( ;\,y) (55) ∎ This shows that δ is controlled by how much the noise ε leaks information about each label separately. If the model successfully disentangles μcμ^c and μyμ^y from ε , both terms vanish and δ→0δ→ 0, recovering the ideal case: I(v,c,y)≈I(μc,c)+I(μy,y)I(v;\,c,y)≈ I(μ^c;\,c)+I(μ^y;\,y) (56) Thus, under the assumption that such orthogonal centroids can be learned (see Figure 5), the ideal mutual information equality in the practical case is consistent with the TDIR proposition up to an upper bound δ characterized by the entanglement of training data. μcμ^cε μyμ^yccvvyyp(c∣μc)p(c μ^c)p(y∣μy)p(y μ^y)category centroidnoisetemporal centroidcategoryrepresentationyear⟂ ⟂ =μc+μy+εv=μ^c+μ^y+ (v∣μc)p(v μ^c)p(v∣μy)p(v μ^y) Figure 5: Resulting graphical model from Proposition 2 The joint optimization of ℒcL^c and ℒyL^y implicitly enforces the orthogonality condition ⟨μc,μy⟩=0 μ^c,\,μ^y =0 required by the TDIR definition, without imposing it as an explicit constraint. The proof relies on a preliminary result about the structure of the year proxies, which we establish first. A.4 TDIR as an Emergent Property of the Joint Loss With all the properties and propositions derived above, we can then state that the following theorem holds: Theorem 3 (Emergent TDIR). Let ℒcL^c and ℒyL^y be the proxy loss functions defined over an embedding space ℝdR^d. Then: Part I (exact case). If I(ϵ,c,y)=0I(ε\,;\,c,y)=0, joint minimization of ℒ=ℒc+ℒyL=L^c+L^y implies TDIR: minθℒc+ℒy⟹p(μc,μy,c,y)=p(μc,c)⋅p(μy,y) _θ\,L^c+L^y p(μ^c,μ^y,c,y)=p(μ^c,c)· p(μ^y,y) (57) Part I (asymptotic case). If I(ϵ,c,y)≠0I(ε\,;\,c,y)≠ 0, joint minimization implies asymptotic TDIR: minθℒc+ℒy⟹⟨μc,μy⟩→O(I(ϵ,c,y)‖μy‖2)∀c,y _θ\,L^c+L^y μ^c,\,μ^y → O\! ( I(ε\,;\,c,y)\|μ^y\|^2 ) ∀\,c,\,y (58) Proof. Part I: I(ϵ,c,y)=0⟹I(ε\,;\,c,y)=0 TDIR. The category proxy loss is minimized when the score v⊤μcv μ^c is maximal and uniform across all years. Since vAc=μc+μyAv_A^c=μ^c+μ^y_A and vBc=μc+μyBv_B^c=μ^c+μ^y_B share the same category but differ in year, the minimum requires: vA⊤μc=vB⊤μc∀yA,yB⟹⟨μyA,μc⟩=⟨μyB,μc⟩∀yA,yBv_A μ^c=v_B μ^c ∀\,y_A,y_B μ^y_A,μ^c = μ^y_B,μ^c ∀\,y_A,y_B hence I(μc,y)=0I(μ^c\,;\,y)=0. For the temporal loss, the shared displacement vectors kyk^y impose, for any two categories c≠c′c≠ c sharing the same years: rAc+kBy≈μyB,rAc′+kBy≈μyBr_A^c+k^y_B≈μ^y_B, r_A^c +k^y_B≈μ^y_B Subtracting yields rAc≈rAc′r_A^c≈ r_A^c , so the score rAc⊤μyr_A^c μ^y is independent of the category, hence I(μy,c)=0I(μ^y\,;\,c)=0. Applying the chain rule of mutual information: I(μc,μy,c,y)=I(μc,c)+I(μy;c∣μc)⏟= 0+I(μc;y∣μy,c)⏟= 0+I(μy,y)=I(μc,c)+I(μy,y)I(μ^c,μ^y\,;\,c,y)=I(μ^c\,;\,c)+ I(μ^y\,;\,c μ^c)_=\,0+ I(μ^c\,;\,y μ^y,c)_=\,0+I(μ^y\,;\,y)=I(μ^c\,;\,c)+I(μ^y\,;\,y) This equality is equivalent, via the KL divergence and Fubini’s theorem, to the factorization p(μc,μy,c,y)=p(μc,c)⋅p(μy,y)p(μ^c,μ^y,c,y)=p(μ^c,c)· p(μ^y,y). Part I: I(ϵ,c,y)≠0⟹I(ε\,;\,c,y)≠ 0 asymptotic TDIR. The proxy loss ℒ(u,μ∗,)=−logσ(u⊤μ∗)L(u,μ^*;P)=- \,σ(u μ^*) is a softmax cross-entropy, which is smooth and strictly convex in a neighborhood of its minimum. A second-order Taylor expansion around the minimizer yields a local quadratic approximation of the form ‖u−μ∗‖2+O(‖u−μ∗‖3)\|u-μ^*\|^2+O(\|u-μ^*\|^3). We therefore adopt the squared-error surrogates as a standard local approximation: ℒc(T) ^c(T) ∝‖vA−μc‖2+‖vB−μc‖2, \; \;\|v_A-μ^c\|^2+\|v_B-μ^c\|^2, (59) ℒy(T) ^y(T) ∝‖r^A−μyA‖2+‖r^B−μyB‖2+‖r^A→B−μyB‖2+‖r^B→A−μyA‖2. \; \;\| r_A-μ^y_A\|^2+\| r_B-μ^y_B\|^2+\| r_A→ B-μ^y_B\|^2+\| r_B→ A-μ^y_A\|^2. (60) This approximation is tight near convergence, where embeddings cluster around their respective proxies, and the conclusions below hold to the same asymptotic order as the Taylor remainder. Decompose μcμ^c over the temporal subspace as μc=μ~c+αμyA+βμyBμ^c= μ^c+α\,μ^y_A+β\,μ^y_B with μ~c⟂μy μ^c μ^y. Under v=μc+μy+ϵv=μ^c+μ^y+ε, the year residuals are: rAc=(1−α)μyA−βμyB+ϵAc,rBc=(1−β)μyB−αμyA+ϵBcr_A^c=(1-α)μ^y_A-βμ^y_B+ _A^c, r_B^c=(1-β)μ^y_B-αμ^y_A+ _B^c Substituting into the loss surrogates to ℒyL^y and expanding: ℒy∝ ^y +‖ϵAc‖2+‖ϵBc‖2−2α⟨μyA,ϵAc+ϵBc⟩−2β⟨μyB,ϵAc+ϵBc⟩+ +\| _A^c\|^2+\| _B^c\|^2-2α μ^y_A, _A^c+ _B^c -2β μ^y_B, _A^c+ _B^c + +2α2‖μyA‖2+2β2‖μyB‖2+4αβ⟨μyA,μyB⟩⏟Q(α,β)≥0 + 2α^2\|μ^y_A\|^2+2β^2\|μ^y_B\|^2+4αβ μ^y_A,μ^y_B _Q(α,β)≥ 0 The quadratic form Q(α,β)Q(α,β) is positive semidefinite by Cauchy-Schwarz. The cross terms displace the minimum from (0,0)(0,0) to: α∗=⟨μyA,ϵAc+ϵBc⟩2‖μyA‖2+O(ϵ2),β∗=⟨μyB,ϵAc+ϵBc⟩2‖μyB‖2+O(ϵ2)α^*= μ^y_A,\, _A^c+ _B^c 2\|μ^y_A\|^2+O(ε^2), β^*= μ^y_B,\, _A^c+ _B^c 2\|μ^y_B\|^2+O(ε^2) For small correlations, the covariance between ϵε and μyμ^y satisfies Cov(ϵ,μy)≈‖ϵ‖⋅‖μy‖⋅2I(ϵ,c,y)Cov(ε,μ^y)≈\|ε\|·\|μ^y\|· 2\,I(ε\,;\,c,y), which gives: ⟨μc,μy⟩=α∗‖μy‖2+O(ϵ2)=O(I(ϵ,c,y)‖μy‖2) μ^c,\,μ^y =α^*\|μ^y\|^2+O(ε^2)=O\! ( I(ε\,;\,c,y)\|μ^y\|^2 ) Global consistency follows from the shared kyk^y: a single kByk^y_B cannot simultaneously compensate projections α≠α′α≠α across categories unless (α−α′)μyA=ϵAc′−ϵAc(α-α )μ^y_A= _A^c - _A^c, which cannot hold for all category pairs since ϵε is not a control variable of the optimizer. The unique globally consistent solution is therefore: ⟨μc,μy⟩→O(I(ϵ,c,y)‖μy‖2)∀c,y∎ μ^c,\,μ^y → O\! ( I(ε\,;\,c,y)\|μ^y\|^2 ) ∀\,c,y Corollary 4 (Swapping Trick Enforces Cross-Category Orthogonality). Under the conditions of Theorem 3, the shared displacement vectors kyk^y enforce ⟨μc,μy⟩→0 μ^c,μ^y → 0 uniformly across all categories. Since the same kByk^y_B is used when translating residuals from any category, satisfying the swapping constraint simultaneously for two categories c≠c′c≠ c with residual projections α≠α′α≠α onto μyμ^y would require: (α−α′)μyA=ϵAc′−ϵAc∀c,c′.(α-α )\,μ^y_A\;=\; _A^c - _A^c ∀\;c,\,c . (61) This cannot hold globally, since ϵε is not a control variable of the optimizer. The unique globally consistent solution is therefore α=α′=0α=α =0 for all c,c′c,\,c , giving: Projμy(μc)→ 0∀c,y.Proj_μ^y(μ^c)\;→\;0 ∀\;c,\,y. (62) Appendix B Photographic material calculation To estimate the proportion of photographic material in large archival collections, we surveyed three major public archives. Archives Nationales (France). Using the publicly available digitisation corpus dataset [Archives Nationales, France()], a small available sample drawn from the 2025 digitisation batch was analysed. Of the total documents, 123 were non-textual, representing 23% of the sample (123/0.23≈535123/0.23≈ 535 documents in total). Among those non-textual items, 68 were strictly photographic, accounting for approximately 13% of the full sample. Library of Congress (USA). The collection can be divided into digitised and non-digitised holdings [Library of Congress()]. Among the non-digitised material, the catalogue lists 16M books, 3M newspapers, 1M periodicals, 0.5M manuscripts, 1.3M photographic and pictorial items, and 450M audiovisual items. Among the digitised material, it lists 3M newspapers, 0.7M books, 0.5M manuscripts, 0.5M periodicals, and 1.2M photographs. Focusing on the textual-adjacent holdings (i.e. excluding the bulk audiovisual collection, which is a category of its own), photographic material represents: 1.3+1.216+3+1+0.5+1.3+3+0.7+0.5+0.5+1.2≈2.527.7≈%. 1.3+1.216+3+1+0.5+1.3+3+0.7+0.5+0.5+1.2≈ 2.527.7 9\%. Portal de Archivos Españoles (Spain). The portal reports approximately 40 million digital objects in total [Ministerio de Cultura, España()]. According to [Ministerio de Cultura, España(2025)], roughly 2 million of these are classified as photographic objects (excluding videos, posters, and other non-textual formats), yielding a photographic fraction of approximately 2/40=%2/40=5\%. Appendix C Object-Centric Analysis The per-object analysis reveals two complementary sources of error in the composed retrieval framework, illustrated in Figure 6. When considering the reference image (the image from which temporal information is extracted), objects with low Year MAE such as tires, shorts, glasses, and human faces act as reliable temporal vessels: their embeddings carry a strong and unambiguous date signal, suggesting minimal entanglement between their categorical and temporal subspaces. Conversely, objects such as wheels, vehicle registration plates, dresses, buildings, and men exhibit higher reference error, indicating that their visual representations encode date information less cleanly, likely due to higher intra-class visual variance across decades. On the target side (the object whose category should be preserved after temporal transplantation), objects like posters, girls, trains, houses, and dresses are retrieved with relatively low temporal error, while ladders, human faces, tires, wheels, and hats show the highest MAE. This asymmetry suggests that temporally ambiguous objects are doubly penalised: they are poor sources of date information and poor recipients of temporal injection. Figure 6: Average Year Error decomposed by reference object (top) and target object (bottom) in the Image+Image inference mode. Reference objects with low error are reliable vessels for date injection, indicating clean category representations with low temporal entanglement. Target objects with high error are temporally ambiguous, making them difficult to anchor to a specific period regardless of the source image used. Figure 7 reports per-object precision across all three inference modes using radar plots, for both object retrieval (top row) and date estimation (bottom row). In the object retrieval setting, the largest errors concentrate on categories that are visually sub-specific or easily confused with related classes: glasses, vehicle registration plates, and items of clothing such as socks and footwear tend to share visual features with broader categories, creating spurious correlations that degrade categorical fidelity. For date estimation (bottom row), the Label+Label mode achieves strong and consistent performance across virtually all object categories, confirming that the proxy-based subspace structure encodes temporal information reliably when labels are directly available. However, when queries are issued through images, performance degrades selectively for objects that lack distinctive temporal signatures, such as generic architectural elements or accessories, producing a marked and category-specific drop in precision. Figure 7: Per-object precision radar plots for object retrieval (top row) and date estimation (bottom row) across Label+Label, Image+Label, and Image+Image inference modes (left to right). Object retrieval errors concentrate on visually sub-specific categories such as glasses, vehicle registration plates, and footwear, which share features across class boundaries. For date estimation, Label+Label achieves consistently strong performance, while image-based modes reveal a pronounced and category-dependent performance gap for temporally ambiguous objects. The temporal behaviour of each inference mode is further analysed in Figure 8. Date estimation precision degrades toward earlier decades for both Image+Label and Image+Image modes, consistent with the sparser representation of pre-1950 material in the DEW dataset. Notably, the Image+Image mode exhibits a reversal of the general trend observed in label-based settings: whereas date precision typically exceeds object precision under Label+Label and Image+Label queries, the Image+Image mode yields stronger object than date precision. This suggests that the categorical subspace is more robust to temporal transplantation than the temporal subspace is to cross-image injection, a finding aligned with the theoretical guarantee of Proposition 3, which ensures category preservation under orthogonal decomposition but imposes no analogous bound on the fidelity of the transplanted temporal signal. Figure 8: Left: date estimation precision per source year across inference modes. Both Image+Label and Image+Image exhibit degraded performance for earlier decades, with Image+Label remaining more stable across the temporal range. Right: scalar summary of object and date precision per inference mode. In Image+Image retrieval, the relationship between object and date precision inverts relative to label-based modes, indicating that temporal injection is a stricter operation than category preservation under the learned decomposition. Figures 9 and 10 decompose the performance drop from Image+Label to Image+Image inference on a per-object basis, for date estimation and object retrieval respectively. Object retrieval degradation is more uniformly distributed across categories, suggesting that the category subspace remains largely stable under temporal transplantation for most classes, with isolated exceptions corresponding to categories with high inter-class visual overlap. Figure 9: Per-object drop in date retrieval precision. Figure 10: Per-object drop in object retrieval precision. Appendix D CLIP Baseline Both CLIP [Radford et al.(2021)Radford, Kim, Hallacy, Ramesh, Goh, Agarwal, Sastry, Askell, Mishkin, Clark, et al.] baselines use OpenCLIP [Ilharco et al.(2021)Ilharco, Wortsman, Wightman, Gordon, Carlini, Taori, Dave, Shankar, Namkoong, Miller, Hajishirzi, Farhadi, and Schmidt] ViT-B/32 [Dosovitskiy et al.(2020)Dosovitskiy, Beyer, Kolesnikov, Weissenborn, Zhai, Unterthiner, Dehghani, Minderer, Heigold, Gelly, et al.] pre-trained on LAION-2B [Schuhmann(2022)] as a frozen encoder, with no fine-tuning on DEW or on any historical imagery. All image embeddings are ℓ2 _2-normalised to unit norm prior to any arithmetic operation or nearest-neighbour query, following the same convention as TDIR. A single Annoy index 44 4 https://github.com/spotify/annoy of Ntrees=50N_trees=50 trees is built over the CLIP vision embeddings of the full training split (identical images to those used for TDIR training), so that retrieval pool and gallery are exactly matched across methods. Nearest-neighbour search is performed in Euclidean space over the unit hypersphere, which is equivalent to cosine similarity for ℓ2 _2-normalised vectors. The top-K=10K=10 neighbours are returned for every query, and precision is computed identically to Section 4.3. D.1 CLIP Arithmetic Baseline The arithmetic baseline replaces the learned proxy vectors μcμ^c and μyμ^y of TDIR with CLIP text embeddings, exploiting the joint vision–language alignment acquired during pre-training. For each object category c we encode the prompt ”A photo of c”, and for each target year y (bucketed in 5-year lustrum bins from 1930 to 1999, yielding 14 bins) we encode ”A photo in year”. All text embeddings are ℓ2 _2-normalised after encoding, producing a category proxy matrix Tc∈ℝC×dT^c ^C× d and a year proxy matrix Ty∈ℝY×dT^y ^Y× d, with d=512d=512 for ViT-B/32. These matrices serve as drop-in substitutes for PcP^c and PyP^y in all three inference modes of Section 3.5, with no modification to the retrieval procedure. Label+Label. The query vector is the un-normalised sum q=Tcc+Tyyq=T^c_c+T^y_y, directly mirroring the proxy-sum query q=μc+μyq=μ^c+μ^y of TDIR. Image+Label. The query is q=V(xc)+Tyyq=V(x_c)+T^y_y, where V(xc)V(x_c) is the unit-norm CLIP vision embedding of the category image. Image+Image (zero-label). The reference image xyx_y is sampled uniformly from the training index. Its inferred category c c is obtained by argmax cosine similarity against all category text embeddings, c^=argmaxc′V(xy)⊤Tc′c c= *arg\,max_c V(x_y) T^c_c . The temporal residual is then extracted as ry=V(xy)−Tc^cr_y=V(x_y)-T^c_ c, and the final query is q=V(xc)+ryq=V(x_c)+r_y. The poor date precision reported in Table 2 for this baseline is informative: despite CLIP embeddings carrying rankable temporal information [Sonthalia et al.(2026)Sonthalia, Uselis, and Oh], the arithmetic composition fails to disentangle temporal and categorical signals because CLIP was never trained to produce orthogonal category and year subspaces. The residual ryr_y retains substantial categorical content, and the year text prompts encode only a weak and diffuse notion of historical period that does not translate into a geometrically consistent displacement on the visual embedding manifold. This confirms that temporal composability is not an emergent property of vision–language pre-training, but requires the explicit geometric structure imposed by TDIR. D.2 CLIP Prompting Baseline The prompting baseline replaces vector arithmetic with a two-stage retrieve-and-rerank strategy. Rather than composing both signals into a single query vector, retrieval is first performed using the primary modality alone, and the resulting candidate pool is then re-ranked by cosine similarity to the secondary modality signal. This avoids relying on cross-modal vector addition and instead lets each signal operate independently on the unit hypersphere. Concretely, a candidate pool of size P=100P=100 is retrieved by Annoy nearest-neighbour search using the primary modality embedding. The stored embeddings of all pool candidates are then re-scored by their cosine similarity to the secondary modality embedding, and the top-K=10K=10 candidates after re-scoring are returned. Label+Label. The candidate pool is retrieved by TccT^c_c alone; pool candidates are then re-ranked by cosine similarity to TyyT^y_y. Image+Label. The candidate pool is retrieved by V(xc)V(x_c); re-ranking uses TyyT^y_y. The pool is computed once per test image and reused for all 14 year bins, amortising the retrieval cost across the full label set. Image+Image. The candidate pool is retrieved by V(xc)V(x_c); re-ranking uses V(xy)V(x_y) directly, replacing the residual ryr_y with the raw reference image embedding. This sidesteps the category-inference step and provides a stronger re-ranking signal when category and date co-vary in the reference image. As shown in Table 2, the prompting baseline recovers competitive object precision across all three modes, confirming that CLIP vision embeddings carry strong categorical signal. Date precision, however, remains substantially below TDIR for all inference modes even under reranking. This gap persists because reranking by TyyT^y_y or V(xy)V(x_y) can only promote candidates that already appear in the pool retrieved by the category signal, and nothing in the CLIP representation space guarantees that temporally relevant images rank highly under a category-only query. The contrast between the two baselines and TDIR thus isolates the contribution of the learned orthogonal decomposition: without it, object and temporal retrieval remain fundamentally entangled regardless of the composition strategy employed. Appendix E Frequently Asked Questions (FAQs) In this section, readers will find a series of answers to questions raised during peer review, discussions, and others. We hope they cana be useful in interpreting the data and results presented in the paper, as well as the method itself. E.1 Joint metric: Def. 1’s independence assumption holds empirically. We added Precision@K requiring both category and date correct simultaneously to all three retrieval modes, letting us test directly (rather than assume) whether treating category and date as independent (Def. 1) actually holds at inference. On Label+Label (K=10, a stratified 7 407-crop pool): Obj P@10 = 0.488 Date P@10 = 0.718 Joint P@10 = 0.349 against 0.488×0.718=0.3500.488× 0.718=0.350 predicted under independence, a 0.4%0.4\% relative gap well within noise. Object- and date-correctness are therefore empirically independent: neither axis is traded off against, nor rides for free on, the other, exactly the behaviour Def. 1 and Thm. 1’s joint optimisation are designed to produce. E.2 Proxy geometry corroborates this at the representation level. We further computed the full cosine matrix between the 50 category and 14 year proxies of the reported ConvNeXt+TDIR checkpoint (its cached index reproduces Table 2/3 exactly). Cross-block mean |cos|=0.204| |=0.204, close to the 0.0250.025 floor for random unit vectors in ℝ1024R^1024 at this scale. Fig. 1 shows both regimes side by side across all 14 year proxies: 27/50 categories (e.g Window, Bus) sit at that floor throughout, genuinely orthogonal as Thm. 1 predicts in the exact case. A smaller set of ∼ 15 categories with a plausible intrinsic appearance-era link (tie, belt, watch, sock, footwear,…) shows the bounded, data-dependent residual δ that Prop. 2 anticipates for the realistic case: e.g Footwear peaks at +0.87+0.87 in 1930 and is 3–6× weaker by 1965 on. This residual tracks training-set size almost exactly (per-year mean |cos|| | vs. log sample count: r=−0.90r=-0.90, p<10−4p<10^-4; 1930: 14k crops →0.44→ 0.44; 1985: 108k crops →0.11→ 0.11), giving δ a precise, measured, localized explanation rather than an abstract constant. Figure 11: Category-year cosine similarity across all 14 year proxies for the two most orthogonal (Window, Bus) and the most entangled (Footwear) of the 50 categories. E.3 Baselines and Previous Work Modern SSL pretraining (DINOv3). Because DINOv3 is not a specific architecture but a pretraining regime, we have loaded DINOv3 weights into the same ConvNeXt-Base backbone used throughout (Table 2) and train under the identical 50-category protocol. At epoch 1 already, cross-block mean |cos|=0.173| |=0.173, already at or below the fully-converged ImageNet-supervised value (0.2040.204), with the same ordinal decay replicating: the decomposition is therefore not an artifact of ImageNet-supervised pretraining. On the suggested [Molina et al.(2022)Molina, Gomez, Ramos Terrades, and Lladós, Net et al.(2024)Net, Hernández, Molina, and Gómez] baselines. We are well aware of both [Molina et al.(2022)Molina, Gomez, Ramos Terrades, and Lladós]’s smooth-nDCG loss needs a relevance function derived from a numerical variable: it can replace our proxy loss on the year component, but never on the category, so it does not serve as a composed-retrieval loss. However, further results may incorporate such loss as date proxy loss, but the method itself is not capable of composing concepts. And, on the other hand, [Net et al.(2024)Net, Hernández, Molina, and Gómez] is incompatible with our setup by any means: it runs object detection over multi-object images and requires a dedicated encoder per object label to predict a date (in [Net et al.(2024)Net, Hernández, Molina, and Gómez] the category is not learned, it is passed as input). E.4 Ordinal structure & unseen years. Cosine similarity between two year proxies decays smoothly with their temporal gap, although only pairwise transitivity (Def. 2) was ever enforced at training time: Leave-one-out check: a year’s true proxy is consistently closer to the midpoint of its two neighbours than to either neighbour alone (e.g cos=0.66 =0.66 vs. 0.570.57 at 1935; all 12 interior years), so the subspace already extrapolates locally. Objects with weak temporal signature. The categories with the lowest year-entanglement above (window, bus, door, boat, tower) are the slow-changing, architecture-type categories that one may flag as barely changing across decades. This is actually consistent, not contradictory, since their proxies correctly encode little date information because their appearance carries little.