Paper deep dive
Temporal Preference Optimization for Unsupervised Retrieval
HyunJin Kim, Jaejun Shim, Young Jin Kim, JinYeong Bak
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 6/20/2026, 10:29:53 AM
Summary
The paper introduces TPOUR (Temporal Preference Optimization for Unsupervised Retriever), a novel training framework designed to address temporal misalignment in unsupervised dense retrievers. While standard unsupervised retrievers optimize solely for semantic similarity, TPOUR incorporates Temporal Retrieval Preference Optimization (TRPO) to guide models to favor temporally aligned documents. The method utilizes time vector extraction and interpolation, allowing the retriever to generalize to unseen intermediate and future time periods without retraining. Experimental results demonstrate that TPOUR improves performance on temporal information retrieval (T-IR) tasks, such as SituatedQA and RealTimeQA, and can even be used for document timestamp prediction via a mixture-of-TPOUR approach.
Entities (8)
Relation Signals (4)
TPOUR → builtupon → MoCo
confidence 100% · Built upon MoCo, TPOUR jointly learns semantic similarity and temporal relevance...
TPOUR → evaluatedon → SituatedQA
confidence 100% · We evaluate whether TPOUR-trained retrievers learn temporally aligned representations... on the SituatedQA dataset.
TPOUR → improves → Contriever
confidence 100% · TPOUR Contriever improves average nDCG@5 by +4.04 (+12.15%) on explicit and +4.98 (+15.21%) on implicit queries.
TPOUR → uses → TRPO
confidence 100% · We propose TPOUR... which uses our novel training method Temporal Retrieval Preference Optimization (TRPO).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Unsupervised dense retrievers offer scalability by learning semantic similarity from unlabeled documents via contrastive learning, but they struggle to capture the temporal relevance, retrieving semantically related but temporally misaligned documents-an important aspect when a document collection spans multiple time periods (e.g., retrieving documents from 2018-2025 for "Who is the president in 2019?" introduces temporal ambiguity). Existing methods rely on supervised training with explicit timestamps, which are not always feasible. We propose TPOUR (Temporal Preference Optimization for Unsupervised Retriever), which uses our novel training method Temporal Retrieval Preference Optimization (TRPO). TRPO reinterprets preference learning in the temporal dimension, guiding the retriever to favor temporally aligned documents. TPOUR further generalizes to unseen time periods via interpolation in a learned time embedding, enabling continuous temporal alignment. Experiments on temporal information retrieval (T-IR), TPOUR outperforms both unsupervised and supervised baselines. Compared to Qwen-Embedding-8B, despite being about 72.7x smaller, TPOUR Contriever improves average nDCG@5 by +4.04 (+12.15%) on explicit and +4.98 (+15.21%) on implicit queries. We provide our code at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2606.17664v1
- Canonical: https://arxiv.org/abs/2606.17664v1
Trouble viewing inline? Open PDF directly →
Full Text
130,475 characters extracted from source content.
Expand or collapse full text
Temporal Preference Optimization for Unsupervised Retrieval HyunJin Kim 1 Jaejun Shim 1 Young Jin Kim 2∗ JinYeong Bak 1∗ Abstract Unsupervised dense retrievers offer scalability by learning semantic similarity from unlabeled doc- uments via contrastive learning, but they strug- gle to capture the temporal relevance, retrieving semantically related but temporally misaligned documents–an important aspect when a document collection spans multiple time periods (e.g. re- trieving documents from 2018-2025 for “Who is the president in 2019?” introduces temporal ambiguity.). Existing methods rely on super- vised training with explicit timestamps, which are not always feasible. We proposeTPOUR (Temporal Preference Optimization for Unsuper- vised Retriever), which uses our novel training method Temporal Retrieval Preference Optimiza- tion (TRPO). TRPO reinterprets preference learn- ing in the temporal dimension, guiding the re- triever to favor temporally aligned documents. TPOURfurther generalizes to unseen time peri- ods via interpolation in a learned time embedding, enabling continuous temporal alignment. Exper- iments on temporal information retrieval (T-IR), TPOURoutperforms both unsupervised and super- vised baselines. Compared to Qwen-Embedding- 8B, despite being about 72.7×smaller,TPOUR Contriever improves average nDCG@5 by +4.04 (+12.15%) on explicit and +4.98 (+15.21%) on implicit queries. We provide our code athttps: //github.com/agwaBom/TPOUR. 1. Introduction Document retrieval is the process of identifying relevant documents from document collections (Gao et al., 2024; Zhao et al., 2024a;b; Zhu et al., 2025; Li et al., 2025). It is widely used for various applications, including search en- gines (Brin & Page, 1998; Li et al., 2025), recommendation 1 Sungkyunkwan University, Suwon, South Korea 2 Microsoft, Redmond,USA. Correspondence to:Young Jin Kim <youki@microsoft.com>, JinYeong Bak <jy.bak@skku.edu>. Proceedings of the43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s). systems (Bobadilla et al., 2013; Zhang et al., 2019; Singh, 2023; Li et al., 2024), question answering (Karpukhin et al., 2020; Zhang et al., 2023a;b), and retrieval-augmented gen- eration (Lewis et al., 2020; Zhao et al., 2024a; Fan et al., 2024; Kwon et al., 2025). Retrieval training generally falls into supervised and unsupervised methods. Supervised methods utilize labeled query-document pairs (Karpukhin et al., 2020), whereas unsupervised methods leverage term- frequency (Robertson & Zaragoza, 2009) or contrastive learning from unlabeled data (Izacard et al., 2022). Despite advancements in retrieval research, most retrieval systems overlook temporal misalignment (i.e., mismatch between the temporal context of user queries and the times- tamps of retrieved documents). Temporal retrieval aims to address this limitation by incorporating temporal context into the retriever. As shown in Fig. 1, queries may contain explicit (e.g., “in 2019”) or implicit (e.g., “this year”) tempo- ral information. While explicit references clearly anchor the query in time, implicit ones require interpretation. We adopt an approach that trains the retriever to prefer documents from a specific time period and interpret implicit queries accordingly. For example, a retriever trained on 2018 data would interpret “this year” as referring to 2018, aligning implicit temporal expressions with its training period. Tem- poral retrieval is important in domains such as news (Litty K Mathews, 2012; Wang et al., 2012; Luu et al., 2022) and law (Schilder & McCulloh, 2005), where the relevance of information depends on its publication date. For instance, the query “What was the minimum wage law in effect in 2019?” should retrieve the regulation in effect at that time. However, existing retrieval methods often neglect temporal signals, particularly when timestamps are implicit rather than explicitly stated in the query. For instance, consider the query “Who is the current president?”, which implicitly requires an answer at the time the query is raised, despite the absence of an explicit timestamp. Time-unaware retrievers such as Contriever (Izacard et al., 2022) or DPR (Karpukhin et al., 2020) are trained to maximize semantic similarity, and thus often retrieve documents that are semantically relevant but temporally unaligned. Fig. 1 illustrates this limitation–a time-unaware retriever fails to distinguish tem- porally aligned documents from temporally misaligned ones when relying solely on semantic similarity. 1 arXiv:2606.17664v1 [cs.IR] 16 Jun 2026 Temporal Preference Optimization for Unsupervised Retrieval Query: Who was the champion in women’s singles at Roland Garros (in 2019, this year)? Index: 2 Ashleigh Barty won the 2019 French Open champion ... Index: 1 Simona Halep won her first major in 2018 French Open TPOUR Time Unaware Index: 4 2015 Worlds Championship is won by SK Telecom T1 ... 2 1 34 Consider semantic similarity and temporal alignment Consider solely on semantic similarity Ranked Index Order: Ranked Index Order: ? VS. 2 1 34 Index: 3 2019 Ballon d’Or. Lionel Messi won sixth award ... Figure 1. Comparison betweenTPOURaligned at 2019 and a time-unaware retriever for queries with explicit (e.g., in 2019) or implicit (e.g., this year) temporal information. Left: A mixed-timestamp document collection containing (i) semantically and temporally aligned documents (green), (i) semantically relevant but temporally misaligned documents (yellow), and (i) irrelevant documents (red). Right: Ranked retrieval results. The time-unaware retriever, trained solely for semantic similarity, struggles to rank the temporally aligned document (green) over the misaligned (yellow). In contrast, the TPOUR-trained retriever prioritizes the temporally aligned document. In practice, addressing temporal misalignment is challeng- ing. On the one hand, supervised approaches may capture temporal relevance, but they require large amounts of la- beled data, making them impractical at scale. On the other hand, unsupervised approaches based on contrastive learn- ing (Shao et al., 2021; Izacard et al., 2022; Wu et al., 2022; Deng et al., 2022) are scalable but solely optimize for se- mantic similarity and ignore temporal relevance. To embed temporal relevance in unsupervised retrieval, we proposeTPOUR(Temporal Preference Optimization for Unsupervised Retriever), which integrates novel train- ing method Temporal Retrieval Preference Optimization (TRPO) with contrastive learning. TRPO incorporates tem- poral preference signal into the retriever, reinterpreting pref- erence learning in the temporal dimension using training signals from document corpora collected at different time periods. Rather than relying solely on semantic similarity, TRPO prioritize temporally aligned documents over mis- aligned ones. Thus,TPOURpreserves semantic similarity while learning temporal relevance, even when explicit time information is missing from the query or document. TPOURdoes not require retraining to adapt to specific time periods. We validate that the time vector, originally pro- posed as a temporal embedding for generative models (Ny- lund et al., 2024), can be applied to encoder-basedTPOUR retriever. By extracting time vectors fromTPOURretrievers fine-tuned on a specific time period and interpolating them, we achieve continuous temporal alignment to intermediate periods without retraining. Our main findings are as follows: 1.Temporal misalignment occurs in existing retrieval. We show that time-unaware retrievers tend to retrieve se- mantically relevant but temporally misaligned documents from a document collection with mixed-timestamps. 2. Integrating preference optimization helps capture temporal awareness. We proposeTPOUR, which learns to prefer temporally aligned over misaligned documents, improving temporal retrieval and enabling timestamp prediction using minimal corpus-level supervision. 3.Time vectors enable continuous temporal generaliza- tion. We validate that time-vector interpolation (Nylund et al., 2024) can be applied toTPOUR-trained retrievers, allowing them to generalize to intermediate time peri- ods without additional training. We further show that extrapolation enables generalization to future time. 4.Temporal awareness reveals time sensitivity in gen- eral retrieval tasks. On the BEIR benchmark,TPOUR uncovers alignment between dataset publication year and optimal retrieval performance, suggesting that temporal modeling improves even general retrieval tasks. 2. Related Work 2.1. Unsupervised Learning for Retrieval Training Unsupervised learning has enabled retrievers to scale with large amounts of unlabeled documents, from early statis- tical methods (Jatowt et al., 2005; 2013; Berberich et al., 2010; Kanhabua & Nørv ̊ ag, 2010; Kanhabua et al., 2012) like BM25 (Robertson & Zaragoza, 2009) to recent neu- ral embedding models (Nussbaum & Duderstadt, 2025). While traditional approaches rely on statistics, unsupervised dense retrievers leverage contrastive learning. In dense re- trieval, DPR (Karpukhin et al., 2020) is a supervised dense retriever trained on labeled query-passage pairs. In con- trast, Contriever (Izacard et al., 2022) utilizes fully unsu- pervised contrastive learning. REALM (Guu et al., 2020) introduces retrieval-augmented masked language model- ing. SimCSE (Gao et al., 2021) applies in-batch contrastive learning for sentence embeddings. E5 (Wang et al., 2024b) extends this with weak supervision over large-scale web data. CPT (Neelakantan et al., 2022) shows that scaling con- trastive learning improves both text and code embeddings. GTE (Li et al., 2023) improves generalization by training on diverse datasets, while M3-Embedding (Chen et al., 2024) uses self-distillation to unify signals from multiple retrieval paradigms. Most recently, Nomic Embed v2 (Nussbaum & Duderstadt, 2025) adopts a sparse mixture-of-experts (MoE) for scalable and efficient general-purpose embedding. 2 Temporal Preference Optimization for Unsupervised Retrieval Contrastive learning is the core of unsupervised retriever training, where a queryQis paired with a positive docu- mentD + and a set of negative documentsD − 1 ,...,D − K . The loss (Eq., 1) is calculated using a similarity function S(·,·)with a query encoderπ q and document (i.e., key) encoderπ k . This loss encourages models to maximize sim- ilarity between a query and its positive document while minimizing similarity to negatives. However, embeddings are solely optimized for semantic similarity. As a result, retrievers such as Contriever (Izacard et al., 2022) degrade in mixed-timestamp document collection settings, failing to distinguish between documents from different time peri- ods. Letq = π q (Q)andd = π k (D)denote the query and document embeddings, respectively. The contrastive loss is: L CE =− log exp S(q,d + ) exp S(q,d + ) + P K i=1 exp S(q,d − i ) (1) Unsupervised retrieval training commonly utilizes either (1) in-batch negative (Lee et al., 2019), or (2) MoCo (Mo- mentum Contrast) (He et al., 2020). The former is effec- tive with large batch sizes, while MoCo simulates large batches with lower memory. In MoCo, the query encoder π q and key encoderπ k are updated during training. Af- ter updatingπ q ’s weightθ q via the contrastive loss in Eq. 1, the key encoder weightθ k is updated via momentum θ k ← m×θ k + (1−m)×θ q . In this work, we adopt MoCo for unsupervised retrieval to train under limited resources. 2.2. Temporal Relevance Modeling Temporal relevance has been explored in language mod- els (Lazaridou et al., 2021; R ̈ ottger & Pierrehumbert, 2021; Rosin et al., 2022; Su et al., 2023; Wang et al., 2023). For instance, (Dhingra et al., 2022) jointly models timestamps with text to improve temporal generalization in language modeling. In temporal information retrieval, recent work in- corporates temporal information for time-aware search (Wu et al., 2024; Abdallah et al., 2025). For example, (Gade et al., 2025) applies retrieval-augmented generation on ex- plicit temporal annotation for both queries and documents, and (Qian et al., 2024) addresses implicit temporal aware- ness through query rewriting over a knowledge graph. Another line of work extracts time vectors from genera- tive language models fine-tuned on data from distinct pe- riods (Nylund et al., 2024). These latent vectors capture temporal context and allow interpolation. They show that adjacent time vectors are close in weight space, enabling generalization to intermediate periods without retraining. We extend time vector extraction from generative language models toTPOUR, enabling continuous temporal alignment of retrievers to unseen intermediate and future periods. 2.3. Direct Preference Optimization RLHF (Reinforcement Learning from Human Feedback) aligns language models with human preferences (Ouyang et al., 2022). It involves training a reward model on human- labeled preferences and optimizing the policyπ θ to max- imize the reward using PPO (Proximal Policy Optimiza- tion) (Schulman et al., 2017) or DPO (Direct Preference Optimization) (Rafailov et al., 2023). L DPO =− logσ β log π θ (y w | x)π ref (y l | x) π θ (y l | x)π ref (y w | x) (2) Building on DPO, we introduce TRPO, which incorporates temporal preferences into unsupervised retrieval. TRPO con- structs preference pairs from document corpora across time and learns to prefer temporally aligned documents without explicit supervision. Unlike DPO, which aligns generation policies using human preferences, TRPO adapts preference optimization to retrieval by replacing log-likelihoods with embedding similarity from unlabeled temporal signals. 3. Temporal Preference Optimization for Unsupervised Retriever 3.1. Incorporating Temporal Preferences into Contrastive Learning We proposeTPOUR(Temporal Preference Optimization for Unsupervised Retriever), a training framework that inte- grates temporal preferences into contrastive learning for unsupervised retrieval. Built upon MoCo,TPOURjointly learns semantic similarity and temporal relevance by com- bining contrastive learning with a preference-based objec- tive from TRPO. This enables the retriever to encode both content relevance and implicit temporal preferences from unlabeled data. We illustrate this with a case study on unla- beled document training in Appendix E. As shown in Fig. 2, the training phase consists of a query documentQ i , a temporally aligned documentD t i , and an unaligned documentD t ′ i . The encoderπ θ encodes these inputs, while a momentum-based reference encoder π ref maintains a queue of negatives for contrastive learn- ing. The training objective combines two losses. The first is a contrastive loss that brings the query closer to its relevant document while distinguishing it from nega- tives, whereS(·,·)is the similarity function. Here, we defineS θ (y w i ) = S(π θ (Q i ),π θ (D t i )), which denotes sim- ilarity with the temporally aligned document (preferred) andS θ (y l i ) = S(π θ (Q i ),π θ (D t ′ i )), the similarity with the unaligned document (less preferred). The valuesS ref (y w j ) andS ref (y l j )correspond to negative pairs from the previous batch queue j, where D − j ∈D t j ,D t ′ j : 3 Temporal Preference Optimization for Unsupervised Retrieval Clemson Tigers football // Once again the teams did battle in the 2018 Sugar Bowl in New Orleans, Louisiana with a trip to the 2018 College Football Playoff National Championship game on the line ... (퐷 ! " ) Document Collected at Time 푡 National Football League Christmas games // On March 19, 2021, the NFL and its broadcast partners agreed to an 11-year contract which will run from 2023 through 2033. As part of the deal, Fox acquired .... (퐷 ! " ! ) Document Collected at Time 푡 # Encoder to Align 푡 Reference Encoder 푠푖푚( ,) 휋 ! 휋 "#$ 휋 ! (퐷 % & ! ) 휋 "#$ (퐷 % & ) 휋 "#$ (퐷 % & ! ) 휋 ! (퐷 % & ) 푠푖푚( ,) =푆 ! (푦 % ' ) =푆 ! (푦 % ( ) 푠푖푚( ,) 푠푖푚( ,) Clemson Tigers football // In the build-up to the 2018 Sugar Bowl, players and coaches from both teams referred to the series as a "respectful" rivalry. The next week, on January 8, 2018, Alabama would win the national... (푄 ! ) Query Document 휋 "#$ (푄) 휋 ! (푄 % ) =푆 "#$ (푦 % ' ) =푆 "#$ (푦 % ( ) Previous Batch Queue 휋 $%& (퐷 ' " ), 휋 $%& (퐷 ' " ! )| 푗=1,...,푖−1 ④Momentum Update ①Calculate similarity between un/aligned document ③Append to Queue ℒ )*)+, =휆ℒ -. 푆 ! 푦 % ' ,푆 "#$ 푦 / ' ,푆 "#$ 푦 / ( +1−휆 ℒ 0123 푆 ! 푦 % ' ,푆 ! 푦 % ( ,푆 "#$ 푦 % ' ,푆 "#$ 푦 % ( ②Calculate ℒ () and ℒ *+,- loss and backpropagate ℒ ./.01 3exp(푆 $%& (푦 ' 2 )) '3! +exp(푆 $%& (푦 ' 4 )) Figure 2. Overview ofTPOUR. Given a queryQ i and two documentsD t i (temporally aligned) andD t ′ i (temporally misaligned), each input is encoded using both the main encoderπ θ and the reference encoderπ ref .1Similarity scores are computed between the query and each document usingπ θ . 2 A contrastive lossL CE , which calculate semantic similarity betweenQ i andD t i , and a TRPO lossL TPRO for preferring temporally aligned documents are calculated to get combined lossL total .3The reference embeddingsπ ref (D t i )andπ ref (D t ′ i ) are added to a queue as negatives for future batches.4The encoderπ θ is updated viaL total , andπ ref is updated via momentum fromπ θ . L CE =− log e S θ (y w i ) e S θ (y w i ) + P j<i n e S ref (y w j ) + e S ref (y l j ) o (3) To model temporal preferences, TRPO aligns the preference gap between the current and reference models, whereS θ (y) andS ref (y)denote scores from the current and reference models given outputy. Given a pairy w i (preferred) and y l i (less preferred), the TRPO loss is defined as Eq. 4. A detailed theoretical basis of TRPO is in Appendix B.1. L TRPO =− logσ β S θ (y w i )− S θ (y l i ) − S ref (y w i )− S ref (y l i ) (4) The total loss is computed asL total = λL CE + (1−λ)L TRPO , whereλ ∈ [0, 1]balances the influence of semantic and temporal signals. The encoderπ θ is optimized usingL total , while the reference encoder weightsθ ref are updated via momentum asθ ref ← m× θ ref + (1− m)× θ, wherem is the momentum coefficient andθis the current weight of π θ . After training,TPOUR-trained retrievers can be used as general-purpose retrieval systems through the standard inference pipeline, as illustrated in Appendix Fig. 6. 3.2. Continuous Temporal Representation Discrete temporal models are inherently limited in modeling continuous time. Since time is inherently continuous, a retriever needs to generalize to queries that fall between the temporal regions covered by separately trained retrievers. To overcome this limitation, we adopt time vector extraction from language modeling (Nylund et al., 2024) and extend it to TPOUR for unsupervised retrieval. We extract time vectors fromTPOUR-trained retrievers fine- tuned on specific time periods (e.g., the years 2018 and 2021). Interpolating between these vectors allows the model to adjust its temporal alignment and generalize to interme- diate periods without retraining. Tab. 3, Fig. 3, and Fig. 4 show the generalization capability through time vector in- terpolation across continuous time shifts. Formally, letθ base denote the base encoder weight andθ t the encoder weight fine-tuned on data from time period t. The time vectorτ t for time periodtis computed as τ t = θ t −θ base , whereτ t captures the temporal shift between the base model and the model adapted to time periodt. To obtain an encoder for an intermediate time periodt mid , given two time vectorsτ t start andτ t end corresponding to the t start (earlier) andt end (later), respectively, we interpolate using a coefficientα ∈ [0, 1], as defined in Eq. 5. Further 4 Temporal Preference Optimization for Unsupervised Retrieval theoretical details are provided in Appendix Sec. B.2. θ t mid = θ base + (1− α)τ t start + ατ t end , where t start ≤ t mid ≤ t end (5) This interpolation allows the model to adjust its temporal alignment without retraining. For example, interpolating between 2018 and 2021 vectors enables retrieval for queries from 2019 or 2020. Tab. 12, Tab. 13, and Tab. 14 in the Appendix show that interpolation improves generalization to intermediate time periods even when the temporal infor- mation is not given in the query. 3.3. Inferring Document Timestamps fromTPOUR In addition to the retrieval,TPOURcan be used to infer a document’s timestamp. Following (Gunasekaran et al., 2023), we formulate timestamp inference as a classification task and introduce a timestamp predictor based on a mixture ofTPOURretrievers (mixture-of-TPOUR). As illustrated in Appendix Fig. 7, the mixture-of-TPOURuses a set of frozen retrieversπ t 1 θ ,...,π t n θ , each specialized for a distinct time periodt i . Given a documentD, each retriever encodesD into a temporally-aware embedding. These embeddings are concatenated and passed to a shared trainable linear classification head to train and predict the timestamp. We compare against a baseline predictor that uses a single frozen retrieverπ θ trained on the full time range. To ensure a fair comparison, we match the number of trainable param- eters by stacking multiple linear layers in the baseline clas- sifier, with the same depth as the number of retrievers in the mixture model, and also compare model with larger param- eter counts. The result shows mixture-of-TPOURachieves superior timestamp prediction performance (Tab. 4). 4. Experiments and Analysis This section presents experiments to answer three main research questions regarding TPOUR: RQ1. DoTPOUR-trained retrievers learn temporally aligned representations? We evaluate whetherTPOUR- trained retrievers retrieve temporally aligned documents and whether interpolation and timestamp prediction reveal embedded temporal representations in the retriever. RQ2. Does temporal awareness improve performance on temporal QA tasks? We assess temporal awareness by evaluating retrieval on temporal QA across time splits and measuring gains in periods via time vector interpolation. RQ3. Can temporal awareness reveal time sensitivity in general retrieval tasks? We conduct a case study on the BEIR benchmark (Thakur et al., 2021), spanning diverse domains and publication years, to assess whether temporal awareness inTPOUR-trained models reveals time sensitivity. Table 1. Evaluation bias test on SituatedQA. To confirm that the dataset construction process is free from bias in benchmark creation, we built a separate gold document collection with Nomic Embed v2 MoE and DPR as the retriever and averaged their results. The performance trends ofTPOURContriever (2018/2021) remain consistent, showing that retriever does not affect benchmark bias. SituatedQA2018/N@52018/N@102021/N@52021/N@10 Contriever31.2834.9435.2138.13 DPR (Dense Passage Retriever)27.2131.1333.8236.37 Nomic Embed v2 MoE31.3034.7434.5236.96 TPOUR Contriever (2018)38.7041.1311.5714.17 TPOUR Contriever (2021)23.1527.6342.9343.62 4.1. Evaluation Benchmarks and Metrics To assessTPOUR, we use two temporal QA datasets, Situ- atedQA and RealTimeQA, for temporal retrieval, and the BEIR for general retrieval tasks. (1) SituatedQA (Zhang & Choi, 2021) is a yearly temporal QA dataset contain- ing 2,795 queries spanning 1700–2021. Since years prior to 2018 each have fewer than 130 queries, we focus on the 2018–2021 subsets, which contain 291, 411, 501, and 491 queries, respectively. (2) RealTimeQA (Kasai et al., 2023) is a monthly temporal QA dataset, providing weekly evaluations from June 2022 to January 2024, with approxi- mately 130 queries per month. For evaluation, we use the queries from January to December 2023. (3) BEIR (Thakur et al., 2021) is a general retrieval benchmark comprising 18 datasets across diverse domains (e.g., medical, financial). We use BEIR to show that temporal awareness reveals time sensitivity in general retrieval tasks. SituatedQA provides only queries and associated answers, while RealTimeQA includes a query, a single associated document, and an answer, which is still insufficient for evaluating retrieval performance, since duplicated docu- ments created or updated at different timestamps are not present. To address this, we construct a custom retrieval benchmark based on these datasets, following the BEIR custom dataset guidelines (Thakur, 2022) to create a tem- poral QA benchmark tailored for retrieval evaluation. Each custom dataset requires a set of documents related to each query. To construct these, we use Contriever (Izacard et al., 2022) to retrieve the top-10 documents per query from a fixed document collection. For instance, when building the document set for queries from the 2018 test set, we use the 2018 Wikipedia document collection, retrieve the top-10 documents using Contriever (Izacard et al., 2022), filter out documents that do not contain the answer, retaining only the answer-containing ones as gold documents. We also per- form an evaluation bias test with a different retriever (DPR, Nomic-Embed v2 MoE) to check whether the performance trends remain, as reported in Tab. 1. We evaluate retrieval performance using normalized dis- counted cumulative gain (nDCG@k, denoted as N@k), which captures relevance and ranking in the top-k. 5 Temporal Preference Optimization for Unsupervised Retrieval Recall@k, the percentage of queries with at least one cor- rect document in the top-k, is reported in Appendix D. For timestamp prediction, we report accuracy, which is the ratio of correct predictions to total examples. 4.2. Training Datasets We construct our training corpus from English Wikipedia database dumps (Johnson et al., 2024) collected at different times to capture temporal differences, retaining newly added or modified document content across the corpus. For the yearly corpus, we use Wikipedia dumps from December 2018 and 2021, which serve as the yearly time span used for SituatedQA. For the monthly corpus, we use dumps from January and December 2023 for RealTimeQA. An additional dump is also used for temporal diversity. To prevent data leakage, we filter out documents that serve as gold documents in SituatedQA and RealTimeQA. Details on training data construction are provided in Appendix A.2 and Tab. 7, and the training setup in Appendix A.3.1. 4.3. Baselines We consider three types of baselines. Standard Retrievers are retrieval models that rank documents primarily based on semantic relevance between the query and document. Temporal-Aware Retrievers are retrieval models that in- corporate temporal signals, such as timestamps or temporal constraints, to better align retrieved documents with the time specified or implied by the query. Large Embedding Models are recent large-scale embedding models trained with broad instruction-following and retrieval objectives, which provide strong general-purpose retrieval performance and can potentially handle temporal intent through their pre- trained representations or prompting. Detailed information about each baseline is provided in Appendix C.2. Standard Retriever: (1) DPR (Karpukhin et al., 2020) is a supervised bi-encoder with 110M parameters, trained with BM25 hard negatives. (2) REALM (Guu et al., 2020) is a 134M parameter retriever that combines retrieval with lan- guage modeling in an end-to-end setup. (3) SimCSE (Gao et al., 2021) is a 110M parameter retriever that learns sentence embeddings via contrastive learning and can be adapted for retrieval. (4) Contriever (Izacard et al., 2022) is an unsupervised retriever with 110M parameters, trained via MoCo-based contrastive learning. Time-Aware Retriever: (1) Berberich et al. (2010) is an early probabilistic model that explored temporal expres- sions represented as tuples. (2) Temporal Contrastive is a temporal-aware contrastive baseline that augments the standard contrastive retrieval objective with temporal super- vision. We include this baseline to examine whether tem- poral alignment can be obtained by directly learning time- based positive and negative documents, without preference- based optimization (i.e. TRPO). For each query, temporally aligned documents are treated as positives and misaligned documents as negatives, yieldingL TempCE , which encour- ages higher similarity between the query and documents that better match the target time. The final objective is L total = λL CE + (1− λ)L TempCE , whereL CE models semantic relevance andL TempCE models temporal align- ment. (3) TimeR 4 (Qian et al., 2024) proposes a time-aware retriever with 113M parameters, trained on temporal knowl- edge graphs. We use their public checkpoint for comparison. Large Embedding Model:(1) Nomic Embed v2 MoE (Nussbaum & Duderstadt, 2025) is a recent general- purpose embedding model with 475M parameters, utiliz- ing a sparse mixture-of-experts architecture. (2) Qwen-3- Embedding-8B (Zhang et al., 2025) is a large-scale embed- ding model with 8B parameters, built upon Qwen3 (Yang et al., 2025). It supports diverse embedding and reranking tasks across multiple domains and languages. We consider retrieval with naive, query rewriting (QR), and time-aware instruction retrieval (TAI), since Qwen-3-Embedding sup- ports instruction-conditioned embeddings. We apply TAI to explicit temporal information, as its target time can be di- rectly incorporated in the instruction. We report our prompt used for TAI in the Appendix Tab. 10. 4.4. Results and Analysis 4.4.1. DO TPOUR-TRAINED RETRIEVERS LEARN TEMPORALLY ALIGNED REPRESENTATIONS? We evaluate whetherTPOURlearns temporally aligned rep- resentations by analyzing the document timestamps distri- bution and timestamp prediction. Fig. 3 shows that inter- polationαsmoothly shifts retrieval distributions toward intermediate time periods. Full distributions for SituatedQA and RealTimeQA are in Appendix Fig. 8 and 9. Notably, TPOUR also captures temporal patterns without explicit su- pervision (Appendix Sec. D.7; Tab. 18). To further assess whetherTPOURencodes temporal in- formation, we evaluate its timestamp prediction accuracy as a classification task, with year prediction as 4 classes (2018–2021) and month prediction as 12 classes. As shown in Tab. 4, the mixture-of-TPOURachieves 76.56% year accu- racy and 27.41% month accuracy, outperforming the base- line predictor built on Contriever (50.18% year, 22.22% month accuracy) with 10,000 training steps. The evalua- tion loss also decreases from 3.13 to 2.66. Furthermore, Mixture-of-TPOURwith 2 encoders (220M) surpasses the larger-capacity Nomic-Embed v2 MoE (305M), achieving a 21.53% improvement on year accuracy and 20.30% on month accuracy. These results indicate thatTPOURem- beddings preserve temporal signals for inference tasks that are both temporally aligned and predictive. The detailed evaluation setup is in Appendix Fig. 7. 6 Temporal Preference Optimization for Unsupervised Retrieval Table 2. Retrieval performance on mixed-timestamp document collections across SituatedQA and RealTimeQA. We compare standard baselines using their public checkpoints against theTPOUR-trained retriever with three training seeds (mean and standard deviation reported with±), denoted as TPOUR Contriever (t). TPOUR Contriever outperforms the baselines across time periods, achieving higher accuracy and stronger generalization regardless of whether queries contain explicit or implicit temporal information. Notably,TPOUR Contriever achieves strong performance on intermediate periods (2019, 2020, and June) without requiring time-specific retraining. SituatedQARealtimeQA 2018201920202021JanuaryJuneDecember Retriever N@5N@10N@5N@10N@5N@10N@5N@10N@5N@10N@5N@10N@5N@10 Query with Explicit Temporal Information Contriever29.3033.3529.6734.4931.2535.7737.8541.0521.7622.3633.0433.1245.9643.99 REALM22.3721.5712.2614.0413.9215.239.3410.0314.6614.2218.2316.8319.5717.83 SimCSE25.1727.5619.6222.5717.8419.7118.2520.9418.4017.9821.4621.7319.1819.56 Supervised:DPR28.6731.2027.5830.7627.9131.2432.7634.6222.9022.5530.2729.0335.6534.49 Berberich et al. (2010)8.649.419.3610.158.489.639.9110.8815.8416.338.618.7422.4724.91 Temporal Contrastive35.00 38.6029.0334.7629.2133.7235.9737.3925.6925.1933.9932.2239.8837.12 TimeR 4 33.6537.7127.6232.2431.0934.7131.3334.9726.4525.3431.3629.478.948.60 Nomic Embed v2 MoE29.6133.6229.6732.4630.7735.1731.0933.7422.3822.1231.4530.6436.9936.28 Qwen3-Embedding-8B30.4533.7732.7735.2636.3140.0635.1737.8522.5423.3033.8633.0041.0639.07 TPOUR Contriever (2018)43.93±0.346.66±0.131.25±0.434.11±1.124.56±2.527.44±2.718.25±3.821.58±3.5— TPOUR Contriever (2021)24.38±3.029.06±3.123.63±3.427.59±3.826.60±1.729.68±1.940.21±0.744.72±6.8— TPOUR Contriever (Jan)—31.78±0.531.37±0.731.96±2.231.05±2.131.01±1.530.51±1.5 TPOUR Contriever (Dec)—9.21±0.79.51±0.327.43±1.426.30±1.448.24±1.645.57±1.0 TPOUR Contriever43.93±0.346.66±0.332.87±0.136.75±0.330.56±1.133.66±0.940.21±0.744.72±6.831.78±0.531.37±0.732.82±0.431.90±0.448.24±1.645.57±1.0 Query with Implicit Temporal Information Contriever29.8934.6030.9636.2031.0034.4333.0637.0827.3828.4832.4632.8340.0338.77 REALM22.4122.0913.2214.7415.3516.2910.4111.0916.4416.1219.8618.3919.9518.30 SimCSE24.3127.2622.6626.4222.6724.3620.7323.6823.0023.0826.3726.7522.9523.83 Supervised:DPR32.7535.1830.3134.4528.4631.7131.1734.2929.2128.3827.8728.3833.8232.45 Berberich et al. (2010)9.149.358.149.377.438.198.058.6813.2113.357.137.7920.4322.91 Temporal Contrastive34.6437.1833.4536.8131.1335.1233.5937.7831.5830.2634.6233.4439.5537.76 TimeR 4 35.5039.5228.4332.9032.0236.1730.8634.2226.9526.088.798.3832.7632.52 Nomic Embed v2 MoE29.2333.2730.6133.8130.5534.8630.4233.0825.7825.8830.6631.3735.8234.94 Qwen3-Embedding-8B28.0732.6033.2733.5235.9139.5733.9136.4524.2524.4032.2532.5741.5138.30 TPOUR Contriever (2018)44.11±0.546.59±0.231.23±0.434.49±0.524.80±2.127.87±2.018.53±3.321.87±2.9— TPOUR Contriever (2021)24.81±2.629.47±2.826.48±1.630.52±1.328.31±1.331.84±1.839.40±1.044.72±6.8— TPOUR Contriever (Jan)—32.03±0.831.67±1.231.97±2.231.37±1.530.48±2.430.25±2.0 TPOUR Contriever (Dec)—10.34±1.310.89±2.228.73±4.227.58±3.649.29±3.446.19±2.1 TPOUR Contriever44.11±0.546.59±0.234.36±0.238.63±0.433.14±0.236.31±0.539.40±1.044.72±6.832.03±0.831.67±1.231.73±0.932.86±0.449.29±3.446.19±2.1 Table 3. Interpolation ofTPOURContriever betweent start and t end periods reduces temporal misalignment in intermediate peri- ods. The result shows (1) interpolation enables generalization in middle time (α = 0.5). And (2) it can surpass directly fine-tuned retriever (Eval-year fine-tuned vs. Best interpolation α). Method SituatedQARealTimeQA nDCG@5nDCG@10nDCG@5nDCG@10 TPOUR Contriever (t start )29.0532.1330.2929.87 TPOUR Contriever (t end )31.7234.3029.3227.97 α = 0.535.7139.2330.4730.36 Best Interpolation α42.4744.5938.7737.30 Eval-year fine-tuned42.4743.96 37.3036.98 Table 4. Performance of the mixture-of-TPOURtimestamp pre- dictor after 10k training steps. The mixture-of-TPOURmodel achieves the lowest evaluation loss (Eval Loss), as well as the high- est year accuracy (Y-Acc) and month accuracy (M-Acc). It also outperforms the larger size Nomic-Embed v2 MoE (305M) when compared to a mixture-of-TPOUR with two encoders (220M). Eval Loss↓M-Acc↑Y-Acc↑ Contriever3.1322.2250.18 Nomic-Embed v2 MoE3.545.5753.03 Mixture-of-TPOUR (2 Encoders)2.7625.8774.56 Mixture-of-TPOUR (10 Encoders)2.6627.4176.56 4.4.2. DOES TEMPORAL AWARENESS IMPROVE PERFORMANCE ON TEMPORAL QA TASKS? We evaluate the impact of temporal awareness on re- trieval using SituatedQA and RealTimeQA. Tab. 2 shows nDCG@5/10 across different test periods. For interpolated TPOURContriever (denoted asTPOURContriever), we ap- ply a heuristic interpolation strategy, selecting the interpo- lated model whoseαcorresponds to thet mid of the test set. For example, for the 2019 set, we use the interpolated model withα = 0.3. TheTPOURContriever consistently outper- forms all baselines. On SituatedQA 2018 (Implicit),TPOUR achieves an nDCG@5 of 44.11, substantially surpassing Contriever (29.89). Similar improvements are observed across later years, including +3.40 nDCG@5 in 2019 and +6.34 in 2021 over Contriever. On RealTimeQA,TPOUR also maintains an advantage across months in January and December. Notably, performance gains remain consistent across time periods, regardless of whether temporal infor- mation is provided explicitly or implicitly in the query. Tab. 5 further comparesTPOURContriever with Qwen3- Embedding-8B, a substantially larger embedding model, un- der query rewriting (QR) and time-aware instruction (TAI) settings. Although QR improves Qwen3-Embedding-8B on the 2018 test set, its gains are not consistent across years. Similarly, TAI improves performance across both years for explicit queries, but still remains belowTPOURContriever. In contrast,TPOURContriever achieves the best perfor- mance across all explicit and implicit settings, improving 7 Temporal Preference Optimization for Unsupervised Retrieval 0.370.280.190.15 0.330.310.220.15 0.310.270.270.15 0.320.280.240.17 2018201920202021 2021 2020 2019 2018 0.300.270.220.21 0.210.280.270.23 0.190.220.310.28 0.200.230.270.30 2018201920202021 0.250.260.230.26 0.150.250.280.32 0.120.180.310.39 0.120.180.270.43 2018201920202021 0.190.250.240.32 0.090.210.280.42 0.070.130.290.51 0.070.130.240.56 2018201920202021 훼 = 0.0훼 = 0.3훼 = 0.7훼 = 1.0 Figure 3. Distribution of retrieved document timestamps with time vector interpolation. Heatmaps show the normalized distribution of retrieved document timestamps in years (x-axis) for each test year (y-axis) on SituatedQA. Each heatmap corresponds to aTPOUR Contriever interpolated between retrievers trained ont start = 2018andt end = 2021, using weightsα, where0.0represents the 2018 and 1.0represents the 2021 model. Retrieved documents are concentrated around the test year when the interpolation weights align, and shift across intermediate years (2019, 2020) as interpolation value changes, showing temporal alignment in intermediate years. 00.51 10 20 30 40 50 00.5100.5100.51 2018201920202021JanJunDec ExplicitImplicitExplicitImplicit nDCG@10 Yearly InterpolationMonthly Interpolation Figure 4. Temporal retrieval performance of interpolatedTPOURContriever. nDCG@10 on Left: SituatedQA (Yearly) and Right: RealTimeQA (Monthly) using interpolatedTPOURContriever betweenπ t start θ andπ t end θ (2018/2021 for SituatedQA, January/December 2023 for RealTimeQA), evaluated with explicit and implicit temporal information in queries. The x-axis indicates the interpolation weightαbetween 2018 and 2021. Each colored line denotes an evaluation set, and star markers (⋆) indicate the interpolation achieving peak performance. Peaks aligning with the corresponding time period show temporal generalization across intermediate periods. Table 5.TPOURContriever outperforms 72.7×larger Qwen3- Embedding-8B variants with query rewriting (QR) and tempo- rally aware instruction (TAI) on explicit and implicit SituatedQA, showing that large embedding models alone cannot fully resolve temporal misalignment without temporal preference optimization. 20182021 RetrieverN@5N@10N@5N@10 Query with Explicit Temporal Information Qwen3-Embedding-8B30.4533.7735.1737.85 Qwen3-Embedding-8B (QR)40.9244.0734.4637.23 Qwen3-Embedding-8B (TAI)32.6935.4838.5140.64 TPOUR Contriever43.9346.6640.2144.72 Query with Implicit Temporal Information Qwen3-Embedding-8B28.0732.6033.9136.45 Qwen3-Embedding-8B (QR)29.1232.8431.7734.48 TPOUR Contriever44.1146.5939.4044.72 nDCG@5 over the strongest Qwen3-Embedding-8B variant by +3.01 in 2018 and +1.70 in 2021 for explicit queries, and by +14.99 in 2018 and +5.49 in 2021 for implicit queries. Tab. 3 shows interpolatedTPOURContriever performance. On SituatedQA, interpolated retrievers achieve an average improvement of +13.4 nDCG@5 over the start-year retriever and +10.8 over the end-year retriever, relative to the best interpolation setting. RealTimeQA shows similar trends, with interpolation improving nDCG@5 by +9.0 points on average compared to retrievers trained on fixed January or December snapshots. Importantly, interpolated retrievers match or outperform retrievers trained directly at the mid- dle time (i.e.,α = 0.5), demonstrating that interpolation enables continuous generalization across time without ex- plicit retraining. Full results across all years and months are provided in Tab. 12, 13, and 14 in the Appendix. Fig. 4 illustrates how interpolation enablesTPOURto adapt to continuous time shifts. Retrieval performance peaks when the interpolation weight aligns with the test times- tamp. For instance, interpolatedTPOURContriever achieves peak nDCG@10 on the 2019 (green line) and 2020 (blue line) test sets in SituatedQA when interpolation is around the intermediate period. Similarly, on RealTimeQA, the interpolated retriever peaks on the June test set (orange line). We also conduct an ablation study on the loss weight λ, which balances semantic and temporal supervision, as shown in Appendix Fig. 10. We find that moderate values of λ (0.7–0.85) yield the optimal performance. 4.4.3. CAN TEMPORAL AWARENESS REVEAL TIME SENSITIVITY IN GENERAL RETRIEVAL? To assess whether temporal awareness can provide insights into general retrieval tasks, we evaluateTPOURon the BEIR benchmark spanning diverse domains and creation years. As shown in Fig. 5 and Appendix Tab. 11, interpolatedTPOUR Contriever between 2018 and 2021, along with interpolation valuesαfor 2021, reveal clear trends. Older datasets (e.g., MS MARCO) perform best whenα = 0.0, while newer 8 Temporal Preference Optimization for Unsupervised Retrieval MS MARCO TREC-COVID NFCorpus NQ HotpotQA FiQA ArguAna Touché-2020 Quora DBPedia SciDocsFEVER Climate-FEVER SciFact 20162017201820192020 0 0.5 1 Dataset Creation Year Best Interpolation ( 훼 ) Figure 5. Best-performing interpolationαfor each BEIR dataset relative to its creation year. Each point denotes a dataset, where αis the interpolation weight for 2021 betweenTPOURContriever (2018) and (2021). The red regression line indicates that datasets prefer retrievers temporally aligned with their publication year. For example, Climate-FEVER (2020) achieves peak performance atα = 0.7. Time-sensitive datasets such as TREC-COVID favor higherα, whereas less sensitive ones (SciFact, SciDocs) perform well with lower weights. Full results are in Appendix Tab. 11. datasets (e.g., TREC-COVID and Climate-FEVER) peak when interpolated toward 2021 (i.e.,α = 1.0). These results show that temporal awareness reveals time sensitivity in retrieval, aligning with dataset years. We conduct a qualitative case study comparing outputs from Contriever andTPOURContriever. As shown in Appendix Tab. 20 and 21,TPOURContriever retrieves documents that are both semantically relevant and temporally aligned with the query. For example, given “When did the Golden State Warriors win the Finals as of 2018,”TPOURContriever returns documents about the 2018 NBA Finals, whereas Contriever retrieves general descriptions of the NBA Finals. Similarly, for “Who has won the most Olympic medals in curling as of 2021,”TPOURContriever retrieves temporally aligned documents, whereas Contriever returns older ones. 4.4.4. EXTRAPOLATING TO FUTURE TIME PERIODS TPOURuses time-vector interpolation to generalize to inter- mediate time periods without additional training. Extending this idea to future or more recent time periods would further improve its practical utility. Thus, we conduct an anal- ysis of time-vector extrapolation for future time periods. Specifically, we construct an extrapolated retriever by com- bining three temporally distinct time vectors extracted from TPOURmodels trained on the 2018, 2021, and 2022 docu- ment dumps. We define the extrapolatedTPOURretriever asθ future = θ base + (1− α)τ t 2018 + α (τ t 2022 − τ t 2021 ), where αcontrols the extrapolation strength. Here,τ t 2018 denotes the base time vector from which extrapolation is performed, whileτ t 2021 andτ t 2022 denote time vectors obtained from later document dumps. The difference vectorτ t 2022 − τ t 2021 cap- tures the temporal direction from an earlier to a later period. Table 6. Results of time vector extrapolation using RealtimeQA (2023, December) test set. The extrapolated model gained using three temporally outdated retrievers (2018/2021/2022) achieves higher performance than temporally outdated checkpoints. ModelN@5N@10 Oracle: TPOUR Contriever (2023)42.4546.15 TPOUR Contriever (2018)21.4520.13 TPOUR Contriever (2021)23.9325.28 TPOUR Contriever (2022)27.7827.99 Extrapolated TPOUR (α = 0.5)30.0030.40 By adding this direction toτ t 2018 , we approximate a future- oriented time vector beyond the observed training periods. Tab. 6 shows time vector extrapolation performance using the RealTimeQA (2023, December) test set. The results show that an appropriate extrapolation strength (α = 0.5) enablesTPOURto approximate future time periods more effectively, outperforming the most recent 2022TPOUR Contriever baseline in N@5 (30.00 vs. 27.78). 5. Conclusion and Future Work We proposeTPOUR, a preference-based training method at the embedding level that injects temporal information into unsupervised dense retrievers. By integrating our TRPO into contrastive learning,TPOURenables retrievers to learn both semantic similarity and temporal preferences from un- labeled data. We show that time-unaware retrievers suffer from temporal misalignment and that training with TRPO improves on temporal retrieval tasks on SituatedQA and Re- alTimeQA. We further show that time vector interpolation allowsTPOUR-trained retrievers to generalize across contin- uous time periods without retraining. Beyond temporal re- trieval,TPOURretrievers also exhibit temporal preferences on the BEIR benchmark, indicating that temporal modeling benefits both time-sensitive and general retrieval tasks. We show thatTPOURimproves temporal retrieval, and sev- eral promising directions remain for future work. (1) Re- laxing the requirement for temporally distributed document collections could broaden applicability. (2) Further analysis of temporal grounding could enhance interpretability across implicit and explicit queries, as the benefits ofTPOURare more pronounced in explicit than in implicit setups. (3) We show that temporal alignment relates to general retrieval. Further studies could expand its usability (e.g., appropriate αselection). Our current setup setsαheuristically based on the test-set time (e.g.,α = 0.3between 2018 and 2021 re- trievers for the 2019 test set). (4) Time vector extrapolation could enableTPOUR-trained retriever to generalize beyond the training period. Our preliminary results (Sec. 4.4.4, Tab. 6) show thatTPOURcan be applied to extrapolation. We provide more details on each aspect of our future work and a more detailed analysis in Appendix F. 9 Temporal Preference Optimization for Unsupervised Retrieval Impact Statement This paper presents a method for improving temporal align- ment in unsupervised information retrieval systems. Im- proved temporal grounding can enhance the reliability of retrieved information. The method uses existing document corpora. As such, it does not directly raise concerns related to privacy or content misuse. Acknowledgments We would like to thank the anonymous reviewers for their helpful questions and comments. This work was partly supported by Institute of Information & communications Technology Planning & Evaluation(IITP) grant funded by the Korea government(MSIT) (RS-2019-I190421, AI Grad- uate School Support Program(Sungkyunkwan University) & RS-2025-02263169, Detection and Prediction of Emerging and Undiscovered Voice Phishing & RS-2024-00398115 , Research on the reliability and coherence of outcomes produced by Generative AI). This work was supported by the Ministry of Education of the Republic of Korea and the National Research Foundation of Korea (NRF-RS-2025- 00523385). References Abdallah, A., Piryani, B., Wallat, J., Anand, A., and Jatowt, A. Tempretriever: Fusion-based temporal dense passage retrieval for time-sensitive questions, 2025. URLhttps: //arxiv.org/abs/2502.21024. Attardi, G. Wikiextractor.https://github.com/ attardi/wikiextractor, 2015. Berberich, K., Bedathur, S., Alonso, O., and Weikum, G. A language modeling approach for temporal information needs. In Proceedings of the 32nd European Confer- ence on Advances in Information Retrieval, ECIR’2010, p. 13–25, Berlin, Heidelberg, 2010. Springer-Verlag. ISBN 3642122744. doi: 10.1007/978-3-642-12275-05. URLhttps://doi.org/10.1007/978-3-642- 12275-0_5. Bobadilla, J., Ortega, F., Hernando, A., and Guti ́ errez, A.Recommender systems survey.Knowledge- Based Systems, 46:109–132, 2013.ISSN 0950- 7051.doi: https://doi.org/10.1016/j.knosys.2013.03. 012. URLhttps://w.sciencedirect.com/ science/article/pii/S0950705113001044. Brin, S. and Page, L.The anatomy of a large-scale hypertextual web search engine. Computer Networks and ISDN Systems, 30(1):107–117, 1998. ISSN 0169-7552. doi: https://doi.org/10.1016/S0169-7552(98)00110-X. URLhttps://w.sciencedirect.com/ science/article/pii/S016975529800110X. Proceedings of the Seventh International World Wide Web Conference. Chen, J., Xiao, S., Zhang, P., Luo, K., Lian, D., and Liu, Z. M3-embedding: Multi-linguality, multi- functionality, multi-granularity text embeddings through self-knowledge distillation. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics: ACL 2024, p. 2318– 2335, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024. findings-acl.137. URLhttps://aclanthology. org/2024.findings-acl.137/. Deng, Z., Zhong, Y., Guo, S., and Huang, W. Insclr: Improv- ing instance retrieval with self-supervision. Proceedings of the AAAI Conference on Artificial Intelligence, 36: 516–524, 06 2022. doi: 10.1609/aaai.v36i1.19930. Dhingra, B., Cole, J. R., Eisenschlos, J. M., Gillick, D., Eisenstein, J., and Cohen, W. W. Time-aware language models as temporal knowledge bases. Transactions of the Association for Computational Linguistics, 10:257– 273, 03 2022. ISSN 2307-387X. doi: 10.1162/tacla 00459. URLhttps://doi.org/10.1162/tacl_ a_00459. Fan, W., Ding, Y., Ning, L., Wang, S., Li, H., Yin, D., Chua, T.-S., and Li, Q. A survey on rag meeting llms: Towards retrieval-augmented large language models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’24, p. 6491–6501, New York, NY, USA, 2024. Association for Computing Machinery. ISBN 9798400704901. doi: 10. 1145/3637528.3671470. URLhttps://doi.org/ 10.1145/3637528.3671470. Gade, A., Jetcheva, J. G., and Trivedi, H. It’s about time: In- corporating temporality in retrieval augmented language models. In 2025 IEEE Conference on Artificial Intelli- gence (CAI), p. 75–82, 2025. doi: 10.1109/CAI64502. 2025.00019. Gao, T., Yao, X., and Chen, D. SimCSE: Simple con- trastive learning of sentence embeddings. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Meth- ods in Natural Language Processing, p. 6894–6910, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main.552. URLhttps:// aclanthology.org/2021.emnlp-main.552. Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Wang, M., and Wang, H. Retrieval-augmented 10 Temporal Preference Optimization for Unsupervised Retrieval generation for large language models: A survey, 2024. URL https://arxiv.org/abs/2312.10997. Goodfellow, I. J. and Vinyals, O. Qualitatively character- izing neural network optimization problems. In Bengio, Y. and LeCun, Y. (eds.), 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceed- ings, 2015. URLhttp://arxiv.org/abs/1412. 6544. Gunasekaran, K. P., Chase Babrich, B., Shirodkar, S., and Hwang, H.Text2time: Transformer-based arti- cle time period prediction. In 2023 IEEE 6th Inter- national Conference on Pattern Recognition and Ar- tificial Intelligence (PRAI), p. 449–455, 2023. doi: 10.1109/PRAI59366.2023.10331985. Guu, K., Lee, K., Tung, Z., Pasupat, P., and Chang, M. Retrieval augmented language model pre-training. In I, H. D. and Singh, A. (eds.), Proceedings of the 37th International Conference on Machine Learn- ing, volume 119 of Proceedings of Machine Learn- ing Research, p. 3929–3938. PMLR, 13–18 Jul 2020. URLhttps://proceedings.mlr.press/ v119/guu20a.html. He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R. Mo- mentum contrast for unsupervised visual representation learning. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 9726–9735, 2020. doi: 10.1109/CVPR42600.2020.00975. Izacard, G., Caron, M., Hosseini, L., Riedel, S., Bo- janowski, P., Joulin, A., and Grave, E. Unsupervised dense information retrieval with contrastive learning. Trans. Mach. Learn. Res., 2022, 2022. URLhttps: //openreview.net/forum?id=jKN1pXi7b0. Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. P., and Wilson, A. G. Averaging weights leads to wider optima and better generalization. In Globerson, A. and Silva, R. (eds.), Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence, UAI 2018, Monterey, California, USA, August 6-10, 2018, p. 876–885. AUAI Press, 2018. URLhttp://auai.org/uai2018/ proceedings/papers/313.pdf. Jatowt, A., Kawai, Y., and Tanaka, K. Temporal rank- ing of search engine results. In Proceedings of the 6th International Conference on Web Information Sys- tems Engineering, WISE’05, p. 43–52, Berlin, Heidel- berg, 2005. Springer-Verlag. ISBN 3540300171. doi: 10.1007/115810624. URLhttps://doi.org/10. 1007/11581062_4. Jatowt, A., Au Yeung, C.-M., and Tanaka, K.Esti- mating document focus time. In Proceedings of the 22nd ACM International Conference on Information & Knowledge Management, CIKM ’13, p. 2273–2278, New York, NY, USA, 2013. Association for Comput- ing Machinery. ISBN 9781450322638. doi: 10.1145/ 2505515.2505655.URLhttps://doi.org/10. 1145/2505515.2505655. Johnson, I., Kaffee, L.-A., and Redi, M. Wikimedia data for AI: a review of wikimedia datasets for NLP tasks and AI-assisted editing. In Lucie-Aim ́ e, L., Fan, A., Gwad- abe, T., Johnson, I., Petroni, F., and van Strien, D. (eds.), Proceedings of the First Workshop on Advancing Natural Language Processing for Wikipedia, p. 91–101, Miami, Florida, USA, November 2024. Association for Com- putational Linguistics. doi: 10.18653/v1/2024.wikinlp- 1.14. URLhttps://aclanthology.org/2024. wikinlp-1.14/. Kanhabua, N. and Nørv ̊ ag, K.Determining time of queries for re-ranking search results. In Proceedings of the 14th European Conference on Research and Ad- vanced Technology for Digital Libraries, ECDL’10, p. 261–272, Berlin, Heidelberg, 2010. Springer-Verlag. ISBN 3642154638. Kanhabua, N., Berberich, K., and Nørv ̊ ag, K. Learning to select a time-aware retrieval model. In Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’12, p. 1099–1100, New York, NY, USA, 2012. Association for Computing Machinery. ISBN 9781450314725. doi: 10. 1145/2348283.2348488. URLhttps://doi.org/ 10.1145/2348283.2348488. Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., and Yih, W.-t. Dense passage retrieval for open-domain question answering. In Webber, B., Cohn, T., He, Y., and Liu, Y. (eds.), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 6769–6781, Online, November 2020. Association for Computational Linguistics. doi: 10.18653/v1/2020.emnlp-main.550. URLhttps:// aclanthology.org/2020.emnlp-main.550. Kasai, J., Sakaguchi, K., takahashi, y., Le Bras, R., Asai, A., Yu, X., Radev, D., Smith, N. A., Choi, Y., and Inui, K.Realtime qa: What's the answer right now?In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Advances in Neural Information Processing Systems, volume 36, p. 49025–49043. Curran Associates, Inc., 2023.URLhttps://proceedings.neurips. c/paper_files/paper/2023/file/ 11 Temporal Preference Optimization for Unsupervised Retrieval 9941624ef7f867a502732b5154d30cb7- Paper-Datasets_and_Benchmarks.pdf. Kwon, M., Bang, J., Hwang, S., Jang, J., and Lee, W. A dynamic-selection-based, retrieval-augmented genera- tion framework: Enhancing multi-document question- answering for commercial applications.Electron- ics, 14(4), 2025.ISSN 2079-9292.doi: 10.3390/ electronics14040659.URLhttps://w.mdpi. com/2079-9292/14/4/659. Lazaridou, A., Kuncoro, A., Gribovskaya, E., Agrawal, D., Liska, A., Terzi, T., Gimenez, M., de Mas- son d'Autume, C., Kocisky, T., Ruder, S., Yogatama, D., Cao, K., Young, S., and Blunsom, P. Mind the gap:Assessing temporal generalization in neural language models. In Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., and Vaughan, J. W. (eds.), Advances in Neural Information Processing Systems, volume 34, p. 29348–29363. Curran Associates, Inc., 2021.URLhttps://proceedings.neurips. c/paper_files/paper/2021/file/ f5bf0ba0a17ef18f9607774722f5698c- Paper.pdf. Lee, K., Chang, M.-W., and Toutanova, K. Latent retrieval for weakly supervised open domain question answering. In Korhonen, A., Traum, D., and M ` arquez, L. (eds.), Proceedings of the 57th Annual Meeting of the Asso- ciation for Computational Linguistics, p. 6086–6096, Florence, Italy, July 2019. Association for Computa- tional Linguistics. doi: 10.18653/v1/P19-1612. URL https://aclanthology.org/P19-1612/. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., K ̈ uttler, H., Lewis, M., Yih, W.-t., Rockt ̈ aschel, T., Riedel, S., and Kiela, D. Retrieval-augmented genera- tion for knowledge-intensive nlp tasks. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., and Lin, H. (eds.), Advances in Neural Information Processing Systems, volume 33, p. 9459–9474. Curran Associates, Inc., 2020.URLhttps://proceedings.neurips. c/paper_files/paper/2020/file/ 6b493230205f780e1bc26945df7481e5- Paper.pdf. Li, X., Jin, J., Zhou, Y., Zhang, Y., Zhang, P., Zhu, Y., and Dou, Z. From matching to generation: A survey on generative information retrieval. ACM Trans. Inf. Syst., March 2025. ISSN 1046-8188. doi: 10.1145/3722552. URL https://doi.org/10.1145/3722552. Li, Y., Liu, K., Satapathy, R., Wang, S., and Cambria, E. Recent developments in recommender systems: A sur- vey [review article]. IEEE Computational Intelligence Magazine, 19(2):78–95, 2024. doi: 10.1109/MCI.2024. 3363984. Li, Z., Zhang, X., Zhang, Y., Long, D., Xie, P., and Zhang, M. Towards general text embeddings with multi-stage contrastive learning, 2023. URLhttps://arxiv. org/abs/2308.03281. Litty K Mathews, S. D. K.A survey on tempo- ral information retrieval systems. International Jour- nal of Computer Applications, 58(4):24–28, Novem- ber 2012.ISSN 0975-8887.doi:10.5120/ 9271-3461.URLhttps://ijcaonline.org/ archives/volume58/number4/9271-3461/. Luu, K., Khashabi, D., Gururangan, S., Mandyam, K., and Smith, N. A. Time waits for no one! analysis and challenges of temporal misalignment. In Carpuat, M., de Marneffe, M.-C., and Meza Ruiz, I. V. (eds.), Pro- ceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguis- tics: Human Language Technologies, p. 5944–5958, Seattle, United States, July 2022. Association for Com- putational Linguistics. doi: 10.18653/v1/2022.naacl- main.435.URLhttps://aclanthology.org/ 2022.naacl-main.435. Ma, X., Gong, Y., He, P., Zhao, H., and Duan, N. Query rewriting in retrieval-augmented large language mod- els. In Bouamor, H., Pino, J., and Bali, K. (eds.), Pro- ceedings of the 2023 Conference on Empirical Meth- ods in Natural Language Processing, p. 5303–5315, Singapore, December 2023. Association for Computa- tional Linguistics. doi: 10.18653/v1/2023.emnlp-main. 322. URLhttps://aclanthology.org/2023. emnlp-main.322/. Neelakantan, A., Xu, T., Puri, R., Radford, A., Han, J. M., Tworek, J., Yuan, Q., Tezak, N., Kim, J. W., Hallacy, C., Heidecke, J., Shyam, P., Power, B., Nekoul, T. E., Sastry, G., Krueger, G., Schnurr, D., Such, F. P., Hsu, K., Thompson, M., Khan, T., Sherbakov, T., Jang, J., Welinder, P., and Weng, L. Text and code embeddings by contrastive pre-training, 2022. URLhttps://arxiv. org/abs/2201.10005. Nussbaum, Z. and Duderstadt, B. Training sparse mixture of experts text embedding models, 2025. URLhttps: //arxiv.org/abs/2502.07972. Nylund, K., Gururangan, S., and Smith, N. Time is en- coded in the weights of finetuned language models. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceed- ings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 2571–2587, Bangkok, Thailand, August 2024. Associa- tion for Computational Linguistics. doi: 10.18653/v1/ 2024.acl-long.141. URLhttps://aclanthology. org/2024.acl-long.141/. 12 Temporal Preference Optimization for Unsupervised Retrieval OpenAI, :, Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., Madry, A., Baker-Whitcomb, A., Beutel, A., Borzunov, A., Carney, A., Chow, A., Kirillov, A., Nichol, A., Paino, A., Renzin, A., Passos, A. T., Kir- illov, A., Christakis, A., Conneau, A., Kamali, A., Jabri, A., Moyer, A., Tam, A., Crookes, A., Tootoochian, A., Tootoonchian, A., Kumar, A., Vallone, A., Karpathy, A., Braunstein, A., Cann, A., Codispoti, A., Galu, A., Kon- drich, A., Tulloch, A., Mishchenko, A., Baek, A., Jiang, A., Pelisse, A., Woodford, A., Gosalia, A., Dhar, A., Pan- tuliano, A., Nayak, A., Oliver, A., Zoph, B., Ghorbani, B., Leimberger, B., Rossen, B., Sokolowsky, B., Wang, B., Zweig, B., Hoover, B., Samic, B., McGrew, B., Spero, B., Giertler, B., Cheng, B., Lightcap, B., Walkin, B., Quinn, B., Guarraci, B., Hsu, B., Kellogg, B., Eastman, B., Lu- garesi, C., Wainwright, C., Bassin, C., Hudson, C., Chu, C., Nelson, C., Li, C., Shern, C. J., Conger, C., Barette, C., Voss, C., Ding, C., Lu, C., Zhang, C., Beaumont, C., Hallacy, C., Koch, C., Gibson, C., Kim, C., Choi, C., McLeavey, C., Hesse, C., Fischer, C., Winter, C., Czar- necki, C., Jarvis, C., Wei, C., Koumouzelis, C., Sherburn, D., Kappler, D., Levin, D., Levy, D., Carr, D., Farhi, D., Mely, D., Robinson, D., Sasaki, D., Jin, D., Valladares, D., Tsipras, D., Li, D., Nguyen, D. P., Findlay, D., Oiwoh, E., Wong, E., Asdar, E., Proehl, E., Yang, E., Antonow, E., Kramer, E., Peterson, E., Sigler, E., Wallace, E., Brevdo, E., Mays, E., Khorasani, F., Such, F. P., Raso, F., Zhang, F., von Lohmann, F., Sulit, F., Goh, G., Oden, G., Salmon, G., Starace, G., Brockman, G., Salman, H., Bao, H., Hu, H., Wong, H., Wang, H., Schmidt, H., Whitney, H., Jun, H., Kirchner, H., de Oliveira Pinto, H. P., Ren, H., Chang, H., Chung, H. W., Kivlichan, I., O’Connell, I., O’Connell, I., Osband, I., Silber, I., Sohl, I., Okuyucu, I., Lan, I., Kostrikov, I., Sutskever, I., Kanitscheider, I., Gulrajani, I., Coxon, J., Menick, J., Pachocki, J., Aung, J., Betker, J., Crooks, J., Lennon, J., Kiros, J., Leike, J., Park, J., Kwon, J., Phang, J., Teplitz, J., Wei, J., Wolfe, J., Chen, J., Harris, J., Varavva, J., Lee, J. G., Shieh, J., Lin, J., Yu, J., Weng, J., Tang, J., Yu, J., Jang, J., Candela, J. Q., Beut- ler, J., Landers, J., Parish, J., Heidecke, J., Schulman, J., Lachman, J., McKay, J., Uesato, J., Ward, J., Kim, J. W., Huizinga, J., Sitkin, J., Kraaijeveld, J., Gross, J., Ka- plan, J., Snyder, J., Achiam, J., Jiao, J., Lee, J., Zhuang, J., Harriman, J., Fricke, K., Hayashi, K., Singhal, K., Shi, K., Karthik, K., Wood, K., Rimbach, K., Hsu, K., Nguyen, K., Gu-Lemberg, K., Button, K., Liu, K., Howe, K., Muthukumar, K., Luther, K., Ahmad, L., Kai, L., Itow, L., Workman, L., Pathak, L., Chen, L., Jing, L., Guy, L., Fedus, L., Zhou, L., Mamitsuka, L., Weng, L., McCal- lum, L., Held, L., Ouyang, L., Feuvrier, L., Zhang, L., Kondraciuk, L., Kaiser, L., Hewitt, L., Metz, L., Doshi, L., Aflak, M., Simens, M., Boyd, M., Thompson, M., Dukhan, M., Chen, M., Gray, M., Hudnall, M., Zhang, M., Aljubeh, M., Litwin, M., Zeng, M., Johnson, M., Shetty, M., Gupta, M., Shah, M., Yatbaz, M., Yang, M. J., Zhong, M., Glaese, M., Chen, M., Janner, M., Lampe, M., Petrov, M., Wu, M., Wang, M., Fradin, M., Pokrass, M., Castro, M., de Castro, M. O. T., Pavlov, M., Brundage, M., Wang, M., Khan, M., Murati, M., Bavarian, M., Lin, M., Yesil- dal, M., Soto, N., Gimelshein, N., Cone, N., Staudacher, N., Summers, N., LaFontaine, N., Chowdhury, N., Ryder, N., Stathas, N., Turley, N., Tezak, N., Felix, N., Kudige, N., Keskar, N., Deutsch, N., Bundick, N., Puckett, N., Nachum, O., Okelola, O., Boiko, O., Murk, O., Jaffe, O., Watkins, O., Godement, O., Campbell-Moore, O., Chao, P., McMillan, P., Belov, P., Su, P., Bak, P., Bakkum, P., Deng, P., Dolan, P., Hoeschele, P., Welinder, P., Tillet, P., Pronin, P., Tillet, P., Dhariwal, P., Yuan, Q., Dias, R., Lim, R., Arora, R., Troll, R., Lin, R., Lopes, R. G., Puri, R., Miyara, R., Leike, R., Gaubert, R., Zamani, R., Wang, R., Donnelly, R., Honsby, R., Smith, R., Sahai, R., Ramchandani, R., Huet, R., Carmichael, R., Zellers, R., Chen, R., Chen, R., Nigmatullin, R., Cheu, R., Jain, S., Altman, S., Schoenholz, S., Toizer, S., Miserendino, S., Agarwal, S., Culver, S., Ethersmith, S., Gray, S., Grove, S., Metzger, S., Hermani, S., Jain, S., Zhao, S., Wu, S., Jomoto, S., Wu, S., Shuaiqi, Xia, Phene, S., Papay, S., Narayanan, S., Coffey, S., Lee, S., Hall, S., Balaji, S., Broda, T., Stramer, T., Xu, T., Gogineni, T., Christian- son, T., Sanders, T., Patwardhan, T., Cunninghman, T., Degry, T., Dimson, T., Raoux, T., Shadwell, T., Zheng, T., Underwood, T., Markov, T., Sherbakov, T., Rubin, T., Stasi, T., Kaftan, T., Heywood, T., Peterson, T., Walters, T., Eloundou, T., Qi, V., Moeller, V., Monaco, V., Kuo, V., Fomenko, V., Chang, W., Zheng, W., Zhou, W., Man- assra, W., Sheu, W., Zaremba, W., Patil, Y., Qian, Y., Kim, Y., Cheng, Y., Zhang, Y., He, Y., Zhang, Y., Jin, Y., Dai, Y., and Malkov, Y. Gpt-4o system card, 2024. URL https://arxiv.org/abs/2410.21276. OpenAI, :, Agarwal, S., Ahmad, L., Ai, J., Altman, S., Ap- plebaum, A., Arbus, E., Arora, R. K., Bai, Y., Baker, B., Bao, H., Barak, B., Bennett, A., Bertao, T., Brett, N., Brevdo, E., Brockman, G., Bubeck, S., Chang, C., Chen, K., Chen, M., Cheung, E., Clark, A., Cook, D., Dukhan, M., Dvorak, C., Fives, K., Fomenko, V., Garipov, T., Georgiev, K., Glaese, M., Gogineni, T., Goucher, A., Gross, L., Guzman, K. G., Hallman, J., Hehir, J., Hei- decke, J., Helyar, A., Hu, H., Huet, R., Huh, J., Jain, S., Johnson, Z., Koch, C., Kofman, I., Kundel, D., Kwon, J., Kyrylov, V., Le, E. Y., Leclerc, G., Lennon, J. P., Lessans, S., Lezcano-Casado, M., Li, Y., Li, Z., Lin, J., Liss, J., Lily, Liu, Liu, J., Lu, K., Lu, C., Martinovic, Z., McCallum, L., McGrath, J., McKinney, S., McLaughlin, A., Mei, S., Mostovoy, S., Mu, T., Myles, G., Neitz, A., Nichol, A., Pachocki, J., Paino, A., Palmie, D., Pantu- liano, A., Parascandolo, G., Park, J., Pathak, L., Paz, C., 13 Temporal Preference Optimization for Unsupervised Retrieval Peran, L., Pimenov, D., Pokrass, M., Proehl, E., Qiu, H., Raila, G., Raso, F., Ren, H., Richardson, K., Robinson, D., Rotsted, B., Salman, H., Sanjeev, S., Schwarzer, M., Sculley, D., Sikchi, H., Simon, K., Singhal, K., Song, Y., Stuckey, D., Sun, Z., Tillet, P., Toizer, S., Tsimpourlas, F., Vyas, N., Wallace, E., Wang, X., Wang, M., Watkins, O., Weil, K., Wendling, A., Whinnery, K., Whitney, C., Wong, H., Yang, L., Yang, Y., Yasunaga, M., Ying, K., Zaremba, W., Zhan, W., Zhang, C., Zhang, B., Zhang, E., and Zhao, S. gpt-oss-120b & gpt-oss-20b model card, 2025. URL https://arxiv.org/abs/2508.10925. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. Training language models to follow instruc- tions with human feedback. In Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., and Oh, A. (eds.), Advances in Neural Information Processing Systems, volume 35, p. 27730–27744. Curran Associates, Inc., 2022.URLhttps://proceedings.neurips. c/paper_files/paper/2022/file/ b1efde53be364a73914f58805a001731- Paper-Conference.pdf. Qian, X., Zhang, Y., Zhao, Y., Zhou, B., Sui, X., Zhang, L., and Song, K.TimeR 4 : Time-aware retrieval- augmented large language models for temporal knowl- edge graph question answering.In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Nat- ural Language Processing, p. 6942–6952, Miami, Florida, USA, November 2024. Association for Com- putational Linguistics. doi: 10.18653/v1/2024.emnlp- main.394.URLhttps://aclanthology.org/ 2024.emnlp-main.394/. Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C.Direct preference opti- mization: Your language model is secretly a reward model.In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Ad- vances in Neural Information Processing Systems, volume 36, p. 53728–53741. Curran Associates, Inc., 2023.URLhttps://proceedings.neurips. c/paper_files/paper/2023/file/ a85b405ed65c6477a4fe8302b5e06ce7- Paper-Conference.pdf. Rame, A., Ahuja, K., Zhang, J., Cord, M., Bottou, L., and Lopez-Paz, D. Model ratatouille: Recycling di- verse models for out-of-distribution generalization. In Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., and Scarlett, J. (eds.), Proceedings of the 40th International Conference on Machine Learn- ing, volume 202 of Proceedings of Machine Learn- ing Research, p. 28656–28679. PMLR, 23–29 Jul 2023. URLhttps://proceedings.mlr.press/ v202/rame23a.html. Robertson, S. and Zaragoza, H. The probabilistic relevance framework: Bm25 and beyond. Found. Trends Inf. Retr., 3(4):333–389, apr 2009. ISSN 1554-0669. doi: 10.1561/ 1500000019. URLhttps://doi.org/10.1561/ 1500000019. Rosin, G. D., Guy, I., and Radinsky, K. Time mask- ing for temporal language models. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining, WSDM ’22, p. 833–841, New York, NY, USA, 2022. Association for Comput- ing Machinery. ISBN 9781450391320. doi: 10.1145/ 3488560.3498529.URLhttps://doi.org/10. 1145/3488560.3498529. R ̈ ottger, P. and Pierrehumbert, J. Temporal adaptation of BERT and performance on downstream document clas- sification: Insights from social media. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Find- ings of the Association for Computational Linguistics: EMNLP 2021, p. 2400–2412, Punta Cana, Domini- can Republic, November 2021. Association for Com- putational Linguistics. doi: 10.18653/v1/2021.findings- emnlp.206. URLhttps://aclanthology.org/ 2021.findings-emnlp.206/. Schilder, F. and McCulloh, A. Temporal information extrac- tion from legal documents. In Katz, G., Pustejovsky, J., and Schilder, F. (eds.), Annotating, Extracting and Rea- soning about Time and Events, volume 5151 of Dagstuhl Seminar Proceedings (DagSemProc), p. 1–9, Dagstuhl, Germany, 2005. Schloss Dagstuhl – Leibniz-Zentrum f ̈ ur Informatik. doi: 10.4230/DagSemProc.05151.9. URL https://drops.dagstuhl.de/entities/ document/10.4230/DagSemProc.05151.9. Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O.Proximal policy optimization algo- rithms, 2017.URLhttps://arxiv.org/abs/ 1707.06347. Scrapinghub. dateparser: python parser for human readable dates.https://github.com/scrapinghub/ dateparser. Accessed: 2025-09-22. Shao, J., Wen, X., Zhao, B., and Xue, X. Temporal con- text aggregation for video retrieval with contrastive learn- ing. In 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), p. 3267–3277, 2021. doi: 10.1109/WACV48630.2021.00331. 14 Temporal Preference Optimization for Unsupervised Retrieval Singh, J. Combining machine learning and rag models for enhanced data retrieval: Applications in search engines, enterprise data systems, and recommendations. Journal of Computational Intelligence and Robotics, 3(1):163–204, Mar. 2023. URLhttps://thesciencebrigade. com/jcir/article/view/421. Su, Z., Li, J., Zhang, Z., Zhou, Z., and Zhang, M. Effi- cient continue training of temporal language model with structural information. In Bouamor, H., Pino, J., and Bali, K. (eds.), Findings of the Association for Com- putational Linguistics: EMNLP 2023, p. 6315–6329, Singapore, December 2023. Association for Compu- tational Linguistics. doi: 10.18653/v1/2023.findings- emnlp.418. URLhttps://aclanthology.org/ 2023.findings-emnlp.418/. Thakur,N.Loadingyourcustomdataset. https://github.com/beir-cellar/beir/ wiki/Load-your-custom-dataset,30Jun 2022. Accessed: 2026-01-27. Thakur, N., Reimers, N., R ̈ uckl ́ e, A., Srivastava, A., and Gurevych, I. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Thirty-fifth Conference on Neural Information Process- ing Systems Datasets and Benchmarks Track (Round 2), 2021. URLhttps://openreview.net/forum? id=wCu6T5xFjeJ. Wang, C., Jiang, Y., Yang, C., Liu, H., and Chen, Y. Beyond reverse KL: generalizing direct preference optimization with diverse divergence constraints. In The Twelfth Inter- national Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024a. URLhttps://openreview.net/forum? id=2cRzmWXK9N. Wang, H., Dong, A., Li, L., Chang, Y., and Gabrilovich, E. Joint relevance and freshness learning from click- throughs for news search. In Proceedings of the 21st International Conference on World Wide Web, W ’12, p. 579–588, New York, NY, USA, 2012. Associa- tion for Computing Machinery. ISBN 9781450312295. doi: 10.1145/2187836.2187915. URLhttps://doi. org/10.1145/2187836.2187915. Wang, J., Jatowt, A., Yoshikawa, M., and Cai, Y. Bitime- bert: Extending pre-trained language representations with bi-temporal information. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR ’23, p. 812–821, New York, NY, USA, 2023. Associa- tion for Computing Machinery. ISBN 9781450394086. doi: 10.1145/3539618.3591686. URLhttps://doi. org/10.1145/3539618.3591686. Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., and Wei, F. Text embeddings by weakly- supervised contrastive pre-training, 2024b. URLhttps: //arxiv.org/abs/2212.03533. Wu, B., Zhang, Z., Wang, J., and Zhao, H. Sentence- aware contrastive learning for open-domain passage re- trieval.In Muresan, S., Nakov, P., and Villavicen- cio, A. (eds.), Proceedings of the 60th Annual Meet- ing of the Association for Computational Linguistics (Volume 1: Long Papers), p. 1062–1074, Dublin, Ire- land, May 2022. Association for Computational Linguis- tics. doi: 10.18653/v1/2022.acl-long.76. URLhttps: //aclanthology.org/2022.acl-long.76/. Wu, F., Liu, L., He, W., Liu, Z., Zhang, Z., Wang, H., and Wang, M. Time-sensitve retrieval-augmented gen- eration for question answering. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, CIKM ’24, p. 2544–2553, New York, NY, USA, 2024. Association for Comput- ing Machinery. ISBN 9798400704369. doi: 10.1145/ 3627673.3679800.URLhttps://doi.org/10. 1145/3627673.3679800. Xiong, W., Dong, H., Ye, C., Wang, Z., Zhong, H., Ji, H., Jiang, N., and Zhang, T. Iterative preference learn- ing from human feedback: Bridging theory and prac- tice for RLHF under KL-constraint.In Salakhutdi- nov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., and Berkenkamp, F. (eds.), Proceedings of the 41st International Conference on Machine Learn- ing, volume 235 of Proceedings of Machine Learn- ing Research, p. 54715–54754. PMLR, 21–27 Jul 2024. URLhttps://proceedings.mlr.press/ v235/xiong24a.html. Yang, A., Li, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Gao, C., Huang, C., Lv, C., Zheng, C., Liu, D., Zhou, F., Huang, F., Hu, F., Ge, H., Wei, H., Lin, H., Tang, J., Yang, J., Tu, J., Zhang, J., Yang, J., Yang, J., Zhou, J., Zhou, J., Lin, J., Dang, K., Bao, K., Yang, K., Yu, L., Deng, L., Li, M., Xue, M., Li, M., Zhang, P., Wang, P., Zhu, Q., Men, R., Gao, R., Liu, S., Luo, S., Li, T., Tang, T., Yin, W., Ren, X., Wang, X., Zhang, X., Ren, X., Fan, Y., Su, Y., Zhang, Y., Zhang, Y., Wan, Y., Liu, Y., Wang, Z., Cui, Z., Zhang, Z., Zhou, Z., and Qiu, Z. Qwen3 technical report, 2025. URLhttps: //arxiv.org/abs/2505.09388. Zhang, L., Zhang, J., Ke, X., Li, H., Huang, X., Shao, Z., Cao, S., and Lv, X. A survey on complex factual question answering. AI Open, 4:1–12, 2023a. ISSN 2666-6510. doi: https://doi.org/10.1016/j.aiopen.2022.12. 003. URLhttps://w.sciencedirect.com/ science/article/pii/S2666651022000249. 15 Temporal Preference Optimization for Unsupervised Retrieval Zhang, M. and Choi, E. SituatedQA: Incorporating extra- linguistic contexts into QA. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, p. 7371–7387, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021. emnlp-main.586. URLhttps://aclanthology. org/2021.emnlp-main.586. Zhang, Q., Chen, S., Xu, D., Cao, Q., Chen, X., Cohn, T., and Fang, M. A survey for efficient open domain question answering. In Rogers, A., Boyd-Graber, J., and Okazaki, N. (eds.), Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 14447–14465, Toronto, Canada, July 2023b. Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.808. URLhttps:// aclanthology.org/2023.acl-long.808/. Zhang, S., Yao, L., Sun, A., and Tay, Y. Deep learning based recommender system: A survey and new perspectives. ACM Comput. Surv., 52(1), February 2019. ISSN 0360- 0300. doi: 10.1145/3285029. URLhttps://doi. org/10.1145/3285029. Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. Qwen3 embedding: Advancing text embedding and reranking through foundation models, 2025. URL https://arxiv.org/abs/2506.05176. Zhao, P., Zhang, H., Yu, Q., Wang, Z., Geng, Y., Fu, F., Yang, L., Zhang, W., Jiang, J., and Cui, B. Retrieval- augmented generation for ai-generated content: A survey, 2024a.URLhttps://arxiv.org/abs/2402. 19473. Zhao, W. X., Liu, J., Ren, R., and Wen, J.-R. Dense text retrieval based on pretrained language models: A survey. ACM Trans. Inf. Syst., 42(4), February 2024b. ISSN 1046-8188. doi: 10.1145/3637870. URLhttps:// doi.org/10.1145/3637870. Zhu, Y., Yuan, H., Wang, S., Liu, J., Liu, W., Deng, C., Chen, H., Liu, Z., Dou, Z., and Wen, J.-R. Large lan- guage models for information retrieval: A survey. ACM Trans. Inf. Syst., September 2025. ISSN 1046-8188. doi: 10.1145/3748304. URLhttps://doi.org/10. 1145/3748304. Just Accepted. 16 Temporal Preference Optimization for Unsupervised Retrieval A. Reproducibility Statements A.1. System Architecture and Inference Process What was at stake when the Clemson Tigers faced off in the Sugar Bowl now? (푄 ! ) Query 푡Aligned Encoder 휋 ! Document Index Clemson Tigers football // Once again the teams did battle in the 2018 Sugar Bowl in New Orleans, Louisiana with a trip to the 2018 College Football Playoff National Championship game on the line ... Document Collected at Time 푡,푡 " 휋 ! Pre-compute Store 휋 ! (푄 " ) ! Clemson Tigers football // Once again the teams did battle in the 2018 Sugar Bowl in New Orleans, Louisiana with a trip to the 2018 College Football Playoff National Championship game on the line ... Top 1 Retrieved Document 푡Aligned Encoder Figure 6. An illustration ofTPOURinference. Like standard retrieval, we use the trained encoderπ θ to pre-compute representations for all documents at mixed-timestampstandt ′ , which are then stored in the document index. At inference, a queryQ i is encoded asπ θ (Q i ), and retrieves the document from the index with the highest similarity to the query. The retrieved document is both semantically relevant and temporally aligned with the query. 휋 ! " ! Clemson Tigers football // Once again the teams did battle in the 2018 Sugar Bowl in New Orleans, Louisiana with a trip to the 2018 College Football Playoff National Championship game on the line ... Documents 퐷with random timestamp 휋 ! " " ! ! ! 휋 ! " # ! ! ! Mixture-of-TPOUR Encoder 휋 ! " " (퐷) 휋 ! " # (퐷) 휋 ! " ! (퐷) Linear Predicted Timestamp 푡 휋 ! Clemson Tigers football // Once again the teams did battle in the 2018 Sugar Bowl in New Orleans, Louisiana with a trip to the 2018 College Football Playoff National Championship game on the line ... Documents 퐷with random timestamp Single Baseline Encoder 휋 ! (퐷) Linear Predicted Timestamp 푡 Linear Linear Baseline Timestamp Predictor Mixture-of-TPOUR Timestamp Predictor Figure 7. An illustration of the Baseline and the mixture-of-TPOURTimestamp Predictor under a setup where the linear classifier has the same number of parameters. Given a document, the baseline model (upper) uses a single encoder to generate a representation, which is then passed to a linear classifier to predict the timestamp. In contrast, the mixture-of-TPOUR(lower) uses a set of frozen retrievers π t 1 θ ,...,π t n θ , each specialized for a different time period, to produce temporally-aware embeddings. These are concatenated and fed into a linear classification layer to predict the most likely timestamp. For a fair comparison, we matched the total number of trainable parameters by stacking multiple linear layers in the baseline predictor equal to the number of TPOUR encoders. 17 Temporal Preference Optimization for Unsupervised Retrieval A.2. Training Dataset Construction Procedure We extract document texts from each dump using Wikiextractor (Attardi, 2015). As summarized in Tab. 7, we first filter out short documents (>50 words), which are mostly hyperlink pages with no content. We then identify overlapping documents across timestamps (Intersection) and retain only those with content changes (Filtered Intersection). Finally, we include timestamp-specific unique documents (Unique) to build the final dataset (Final), ensuring that each timestamp-specific collection contains meaningful temporal differences. Lastly, we remove all documents that appear in the test sets to prevent any data leakage during evaluation. The resulting dataset comprises temporally distinct document collections from each Wikipedia dump, with minimal explicit mentions of the target year. As shown in Tab. 8, fewer than 2.5% of documents contain the target year explicitly within their content. Table 7. Statistics of Wikipedia dumps used for monthly and yearly training & evaluation. (Original) Starting from the full set of documents (>50 words), we filter out those with fewer than 50 words. (Intersection) We then identify overlapping documents across timestamps (Filtered Intersection), further filter for documents that changed between each dump set, and (Unique) add unique documents that are created only at the specific dump set (Final) to obtain the final dataset. # Docs - Monthly Dump SetOriginal >50 wordsIntersectionFiltered IntersectionUniqueFinal 2023-01-0116,228,2284,876,6824,842,453736,52734,229770,756 2023-07-0116,505,5314,963,0324,842,453736,527120,579857,106 2023-12-2016,619,6445,011,0404,842,453736,527168,587905,114 # Docs - Yearly 2018-12-2013,717,0224,021,0803,888,1232,674,468132,9572,807,425 2021-12-2015,567,2194,706,7053,888,1232,674,468818,5823,493,050 2023-12-2016,619,6445,011,0403,888,1232,674,4681,122,9173,797,385 Table 8. Percentage of documents in each Wikipedia dump that contain an explicit mention of the corresponding collection year. As shown, the majority of documents (>97%) do not include lexical references to the target year, reinforcing thatTPOURlearns temporal preferences from semantic drift across documents collected at different times, rather than from explicit timestamp information. Dump SetTarget Year in Document (%)Target Year not in Document (%) 2018-12-202.5097.50 2021-12-201.8998.01 2023-12-201.1098.90 A.3. Training & Evaluation Environment A.3.1. TRAINING CONFIGURATION We fully fine-tuneTPOURusing Contriever (Izacard et al., 2022) as the base model (TPOURContriever) on a single NVIDIA A100 (80GB) GPU, an AMD EPYC 7763 64-core CPU, and 200GB of memory. The hyperparameters used forTPOUR training are listed in Tab. 9. We use a learning rate of1e−6, 4,000 warmup steps, and a MoCo queue of length 131,060. The temporal preference objective is combined with the contrastive loss usingλ = 0.925. We apply token deletion augmentation with a probability of 10%, and chunk input texts to a maximum length of 256 tokens. All normalization options are disabled to preserve the original text form. We use three different random seeds to train TPOUR Contriever. 18 Temporal Preference Optimization for Unsupervised Retrieval Table 9. Hyperparameters used for training the TPOUR Contriever. HyperparameterValue Contrastive / TRPO Loss Weight (λ)0.925 Temperature (T )0.05 OptimizerAdamW AdamW β 1 , β 2 , ε0.9, 0.98, 1e−6 Learning Rate1e−6 Scheduler / Warmup StepsLinear / 4000 Batch Size10 MoCo Queue Size131,060 Momentum (m)0.9999 Projection Size768 Dropout Rate0.1 Chunk Length256 Text AugmentationDeletion (prob = 0.1) Normalization (Query / Doc / Text)False / False / False Training Steps100,000 Data Augmentation (Random Cropping / Delete)False / True (10%) A.3.2. COMPUTATIONAL COST TPOUR Contriever is based on BERT-base-uncased (110M parameters, 440MB). For interpolation or mixture-of-TPOUR experiments, we train twoTPOURContrievers (2018 and 2021), each taking 4.5 GPU hours on a single A100. These are then interpolated to produce 10 time-specific models, resulting in a total storage of 4.4GB. For the mixture-of-TPOUR predictor, the training requires 16 GPU hours (10 hours for the baseline). The predictor uses 0.12M trainable parameters, of about 600KB in size, while all TPOUR retrievers remain frozen. B. Theoretical Basis ofTPOUR B.1. Temporal Retrieval Preference Optimization Direct Preference Optimization (DPO) (Rafailov et al., 2023) forms a preference pair given a prompt (x), preferred (y w ) and less preferred response (y l ) as(x,y w ,y l ). Like DPO, Temporal Retrieval Preference Optimization (TRPO) forms (Q,D t ,D t ′ )as a pairwise preference pair over a timestamped document corpus given a queryQ, time-aligned document (D t ), and unaligned document (D t ′ ). The goal of TRPO is to prefer temporally aligned document (D t ) over the misaligned (D t ′ ) given a query (Q) and minimizing the TRPO loss function (L TRPO ) in Eq. 4. L TRPO is based on the Bradley-Terry model. While DPO aligns model score with human-labeled preference, TRPO aligns score with temporal relevance with an implicit signal derived from corpus-level differences (preferringD t overD t ′ ). In this view, TRPO requires working under the following three conditions. 1.Temporal preference margin. There must be a certain temporal preference gap (i.e., margin) between aligned and misaligned documentE[S(Q,D t )− S(Q,D t ′ )] > δ whent ′ ̸= twhereδis a minimum gap required. If the actual document update with temporal change is too small relative to noise, TRPO learning could be unstable. To handle this issue, we comprise a temporally distinct document collection by filtering the dataset in Appendix A.2. 2.Similar semantic across corpora. Aligned and misaligned temporal corpora should cover a similar set of topics, so semantic similarity may remain high and the only difference is the timestamp and the document content at that timestamp. 3.Model capacity. Encoder should have sufficient capacity to represent latent temporal signal as well as semantic similarity. Under these conditions, TRPO encourages the model to rank temporally aligned documents higher. The resulting scoring functionS θ is expected to approximate one that reflects temporal alignment between query and document. This mirrors the theoretical guarantees for DPO by replacing “generation quality” with “temporal relevance” as the underlying reward (Wang et al., 2024a; Xiong et al., 2024). To sum up, TRPO is a preference alignment variant, where preferences are defined by temporal grounding between a versioned corpus. This generalizes preference learning to the temporal dimension. 19 Temporal Preference Optimization for Unsupervised Retrieval B.2. Time Vector Interpolation The assumption that time vectors (i.e., model parameters trained on temporally adjacent corpora) are close in weight space is supported both empirically in Fig. 3, where retrieved documents change smoothly across interpolated models, showing continuity in the learned representation space. Theoretically, time vector interpolation is supported in two parts. Distributional similarity leads to weight-space proximity. LetP t andP t ′ be training distributions at timetandt ′ . If P t ≈ P t ′ (e.g., under lowKL(P t |P t ′ ) ), then under gradient descent, the learned parametersθ t ≈ θ t ′ will be nearby in weight space. The idea is formalized by (Goodfellow & Vinyals, 2015) and aligns with our setup, where temporally adjacent corpora (e.g., 2018 vs. 2019) are close in weight space. For adjacent periods, only temporal preferences differ, while the training data come from similar Wikipedia distributions. Interpolation preserves generalization. Prior work has shown that models trained on related tasks or distributions often lie in connected regions of the loss landscape (Izmailov et al., 2018; Rame et al., 2023). In our setting,θ t andθ t ′ are trained on temporally adjacent corpora (e.g., 2018 vs. 2019), which tend to share topical and linguistic structures, yielding a time vectorτ. As shown by (Izmailov et al., 2018), linearly interpolating such weight vectors,ατ t + (1− α)τ t ′ , often produces low-loss solutions if the endpoints lie in a shared basin. This smoothness in weight space supports generalization and has been used in practice via stochastic weight averaging (SWA). Also, (Rame et al., 2023) shows that interpolating across models trained on diverse but related domains can produce generalizable models that outperform the individual components. Analogously, we treat time as an axis of distributional change, and our interpolation procedure leverages this continuity to produce retrievers that generalize to intermediate periods. C. Related Work in Information Retrieval C.1. Early Work on Temporal Alignment in Information Retrieval (Berberich et al., 2010) explored the inherent uncertainty of temporal expressions and proposed representing them as tuples, integrating this representation into a probabilistic language modeling framework for information retrieval. (Jatowt et al., 2005) proposed a re-ranking method that utilizes archived web snapshots to prioritize documents based on content freshness and relevance. They also introduced the concept of document focus time, which refers to the temporal period indicated by the document content and is distinct from its creation time. Additionally, they proposed a method to automatically estimate this temporal reference using large news collections and external knowledge bases (Jatowt et al., 2013). (Kanhabua & Nørv ̊ ag, 2010) developed methods for determining the time of implicit temporal queries by leveraging temporal language models trained on timestamped corpora. They further proposed the first machine learning framework capable of automatically selecting the most effective temporal ranking strategy for a given query (Kanhabua et al., 2012). C.2. Baseline Models DPR (Dense Passage Retrieval) (Karpukhin et al., 2020) is a supervised dense retriever trained on query-passage pairs using a bi-encoder architecture. It optimizes retrieval by maximizing similarity between queries and relevant passages while minimizing similarity to negative samples. DPR is trained using hard negatives from BM25 to improve retrieval quality. Contriever (Izacard et al., 2022) is a self-supervised dense retriever trained with contrastive learning, removing the need for labeled query-document pairs. It constructs high-quality negative samples using a momentum encoder, enabling scalable pretraining on large unlabeled corpora. REALM (Retrieval-Augmented Language Model) (Guu et al., 2020) jointly trains a dense retriever and a language model in an end-to-end manner. During pretraining, the retriever is updated to select relevant documents that improve language model performance. This integration enables the model to dynamically leverage external knowledge, making it particularly effective for knowledge-intensive NLP tasks such as open-domain QA. SimCSE (Gao et al., 2021) is a sentence embedding model trained using contrastive learning in both supervised and unsupervised settings. The unsupervised variant leverages dropout as noise, while the supervised variant uses natural language inference (NLI) data. Though not originally intended for retrieval, SimCSE embeddings can be used for dense retrieval by comparing query and document representations in a shared semantic space. Temporal Language Modeling (Berberich et al., 2010) is a retrieval framework that integrates temporal expressions into language models by modeling their inherent uncertainty. The proposed uncertainty-aware model represents temporal 20 Temporal Preference Optimization for Unsupervised Retrieval expressions as interval distributions and measures temporal relevance via overlap between query and document intervals. Temporal Contrastive is a temporally-aware contrastive baseline that augments the standard contrastive retrieval objective with temporal supervision. We include this baseline to examine whether temporal alignment can be obtained by directly constructing time-based positive and negative pairs, without preference-based optimization. For each query, temporally aligned documents are treated as positives and temporally misaligned documents as negatives, yieldingL TempCE , which encourages higher similarity between the query and documents that better match the target time. The final objective is L total = λL CE + (1− λ)L TempCE , whereL CE models semantic relevance andL TempCE models temporal alignment. TimeR 4 (Qian et al., 2024) is a retrieval-augmented generation framework for temporal knowledge graph question answering. It includes a time-aware dense retriever trained with contrastive learning to capture both semantic and temporal constraints. In our experiments, we use its retriever component for evaluation. Nomic Embed v2 MoE (Nussbaum & Duderstadt, 2025) is a sparse Mixture-of-Experts (MoE) embedding model developed for efficient and scalable dense retrieval. It activates a small subset of expert networks per input, balancing high capacity with low inference cost. Trained using hard negative mining and consistency filtering, it achieves competitive retrieval performance compared to fully dense models. As a general-purpose model, it is open-sourced and designed to perform well across various domains and tasks without extensive fine-tuning. Qwen3-Embedding-8B (Zhang et al., 2025) is a large-scale embedding model with 8B parameters, built upon Qwen3 (Yang et al., 2025). It supports diverse embedding and reranking tasks across multiple domains and languages. We consider three retrieval setups. (1) Naive retrieval: As in conventional retrieval methods, we directly use the original query to retrieve relevant documents. (2) Query rewriting retrieval: This method uses a large language model to rewrite the query to make the temporal intent more explicit. Specifically, we use GPT-OSS-20B (OpenAI et al., 2025) as a frozen rewriter, following prior work (Ma et al., 2023). (3) time-aware instruction retrieval: Since Qwen3-Embedding-8B supports instruction-conditioned embeddings, we apply TAI to queries with explicit temporal information, where the target time can be directly encoded in the instruction prompt shown in Tab. 10. Table 10. Instruction-aware prompting template for temporal retrieval with Qwen-3-Embedding-8B.QUERY denotes a placeholder. Prompt Content You are a retrieval system that selects documents relevant to a query. Temporal requirement: - Preference: prioritize documents whose content reflects knowledge available at that time. - Avoid documents containing information published after the target time unless explicitly requested. Instructions: 1. Identify the temporal intent of the query. 2. Filter or downweight documents that violate the temporal constraint. 3. Rank documents by both semantic relevance and temporal alignment. 4. Prefer documents whose timestamps are closest to, but not exceeding, the target time. Query: QUERY 21 Temporal Preference Optimization for Unsupervised Retrieval D. Additional Experimental Results & Analysis D.1. Full Results on BEIR Benchmark Table 11. Retrieval performance (nDCG@10) on the BEIR benchmark, with dataset publication years shown below each dataset name. Each benchmark exhibits specific temporal preferences that mostly align with its creation date, suggesting thatTPOURcan improve general retrieval performance by adapting to the temporal characteristics of different datasets. Interpolation BEIR Benchmark Datasets MS MARCO TREC-COVID NFCorpus NQ HotpotQA FiQA ArguAna Touch ́ e-2020 Quora DBPedia SciDocs FEVER Climate-FEVER SciFact 2018 2021 2017 2020 2016 2019 2018 2018 2018 2020 2016 2017 2020 2018 2020 2020 1 0 45.56 21.44 25.52 23.90 40.88 19.22 43.65 9.39 81.98 30.11 13.51 45.63 19.75 60.68 0.9 0.1 45.30 22.57 25.98 24.23 41.52 19.75 43.69 9.83 82.01 30.42 13.74 46.37 20.49 60.95 0.8 0.2 45.14 22.60 26.36 24.52 42.00 20.31 43.57 10.44 82.15 30.67 13.86 46.72 21.08 61.12 0.7 0.3 44.98 23.11 26.91 24.82 42.41 20.51 43.56 10.80 82.15 30.62 13.88 46.85 21.62 61.34 0.6 0.4 44.33 23.42 27.18 24.82 42.60 20.54 43.37 10.92 82.22 30.52 13.82 46.69 22.11 61.57 0.5 0.5 43.84 24.16 27.17 24.80 42.61 20.56 43.52 10.91 82.23 30.35 13.77 46.19 22.59 61.73 0.4 0.6 43.40 24.98 27.03 24.51 42.40 20.25 43.53 10.95 82.21 29.74 13.61 45.33 22.62 61.73 0.3 0.7 42.78 25.07 26.90 24.21 41.92 20.16 43.48 10.79 82.14 29.31 13.39 44.07 22.68 61.26 0.2 0.8 41.26 25.74 26.64 23.82 41.30 19.78 43.38 10.85 82.00 28.74 13.08 42.50 22.68 60.56 0.1 0.9 40.90 26.35 26.27 23.08 40.38 19.15 42.99 10.29 81.84 28.10 12.81 40.55 22.44 59.58 0 1 39.95 27.08 25.78 22.58 39.27 18.29 42.69 9.99 81.67 27.16 12.43 38.49 22.13 59.35 2021 2023 1 0 39.95 27.08 25.78 22.58 39.27 18.29 42.69 9.99 81.67 27.16 12.43 38.49 22.13 59.35 0.9 0.1 40.78 26.95 26.40 22.96 40.06 18.71 42.80 10.06 81.85 27.84 12.81 39.88 22.15 59.59 0.8 0.2 41.84 27.03 26.72 23.19 40.59 19.13 42.96 10.28 81.94 28.36 13.09 40.86 22.17 59.99 0.7 0.3 41.89 28.13 26.96 23.40 40.90 19.38 42.89 10.74 81.99 28.54 13.20 41.72 22.14 60.53 0.6 0.4 42.25 28.72 26.78 23.48 41.05 19.51 42.92 10.92 81.98 28.58 13.38 42.17 21.89 60.63 0.5 0.5 42.43 28.93 26.75 23.48 41.12 19.40 43.09 10.99 81.99 28.71 13.25 42.23 21.62 60.55 0.4 0.6 42.26 28.12 26.43 23.25 40.97 19.01 43.12 11.41 81.98 28.69 13.18 41.98 21.16 60.49 0.3 0.7 42.43 28.51 25.96 22.89 40.58 18.77 43.11 10.92 81.92 28.23 13.08 41.44 20.38 60.20 0.2 0.8 42.63 28.91 25.12 22.56 40.11 18.58 43.26 10.61 81.71 27.67 12.82 40.79 19.65 59.27 0.1 0.9 42.21 28.80 24.38 22.20 39.50 18.06 43.13 10.09 81.44 27.38 12.44 39.98 18.79 58.73 0 1 41.31 29.11 23.54 21.68 38.60 17.35 43.18 10.38 81.14 27.04 12.18 38.90 18.00 57.81 22 Temporal Preference Optimization for Unsupervised Retrieval D.2. Full Results on Interpolation Table 12.TPOURyearly transition in performance with interpolation on SituatedQA. The color saturation indicates the relative performance, with darker green representing higher scores within each column. The table shows the impact of time vector interpolation on retrieval performance across different time periods, where the highest scores are achieved at their corresponding evaluation times. Gradual changes in performance are observed as the interpolation values shift. InterpolationnDCG@5nDCG@10Recall@5Recall@10 20182021 2018201920202021201820192020202120182019202020212018201920202021 1044.1030.8321.6614.0046.5732.8424.3017.6834.4626.0421.2812.8446.1834.6929.8920.38 0.90.143.3833.5726.1918.8646.5536.6429.2822.1534.6628.3126.1517.6447.0439.3836.0925.85 0.80.242.3933.9229.7524.1445.8336.9533.1027.7135.2830.4428.3523.0447.0542.1039.4833.31 0.70.3 40.7232.8431.1428.9244.7336.4535.0732.4234.2829.7230.7226.8246.9541.6843.2937.43 0.60.437.9330.9432.6131.2941.4935.1336.0735.1233.7229.2032.1427.4645.9242.1043.3839.55 0.50.535.7928.8832.5335.6339.7133.8335.6738.7433.0028.4631.6729.2546.0142.4142.8340.85 0.40.634.0427.5531.4138.5238.0432.2934.4840.6632.2127.6831.9230.4345.2140.5641.7441.39 0.30.7 32.2426.7729.6739.8036.8931.4733.1942.1629.8726.7930.0630.0944.5439.4541.0841.56 0.20.828.7924.5728.1840.5834.2928.8332.1241.9826.2025.0926.7130.3342.3237.0939.5239.78 0.10.925.6721.7426.0540.8430.6426.2829.3242.6023.7820.8724.7528.7737.5233.8335.2839.55 0122.6119.7624.6240.8327.2523.2227.4441.3422.3719.1022.8427.0533.1729.1532.9936.99 Table 13. Yearly transition inTPOURperformance with interpolation on SituatedQA, when time information is given implicitly in the query. InterpolationnDCG@5nDCG@10Recall@5Recall@10 201820212018201920202021201820192020202120182019202020212018201920202021 1044.6330.7822.1414.8346.3733.9925.5718.5736.1525.6621.2613.0846.7036.9031.2822.67 0.90.142.6934.0426.5718.6845.4837.1529.9022.8236.2229.6727.1617.4847.0840.5637.0427.51 0.80.240.4234.3127.9221.7944.8437.8931.2726.0533.9730.8928.5421.2348.0342.7637.9832.49 0.70.3 38.6434.6229.9425.2942.3238.7333.4828.9934.6131.6030.6824.6147.0443.9541.1835.29 0.60.437.6234.6529.8527.9342.1238.1533.9931.6534.4532.1231.1526.1647.8043.6642.2137.26 0.50.535.1433.8131.9830.0240.2037.1035.7134.1132.4632.9632.2227.5647.1743.4143.5339.06 0.40.633.2033.2132.9331.9838.0936.1836.2135.6231.6433.0133.4728.5845.8043.0843.4639.34 0.30.732.2432.0633.1434.3636.0335.5036.9037.1731.3431.8634.2129.3642.5443.3244.3539.54 0.20.828.5130.9032.5736.2634.2834.7236.6338.7926.6930.4733.1130.9642.1342.9845.0640.19 0.10.9 26.1529.6830.3138.4730.6833.6634.9440.4125.7528.3932.8230.8237.1941.1544.4140.38 0123.8928.2829.7538.4028.4931.9933.9141.3322.5327.0231.9029.7635.4739.1542.5842.24 Table 14.TPOURmonthly performance with interpolation on RealTimeQA, when time information is given explicitly (left) or implicitly (right) in the query. Results show temporal alignment in January (Jan), June (Jun), and December (Dec). nDCG@5nDCG@10nDCG@5nDCG@10 InterpolationTest Month (Explicit)Test Month (Implicit) JanDecJanJunDecJanJunDecJanJunDecJanJunDec 1032.0829.4129.3632.1828.6428.7832.8229.4527.7833.0729.6128.00 0.90.128.8028.6734.3630.1928.8933.7733.5330.0130.1332.4630.1528.80 0.80.225.2629.9840.2025.7329.6738.8533.5529.8831.5932.0730.0829.62 0.70.321.9630.0144.4321.5429.6441.5832.0330.1933.6430.5630.3432.39 0.60.4 17.5830.7947.2318.3330.5144.3429.3029.9137.8328.7630.9035.53 0.50.516.0032.3949.8716.0931.7246.0625.3230.6839.7225.5032.1738.05 0.40.6 14.4132.8849.9914.7831.4947.3620.1931.1642.6723.1232.3440.73 0.30.7 12.9531.7749.9413.6531.9447.6317.2032.0245.1220.8632.6942.90 0.20.810.8631.8050.9812.5330.9048.5415.2433.5748.1018.5532.7645.03 0.10.9 9.8230.9851.1311.2129.9947.4413.0734.0450.8015.7532.5546.54 018.4129.5749.989.3327.8446.7511.8033.4753.1213.4531.6848.59 23 Temporal Preference Optimization for Unsupervised Retrieval D.3. Timestamp Distribution of Retrieved Documents 0.370.280.190.15 0.330.310.220.15 0.310.270.270.15 0.320.280.240.17 2021 2020 2019 2018 0.270.260.230.25 0.170.260.280.29 0.140.190.310.36 0.140.190.280.39 0.350.280.200.17 0.330.300.210.16 0.320.260.250.17 0.330.270.220.18 0.260.260.230.25 0.230.270.250.25 0.230.230.280.26 0.240.240.240.28 0.340.280.200.18 0.270.300.240.18 0.260.250.290.20 0.270.260.250.22 2021 2020 2019 2018 0.250.260.230.26 0.150.250.280.32 0.120.180.310.39 0.120.180.270.43 0.320.280.210.20 0.290.300.230.19 0.290.250.260.19 0.300.260.230.21 0.250.260.230.26 0.220.270.250.26 0.210.220.290.28 0.220.240.240.30 0.310.270.210.20 0.240.290.260.21 0.220.240.300.24 0.220.240.270.26 2021 2020 2019 2018 0.230.260.230.28 0.130.240.280.35 0.110.160.300.43 0.100.160.260.47 0.300.270.220.21 0.270.290.240.21 0.270.240.270.21 0.280.260.230.23 0.230.260.230.28 0.210.260.250.28 0.200.220.290.30 0.200.230.250.32 0.300.270.220.21 0.210.280.270.23 0.190.220.310.28 0.200.230.270.30 2021 2020 2019 2018 0.210.250.230.30 0.110.230.280.38 0.090.140.300.47 0.080.150.250.52 0.290.270.220.22 0.260.280.240.22 0.260.240.280.23 0.270.250.240.24 0.220.250.230.29 0.190.260.260.30 0.170.210.290.33 0.180.220.250.35 0.300.270.220.22 0.200.270.270.25 0.180.210.310.30 0.180.220.280.33 2021 2020 2019 2018 0.190.250.240.32 0.090.210.280.42 0.070.130.290.51 0.070.130.240.56 0.280.260.220.23 0.250.280.240.22 0.250.240.280.23 0.260.250.240.25 0.200.250.230.31 0.170.240.260.32 0.160.200.290.36 0.160.210.250.38 0.280.260.220.23 0.180.270.280.27 0.160.200.310.33 0.160.210.280.36 2018201920202021 2021 2020 2019 2018 0.270.260.230.24 0.240.280.250.23 0.240.230.280.25 0.250.240.240.27 2018201920202021 훼 = 0.0훼 = 0.6훼 = 0.0훼 = 0.6 훼 = 0.1훼 = 0.7훼 = 0.1훼 = 0.7 훼 = 0.2훼 = 0.8훼 = 0.2훼 = 0.8 훼 = 0.3훼 = 0.9훼 = 0.3훼 = 0.9 훼 = 0.4훼 = 1.0훼 = 0.4훼 = 1.0 훼 = 0.5훼 = 0.5 ExplicitImplicit Figure 8. Normalized count of retrieved documents per year (X-axis) given the test set year (Y-axis) on SituatedQA, with queries containing explicit (Explicit) or implicit (Implicit) temporal information, when interpolated between 2018 (α = 0.0) and 2021 (α = 1.0). 24 Temporal Preference Optimization for Unsupervised Retrieval 0.370.330.30 0.340.360.29 0.370.300.32 Dec Jun Jan 0.240.350.41 0.230.370.40 0.210.300.49 0.390.320.29 0.370.350.28 0.410.310.28 0.310.330.35 0.300.370.33 0.310.310.39 0.330.340.33 0.310.370.33 0.330.310.37 Dec Jun Jan 0.230.360.41 0.230.370.40 0.200.300.50 0.380.320.30 0.360.350.28 0.400.310.29 0.300.330.36 0.290.370.34 0.290.310.40 0.290.340.37 0.280.370.36 0.280.300.42 Dec Jun Jan 0.210.360.42 0.210.380.41 0.180.310.51 0.370.320.31 0.350.360.29 0.380.310.31 0.290.340.37 0.270.370.35 0.270.310.42 0.260.350.39 0.250.370.38 0.250.310.44 Dec Jun Jan 0.200.370.44 0.190.380.42 0.160.310.54 0.360.320.32 0.330.360.30 0.360.310.33 0.270.350.38 0.250.370.37 0.240.310.45 0.250.350.40 0.240.370.39 0.230.310.46 Dec Jun Jan 0.170.360.47 0.170.390.44 0.120.310.57 0.340.330.33 0.320.360.31 0.340.310.35 0.240.360.40 0.240.380.38 0.210.310.48 0.240.350.41 0.240.370.40 0.220.300.48 JanJunDec Dec Jun Jan 0.330.330.34 0.310.370.32 0.320.310.37 JanJunDec 훼 = 0.0훼 = 0.6훼 = 0.0훼 = 0.6 훼 = 0.1훼 = 0.7훼 = 0.1훼 = 0.7 훼 = 0.2훼 = 0.8훼 = 0.2훼 = 0.8 훼 = 0.3훼 = 0.9훼 = 0.3훼 = 0.9 훼 = 0.4훼 = 1.0훼 = 0.4훼 = 1.0 훼 = 0.5훼 = 0.5 ExplicitImplicit Figure 9. Normalized count of retrieved documents per year (X-axis) given the test set year (Y-axis) on RealTimeQA, with queries containing explicit (Explicit) or implicit (Implicit) temporal information, when interpolated between January (α = 0.0) and December (α = 1.0). 25 Temporal Preference Optimization for Unsupervised Retrieval D.4. Lambda Interpolation Figure 10. Ablation ofλ, the interpolation ratio betweenL TRPO (λ = 0.0)L CE (λ = 1.0), forTPOURContriever 2018 and 2021, evaluated on SituatedQA 2018 and 2021 respectively. Performance improves significantly with moderateλvalues, showing that combining semantic and temporal supervision is more effective than relying solely on either. Dashed lines atλ = 1.0indicate performance using contrastive-only training. Vertical arrows show the performance gap compared to TPOUR ’s peak setting for each year. D.5. Queue Size Ablation Table 15. Effect of contrastive queue size on temporal retrieval performance forTPOUR-trained retrievers. We vary the queue size used in contrastive training from 100 to 256k while keeping all other training settings fixed. Performance first improves as the queue grows, indicating that a larger set of in-batch negatives helps the model learn stronger temporal preference signals. The optimal queue size is between 4k (50.20) to 16k (44.32). Model\ Queue Size1005001k2k4k16k64k128k256k TPOUR Contriever (2018)46.8748.5648.8850.0150.2049.3647.7847.3846.96 TPOUR Contriever (2021)37.2639.9740.1941.1441.8344.3242.9141.9941.15 D.6. Retrieved Document-Year Distribution Over Training To verify thatTPOURlearns to distinguish content updates over time (beyond matching explicit temporal markers), we track how the distribution of retrieved document years changes throughout training. Concretely, at several checkpoints (0k–100k steps), we retrieve documents for a fixed evaluation set and compute the fraction of retrieved documents belonging to each snapshot year (normalized so each column sums to 100%). If the model learns temporal alignment from implicit semantic shifts across versions, the retrieved-year distribution should progressively concentrate around the target snapshot time. Across training, the retrieved-year distribution shifts toward the snapshot time each model is trained to prefer.TPOUR Contriever (2018) increases the retrieved 2018 documents (25.8→31.0) while decreasing later years, whereasTPOUR Contriever (2021) increasingly concentrates on 2021 documents (37.2→42.1) while reducing earlier years. Table 16. Retrieved document-year distribution (%, normalized) over training steps for TPOUR Contriever (2018). Doc Year\ Step0k20k40k60k80k100k 201825.829.930.630.830.931.0 201924.826.026.126.226.226.2 202023.622.121.921.921.821.4 202125.721.921.321.121.121.4 26 Temporal Preference Optimization for Unsupervised Retrieval Table 17. Retrieved document-year distribution (%, normalized) over training steps for TPOUR Contriever (2021). Doc Year\ Step0k20k40k60k80k100k 201817.116.816.715.314.814.0 201919.520.419.718.818.418.0 202026.326.326.326.225.725.8 202137.236.537.239.841.142.1 D.7. Seasonal Preference ofTPOUR-trained Retriever We analyze preference on temporal patterns usingTPOUR. WhileTPOURdoes not explicitly train to capture temporal patterns (e.g., seasonal recurrences), it learns to align with the document distribution observed in corpora, which may naturally encode temporal patterns. Specifically, we investigate document distribution across a monthly set from twoTPOURretrievers (January and June, 2023). The result of document distribution, computed as the ratio of retrieved to total documents per month, is in Tab. 18. We observe the January retriever favors winter months, while the June retriever favors summer months across years. This shows TPOUR’s sensitivity to seasonal patterns without explicit supervision. Table 18. Monthly document distribution ofTPOUR-trained retrievers. We report monthly retrieval frequencies for two retrievers trained at different checkpoints (January 2023 and June 2023). The January retriever exhibits stronger alignment with winter months (e.g., December–February), while the June retriever favors summer months (e.g., May–August). This showsTPOURcan internalize seasonal patterns present in the training corpus without being explicitly trained for temporal recurrences. YearJanFebMarAprMayJunJulAugSepOctNovDec TPOUR Contriever (January, 2023) 20220.450.4480.4030.4380.4190.3650.440.5170.5360.5340.5240.656 2023 0.6280.5510.4750.430.4660.5110.4520.4050.4380.4640.4630.513 Total1.0780.9990.8780.8680.8850.8760.8920.9220.9740.9980.9871.169 TPOUR Contriever (June, 2023) 20220.270.3330.2450.3260.3760.4720.4580.4190.3870.3830.3990.402 2023 0.4390.3730.4240.3940.4590.5590.5630.3680.4310.5050.4780.449 Total0.7090.7060.6690.720.8351.0311.0210.7870.8180.8880.8770.851 E. Qualitative Case Studies E.1. Temporal Preference Learning Without Explicit Time Expressions To illustrate howTPOURcaptures temporal preferences without explicit timestamp expressions, we present a qualitative case study using the Wikipedia article Office 1 Superstore. This example shows how semantic changes across document versions serve as implicit temporal signals. Tab. 19 compares three versions of the same document from the 2018, 2021, and 2023 Wikipedia dumps used inTPOUR’s training set. The 2018 version describes contraction following the 2008 economic crisis, including market exits and a shift to e-commerce. The 2021 version reflects a structural change, emphasizing the 2018 acquisition by Panda Cooperation. By 2023, the company is portrayed as having re-expanded globally under Panda’s ownership. Notably, none of these documents contain explicit temporal information such as year strings. The distinctions arise solely from semantic content.TPOUR’s preference-based training setup contrasts such temporally distinct documents, enabling the model to learn implicit temporal alignment cues. As shown in Tab. 8, fewer than 2.5% of training documents include explicit year references, underscoring the importance of implicit signals in learning temporal preferences. 27 Temporal Preference Optimization for Unsupervised Retrieval Table 19. Three versions of the same document are used inTPOURtraining. Although no explicit timestamp strings appear in the document content, the semantic update—retrenchment (2018), ownership transfer (2021), and re-expansion (2023)—shows real-world temporal progression. TPOUR leverages such a document update to learn temporal preference without explicit supervision. TimestampTraining Document Example (Title: Office 1 Superstore) 2018-09-16Office 1 Superstores International Inc. (OFFICE 1) was founded in 1994 as a franchise retail chain selling office products and supplies, including office furniture and electronics. The company is headquartered in West Palm Beach, Florida, with international operations run from a central office and warehouse in Sofia, Bulgaria. The company uses multiple channels of distribution to reach customers, including retail stores, telemarketing, direct mail, e-commerce, and contract sales. OFFICE 1 expanded its operations through master franchises in Europe, Asia, Africa, Latin America, and the Caribbean, and at its peak had stores in 25 countries. Post the 2008 economic crisis, the company retrenched and closed vulnerable markets such as Italy, Slovenia, and Iceland, shifting focus to e-commerce. It entered France (2010) and Germany (2011) through joint ventures. 2021-11-06Office 1 International Inc. (Office 1) is an international franchise company established in Florida, USA, and present in three countries—Bulgaria, France, and Greece. On February 20, 2018, Panda Cooperation officially acquired all trademark rights of the Office 1 Superstore portfolio. From a major franchisee in Bulgaria, Panda Cooperation became the sole owner and representative of Office 1 brands worldwide. In 1998, Panda had received a master franchise for Bulgaria, and by 2021, Office 1 Superstore was the largest office supply chain in Bulgaria, serving over 130,000 business clients. 2023-11-06Office 1 International Inc. (Office 1) is an international franchise company established in Florida, USA, and currently present in 27 countries including Bulgaria, France, and Greece, with over 600 locations. Office 1 was founded in 1989 by Mark Baccash. Panda Cooperation, having acquired all Office 1 trademark rights in 2018, remains the sole global owner and operator. Office 1 maintains an extensive store network in Bulgaria and has expanded its online presence through multiple social media accounts. E.2. Comparative Analysis of Retrieved Documents Table 20. Retrieved documents comparison betweenTPOURContriever (2021) and Contriever for three example queries. The text containing the correct answers is highlighted in bold. ModelRankDocumentTimestamp Query: Who has won the most Olympic medals in curling as of 2021? TPOUR Contriever Top 1 Brad Gushue // [...] Defeating Edin in the final. [...] Defeating Scotland’s Bruce Mouat in the final. [...] 2021-11-30 Top 2 United States Curling Association // [...] Skip John Shuster’s team won the gold medal. John Shuster [...] 2021-11-30 Contriever Top 1 Canada at the Olympics // [...] Jones, Kaitlyn Lawes, Jill Officer, Dawn McEwen and spare Kirsten Wall went unbeaten [...] 2018-12-06 Top 2 Canada at the Olympics // [...] Jones, Kaitlyn Lawes, Jill Officer, Dawn McEwen and spare Kirsten Wall went unbeaten [...] 2020-11-27 Query: Who is the No. 1 ranked tennis player in the world as of 2021? TPOUR Contriever Top 1Juan Martin del Potro // [...] Lost his quarterfinal against world number 1 Novak Djokovic [...] 2021-12-11 Top 2Tennis in Spain // [...] Tying him with Federer and Novak Djokovic. [...] 2021-11-09 Contriever Top 1Tennis // [...] Novak Djokovic, a rival of both Nadal and Federer, is also [...] 2020-12-06 Top 2Alexander Zverev // [...] Novak Djokovic has said, ”Hopefully, he can surpass me.” [...] 2018-12-12 Query: What is the current macOS operating system as of 2021? TPOUR Contriever Top 1macOS // [...] macOS Monterey was presented as version 12 in 2021. [...] 2021-12-05 Top 2macOS Server // [...] macOS 12 (Server 5.12) [...] Operates on macOS Monterey (12) and later. [...] 2021-12-15 Contriever Top 1 Personal Computer // [...] macOS is a Unix-based graphical operating system, and [...] 2018-12-15 Top 2 macOS // [...] macOS Monterey was presented as version 12 in 2021. [...] 2021-12-05 28 Temporal Preference Optimization for Unsupervised Retrieval Table 21. Retrieved documents comparison betweenTPOURContriever (2018) and Contriever for three example queries. The text containing the correct answers is highlighted in bold. ModelRankDocumentTimestamp Query: When did the Golden State Warriors win the Finals as of 2018 TPOUR Contriever Top 1 Willie Green // [...] defeated the Cleveland Cavaliers in four games of the 2018 NBA Finals. [...] 2018-11-25 Top 2Jarron Collins // [...] Collins won his third championship in four years when the Warriors defeated the Cleveland Cavaliers in the 2018 NBA Finals. [...] 2019-12-27 Contriever Top 1 National Basketball Association Criticisms and Controversies // [...] Some NBA fans have accused the league of conspiring to have large-market teams [...] 2019-12-30 Top 2NBA Finals // [...] The Warriors swept the Cavaliers 4-0 [...]2020-12-11 Query: What NFL player has the most NFL rings as of 2018 TPOUR Contriever Top 1NFL Top 100 Players of 2018 // [...] It ended with reigning NFL MVP Tom Brady being ranked #1 [...] 2018-12-07 Top 2 Jeff Stoutland // [...] Stoutland won his first Super Bowl ring when the Eagles defeated the New England Patriots in Super Bowl LII. [...] 2020-12-19 Contriever Top 1 Super Bowl Ring // [...] The New England Patriots’ Super Bowl XLIX rings reportedly cost $36,500 each [...] 2019-12-30 Top 2 Super Bowl Ring // [...] Super Bowl LI ring has 283 diamonds, to commemorate their comeback [...] 2020-12-19 Query: When did the Philadelphia Eagles play in the Super Bowl last as of February 23, 2018 TPOUR Contriever Top 1 Curse of Billy Penn // [...] On February 4, 2018, the Philadelphia Eagles defeated the New England Patriots in Super Bowl LII 41-33 [...] 2018-12-07 Top 2Jeff Stoutland // [...] Stoutland won his first Super Bowl ring when the Eagles defeated the New England Patriots in Super Bowl LII. [...] 2020-12-19 Contriever Top 12018 Philadelphia Eagles Season // [...] A new Super Bowl champion would be crowned. [...] 2020-12-19 Top 2Sports-Related Curses // [...] The Eagles accumulated a lot of playoff heartbreak, including 2 Super Bowl losses [...] 2020-12-19 F. Discussion on Future Work F.1. Relaxation of Temporally Distributed Corpora As noted in Sec. 5,TPOURrequires temporally distributed corpora (e.g., Wikipedia dumps). Each dump is treated as a snapshot of world knowledge at a specific point in time (Jatowt et al., 2005). While documents may mention events from various eras, their dominant temporal context aligns with the collection period (e.g., the phrase “last week” in a 2020 dump naturally grounds to that year). This assumption allows TPOUR to induce temporal preferences at the corpus level without requiring document-level timestamp supervision. Such versioned corpora may not always be available in practice. However, we believe that utilizing coarse-grained temporal signals is a promising future direction. Coarse-grained temporal signals often exist in other domains. For example, user-generated content typically carries internal timestamps (e.g., server logs or metadata), even if not explicitly exposed. The central insight ofTPOURis that even minimal corpus-level temporal signals can be sufficient to induce temporal awareness in retrievers, without relying on explicit document-level timestamps. Moreover, document-level annotations, while useful, are often noisy, missing, or inconsistent due to edits, revisions, or formatting errors (Dhingra et al., 2022). F.2. Analysis of Temporal Grounding In practice, temporal grounding is expected to occur at the time of querying (or inference), reflecting the user’s current context for implicit queries. We first conducted a preliminary experiment to test whether aTPOUR-trained retriever optimized to predict more recent times can surpass general retriever baselines (e.g., Contriever, Nomic Embed v2 MoE). To empirically validate this assumption, we evaluated theTPOUR-trained Contriever (2021) on RealtimeQA (2023) by aggregating all 29 Temporal Preference Optimization for Unsupervised Retrieval monthly test sets from RealtimeQA. Tab. 22 shows that theTPOUR-trained Contriever (2021) outperforms general retrievers (e.g., Contriever and Nomic Embed v2 MoE) when the test set contains 2023-related queries. This shows thatTPOURcan train retrievers to handle recent queries better than general-purpose retrievers. To further analyze the impact of temporal grounding, we categorized RealTimeQA queries along two different axes. We used GPT-4o (OpenAI et al., 2024) to assign each of the 1,428 queries to both a (1) Temporal Category and a (2) Topic Category. We then manually reviewed all queries to ensure accurate classification. Queries from underrepresented topic categories (fewer than 30 examples) were grouped under “Others” to stabilize analysis. Detailed information on each category is shown at Tab. 23 and Tab. 24. Given these queries assigned to each temporal/topic category, we evaluated NDCG@5 (N@5) performance across categories. Here,∆represents the score difference between TPOUR Contriever (2021) and baseline Contriever. Tab. 25 and 26. The temporal category results show an interesting insight.TPOURContriever (2021) is especially effective on “Timeless” temporal queries, with smaller improvements for “Distant Past” queries. In terms of topic category, timely categories such as “Sports” and “Business” benefited the most, while “Health” and “Environment” showed relatively smaller performance gains over Contriever. Given the per-query∆, we further investigate a case study to examine which examplesTPOURContriever (2021) performs better on compared to Contriever in Tab. 27. It shows thatTPOURContriever (2021) outperforms Contriever on queries requiring temporal grounding by retrieving contextually and temporally aligned documents. Table 22. Performance of theTPOUR-trained retriever aligned to recent time (TPOURContriever (2021)), which surpasses general retrievers (e.g., Contriever and Nomic Embed v2 MoE). RealtimeQA (2023)N@5N@10 Contriever44.3945.25 Nomic Embed v2 MoE35.2035.88 TPOUR Contriever (2018)22.4823.92 TPOUR Contriever (2021)48.4351.22 Table 23. Topic categories. RealTimeQA (2023) queries are categorized into topical domains such as Sports, Business and Health. Queries from underrepresented domains are grouped under Others. Category# Queries Sports122 Business119 International224 Entertainment114 Politics217 Environment61 Health99 Others472 Table 24. Temporal categories. RealTimeQA (2023) queries are categorized as Timeless, Recent Past, Immediate, or Distant Past based on their temporal information. Most queries fall into the “Timeless” category, which requires retrieving temporally up-to-date documents. Category# QueriesDescription / Example Timeless687The query does not mention time, but requires up-to-date documents. e.g., “Which Covid-19 variant of Omicron become the most dominant in US?” Recent Past (≤ 1 year)165Explicitly references events from the recent past (e.g., “last year”). e.g., “How many flights on private jets were made globally last year?” Immediate504Refers to ongoing or very recent events (e.g., “this week”). e.g., “The U.S. embassy in which country was evacuated this week?” Distant Past (> 1 year)72Refers to events that occurred more than a year ago (e.g., “after the 2020”). e.g., “Dominion Voting Systems settled with which TV network in a defamation lawsuit over the broadcast of lies after the 2020 presidential election?” 30 Temporal Preference Optimization for Unsupervised Retrieval Table 25. Temporal category performance.TPOURContriever (2021) is effective on “Timeless” (+7.36) compared to baseline Contriever, while showing smaller gains on “Distant Past” (+2.23). ModelTimelessRecent Past (≤1 year)ImmediateDistant Past (>1 year) Contriever46.0844.5043.1450.81 Nomic Embed v2 MoE37.5035.1632.6940.04 TPOUR Contriever (2018)23.1723.0521.3023.72 TPOUR Contriever (2021)53.4450.5348.0553.04 ∆ over Contriever+7.36+6.03+4.91+2.23 Table 26. Topic category performance. Timely categories such as “Sports” (+10.81) and “Business” (+7.39) benefited the most from TPOURContriever (2021), while “Health” (+3.06) and “Environment” (+4.71) showed relatively smaller gains over baseline Contriever. ModelSportsBusinessInternationalEntertainmentPoliticsEnvironmentHealthOthers Contriever42.6441.9845.0745.1845.2944.6451.1146.07 Nomic Embed v2 MoE34.4533.3636.7633.7237.9635.7835.2937.28 TPOUR Contriever (2018)19.1624.3720.7920.3121.9225.3929.8123.32 TPOUR Contriever (2021)53.4549.3751.7951.2650.4849.3554.1751.88 ∆ over Contriever+10.81+7.39+6.72+6.08+5.19+4.71+3.06+5.81 Table 27. Example queries across temporal and topic categories. Each example illustrates howTPOURContriever (2021) outperforms baseline Contriever, with improvements ranging from +48.52 (Environment, Recent Past) to +86.88 (Entertainment, Immediate). QueryTemporal CategoryTopic Category∆ over Contriever The Biden administration is monitoring a potentially major labor strike brewing in which industry? TimelessPolitics+69.92 Newly released figures show that the amount of elec- tricity produced by which type of renewable energy hit a record high in Britain last year? Recent PastEnvironment+48.52 The nominees for the 75th Emmy Awards televi- sion’s top honor were announced this week. Which show received the most nominations? ImmediateEntertainment+86.88 China’s birth rate declined for the first time in decades in 2022. It has been the world’s most popu- lous nation since at least when? Distant PastInternational+78.60 F.3. Appropriate α Selection Determining the optimal interpolation weightαis a non-trivial problem. We assume temporal grounding for each query, determined by either explicit or implicit temporal intent. This offers an advantage over using a single “global” retriever to handle queries from multiple time periods. Reduced training burden. Avoids forcing a single model to learn both semantic and temporal alignment simultaneously. Temporal sensitivity. A global retriever must balance signals across many time periods, which can weaken or distort its sensitivity for specific periods. Modularity. We can decouple the problem into two subproblems. (1) Router to detect a query’s temporal intent and (2) Retriever to retrieve temporally aligned documents. Interpretability. Interpolation weights α make it easy to trace how retrieval preferences shift across time. For explicit temporal queries (e.g., “in 2019”), tools like dateparser (Scrapinghub) can be used to extract the timestamp, which directly mapsαto select or interpolate amongTPOURretrievers. For implicit temporal queries, we distinguish two types: (1) Queries referring to the current time (e.g., “Who is the current prime minister?”, “What time is it?”). In such cases, defaulting to the most recentTPOURretriever is a viable approach, under the assumption that users intend to refer to the present.TPOURContriever (2021), despite being trained two years earlier, still outperforms general-purpose retrievers on the RealTimeQA (2023) benchmark, as shown in Tab. 22. (2) Queries implying a specific but unstated time (e.g., “When was the 21st conference held?”). In these cases, training and using a query intent classifier to predict the optimalαis feasible. (Wu et al., 2024) has already demonstrated that predicting query timestamps is possible, achieving 96% test accuracy. 31 Temporal Preference Optimization for Unsupervised Retrieval G. Notations Table 28. Definitions of notations used in the above formalizations. SymbolDefinition QQuery text DDocument text D + Positive document D − Negative document D t Temporally aligned document D t ′ Temporally misaligned document S(·,·)Similarity function S θ (y w )Abbreviated form of S(π θ (Q),π θ (D t )) S θ (y l )Abbreviated form of S(π θ (Q),π θ (D t ′ )) π q Query encoder π k Document encoder π ref Reference policy (encoder) π θ Training target policy (encoder) π t θ Training target policy (encoder) that aligned at time t L(·)Loss function L TRPO (·)TRPO loss L CE (·)Contrastive loss L total (·)Total loss mMomentum hyperparameter λL TRPO andL CE balance hyperparameter αTime vector interpolation hyperparameter θ q Query (policy) encoder weight θ k Document (policy) encoder weight θ ref Reference policy weight θTraining target policy (π θ ) weight θ base Base pretrained encoder weight θ t The encoder weight fine-tuned on data from time period t y w Preferred output y l Less preferred output xPrompt input σ(·)Sigmoid function βDPO temperature parameter TContrastive loss temperature parameter τ t Time vector for time t t start Start time period t mid Middle time period t end End time period 32