Paper deep dive
STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering
Xinlong Dai, Jinchuan Zhang, Lei Gao, Xinzhe Hu, Yuefeng He, Hui Gao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/23/2026, 1:43:12 AM
Summary
The paper introduces STAIR (Semantic-Temporal Automaton for Interpretable Reasoning), a neuro-symbolic system for Temporal Question Answering (TQA). STAIR separates semantic interpretation from temporal inference using a rule-first architecture. It employs an answer-free LLM adapter for semantic normalization of complex queries and a deterministic temporal automaton for precise, verifiable evidence selection. This approach reduces probabilistic errors and improves interpretability, achieving significant F1 improvements over baselines like NeSTR on datasets such as TimeQA and TempReason.
Entities (8)
Relation Signals (7)
STAIR → evaluatedon → TimeQA
confidence 98% · Across the TimeQA-Easy, TimeQA-Hard... datasets
STAIR → evaluatedon → TempReason
confidence 98% · TempReason-L2, and TempReason-L3 datasets
STAIR → usescomponent → Temporal Automaton Selector
confidence 95% · The Temporal Automaton Selector (TAS) first attempts rule-only execution.
STAIR → usescomponent → Semantic Adapter
confidence 95% · an answer-free LLM adapter maps complex question formulations to normalized temporal intents
STAIR → outperforms → NeSTR
confidence 90% · STAIR consistently outperforms strong baselines... achieving average F1 improvements
STAIR → utilizesllm → GPT-4o-mini
confidence 90% · and GPT-4o-mini models, respectively.
STAIR → utilizesllm → Qwen2.5-7B
confidence 90% · when utilizing the Qwen2.5-7B... models
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific training. However, existing prompt-based neuro-symbolic systems continue to rely on LLMs for both semantic interpretation and exact temporal inference. Consequently, discrete decisions regarding intervals, time anchors, and ordered states remain vulnerable to probabilistic errors and difficult to verify. We present STAIR, a \textbf{S}emantic-\textbf{T}emporal \textbf{A}utomaton for \textbf{I}nterpretable \textbf{R}easoning. STAIR separates semantic interpretation from precise temporal inference: an answer-free LLM adapter maps complex question formulations to normalized temporal intents, while a deterministic temporal automaton with finite control and guarded transitions executes the corresponding policies over canonicalized evidence. Following a rule-first design, STAIR resolves standard questions without invoking an LLM and applies semantic adaptation only when the rule path fails to produce an executable intent. This approach reduces free-form reasoning, making temporal decisions verifiable and interpretable. Specifically, guarded execution supports precise point-time containment and before/after selection, while semantic adaptation handles non-exact intervals and time-anchored queries. Across the TimeQA-Easy, TimeQA-Hard, TempReason-L2, and TempReason-L3 datasets, STAIR consistently outperforms strong baselines in the TQA task using matched model settings, achieving average F1 improvements of 16.57\% and 3.10\% when utilizing the Qwen2.5-7B and GPT-4o-mini models, respectively. Furthermore, ablations and diagnostic analyses demonstrate that STAIR excels at handling both boundary-sensitive and order-sensitive queries, while its guarded execution and semantic adaptation ensure precise point-time reasoning and inexact intervals, respectively.
Tags
Links
- Source: https://arxiv.org/abs/2608.16224v1
- Canonical: https://arxiv.org/abs/2608.16224v1
Trouble viewing inline? Open PDF directly →
Full Text
79,895 characters extracted from source content.
Expand or collapse full text
STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering Xinlong Dai, Jinchuan Zhang, Lei Gao, Xinzhe Hu, Yuefeng He, Hui Gao University of Electronic Science and Technology of China (UESTC) Correspondence: jc.zhang@uestc.edu.cn Abstract By leveraging large-scale pretraining, LLMs can interpret di- verse temporal expressions and question formulations without task-specific training. However, existing prompt-based neuro- symbolic systems continue to rely on LLMs for both seman- tic interpretation and exact temporal inference. Consequently, discrete decisions regarding intervals, time anchors, and or- dered states remain vulnerable to probabilistic errors and dif- ficult to verify. We present STAIR, a Semantic-Temporal Automaton for Interpretable Reasoning. STAIR separates semantic interpretation from precise temporal inference: an answer-free LLM adapter maps complex question formula- tions to normalized temporal intents, while a deterministic temporal automaton with finite control and guarded transi- tions executes the corresponding policies over canonicalized evidence. Following a rule-first design, STAIR resolves stan- dard questions without invoking an LLM and applies seman- tic adaptation only when the rule path fails to produce an executable intent. This approach reduces free-form reason- ing, making temporal decisions verifiable and interpretable. Specifically, guarded execution supports precise point-time containment and before/after selection, while semantic adap- tation handles non-exact intervals and time-anchored queries. Across the TimeQA-Easy, TimeQA-Hard, TempReason-L2, and TempReason-L3 datasets, STAIR consistently outper- forms strong baselines in the TQA task using matched model settings, achieving average F1 improvements of 16.57% and 3.10% when utilizing the Qwen2.5-7B and GPT-4o-mini mod- els, respectively. Furthermore, ablations and diagnostic analy- ses demonstrate that STAIR excels at handling both boundary- sensitive and order-sensitive queries, while its guarded exe- cution and semantic adaptation ensure precise point-time rea- soning and inexact intervals, respectively. Introduction Temporal question answering (TQA) requires models to an- swer questions over time-indexed evidence. To provide re- liable and high-confidence answers, the reasoning process requires both the semantic interpretation of diverse question formulations and precise, discrete selection over temporally ordered evidence. TQA datasets generally share an input-output structure consisting of a question Q, a temporal context C, and an an- swer A. The temporal context frequently contains multiple time-indexed facts that can be normalized as (s,r,o,t s ,t e ). Even within the same context C, different questions Q may Question:Which political party did Louise Mensch belong to between Jul 1997 and Sep 1997 (a) LLM-based (b) Rule-based (c) STAIR Probabilistic Inference Predicted Answer Labour Party Poor discrete selection RULES Feature Extraction •Entity A •Relation R •Entity B •Start t1 •End t2 Output No exact match(fail) 1995 – 1996 : Louise Mensch‛s party is (Conservative Party) Answer Conservatives Canonical Facts belong_to(Louise Mensch, Conservatives, 1997-07-01, 1998-06-30) Structured Intent Interval(start=풕 풔 ,end=풕 풆 , policy=overlap) Deterministic Temporal Automaton 19971998 Conservatives overlap = max Inference Breakdown Context Target Range [Jul 1997,Sep 1997] Strict Temporal Matching Query: Range=[Jul 1997,Sep 1997] Fact 2: Jan 1996-Dec 1997 Fact 3: Jan 1997-Dec 1998 NLP 1996-1997 (Labour) Select Range Coarse Overlooks months Lack flexible mappingLimited Syntactic Coverage Preprocessing S Facts S ଵ Intent ... Select Answer S ସ S ହ •Fact1 •Fact2 •Fact3 Labour Party Conservatives Conservative Party Other 0.42 0.38 0.12 0.08 Candidate Answers Uncalibrated confidence Temporal Constraints 1996 –1997 : Louise Mensch‛s party is (Labour Party) 1997 – 1998 : Louise Mensch‛s party is (Conservatives) Figure 1: TQA setting that motivates STAIR. Semantic pars- ing normalizes diverse temporal question forms, whereas deterministic execution performs temporal selection. impose distinct temporal constraints, requiring boundary- aligned interval matching, non-exact interval overlap, point- in-interval containment, or before/after selection relative to temporal or entity anchors. These operations expose two re- curring challenges: boundary-sensitive reasoning over inter- val overlap and point containment, and order-sensitive rea- soning over predecessor or successor states. Time-anchored before/after questions involve both, because they require in- terpreting a temporal boundary before selecting an ordered state. Figure 1 illustrates this challenge using a non-boundary- aligned interval query. An LLM can interpret the meaning of the query but may select an incorrect adjacent state through probabilistic inference. A strict rule-based matcher provides reproducible execution, yet fails when the query boundaries do not exactly align with those of the supporting fact. The underlying difficulty is thus a fundamental mismatch: flexible semantic interpretation must accommodate diverse surface forms, whereas exact temporal execution requires applying an unambiguous policy to select the correct evidence. This contrast directly motivates our approach: separating semantic normalization from deterministic temporal execution. arXiv:2608.16224v1 [cs.CL] 17 Aug 2026 020406080100 TimeQA-Easy TimeQA-Hard TempReason-L2 TempReason-L3 76.91 27.65 76.91 91.46 Proportion of examples (%) Rule-only resolvedLLM intervention required Figure 2: Routing analysis of the rule-only temporal automa- ton across benchmarks. Each bar decomposes the dataset into examples deterministically resolved by rule-only automaton of STAIR and examples requiring LLM intervention. Recent prompting methods decompose temporal reason- ing into symbolic representation, inference, verification, re- flection, and answer generation. NeSTR combines symbolic temporal representations with LLM inference and feedback- based correction (Liang et al. 2026), while TISER constructs and revises timelines through self-reflection (Bazaga et al. 2025). Although these methods improve over direct prompt- ing, the LLM still performs the decisive temporal operations and generates the final answer. Consequently, even a coher- ent reasoning trace may select an incorrect interval bound- ary, over-retrieve before/after states, or alter the answer span. Thus, we conclude the key limitation is not merely the lack of intermediate representations, but the use of free-form gener- ation for discrete temporal decisions that admit explicit and verifiable execution. Figure 2 reveals an important asymmetry: most temporal questions already admit deterministic execution once their facts and operators are canonicalized. The rule-only path resolves 76.91% of TimeQA-Easy, 76.91% of TempReason- L2, and 91.46% of TempReason-L3. These instances mainly involve explicit from ... to ... intervals, point-time containment, or before/after relations with entity anchors. In contrast, rule-only coverage on TimeQA-Hard is only 27.65%, because many questions involve non-exact inter- vals, time-valued anchors, or formulations that do not di- rectly expose an executable operator. For example,between Apr 1987 and Nov 1988, after Jan 1996, and before Jan 1999 require interval-policy selection, an- chor typing, or structural normalization before determinis- tic execution becomes applicable. These findings suggest that LLM intervention is primarily needed to interpret non- canonical temporal expressions and infer the intended tempo- ral constraints, rather than to make every temporal decision. Motivated by these observations, we propose STAIR, a Semantic-Temporal Automaton for Interpretable Reasoning. STAIR adopts a rule-first architecture separating semantic interpretation from deterministic temporal execution. It first canonicalizes temporal facts and maps recognizable ques- tions to finite executable intents, resolved by a temporal automaton via explicit policies, finite control, and guarded transitions. If rule-based parsing fails, an answer-free seman- tic adapter maps the question into this intent space, while programmatic validation and constrained repair ensure exe- cutability. Guided by dual-process reasoning (Kahneman 2011), cog- nitive offloading (Risko and Gilbert 2016), and symbolic interval models (Allen 1983), STAIR operationalizes the principle of LLM-as-parser, Automaton-as-reasoner. On the main reasoning path, the LLM is invoked only when diffi- cult temporal expressions require semantic normalization, whereas the deterministic automaton performs temporal comparison, selects evidence, and extracts the answer. Moreover, STAIR employ a procedural interpretability rather than post-hoc. For each instance resolved by the deter- ministic temporal automaton, the system exposes the canoni- cal facts, normalized intent, activated temporal policy, guard outcomes, selected evidence, and answer provenance. Fail- ures within the rule path, semantic repairs, and invocations of the final fallback are explicitly recorded. Consequently, the reasoning trace reflects the actual computation performed by the system rather than a natural-language rationale generated after prediction. • We propose STAIR, a rule-first Semantic-Temporal Au- tomaton for Interpretable Reasoning that instantiates the principle of LLM-as-parser, Automaton-as-reasoner through answer-free semantic parsing and guarded deter- ministic temporal execution. • We introduce a hard-only semantic adapter that converts difficult temporal expressions, including non-exact inter- vals and time-anchored before/after questions, into intents that the selector can execute without allowing the LLM to choose the final answer. • STAIR is evaluated on TimeQA, TempReason and Cron- Questions dataset. Component-level ablations, fallback analysis, and rule-coverage diagnostics further charac- terize when semantic adaptation improves deterministic temporal reasoning. Related Work TQA Benchmarks and Prior Systems. Temporal reasoning has been studied through both relation-extraction resources and question-answering benchmarks. TempEval focuses on identifying temporal relations among events and time expres- sions (Verhagen et al. 2007). TempQuestions, TimeQA, and TempReason extend this setting to question answering over time-indexed facts, evolving entities, and multi-level tempo- ral relations (Jia et al. 2018a; Chen, Wang, and Wang 2021; Tan, Ng, and Bing 2023, 2024). More recent benchmarks, including TRAM, TimeBench, ChronoSense, and UnSeen- TimeQA, expose persistent limitations in event ordering, in- terval reasoning, and memorization-free temporal general- ization (Wang and Zhao 2024; Chu et al. 2024; Islakoglu and Kalo 2025; Uddin et al. 2025). Prior systems improve TQA through stronger reading- comprehension models (Zaheer et al. 2020; Raffel et al. 2020; Izacard and Grave 2021), temporal pretraining and structured knowledge representations (Yang et al. 2023; Jia et al. 2018b; Shang et al. 2022; Mavromatis et al. 2022), or prompting and programmatic reasoning (Li et al. 2024; Zhu et al. 2023; Xiong et al. 2024; Wu et al. 2024). These approaches im- prove temporal modeling but often require supervised adap- tation, task-specific graph construction, or LLM-mediated inference. STAIR instead focuses on zero-shot TQA with an explicit temporal executor. Inference-Time and Neuro-Symbolic Reasoning. Inference-time reasoning methods improve LLM delibera- tion through intermediate rationales and sampled reasoning paths (Wei et al. 2022; Wang et al. 2023), search and action- based reasoning (Yao et al. 2023, 2022), or feedback and additional test-time computation (Shinn et al. 2023; Snell et al. 2025; Guo et al. 2025). In temporal QA, TISER re- vises constructed timelines through self-reflection, whereas NeSTR combines symbolic representations with abductive LLM reasoning. These systems organize temporal evidence and reasoning, but the LLM remains responsible for apply- ing temporal relations and producing the final answer. Con- sequently, a coherent reasoning trace may still select an in- correct interval boundary or ordered state. STAIR differs by restricting the LLM to semantic normalization and executing the final temporal operation programmatically. Programmatic Execution and Cognitive Perspectives. Program-aided language modeling externalizes operations that require exact and reproducible execution (Gao et al. 2023). This division of labor is also consistent with dual- process reasoning and cognitive offloading, which motivate separating flexible interpretation from deliberate symbolic manipulation (Kahneman 2011; Risko and Gilbert 2016). Classical interval formalisms provide a basis for precise tem- poral comparison (Allen 1983). STAIR operationalizes these perspectives through a constrained semantic interface and an explicit temporal executor. Method Figure 3 presents an overview of STAIR. Given a question Q and a temporal context C, the system canonicalizes the con- text into temporal facts, maps the question to an executable intent, and ultimately aggregates the object spans selected by the automaton to produce the final answer A. The Temporal Automaton Selector (TAS) first attempts rule-only execution. If this path fails, the hard-structure de- tector routes supported difficult cases to an answer-free se- mantic adapter, whose output is validated before TAS exe- cutes again. A failed guard returns a typed reason for repair and deterministic reselection. Only when this process re- mains unsuccessful does STAIR invoke a 4-stage LLM fall- back with separate agents for symbolic representation, tem- poral inference, consistency checking, and reflection/final answer generation. Rule-First Principle. Rule-first execution defines STAIR’s default inference regime. For directly recogniz- able operators such as from t1 to t2, in t, before entity, and after entity, the context is converted into canonical facts and the question is mapped to an exe- cutable intent. The automaton then applies the correspond- ing policy without invoking an LLM. The LLM is introduced only when the rule path fails and cannot override a success- ful deterministic result, making each route explicit: rule-only execution, semantic adaptation, or final fallback. Canonical Fact Construction STAIR represents each temporal fact as a canonical tuple (s,r,o,t s ,t e ), where s, r, and o denote the subject, re- lation, and object, and t s ,t e denote normalized start and end times. The system retains provenance metadata such as extraction source. Fact construction combines deter- ministic rule parsing with constrained LLM-based sym- bolic parsing. Rule parsers handle semi-structured records such as t s -t e : subject’s relation is object and sentence-style descriptions such as subject works for object from t s to t e . LLM-assisted symbolic parsing maps context statements to the same canonical schema. The module normalizes temporal values, rejects incom- plete records, and merges facts with the same normalized key while preserving their provenance metadata. These ex- plicit representations make subsequent reasoning auditable: selected evidence can be traced to its source, and temporal policies operate over inspectable structured records rather than latent model states. Hard-Structure Detector and Semantic Interface Hard-Structure Detector. After rule-only execution fails, the Hard-structure Detector identifies cases requiring se- mantic normalization, including non-exact intervals (e.g., between t1 and t2), time-valued before/after an- chors, unresolved date formats, and questions that remain unparsed despite available canonical facts. Restricting adap- tation to these cases avoids unnecessary LLM calls and se- mantic drift. Semantic Adapter. For a difficult question Q, the adapter receives the question and canonical facts and produces a raw structured intent I 0 specifying a temporal oper- ator, arguments, and execution policy. Supported intent types include interval, point, before_anchor, after_anchor, first, and last. The adapter is answer-free: it specifies execution but cannot select evidence or generate an answer. Program Validation and Normalization. We use I 0 for the raw adapter output and I for the validated intent passed to TAS. Field aliases and date formats are canonicalized, while unsupported intents, missing temporal arguments, in- valid dates, and answer-like content are rejected. Repaired facts or intents must pass the same validation procedure be- fore deterministic reselection. Temporal Automaton Selector Automaton Formalization. TAS is a guarded extended finite-state machine M = (S, Ξ,δ,S 0 ,S term ),(1) whereS = S 0 ,...,S 5 ,S ⊥ ,S term = S 5 ,S ⊥ , and Ξ is the configuration space. Here S 0 is the initial state, S ⊥ is the Input Question Which political party did Louise Mensch belong to between Jul 1997 and Sep 1997 Temporal Contexts 1995 –1996 : Louise Mensch‛s party is (Conservative Party) 1996 –1997 : Louise Mensch‛s party is (Labour Party) 1997 –1998 : Louise Mensch‛s party is (Conservatives ) Labour Party Conservative Party Answer NeSTR STAIR •Prompt-only •interval boundaries (subject, relation, object, start, end) 푓: belong_to (Louise Mensch, Conservatives, 1997-07-01, 1998-06-30) L. M. Symbolic parsing Rule parsing start Con Jul 1997 Jun 1998 Hard-Structure Detector •Between 1987 and Nov 1988 •Before Jan 1999 •After Jan 1996 Question Context Facts Canonical facts construct ① Deterministic temporal reasoning factsintent select group apply policy aggregate answer Selector trace S2: subject = Louise Mensch, relation = belong_to S3: candidates ○1995-07-01 to 1996-06-30 → no overlap ○1996-07-01 to 1997-06-30 → no overlap ●1997-07-01 to 1998-06-30 → overlap: 3 S4: aggregate → Conservatives S5: emit answer Temporal Automaton Selector ③ Query dates Fact dates 1995-07 1996-07 1997-07 1998-07 Q Con Labour Con selected max-overlap Temporal Policies Interval: Exact/max- min/avg Point: Containment Anchor: before/after Earliest/ Latest Validation & Normalization Schema Check Time Repair Invalid Reject LLM Semantic Interface For hard-structure Semantic Normalization Intent Extraction Time-anchor Typing Optional Fact Repair Semantic adapter❄ ② Structure intent 퐼 No answer output Type: interval, Start: 1997-07-01 End: 1997-09-30 Policy: overlap Normalized intent Interval(1997-07-01, 1997-09-30,overlap) Block answer leakage Easy / Hard Structure Hard-Structure Branch Figure 3: Overview of STAIR. The system first attempts rule-only execution. Difficult cases undergo semantic adaptation and validation before the Temporal Automaton Selector performs guarded temporal reasoning over canonical facts and exposes the activated policy, selected evidence, and emitted answer. typed failure state, and states S 1 through S 5 record the suc- cessful completion of fact validation, intent validation, fact- group selection, policy execution, and answer aggregation, respectively. The transition function is δ :S× Ξ→S× Ξ. A configuration is ξ = (F,I,G,F ⋆ ,A), where F is the set of canonical facts, I = (s q ,r q ,τ,α,π) is a validated intent, s q and r q are the target subject and relation, τ is the intent type, α is an optional entity- or time-valued anchor, π is the execution policy, G is the selected fact group, F ⋆ is the selected evidence set, and A is the answer buffer whose terminal content is returned as the final answer. For k ∈ 0,..., 4, δ(S k ,ξ) = (S k+1 ,u k (ξ)), g k (ξ) = 1, (S ⊥ ,ξ),g k (ξ) = 0, (2) where g k : Ξ → 0, 1 is the guard at step k and u k : Ξ→ Ξ is the corresponding deterministic update. The guards successively verify canonical facts, the validated intent, fact- group selection, policy output, and answer aggregation. A successful execution follows S 0 → S 1 → S 2 → S 3 → S 4 → S 5 , whereas a failed guard enters S ⊥ and returns a typed failure reason to the semantic-repair controller. The finite-state execution trace makes each prediction au- ditable through the evaluated guards, activated temporal pol- icy, and selected evidence. If execution reaches S ⊥ , the first failed guard identifies the stage requiring repair. Fact Grouping and State Chains. Given a set of facts F and the target key (s q ,r q ) from the validated intent, TAS selects G s q ,r q =f i ∈ F | s i = s q , r i = r q .(3) Here f i = (s i ,r i ,o i ,t s i ,t e i ) is a canonical fact, and the se- lected group in the automaton configuration is G = G s q ,r q . Each group defines a local temporal state chain. For be- fore/after questions, TAS follows this chain to the nearest predecessor or successor rather than returning every fact on the corresponding side of the anchor. Temporal Policies. Let t s f and t e f denote the normalized boundaries of a candidate fact f, q = [q s ,q e ] an interval query with normalized boundaries q s and q e , and q t a nor- malized point-time query. For an exact interval intent, TAS selects F ⋆ = f ∈ G s q ,r q | t s f = q s ∧ t e f = q e . For a non-exact interval, it computes ω(f,q) = max 0, min(t e f ,q e )− max(t s f ,q s ) , (4) where ω(f,q) is the overlap length between fact f and query interval q, and retains the facts attaining the largest positive overlap. For a point-time intent, TAS applies containment: F ⋆ =f ∈ G s q ,r q | t s f ≤ q t ≤ t e f . For an entity-anchored intent, TAS locates the anchor fact f a , whose normalized boundaries are t s a and t e a , and selects the nearest temporal predecessor or successor: F ⋆ = argmax f∈G s q ,r q : t e f ≤t s a t e f , before, argmin f∈G s q ,r q : t s f ≥t e a t s f , after. (5) For time-valued anchors, the same predecessor or successor search is applied directly to the normalized anchor time. For first and last intents, TAS selects the earliest and latest facts in the corresponding temporal state chain, respectively. If a policy yields no admissible evidence or multiple distinct answers, the guard at S 3 routes the instance to S ⊥ with a typed failure reason. Direct Answer Emission and Fallback When TAS reaches S 5 , STAIR copies the answer directly from the selected object spans, preventing unsupported enti- ties and surface-form changes. If a guard fails, the returned failure type is used to repair the canonical facts or intent without generating an answer. Every repaired output must be revalidated before TAS is executed again. STAIR invokes the 4-stage LLM fallback only when se- mantic repair and deterministic reselection both fail. Unlike the answer-free repair branch, this final path may generate an answer through its reflection and final-answer stage, preserv- ing coverage for cases outside the current policy inventory. Experiments Experimental Setup Datasets. We evaluate STAIR on the complete test sets of TimeQA-Easy and TimeQA-Hard from TimeQA, as well as TempReason-L2 and TempReason-L3 from TempRea- son. We additionally evaluate cross-source transfer on the operator-supported subset of the official CronQuestions test split (Saxena, Chakrabarti, and Talukdar 2021), converted to the same question–context format. Baselines and Models. The primary comparison is with NeSTR under matched model and benchmark settings. Ta- ble 1 reports the published NeSTR scores; we separately re- produced its original structured prompt and obtained closely aligned results, using this reproduction only as a consistency check. TISER is included as an external matched-model ref- erence using values reported by its authors. We evaluate Qwen2.5-7B, Qwen3-8B, Qwen3-14B, and GPT-4o-mini. Implementation Details. All STAIR runs use the same input contexts, data splits, answer normalization procedure, and evaluation metrics as the NeSTR comparison. For the 4-stage LLM fallback, STAIR adapts the original NeSTR stage descriptions into four agent-specific prompts, invoked only when both semantic repair and deterministic reselection fail. This decomposition is applied only to STAIR; NeSTR remains unmodified. Generation uses temperature 0.1 and a maximum output length of 1024 tokens. STAIR is run independently three times, and we report mean performance in the main paper. Standard deviations are provided in the Appendix. No manual filtering is applied to the four main benchmarks. The experiments are conducted in a PyTorch 2.11.0 with CUDA 13.0 environment running on Ubuntu 24.04, with an NVIDIA GeForce RTX 4090 GPU used for acceleration. Evaluation Metrics. We report Exact Match (EM) and token-level F1 following the standard TQA protocol. EM requires the normalized prediction to match the reference exactly, while F1 assigns partial credit through token overlap, distinguishing exact entity or event selection from partially correct spans. Main Results Table 1 demonstrates that STAIR achieves the highest aver- age Exact Match (EM) and F1 scores across all evaluated models. Relative to NeSTR, STAIR increases the average F1 score by margins ranging from 2.78 points (with GPT- 4o-mini) to 12.71 points (with Qwen2.5-7B), accompanied by EM gains between 5.96 and 17.08 points. Across 16 dis- tinct model-dataset configurations, STAIR improves EM in all cases and F1 in 15. Furthermore, the average F1 score of STAIR varies by only 3.07 points across different mod- els, compared with a variance of 13 points for NeSTR. This significant reduction in variance demonstrates a decreased sensitivity to the underlying capacity of the language model. The most substantial improvements occur on TempReason-L2 and TempReason-L3, achieving aver- age F1 enhancements of 9.48 and 9.62 points over NeSTR, respectively. In contrast, the corresponding gains are 1.89 points on TimeQA-Easy and 3.08 points on TimeQA-Hard. This pattern aligns with the underlying temporal operations. TempReason-L2 and TempReason-L3 primarily contain point-time containment and entity-anchored before/after queries, for which boundary- and order-sensitive decisions map directly to deterministic policies. TimeQA-Hard instead concentrates non-exact intervals and time-anchored before/after queries, which require semantic normalization before the same policies can be executed. Notably, EM improvements consistently surpass F1 gains, indicating that STAIR enhances exact evidence selection rather than merely inflating partial lexical overlap. While performance remains consistent overall, the most notable variations emerge within TimeQA-Hard. For in- stance, when utilizing the Qwen3-14B model on this subset, STAIR increases the EM score from 82.20 to 84.18 but simul- taneously experiences a marginal F1 decrease from 87.30 to 86.59. This localized exception highlights a fundamental and inherent trade-off: imposing deterministic exact selection can occasionally penalize granular token-level overlap on com- plex temporal queries where standard baseline models might generate partially correct, yet overly verbose responses. Component-level Ablation Independent-Run Variation. The Qwen2.5-7B STAIR entries in Tables 1 and 2 are evaluated under the identical running environment. The minor differences in EM and F1 arise from sampling at temperature 0.1, not from a change in configuration. We use TimeQA-Hard for the main ablation study because its non-exact intervals, time-valued anchors, and low rule-only coverage make it the most diagnostic set- ting for evaluating STAIR. Table 2 deconstructs the architecture of STAIR on the TimeQA-Hard dataset to evaluate the individual contribu- tions of the core modules. Note that executor-side ablations are deferred to the Appendix. ModelStrategy TimeQA-Easy TimeQA-Hard TempReason-L2 TempReason-L3Avg EMF1EMF1EMF1EMF1EMF1 Open LLMs Qwen2.5-7B TISER 86.80 92.60 64.30 71.50 61.10 69.80 72.60 77.60 71.20 77.90 NeSTR 85.10 90.20 64.80 71.20 61.50 68.60 73.10 76.70 71.10 76.70 STAIR 93.27 94.25 76.94 80.91 87.61 87.72 94.90 94.77 88.18 89.41 Qwen3-8B TISER 88.80 93.40 77.10 82.50 73.70 78.40 84.30 87.50 80.90 85.40 NeSTR 89.50 94.20 77.70 83.40 79.20 83.50 84.90 87.20 82.80 87.10 STAIR 94.79 95.52 84.32 86.37 88.75 91.31 96.18 95.52 91.01 92.18 Qwen3-14B TISER 90.00 94.30 82.10 87.20 75.50 80.60 81.60 85.20 82.30 86.80 NeSTR 91.10 94.50 82.20 87.30 79.50 84.60 85.10 88.90 84.50 88.80 STAIR 95.01 96.07 84.18 86.59 86.69 90.87 95.96 95.44 90.46 92.25 Closed LLMs GPT-4o-mini TISER 86.70 91.90 74.30 79.90 77.70 84.10 82.30 87.10 80.20 85.80 NeSTR 93.70 96.40 81.70 85.90 80.80 86.40 84.60 90.00 85.20 89.70 STAIR 96.57 97.02 83.97 86.24 89.13 91.10 96.16 95.54 91.46 92.48 Table 1: Exact Match (EM) and token-level F1 results on four temporal reasoning benchmarks under matched model settings.All four benchmarks are evaluated on their complete test sets. NeSTR and TISER values are taken from prior work, while STAIR values are averaged over three independent runs. Boldface indicates the best result for each model and metric. VariantEM F1 Adapter Fallback (%) STAIR-Core68.06 72.62 –72.35 + hard detector68.29 72.75 –72.35 + semantic adapter74.95 80.91 35.3537.00 + validation/repair75.11 81.12 35.4836.87 STAIR complete77.16 81.20 43.7320.50 w/o max-overlap74.01 80.26 43.5320.89 w/o time-anchor typing 76.28 80.19 36.3928.01 w/o 4-stage lLM fallback 70.01 72.47 43.47– Table 2: Interface ablation on TimeQA-Hard with Qwen2.5- 7B. Fallback denotes the final 4-stage LLM fallback rate. The result shows that the hard-structure detector alone provides little benefit because detection does not resolve the underlying interface failure. The main improvement comes from the semantic adapter, which converts previously un- supported question forms into executable intents, increasing F1 from 72.75 to 80.91 while reducing the final fallback rate from 72.35% to 37%. Validation and repair yield only a mod- est additional accuracy gain, but they improve reliability by preventing malformed intents from entering the automaton. The complete system achieves the best EM and F1 to- gether with the lowest fallback rate. The policy ablations reveal two distinct effects: removing max-overlap primarily degrades evidence selection quality, whereas removing time- anchor typing reduces the proportion of questions that can be executed deterministically. Eliminating the 4-stage LLM fall- back causes a substantial performance drop, confirming its role as a coverage mechanism for cases outside the current policy inventory. Overall, the results support the intended division of labor: the LLM maps difficult language into con- strained structures, while the automaton performs the final temporal decision through explicit policies. MethodSamples EMF1 NeSTR13,096 72.00 81.03 STAIR13,096 88.42 87.63 Improvement–+16.42 +6.60 Table 3: Cross-source transfer results on the operator- supported CronQuestions subset. All model configurations are evaluated over three independent runs. Cross-source Transfer Assessment Table 3 evaluates transfer on the successfully converted, operator-supported subset of CronQuestions. From the of- ficial test split of 30,000 instances, we exclude 14,308 ques- tions with unsupported task or answer types and 2,596 type-supported questions whose annotations lack the head field required by the converter to retrieve a subject–relation temporal timeline. This yields 13,096 instances cover- ing simple_entity, entity-answer first_last, and entity-answer before_after. NeSTR and STAIR use ex- actly the same converted questions, answers, and temporal contexts. On this shared subset, STAIR outperforms NeSTR by 16.42 EM points and 6.60 F1 points. The result demon- strates transfer of the current canonical representation and temporal policy inventory across data sources. It should be read as evidence for the converted, operator-supported subset rather than as a claim about unseen temporal operators or the complete CronQuestions test split. Efficiency and Fallback Analysis As shown in Table 4, The efficiency pattern is struc- ture dependent. On TimeQA-Easy, TempReason-L2, and TempReason-L3, STAIR reduces calls, token usage, and wall-clock time; the largest reduction appears on TempReason-L3, where average calls decrease from 1 to 0.21 and time from 3.90s to 0.31s. On TimeQA-Hard, non-exact DatasetSystem Calls Input tok. Output tok. Time TimeQA-Easy NeSTR 1.00548.0288.1 2.66s STAIR 0.64256.052.90.82s TimeQA-Hard NeSTR 1.00543.8351.5 3.13s STAIR 2.12929.3282.0 3.61s TempReason-L2 NeSTR 1.00689.6462.2 3.98s STAIR 0.89545.8100.4 1.44s TempReason-L3 NeSTR 1.00707.3450.3 3.90s STAIR 0.2181.919.80.31s Macro summary−3.65% tokens:−43.87% 2.21× Table 4: Efficiency comparison between the NeSTR baseline and STAIR. Calls and token counts are reported per instance. Results from a single run of Qwen2.5-7B-Instruct. intervals and time-anchored before/after questions require additional adaptation and repair, increasing calls from 1 to 2.12 and wall-clock time from 3.13s to 3.61s. This result reflects an accuracy-efficiency trade-off on structurally diffi- cult inputs: STAIR spends additional computation to convert complex language into validated intents while preserving deterministic final selection. Diagnostic Analysis Category-level Question Structure. Table 5 operational- izes the two high-level temporal challenges into four exe- cutable query categories. Non-exact interval and point-time questions instantiate boundary-sensitive reasoning through interval overlap and point containment, respectively. Time- anchored before/after questions combine boundary-sensitive anchor interpretation with order-sensitive predecessor or suc- cessor selection. STAIR achieves F1 gains of 8.19 points on non-exact inter- vals and 9.28 points on point-time queries. The former relies strongly on semantic adaptation, whereas the latter is handled without adapter intervention, showing that deterministic ex- ecution is beneficial both after semantic normalization and when the temporal operator is explicit. For time-anchored queries, STAIR improves F1 by 6.92 points on before ques- tions and 4.56 points on after questions. The substantially higher fallback rate for time-anchor after indicates that this category more frequently exceeds the coverage of the current semantic interface and policy inventory. Error Analysis. The majority of residual errors arise when contexts resist conversion into canonical facts, query bound- aries intersect multiple plausible intervals, adapter outputs fail validation, or questions require operators outside the predefined policy inventory. Currently unsupported phenom- ena include highly implicit event ordering, duration compar- isons, negation, nested constraints, and multi-hop temporal composition. These limitations reflect the deterministic de- sign scope of STAIR, which requires canonical facts, a finite temporal intent, and directly extractable answers. Instances violating these assumptions trigger rejection; if repair pro- cesses fail, they are routed to the LLM fallback mechanism. (a) Performance by temporal category CategoryPropertyn NeSTRSTAIR∆ EMF1 EMF1EMF1 Non-exact intervalB1285 68.53 75.02 78.60 83.21 +10.06 +8.20 Point-timeB1411 62.37 69.26 74.34 78.54 +11.98 +9.28 Time-anchor before B+O248 76.88 80.57 87.50 87.49 +10.62 +6.93 Time-anchor after B+O134 65.17 73.61 73.88 78.17 +8.71 +4.56 (b) STAIR execution diagnostics CategoryAdapter (%)Fallback (%) Non-exact interval86.1513.62 Point-time0.0022.18 Time-anchor before83.4716.53 Time-anchor after23.8876.12 Table 5: Single-run category-level diagnostics on TimeQA- Hard using Qwen2.5-7B-Instruct. B denotes boundary- sensitive, whereas B+O denotes both boundary-sensitive and order-sensitive. ∆ is calculated as STAIR minus NeSTR. Consequently, while this fallback expands the range of addressable queries, these predictions inherently lack the guard-level execution traces provided by the deterministic TAS module, causing a loss of procedural interpretability. Conclusion We presented STAIR, a rule-first framework separating se- mantic interpretation from precise temporal execution in zero-shot TQA. Regular questions are resolved directly by a deterministic temporal automaton, while an answer-free semantic adapter maps difficult formulations into validated intents executable by the same automaton over canonical- ized evidence. Experiments on four TQA benchmarks and the operator-supported subset of CronQuestions demonstrate consistent improvements over the NeSTR baseline across open-source and proprietary models. Category-level diag- nostics show gains on non-exact intervals, point-time con- tainment, and time-anchored before/after questions, cover- ing both boundary- and order-sensitive reasoning. Ablations attribute these improvements to semantic adaptation, time- anchor typing, guarded temporal policies, and deterministic answer emission. Efficiency results show that rule-first exe- cution substantially reduces model calls on structurally reg- ular datasets, although difficult constraints incur additional adaptation costs. STAIR demonstrates that restricting LLMs to semantic interfaces while delegating discrete temporal de- cisions to interpretable executors provides an effective and reliable approach to temporal question answering. References Allen, J. F. 1983. Maintaining knowledge about temporal intervals. Communications of the ACM, 26(11): 832–843. Bazaga, A.; Blloshmi, R.; Byrne, B.; and de Gispert, A. 2025. Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language Models. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Pro- ceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 28014– 28033. Vienna, Austria: Association for Computational Lin- guistics. ISBN 979-8-89176-251-0. Chen, W.; Wang, X.; and Wang, W. Y. 2021. A Dataset for Answering Time-Sensitive Questions. In 35th Conference on Neural Information Processing Systems (NeurIPS 2021) Track on Datasets and Benchmarks. Chu, Z.; Chen, J.; Chen, Q.; Yu, W.; Wang, H.; Liu, M.; and Qin, B. 2024. TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), 1204–1228. Bangkok, Thailand: Association for Computational Linguis- tics. Gao, L.; Madaan, A.; Zhou, S.; Alon, U.; Liu, P.; Yang, Y.; Callan, J.; and Neubig, G. 2023. PAL: Program-aided Language Models. In Krause, A.; Brunskill, E.; Cho, K.; Engelhardt, B.; Sabato, S.; and Scarlett, J., eds., Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, 10764–10799. PMLR. Guo, D.; Yang, D.; Zhang, H.; Song, J.; Wang, P.; Zhu, Q.; Xu, R.; Zhang, R.; Ma, S.; Bi, X.; et al. 2025. DeepSeek- R1 Incentivizes Reasoning in LLMs through Reinforcement Learning. Nature, 645(8081): 633–638. Islakoglu, D. S.; and Kalo, J.-C. 2025. ChronoSense: Explor- ing Temporal Understanding in Large Language Models with Time Intervals of Events. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), 590–602. Vienna, Austria: Asso- ciation for Computational Linguistics. ISBN 979-8-89176- 252-7. Izacard, G.; and Grave, E. 2021. Leveraging Passage Re- trieval with Generative Models for Open Domain Question Answering. In Merlo, P.; Tiedemann, J.; and Tsarfaty, R., eds., Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, 874–880. Online: Association for Computa- tional Linguistics. Jia, Z.; Abujabal, A.; Saha Roy, R.; Strötgen, J.; and Weikum, G. 2018a. TempQuestions: A Benchmark for Temporal Ques- tion Answering. In Companion Proceedings of the The Web Conference 2018, 1057–1062. Jia, Z.; Abujabal, A.; Saha Roy, R.; Strötgen, J.; and Weikum, G. 2018b. TEQUILA: Temporal Question Answering over Knowledge Bases. In Proceedings of the 27th ACM Interna- tional Conference on Information and Knowledge Manage- ment, CIKM ’18, 1807–1810. New York, NY, USA: Associ- ation for Computing Machinery. ISBN 9781450360142. Kahneman, D. 2011. Thinking, Fast and Slow. Farrar, Straus and Giroux. Li, X.; Cheng, L.; Tan, Q.; Ng, H. T.; Joty, S.; and Bing, L. 2024. Unlocking Temporal Question Answering for Large Language Models with Tailor-Made Reasoning Logic. Liang, F.; Zeng, W.; Zhao, R.; and Zhao, X. 2026. NeSTR: A Neuro-Symbolic Abductive Framework for Temporal Rea- soning in Large Language Models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 31907–31915. Mavromatis, C.; Subramanyam, P. L.; Ioannidis, V. N.; Adeshina, A.; Howard, P. R.; Grinberg, T.; Hakim, N.; and Karypis, G. 2022. TempoQR: Temporal Question Reasoning over Knowledge Graphs. In Proceedings of the AAAI Con- ference on Artificial Intelligence, volume 36, 5825–5833. Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; and Liu, P. J. 2020. Explor- ing the Limits of Transfer Learning with a Unified Text-to- Text Transformer. Journal of Machine Learning Research, 21(140): 1–67. Risko, E. F.; and Gilbert, S. J. 2016. Cognitive offloading. Trends in cognitive sciences, 20(9): 676–688. Saxena, A.; Chakrabarti, S.; and Talukdar, P. 2021. Question Answering Over Temporal Knowledge Graphs. In Zong, C.; Xia, F.; Li, W.; and Navigli, R., eds., Proceedings of the 59th Annual Meeting of the Association for Computational Lin- guistics and the 11th International Joint Conference on Nat- ural Language Processing (Volume 1: Long Papers), 6663– 6676. Online: Association for Computational Linguistics. Shang, C.; Wang, G.; Qi, P.; and Huang, J. 2022. Improv- ing Time Sensitivity for Question Answering over Temporal Knowledge Graphs. In Muresan, S.; Nakov, P.; and Villav- icencio, A., eds., Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8017–8026. Dublin, Ireland: Association for Computational Linguistics. Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; and Yao, S. 2023. Reflexion: language agents with verbal rein- forcement learning. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neu- ral Information Processing Systems, volume 36, 8634–8652. Curran Associates, Inc. Snell, C.; Lee, J.; Xu, K.; and Kumar, A. 2025. Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning. Tan, Q.; Ng, H. T.; and Bing, L. 2023. Towards Bench- marking and Improving the Temporal Reasoning Capability of Large Language Models. In Rogers, A.; Boyd-Graber, J.; and Okazaki, N., eds., Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 14820–14835. Toronto, Canada: Association for Computational Linguistics. Tan, Q.; Ng, H. T.; and Bing, L. 2024. Towards Robust Tem- poral Reasoning of Large Language Models via a Multi-Hop QA Dataset and Pseudo-Instruction Tuning. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Associa- tion for Computational Linguistics: ACL 2024, 6272–6286. Bangkok, Thailand: Association for Computational Linguis- tics. Uddin, M. N.; Saeidi, A.; Handa, D.; Seth, A.; Son, T. C.; Blanco, E.; Corman, S.; and Baral, C. 2025. Un- SeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs’ Memorization. In Che, W.; Nabende, J.; Shutova, E.; and Pilehvar, M. T., eds., Proceedings of the 63rd An- nual Meeting of the Association for Computational Linguis- tics (Volume 1: Long Papers), 1873–1913. Vienna, Austria: Association for Computational Linguistics. ISBN 979-8- 89176-251-0. Verhagen, M.; Gaizauskas, R.; Schilder, F.; Hepple, M.; Katz, G.; and Pustejovsky, J. 2007. Semeval-2007 task 15: Tem- peval temporal relation identification. In Proceedings of the fourth international workshop on semantic evaluations (SemEval-2007), 75–80. Wang, X.; Wei, J.; Schuurmans, D.; Le, Q. V.; Chi, E. H.; Narang, S.; Chowdhery, A.; and Zhou, D. 2023. Self- Consistency Improves Chain of Thought Reasoning in Lan- guage Models. In International Conference on Learning Representations. Wang, Y.; and Zhao, Y. 2024. TRAM: Benchmarking Tem- poral Reasoning for Large Language Models. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Findings of the Associa- tion for Computational Linguistics: ACL 2024, 6389–6415. Bangkok, Thailand: Association for Computational Linguis- tics. Wei, J.; Wang, X.; Schuurmans, D.; Bosma, M.; Ichter, B.; Xia, F.; Chi, E. H.; Le, Q. V.; and Zhou, D. 2022. Chain- of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Sys- tems 35, volume 35, 24824–24837. Curran Associates, Inc. Wu, S.; Li, J.; Zhang, X.; and Feng, Z. 2024. An Event- based Abductive Learning for Hard Time-sensitive Question Answering. In Calzolari, N.; Kan, M.-Y.; Hoste, V.; Lenci, A.; Sakti, S.; and Xue, N., eds., Proceedings of the 2024 Joint International Conference on Computational Linguis- tics, Language Resources and Evaluation (LREC-COLING 2024), 1105–1115. Torino, Italia: ELRA and ICCL. Xiong, S.; Payani, A.; Kompella, R.; and Fekri, F. 2024. Large Language Models Can Learn Temporal Reasoning. In Ku, L.-W.; Martins, A.; and Srikumar, V., eds., Proceedings of the 62nd Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers), 10452–10470. Bangkok, Thailand: Association for Computational Linguis- tics. Yang, S.; Li, X.; Bing, L.; and Lam, W. 2023. Once Upon a Time in Graph: Relative-Time Pretraining for Complex Temporal Reasoning. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empiri- cal Methods in Natural Language Processing, 11879–11895. Singapore: Association for Computational Linguistics. Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; Griffiths, T.; Cao, Y.; and Narasimhan, K. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; and Levine, S., eds., Advances in Neural Information Processing Systems, volume 36, 11809–11822. Curran Associates, Inc. Yao, S.; Zhao, J.; Yu, D.; Shafran, I.; Narasimhan, K. R.; and Cao, Y. 2022. ReAct: Synergizing Reasoning and Acting in Language Models. In NeurIPS 2022 Foundation Models for Decision Making Workshop. Zaheer, M.; Guruganesh, G.; Dubey, K. A.; Ainslie, J.; Al- berti, C.; Ontanon, S.; Pham, P.; Ravula, A.; Wang, Q.; Yang, L.; and Ahmed, A. 2020. Big Bird: Transformers for Longer Sequences. In Larochelle, H.; Ranzato, M.; Hadsell, R.; Balcan, M.; and Lin, H., eds., Advances in Neural Informa- tion Processing Systems, volume 33, 17283–17297. Curran Associates, Inc. Zhu, X.; Yang, C.; Chen, B.; Li, S.; Lou, J.-G.; and Yang, Y. 2023. Question Answering as Programming for Solv- ing Time-Sensitive Questions. In Bouamor, H.; Pino, J.; and Bali, K., eds., Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 12775– 12790. Singapore: Association for Computational Linguis- tics. Appendix STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering Xinlong DaiJinchuan ZhangLei Gao Xinzhe HuYuefeng HeHui Gao University of Electronic Science and Technology of China (UESTC) Correspondence:jc.zhang@uestc.edu.cn A Formal Definition of the Tem- poral Automaton This appendix defines the executable contract of TAS: a finite guarded controller over typed temporal config- urations. Table 1 summarizes the notation used below; symbols shared with the Method section keep the same meaning. A.1 Canonical Fact Definition LetTdenote the totally ordered temporal domain after date canonicalization. Definition 1(Canonical temporal fact).A canonical fact is a tuple f= (s,r,o,t s ,t e ),(1) wheresis the subject,ris the relation,ois the object span to be returned if the fact is selected, andt s ,t e ∈T are the normalized start and end boundaries satisfying t s ≤t e . For a raw temporal contextC, the fact canonicalizer induced by the input grammarG fact is written as Γ fact :C→F, F=f i n i=1 .(2) Facts from deterministic patterns or the constrained sym- bolic parser must share this tuple schema. A fact is admitted only when its entity fields are nonempty, its boundaries are parseable intoT, andt s ≤t e . Duplicate normalized facts may be merged, while provenance meta- data is retained outside the tuple for auditing and is not read by the transition rules. For a target subject–relation pair (s q ,r q ), TAS forms a local temporal state chain G s q ,r q =f i ∈F|s i =s q ∧r i =r q ,(3) ordered by (t s i ,t e i ). All deterministic policies operate on this local chain, which prevents before/after selection from crossing relations. A.2 Intent Space The query interface converts a question into a finite, answer-free intent. The supported intent-type inventory is Y=interval,point,beforeanchor, after anchor,first,last. (4) A validated intent is I= (s q ,r q ,τ,α,π), τ∈Y,(5) where (s q ,r q ) identifies the local fact chain,τis the temporal intent type,αis the typed temporal argument, andπis the deterministic policy selected for execution. The typed argumentαdepends onτ: •Forinterval,α= (q s ,q e ,m), whereq s ,q e ∈ T, q s ≤q e , andm∈ exact,overlapdistinguishes boundary-aligned matching from non-exact interval matching. •Forpoint,α=q t , whereq t ∈ Tis a normalized point-time query. •Forbeforeanchorandafteranchor,α= (a,χ), whereχ∈entity,timespecifies whether the an- chor is another object in the same fact chain or a normalized time value. •Forfirstandlast,α=∅. The query grammarG query and, when needed, the answer-free semantic adapter both target this same intent space: Γ query : (Q,F)→I 0 , V intent (I 0 ,F)→I.(6) HereI 0 is the raw parsed or adapted intent, andV intent re- jects unsupported operators, ill-typed arguments, invalid dates, unresolved target keys, and answer-like adapter outputs. Thus, the LLM-assisted branch may propose an executable temporal interface, but it cannot choose the final evidence or answer. 1 SymbolMeaning C,QRaw temporal context and question. T,≤Normalized temporal domain and its total order. f= (s,r,o,t s ,t e )Canonical fact: subject, relation, object, start time, and end time. F=f i n i=1 Set of admitted canonical facts;ndenotes its size. Γ fact ,Γ query ,V intent Fact canonicalizer, query intent parser, and intent validator/normalizer. I 0 ,IRaw parsed/adapted intent and validated intent. (s q ,r q ,τ,α,π)Target subject, relation, intent type, typed temporal argument, and policy inI. Y,ΠSupported intent-type inventory and deterministic policy inventory. G=G s q ,r q Local subject–relation fact chain used by TAS. F ⋆ ,ASelected evidence set and answer buffer. M,S,Ξ,δTAS automaton, state set, configuration space, and transition function. S 0 ,S 5 ,S ⊥ ,S term Initial state, successful terminal state, typed failure state, and terminal-state set. g k ,u k ,RTransition guard, deterministic update, and auxiliary failure reason. q= [q s ,q e ],q t Interval query and point-time query. [a s ,a e ],f a Anchor interval and resolved anchor fact. ω(f,q),κ T ,emitInterval-overlap score, time-normalization key, and answer-emission operator. uniqueDuplicate-removal operator over normalized answer spans. Table 1: Notation used in the appendix formalization and policy details. A.3 Automaton Definition TAS is a guarded extended finite-state automaton M= (S,Ξ,δ,S 0 ,S term ),(7) where Ξ is the configuration space described below and S=S 0 ,S 1 ,S 2 ,S 3 ,S 4 ,S 5 ,S ⊥ ,S term =S 5 ,S ⊥ . (8) S 0 is the initial state,S 5 is the successful terminal state, andS ⊥ is the typed failure state that triggers semantic repair or the final fallback. The intent-type setYis not an additional state set; it specifies the admissible values ofτinside the intent component of the configuration. A TAS configuration is ξ= (F,I,G,F ⋆ ,A),(9) whereFis the canonical fact set,Iis the validated intent, Gis the local fact chain,F ⋆ is the selected evidence set, andAis the answer buffer. The transition function is δ:S×Ξ→S×Ξ.(10) Each successful nonterminal transition is guarded: δ(S k ,ξ) = ( (S k+1 ,u k (ξ)), g k (ξ) = 1, (S ⊥ ,ξ),g k (ξ) = 0, k∈0,...,4. (11) whereg k is a Boolean guard andu k is the determinis- tic update for that stage. If a guard fails, the selector configuration is left unchanged, and the implementation records the first failure reasonRin the trace. Table 2 lists the guarded execution path. Operationally, S 0 stores the raw input,S 1 stores canonical facts,S 2 stores the validated intent,S 3 binds the local chain and policy,S 4 stores selected evidence, andS 5 emits the an- swer.S ⊥ records a typed failure reason without emitting an answer. A.4 Policy Semantics Letf= (s,r,o,t s ,t e ) be a candidate fact in the selected chainG. The deterministic policy set is finite: Π =π exact ,π overlap ,π point ,π before ,π after ,π first ,π last . (12) For an exact interval queryq= [q s ,q e ], π exact (G,q) =f∈G|t s f =q s ∧t e f =q e .(13) For a non-exact interval query, TAS computes the positive overlap ω(f,q) = max 0,min(t e f ,q e )−max(t s f ,q s ) (14) and selects the maximizers: π overlap (G,q) = argmax f∈G:ω(f,q)>0 ω(f,q).(15) For a point-time query, π point (G,q t ) =f∈G|t s f ≤q t ≤t e f .(16) For before/after intents, the anchor resolver first con- vertsαinto an anchor interval [a s ,a e ]. Ifαis time-valued, thena s =a e =a. Ifαis entity-valued, TAS locates an anchor factf a ∈Gwhose object matches the anchor entity and sets [a s ,a e ] = [t s a ,t e a ]. The predecessor and successor policies are π before (G,α) = argmax f∈G:t e f ≤a s t e f ,(17) π after (G,α) = argmin f∈G:a e ≤t s f t s f .(18) For first/last intents, π first (G) = arg min f∈G t s f , π last (G) = arg max f∈G t e f .(19) After a policy returnsF ⋆ , the answer emitter performs emit(F ⋆ ) = uniqueo f |f∈F ⋆ .(20) Hereuniqueremoves duplicate normalized answer spans. If the resulting set is empty, or if it contains multiple distinct normalized answer spans when a single answer is required, theS 4 guard fails and the configuration enters S ⊥ . This rule prevents the automaton from using an LLM to resolve an ambiguous temporal selection implicitly. 2 Current state Guard conditionDeterministic updateNext state S 0 Canonicalization succeedsSetF= Γ fact (C)S 1 S 1 Intent validates againstFSetI=V intent (Γ query (Q,F),F)S 2 S 2 Target chain and policy are avail- able SetG=G s q ,r q and bindπS 3 S 3 Temporal selection is nonempty and admissible SetF ⋆ =π(G,α)S 4 S 4 Direct answer emission succeedsSetA= emit(F ⋆ )S 5 Any stateThe corresponding guard failsStore the typed reasonRS ⊥ Table 2: Guarded transition structure of TAS. The accepting path isS 0 →S 1 →S 2 →S 3 →S 4 →S 5 . The fallback state records the first failed guard, which is used for repair or final fallback routing. B Temporal Policy Details This section records the TAS policy variants, defaults, and fallback behavior. B.1 Adapter Intent to Program Policy The semantic adapter emits intents rather than answers. Table 3 maps common question patterns to validated TAS policies. B.2 Policy Variants and Defaults Table 4 summarizes the implemented policy variants. TAS emits an answer only when the selected facts imply an admissible answer set. Interval policy.For anintervalintent, TAS first attempts exact boundary matching. If no exact match exists andpolicy=overlap, it selects the facts with the largest positive overlap. If max-overlap is disabled, all overlapping facts are returned to the emission guard. Before policy.For an entity-valued before intent, TAS resolves the anchor factf a inG s q ,r q and selects candi- dates witht e f ≤t s a . The defaultimmediatepolicy returns the nearest predecessor, whileallpreviousreturns all earlier facts. Time-valued anchors use the same predeces- sor rule against the normalized anchor time. After policy.For an entity-valued after intent, TAS resolvesf a and selects candidates witht s f ≥t e a . The defaultimmediatepolicy returns the nearest successor, whileallfollowingreturns all later facts. Time-valued anchors compare candidate start times against the nor- malized anchor time. Point policy.For apointintent, TAS selects all facts satisfyingt s f ≤q t ≤t e f . Multiple distinct answers trigger pointqueryambiguousby default. Thealloverlaps option returns all containing facts, andlateststart keeps only the containing facts with the latest start time. First and last policies.firstandlastdo not use command-line policy variants. They select the earliest or latest facts inG s q ,r q and pass ties to the answer-emission guard. B.3Time Normalization and Closed- Interval Comparison The selector compares normalized time keys, denoted by κ T : raw time string→T,(21) whereκ T is the sortable key produced by date normal- ization. A bare year is mapped to a representative mid- point, month-year forms are mapped to the corresponding month, and fully specified dates are mapped to concrete dates. Unparseable fact boundaries are filtered or routed toS ⊥ . All temporal comparisons are closed at the endpoints. Overlap spans are computed after both fact and query intervals are converted byκ T ; only facts with positive overlap are candidates for max-overlap selection. B.4Pre-Transform and Fallback Trans- form Corrections The pre-transform mode ishardonlyby default. It in- vokes the LLM only for hard temporal structures and constrains the output to a normalized intent, not an answer. Its corrections are therefore policy-interface cor- rections: •Forbetween ... and ...questions with month- year granularity, the transform normalizes the intent tointervalwithpolicy=overlap. •For point queries with multiple containing states, the transform may setlateststart, causing TAS to keep the most recent containing state before answer emission. •For date anchors such asbefore Jan 1996or after Jan 1996, the transform labels the anchor asanchortype=time. 3 Question patternAdapter intentProgram policy between t1 and t2 interval(start=t1,end=t2,policy=overlap)Max-overlap interval selec- tion from t1 to t2 interval(start=t1,end=t2,policy=exact)Exact boundary matching in/at t point(time=t)Point containment after t after anchor(anchor=t,anchortype=time)Time-anchor successor before t beforeanchor(anchor=t,anchortype=time)Time-anchor predecessor before entity beforeanchor(anchor=entity, anchor type=entity) Entity-anchor predecessor after entity afteranchor(anchor=entity, anchortype=entity) Entity-anchor successor first first()Earliest state in chain last last()Latest state in chain Table 3: Mapping from semantic adapter intents to deterministic TAS policies. The adapter chooses the policy descriptor, but TAS still performs final evidence selection over canonical facts. Intent typeDefault policyOptional behaviorFailure or tie behavior interval exact, withoverlapwhen re- quested Max-overlap may be disabled to return all overlapping facts Failure occurs if exact matching fails andpolicy=overlapis absent, or if no positive overlap exists. before anchor immediatepredecessorallpreviousreturns all earlier facts Missing anchor, empty predecessor set, or ambiguous answer may trigger fallback. afteranchor immediatesuccessorallfollowingreturns all later facts Missing anchor, empty successor set, or ambiguous answer may trigger fall- back. point singleorambiguous alloverlaps;lateststart may be injected by transform Multiple distinct answers trigger pointqueryambiguousby default. first/lastEarliest/latest start stateNo command-line policy switchTied earliest/latest facts are returned together and then checked by emis- sion. Table 4: Default and optional temporal policy behavior in TAS. The fallback transform follows the same boundary: it may rewrite a failed or ambiguous temporal structure into a validated intent and policy descriptor, but it returns the instance to TAS whenever possible. Only when validated repair and reselection fail does STAIR enter the final LLM fallback path. Overall, the policy design remains conservative. In- terval queries support exact and overlap policies, be- fore/after queries support immediate and all-previous or all-following variants, and point queries fall back when multiple distinct answers are selected under the default policy. LLM transformation is limited to structure and policy rewriting; final temporal selection remains a TAS operation. C Semantic Adapter Schema The semantic adapter emits an answer-free JSON inter- face for hard temporal questions before TAS validation and execution. C.1 Answer-Free Adapter Contract The adapter prompt enforces the following contract:Do not answer the question. Its JSON output is limited to structure normalization, intent generation, policy selec- tion, and schema-compatible fact repair. The adapter may instantiateI 0 in Appendix A.2, but it does not instantiateF ⋆ orA: I 0 V intent −→I TAS −→F ⋆ emit −→A. Answer-like fields such asanswerandfinalanswer are therefore outside the adapter schema and are rejected by validation. C.2 JSON Schema Table 5 lists the top-level fields accepted from the seman- tic adapter. Theintentobject contains the fields in Table 6. Some fields are type-specific. For example,startandendare required for interval intents, whereastimeis required for point intents. 4 FieldStatusMeaning and validation role transformedquestionOptionalNormalized question text used for traceability or parser repair; it is not an answer and is never copied toA. intentRequiredStructured temporal interface specifying the operator, temporal ar- guments, and policy descriptor. This field is converted intoI 0 and then validated intoI. factsOptionalSchema-compatible canonical facts recovered from the raw context. Every record must satisfy the canonical fact definition before entering F. syntheticfactsOptionalSchema-compatible facts added only to repair a missing but context- supported temporal record. These records are candidates for valida- tion, not selected evidence. reasonOptional Short explanation of the structural transformation, used for auditing. It has no effect on temporal selection or answer emission. Table 5: Top-level answer-free JSON schema emitted by the semantic adapter. Intent fieldAllowed valuesMeaning type interval,point, before anchor,afteranchor, first event,lastevent Temporal operator requested by the adapter. The val- idator mapsfirsteventandlasteventto the TAS intent typesfirstandlast. subject,relationStrings or omitted if inheritedTarget key fields used to populate (s q ,r q ). If omitted, validation must recover them from the rule parser or canonical facts. start,endNormalized or parseable time stringsInterval boundaries used whentype=interval. timeNormalized or parseable time stringPoint-time argument used whentype=point. anchorEntity string or time stringAnchor argument forbeforeanchorandafteranchor. anchortype entityortimeType annotation for the anchor. Date anchors such as Jan 1996are marked astime. policy exact,overlap,latest start, all overlaps,immediate, allprevious,allfollowing Execution policy descriptor. The descriptor selects a TAS policy variant; it does not select a fact or answer. Table 6: Intent-level fields in the semantic adapter schema. C.3 Example ForWho was the employer after X?, the adapter emits the structure in Figure 1 rather than the employer name. C.4 Validation Before TAS The adapter output is accepted only after programmatic validation. The validator performs four checks before TAS consumes the result: •Schema check.The output must be parseable JSON with a supportedintent.type. Unsupported operators, malformed objects, and answer-like fields are rejected. •Temporal argument check.Required fields such asstart/end,time, oranchormust be present for the corresponding intent type, and time strings must be parseable byκ T when they are time-valued. Input question: Who was the employer after X? Adapter output: "transformed_question": "Who was the employer immediately after X?",,→ "intent": "type": "after_anchor", "relation": "employer", "anchor": "X", "anchor_type": "entity", "policy": "immediate" , "facts": [], "synthetic_facts": [], "reason": "The question asks for the successor state after an entity anchor.",→ Figure 1: Example of an answer-free semantic adapter output. No final employer name appears in the JSON; TAS performs the successor selection. • Target-key check.The validator must populate (s q ,r q ) from the adapter output, the rule parser, or 5 the canonical fact set. If the local chainG s q ,r q cannot be formed, execution entersS ⊥ . •Fact admission check.Any records infactsor syntheticfactsmust satisfy the canonical fact schema and temporal ordering constraints before merging intoF. After validation,Iis passed to TAS in the same form as an intent produced by the deterministic query parser. This boundary permits schema repair while keeping tem- poral comparison, evidence selection, and answer aggre- gation symbolic. DCronQuestions Conversion and Evaluation Protocol This section specifies the CronQuestions conversion pro- tocol used for the operator-supported transfer evaluation. D.1 Source Data and Temporal KG We use the official CronQuestions test split and associated temporal KG: data/external/CronQuestions/wikidata_big/questions/test.pickle data/external/CronQuestions/wikidata_big/kg/full.txt data/external/CronQuestions/wikidata_big/kg/wd_id2entity_text.txt data/external/CronQuestions/wikidata_big/kg/wd_id2relation_text. ⌋ txt,→ The test split contains 30,000 questions. The KG contains 328,635 temporal facts, 125,726 entity labels, and 203 relation labels. Each row has the format subject relation object start end Q25559009 P39 Q41582555 1847 1852 and is interpreted as a temporal fact (s,r,o,[t s ,t e ]). The converter maps QIDs and PIDs to natural-language labels before writing the TAS context. D.2Supported Question Types and Fil- tering Undertaskmode=currenttas, an item is type- supported only ifanswertype=entityand its type is simpleentity,firstlast, orbeforeafter. Table 7 reports the resulting deterministic filtering stages. Table 8 gives the distribution of the 2,596 type- supported but unconverted questions assigned to missinghead. missingheadmeans that the item passes the type filter but lacksannotation.head. The converter retrieves forward timelines by (head,relation), so the temporal context cannot be constructed. Many such cases provide tailbut nothead: Who was the first to hold the position of Member of the Landtag of Lower Saxony? "tail": "...", "adj": "first" Thus,missingheadis not a post-hoc filter for answer availability, missing KG facts, missing labels, or manual data cleaning. D.3 Temporal Context Construction For a standard forward item, the converter readshead andrelationand retrieves all KG facts of the form (head,relation,o,t s ,t e ). It then maps IDs to labels, sorts facts by time, and writes a colon-style timeline. A point-time fact uses identical boundaries, such as2009 - 2009. Facts with the same start, end, subject, and relation are grouped into one context line. The materializedtemporalcontextfield is a string. The formatter emits lines of the form start - end : subject's relation is ( object ) and joins multiple lines with.. For example: 2009 - 2009 : Richard Trythall's award received is ( Prix Italia ),→ The question text usesparaphrases[0]when available. D.4 Inverse and Before/After Queries Somebeforeafterquestions require a holder time- line. When the relation isP39and the wording con- tains holder-style cues, the converter retrieves facts by (relation,tail) and reverses them: 1990 - 1997 : Prime Minister of New Zealand's holder is ( Jim Bolger ).,→ 1997 - 1999 : Prime Minister of New Zealand's holder is ( Jenny Shipley ),→ This construction makes adjacency-based before/after policies return the person entity. A general inverse con- struction for allsimpleentityandfirstlastques- tions is not implemented, which leaves some inverse items in themissingheadbucket. D.5 Answer Construction Gold answers are mapped from Wikidata IDs through wdid2entitytext.txtand written as labels inanswer. The converter does not require the gold answer to ap- pear in the constructed context; ifexpectedanswerids cannot be inferred from matched facts, it falls back to the original answer IDs and still writes the sample. The converted subset is therefore not selected by checking whether the answer is present in the context. D.6 Shared Evaluation Protocol NeSTR and STAIR read the same materialized JSONL file, including identical question IDs, questions, answers, contexts, sample order, prompts, and metrics. This pre- vents method-specific conversion or sample-ordering dif- ferences from affecting the comparison. The retained and 6 StageRetained Excluded Reason Official test split30,000– Original CronQuestions test set Type-supported subset15,69214,308Unsupported task type or unsupported answer type Successfully converted subset13,0962,596Type-supported item lacks the annota- tion field needed to retrieve a timeline Table 7: Deterministic filtering stages for the CronQuestions test split. Question typemissinghead simpleentity1,530 firstlast666 beforeafter400 Total2,596 Table 8: Breakdown of type-supported questions skipped becauseannotation.headis absent. skipped counts in Table 7 are produced by the converter without manual annotation. D.7 Scope and Limitations This is not a complete CronQuestions test-set evalua- tion. It covers the operator-supported and successfully converted subset; time-answer questions,timejoin, un- supported operators, and general inverse timelines remain outside the current scope. EAdditional Main Result Details Table 9 reorganizes the reproduced baseline results and the three-run STAIR statistics into the same layout as the main result table. TheNeSTR (repr.)rows are re- produced under our pipeline as a consistency check. FFull Experimental Configura- tion Table 10 records the shared runtime configuration used by the experiments. F.1 Shared Settings G Prompt Templates The prompt-only baseline uses one NeSTR-style prompt. The prompt requires the model to represent the temporal context symbolically, perform inference, check consis- tency, optionally reflect, and output the final result in an <answer>tag. The same template is used across datasets afterquestionandtemporalcontextare filled. STAIR uses the same prompt family only as the fi- nal fallback, after TAS and validated semantic repair fail to produce a deterministic answer. The fallback is decomposed into four calls: symbolic representation, con- strained inference, consistency checking, and final re- flection/answer emission. This decomposition keeps the fallback comparable to the baseline and makes each in- termediate step auditable. All prompts require entity strings to preserve the spelling, accents, and encoding used in the temporal context. Adapter prompts remain answer-free; only the residual fallback prompt may emit a final answer. H Case Study This section provides compact examples of TAS behavior. The purpose is to show the trace format and one boundary limitation, rather than to add new quantitative analysis. H.1 Regular Interval Match Question: Which team did Attaphol Buspakom play for from 1985 to 1989? Temporal context: 1985 - 1989 : Attaphol Buspakom's team are ( Port F.C ), ( Thailand national football team ).,→ 1989 - 1991 : Attaphol Buspakom's team are ( Pahang FA ), ( Thailand national football team ).,→ Gold answer: Port F.C Thailand national football team This case is handled by the rule-based path. TAS canon- icalizes the first line into two facts, parses the question as an exact interval intent over 1985–1989, and selects the facts with matching boundaries. The answer is emitted directly without LLM temporal inference. intent: interval(start=1985, end=1989, policy=exact) selected: - team(Attaphol Buspakom, Port F.C, 1985, 1989) - team(Attaphol Buspakom, Thailand national football team, 1985, 1989),→ answer: Port F.C Thailand national football team H.2 Hard Non-Exact Interval Match Question: Which political party did Louise Mensch belong to between Jul 1997 and Sep 1997?,→ Temporal context: 1995 - 1996 : Louise Mensch's party is ( Conservative Party ). 1996 - 1997 : Louise Mensch's party is ( Labour Party ). 1997 - 1998 : Louise Mensch's party is ( Conservatives ). 7 Table 9: Additional TQA results in the main-table layout. Values with subscripts are mean ±std over three independent runs; † , TISER, and unmarked NeSTR rows were not rerun in this setting. ModelStrategy TimeQA-EasyTimeQA-HardTempReason-L2TempReason-L3Avg EMF1EMF1EMF1EMF1EMF1 Prior Methods BigBird † Vanilla51.271.659.568.132.750.928.846.843.159.4 FiD † Vanilla60.567.946.854.6– T5-large † Vanilla63.171.659.568.132.750.928.846.846.059.3 Temp-T5 † Vanilla–31.849.626.143.0– REMEMO-large † Vanilla63.772.360.569.337.454.933.449.348.861.5 TG-LLM † CoT66.469.163.166.442.452.235.646.951.958.7 ReAct † Few-shot45.055.128.334.439.345.642.750.838.846.5 QAaP † Few-shot46.354.441.755.343.750.145.348.344.352.0 Event-ALT † –63.073.861.770.455.362.858.059.559.566.6 Open-source LLMs Qwen2.5-7B TISER86.8092.6064.3071.5061.1069.8072.6077.6071.2077.90 NeSTR85.1090.2064.8071.2061.5068.6073.1076.7071.1076.70 NeSTR (repr.) 86.73 ±0.30 92.17 ±0.24 66.23 ±1.16 72.76 ±0.97 59.36 ±0.56 66.08 ±0.59 69.63 ±0.11 73.80 ±0.19 70.49 ±0.34 76.20 ±0.27 STAIR93.27 ±0.13 94.25 ±0.08 76.94 ±0.27 80.91 ±0.14 87.61 ±0.06 87.72 ±0.09 94.90 ±0.03 94.77 ±0.02 88.18 ±0.05 89.41 ±0.02 Qwen3-8B TISER88.8093.4077.1082.5073.7078.4084.3087.5080.9085.40 NeSTR89.5094.2077.7083.4079.2083.5084.9087.2082.8087.10 NeSTR (repr.) 90.67 ±0.17 94.51 ±0.07 78.58 ±0.54 83.65 ±0.40 79.32 ±0.08 81.38 ±0.05 85.43 ±0.31 85.71 ±0.20 83.50 ±0.19 86.31 ±0.14 STAIR94.79 ±0.03 95.52 ±0.02 84.32 ±0.02 86.37 ±0.05 88.75 ±0.08 91.31 ±0.04 96.18 ±0.02 95.52 ±0.01 91.01 ±0.01 92.18 ±0.02 Qwen3-14B TISER90.0094.3082.1087.2075.5080.6081.6085.2082.3086.80 NeSTR91.1094.5082.2087.3079.5084.6085.1088.9084.5088.80 NeSTR (repr.) 92.56 ±0.15 94.88 ±0.14 83.27 ±0.17 86.41 ±0.13 78.70 ±0.23 80.71 ±0.27 86.76 ±0.06 86.56 ±0.15 85.32 ±0.03 87.14 ±0.03 STAIR95.01 ±0.02 96.07 ±0.06 84.18 ±0.14 86.59 ±0.12 86.69 ±0.03 90.87 ±0.02 95.96 ±0.02 95.44 ±0.01 90.46 ±0.03 92.25 ±0.03 Closed-source LLMs GPT-4o-mini TISER86.7091.9074.3079.9077.7084.1082.3087.1080.2085.80 NeSTR93.7096.4081.7085.9080.8086.4084.6090.0085.2089.70 NeSTR (repr.) 94.99 ±0.15 96.84 ±0.06 82.41 ±0.15 85.46 ±0.17 81.77 ±0.16 85.76 ±0.12 89.93 ±0.26 90.26 ±0.21 87.28 ±0.05 89.58 ±0.04 STAIR96.57 ±0.05 97.02 ±0.02 83.97 ±0.23 86.24 ±0.24 89.13 ±0.03 91.10 ±0.02 96.16 ±0.00 95.54 ±0.01 91.46 ±0.07 92.48 ±0.06 Table 10: Shared experimental configuration. ItemValue Workflowtas. BenchmarksTimeQA-Easy, TimeQA-Hard, TempReason-L2, TempReason-L3, and the converted CronQuestions subset in Appendix D. Sample sizesample=0, i.e., full selected splits. Decoding Temperature 0.1; maximum output length 1024 tokens. The script argument ismaxnewtokens; API and Ollama backends bind it to their output-length controls. Main runs and seeds Main STAIR and CronQuestions model configurations use three independent runs with seeds 42, 43, and 44. Single-run ablations use seed 42 unless stated otherwise. Framework and hardwarePyTorch 2.11.0 with CUDA 13.0 on Ubuntu 24.04; NVIDIA GeForce RTX 4090 GPU. Gold answer: Conservatives In this case, the query is month-level whereas the con- text is year-level. The semantic adapter does not answer the question; it only rewrites the temporal interface: "intent": "type": "interval", "start": "1997-07-01", "end": "1997-09-30", "policy": "overlap" TAS then applies overlap selection over canonical facts and retains the interval that covers the queried months. 1995-07-01 to 1996-06-30 -> no_match 1996-07-01 to 1997-06-30 -> no_match 1997-07-01 to 1998-06-30 -> overlap answer: Conservatives H.3 Year-Boundary Ambiguity The following failure case illustrates a current boundary limitation. Question: Oliver Bulleid was an employee for whom in Feb 1908? Temporal context: 1901 - 1908 : Oliver Bulleid's employer is ( Great Northern Railway ).,→ 1908 - 1912 : Oliver Bulleid's employer is ( Westinghouse Electric Corporation ).,→ 1912 - Dec 1922 : Oliver Bulleid's employer is ( GNR ). 1923 - 1937 : Oliver Bulleid's employer is ( new London and North Eastern Railway ).,→ ... Gold answer: Westinghouse Electric Corporation TAS output: Great Northern Railway EM = 0, F1 = 0.0 The TAS trace shows that this failure is not an arbi- trary LLM generation error. It is a deterministic boundary decision made by the selector. 8 intent: point(time = Feb 1908, policy = single_or_ambiguous) temporal_tests: - employer(Oliver Bulleid, Great Northern Railway, 1901, 1908) -> contains_point,→ - employer(Oliver Bulleid, Westinghouse Electric Corporation, 1908, 1912) -> no_match,→ candidate_answer: Great Northern Railway The failure is caused by a granularity mismatch. The facts provide year-level boundaries, whereas the question asks for a month-level point. TAS mapsFeb 1908to the preceding interval and excludes the adjacent1908--1912 interval, although the gold annotation treats the latter as valid. The same failure was observed across multiple backbones: results/api/gpt-4o-mini-2024-07-18_timeqa_hard_tas_pre_transform ⌋ _full_20260609_053712.json,→ results/api/qwen2.5_7b-instruct_timeqa_hard_tas_pre_transform_fu ⌋ l_20260613_052455.json,→ results/api/qwen3_14b_timeqa_hard_tas_pre_transform_full_2026061 ⌋ 3_155159.json,→ This case suggests a concrete extension: a boundary- uncertain state could retain both adjacent candidates when a month-level query falls inside a shared year bound- ary, and the ambiguity could then be resolved by a finer- grained rule or a validated semantic normalizer. I Executor Ablation Details Table 11 reports ablations of the STAIR-Core executor. This configuration omits the hard detector and seman- tic adapter, thereby isolating design choices within the deterministic executor. Table 11: Executor ablation of STAIR-Core, which ex- cludes the hard detector and semantic adapter. BackboneVariantEMF1∆F1 GPT-4o-mini STAIR-Core executor91.1592.10– without deterministic answer83.9787.63-4.47 without before/after adjacency72.9082.19-9.92 symbolic facts only86.9689.83-2.28 rule facts only90.4891.96-0.15 Qwen2.5-7B STAIR-Core executor85.7987.20– without deterministic answer70.5774.66-12.54 without before/after adjacency68.1377.58-9.62 symbolic facts only74.5778.73-8.47 rule facts only86.3687.61+0.41 Removing deterministic answer construction decreases F1 by 4.47 points for GPT-4o-mini and 12.54 points for Qwen2.5-7B. Removing before/after adjacency also causes a substantial decline, whereas changing the fact source has a smaller effect. These results indicate that the main gains of the executor come from deterministic selection and direct answer emission. This ablation is intended as an executor-local diag- nostic rather than a second main result table. Without direct answer construction, the system may still identify temporal evidence but must rely on a generative step to verbalize the answer, which introduces avoidable ag- gregation errors. Without before/after adjacency, entity- anchored questions lose the ordered-chain constraint that distinguishes a predecessor or successor from any tem- porally compatible fact. The fact-source variants further show that, once an admissible local chain is available, the decisive factor is the deterministic policy applied to that chain. 9