Paper deep dive
FinanceHarness: Autonomous Financial Deep Research Framework
Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying CuiZhu, Vishy Tirumalashetty, Wei Wang, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/1/2026, 2:26:33 AM
Summary
The paper introduces FinanceHarness, an autonomous framework for financial deep research, and FinanceGym, a point-in-time benchmark. FinanceHarness automates the research loop using specialized tools and workflows, while FinanceGym provides expert-validated, thesis-driven research questions with pre-cutoff and post-cutoff rubrics to evaluate agents' ability to retrieve historical evidence and forecast future events without information leakage.
Entities (6)
Relation Signals (4)
FinanceHarness â developedby â Google Cloud AI Research
confidence 95% ¡ Authors are affiliated with Google Cloud AI Research
FinanceHarness â uses â FinanceGym
confidence 95% ¡ FinanceHarness runs agents against the PIT sandbox and scores them with the FinanceGym rubrics
FinanceGym â contains â Point-in-Time (PIT) Search Sandbox
confidence 92% ¡ Our benchmark infrastructure consists of a point-in-time (PIT) financial search sandbox and FinanceGym
FinanceGym â evaluates â LLM
confidence 90% ¡ Even leading LLMs and agents score below 40% on the rubrics
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2607.27853v1
- Canonical: https://arxiv.org/abs/2607.27853v1
Trouble viewing inline? Open PDF directly â
Full Text
144,811 characters extracted from source content.
Expand or collapse full text
2026-07-24 FinanceHarness: Autonomous Financial Deep Research Framework Yijia Xiao *1,2 , Rujun Han 1 , Yanfei Chen 1 , Zifeng Wang 1 , Ke Jiang 1 , Zhongying CuiZhu 1 , Vishy Tirumalashetty 1 , Wei Wang 2 , Burak Gokturk 1 , Tomas Pfister 1 and Chen-Yu Lee 1 1 Google Cloud AI Research, 2 University of California, Los Angeles Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. Even leading LLMs and agents score below 40% on the rubrics, showing that FinanceGym is challenging and leaves substantial headroom. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%. FinanceHarness is available at https://github.com/Yijia-Xiao/FinanceHarness. 1. Introduction Deep research is now among the most widely adopted agentic products in the industry (Google, 2026; OpenAI, 2025; Perplexity, 2025; Team et al., 2025b). Equipped with leading LLMs and a carefully designed agent harness, deep research conducts comprehensive searches that can answer challenging multi-hop reasoning questions and synthesize insightful long-form reports from multiple sources and topics (Han et al., 2025; Team et al., 2025a, 2026, 2025b). Despite the impressive progress, a general deep research harness falls short for highly specialized domains (Wu et al., 2023; Xie et al., 2023, 2024). In this work, we target financial deep research, which requires capabilities such as evidence gathering, cross-source validation, and report synthesis that go beyond a generic deep research skillset. It further requires the system to grasp the relationships among financial entities and events, and anticipate their evolution over time (Ding et al., 2015; Xu and Cohen, 2018). For example, an equity analyst investigating a semiconductor company needs to examine historical earnings reports, analyze current competitor dynamics, and predict future demand and macroeconomic conditions to derive a reasonable rating for the stock. A general deep research system without financial tools and expertise to analyze earnings reports and economic events may fail to identify specific information for such complex analysis. Moreover, despite the abundance of recent deep research benchmarks (Coelho et al., 2025; Du et al., 2025; Li et al., 2026a; Sharma et al., 2025; Wang et al., 2025), none of them provide scalable evaluation data for financial research. Existing deep research benchmarks focus mostly on testing past or current conditions. Professional financial research reports typically need to reason about future events (post-cutoff reasoning), which requires point-in-time (PIT) evaluation to avoid leakage of future information. On the other hand, existing financial NLP benchmarks target isolated capabilities (sentiment, entity recognition, filing QA) (Islam et al., 2023; Shah et al., 2022; Yang et al., 2023), * This work was done while Yijia Xiao was a Student Researcher at Google Cloud AI Research. Corresponding author(s): yijia.xiao@cs.ucla.edu, rujunh@google.com, weiwang@cs.ucla.edu, chenyulee@google.com arXiv:2607.27853v1 [cs.CL] 30 Jul 2026 FinanceHarness: Autonomous Financial Deep Research Framework which are insufficient to measure the quality of comprehensive, temporally constrained financial reports. We address these gaps with a financial deep research suite built on a PIT search sandbox. We construct a large-scale web corpus with real publication dates from the public web, and build a dense retrieval system with cutoff-date access control. This ensures a financial deep research agent can locate finance-related knowledge effectively, while information published beyond each questionâs cutoff date is excluded from the research reports. Leveraging this sandbox, we construct FinanceGym, a large-scale, expert-validated financial deep research benchmark. We build an entity graph from the corpus and sample financial situations from this graph to ground each research question and its rubric. Drawing on input from financial practitioners, we generate an investment thesis and paired pre-cutoff and post-cutoff rubrics for every question. A multi-stage filtering pipeline then removes low-quality records, and expert annotators validate the rest to form the final benchmark. Building on this new environment, we propose FinanceHarness, an expert-knowledge-guided agent harness for financial deep research (Section 4). Following the common definition in the community, FinanceHarness implements a harness as layered services of LLM-facing tools, APIs, execution workflows, and evaluation rubrics (Lee et al., 2026; Ning et al., 2026). It runs agents against the PIT sandbox and scores them with the FinanceGym rubrics, so that evaluation and in-environment optimization share one contract. We summarize our contributions. (1) We build a point-in-time financial search sandbox to serve as the foundation for FinanceGym: an expert-validated financial deep research benchmark of 400 high-quality research questions and rubrics that separate pre-cutoff evidence retrieval from post-cutoff reasoning. (2) We propose FinanceHarness, an expert-knowledge-guided harness that automates environment and data construction, agent execution, and reward modeling for financial deep research under a strict point-in-time contract. (3) We benchmark a wide range of leading LLMs and agents and show that FinanceGym is challenging, with every system scoring below 40%, making it a useful target for advancing financial deep research. 2. Related Work Financial benchmarks. Financial NLP benchmarks have traditionally targeted isolated skills such as sentiment analysis, named entity recognition (Shah et al., 2022), factual question answering over SEC filings (Islam et al., 2023), or broad evaluation of finance-specialized LLMs (Yang et al., 2023). These tasks are valuable, but they do not capture analyst-style financial deep research: long-form reports that connect entities, sectors, events, and time. To bridge this gap, FinanceGym builds questions from a finance entity graph and evaluates reports with a two-tier rubric that separates pre-cutoff evidence retrieval from post-cutoff outcome anticipation. Research agents. Modern research agents build on retrieval-augmented generation (Lewis et al., 2020) and ReAct-style tool use (Yao et al., 2022), but differ in where the research policy lives. Scaffolded agents encode orchestration in code, including Test-Time Diffusion for Deep Research (TTD- DR) (Han et al., 2025), GPT-Researcher, 1 , STORM (Shao et al., 2024), and subagent-based frameworks. Specialized deep-research models instead train the backbone on agentic-search trajectories with a fixed tool stack, as in Tongyi-DeepResearch-30B-A3B (Tongyi-DR) (Team et al., 2025b) and MiroThinker- 1.7-mini (Team et al., 2025a, 2026). Our evaluation includes trained models, a fixed ReAct search 1 https://github.com/assafelovic/gpt-researcher 2 FinanceHarness: Autonomous Financial Deep Research Framework Table 1|Comparison with representative deep-research and search-agent benchmarks.â/â/â denote full, partial, and no support. Reproducible means evaluation uses a fixed corpus rather than live-web access. PIT means retrieval is constrained by a per-question publication-date cutoff. Verifiable means evaluation rests on externally checkable signals such as gold answers, binary rubric items, citation matches, or post-cutoff-date outcomes. BenchmarkDomain Long-form Reproducible PITRubricVerifiable FinanceBench (Islam et al., 2023)FinanceâââGold answerâ GAIA (Mialon et al., 2024)GeneralâââGold answerâ BrowseComp (Wei et al., 2025)GeneralâGold answerâ BrowseComp-Plus (Chen et al., 2025)GeneralâââGold answer + citationâ FRAMES (Krishna et al., 2025)GeneralâââGold answerâ DeepResearchGym (Coelho et al., 2025)GeneralââCitation + LLM-judgeâ DeepResearch Bench (Du et al., 2025) Multi-domainââRACE + FACTâ DeepResearch Bench I (Li et al., 2026a) Multi-domainââ9.4k binary rubricsâ ResearchRubrics (Sharma et al., 2025) Multi-domainââ2.6k expert rubricsâ LiveResearchBench (Wang et al., 2025) Multi-domainââDeepEval checklistâ MiroEval (Ye et al., 2026)Multi-domainââAdaptive + process evalâ DeepScholar-Bench (Patel et al., 2025)AcademicâââAutomated, 3 dimensionsâ FinanceGym (Ours)FinanceâPre-cutoff/post-cutoff rubricsâ wrapper with multiple backbones, and agentic search systems on the same PIT corpus retriever, which lets us separate backbone capability, scaffold contribution, and tool-distribution shift. Within finance, LLM agents have largely targeted trading decisions, through multi-agent frameworks (Xiao et al., 2024) and reasoning models trained with reinforcement learning (Xiao et al., 2025), whereas we target analyst-style research and report generation evaluated with verifiable rubrics. Deep-research evaluation. Recent deep-research benchmarks emphasize long-form, citation- grounded answers and richer rubrics (Coelho et al., 2025; Du et al., 2025; Li et al., 2026a; Sharma et al., 2025; Wang et al., 2025), but most rely on live-web retrieval or a single global snapshot rather than a per-question publication-date cutoff. LLM-as-judge methods offer scalable evaluation (Zheng et al., 2023), yet finance requires criterion-level attribution: a report may retrieve the right historical facts while failing to anticipate the later outcome, or reason plausibly while omitting source-grounded evidence. We therefore use evidence-aware rubric judging with explicit pre-cutoff and post-cutoff rubrics. Table 1 highlights the gap FinanceGym targets. FinanceBench (Islam et al., 2023) is the closest finance-domain benchmark, but evaluates short-form QA over filings rather than analyst- style research. BrowseComp-Plus (Chen et al., 2025) and DeepResearchGym (Coelho et al., 2025) improve reproducibility through fixed corpora, but do not apply per-question publication-date cutoffs. DeepScholar-Bench (Patel et al., 2025) is closest on the PIT axis, but tied to the evolving arXiv index. DeepResearch Bench I (Li et al., 2026a) and ResearchRubrics (Sharma et al., 2025) provide rich rubrics, but do not isolate pre-cutoff-date evidence retrieval from post-cutoff-date outcome anticipation, which is the central difficulty in financial deep research. 3. FinanceGym Construction As mentioned in Section 1, our benchmark infrastructure consists of a point-in-time (PIT) financial search sandbox and FinanceGym, the benchmark constructed on top of it (Figure 1). In this section, we first describe the corpus, retrieval, and temporal access contract that define the search sandbox (§3.1). We then describe how FinanceGym is generated from a finance entity graph, filtered, balanced, and expert-annotated (§3.2). Finally, we define the scoring contract used by the benchmark 3 FinanceHarness: Autonomous Financial Deep Research Framework and by downstream harness training (§3.3). 3.1. Point-in-Time Financial Search Sandbox Source corpus. We built the search sandbox from a large-scale web corpus that we collected from thousands of public web domains. We built the corpus so that all articles carry reliably extracted publication dates: each articleâs date was extracted withhtmldate(Barbaresi, 2020), allowing the environment to enforce PIT access. We extracted clean text withtrafilatura(Barbaresi, 2021) with normalized metadata. The resulting corpus contains 100+ million articles, deliberately preserving the noise and breadth of web-scale financial search rather than reducing the task to a curated filing set. The corpus serves only as the internal testbed for building and evaluating the harness. Users can input any public corpus in our pipeline to produce research questions and rubrics. The released benchmark consists of standalone research questions that are decoupled from the underlying corpus (§3.2). Embedding and storage. Articles are embedded with Qwen3-Embedding-4B (Zhang et al., 2025), producing normalized dense vectors for retrieval. Full texts are stored separately from the vector Figure 1|End-to-end benchmark construction. Web articles flow through a domain filter and FAISS + PIT embedding store into a semantic graph, where motif mining yields the finance entity graph and downstream situations; publication-date cutoff selection then drives unconstrained question generation, followed by curation, balancing, expert annotation, and final evaluation. 4 FinanceHarness: Autonomous Financial Deep Research Framework index so agents can retrieve candidate documents and then fetch source text for citation-grounded synthesis. Search server. We build a FAISS (Douze et al., 2025) IVF-SQ8 index over the normalized embeddings and serve it through a lightweight API. This API defines the environment contract used throughout the paper: agents can search and read historical documents, but they cannot access articles published after the assigned cutoff date. 3.2. FinanceGym: Situation-Driven Benchmark Construction Finance graph construction. We construct a finance entity graph from the corpus by first filtering to finance-relevant sources and then extracting entity-relation triples with Gemini-3.5-Flash (Team et al., 2023) under structured JSON output. Each edge stores a head entity, relation, tail entity, local context, source URL, and publication date. The extractor produces 5.74M raw edges; after filtering generic or malformed nodes, the working graph contains 4.37M edges. These edges span 1.11M unique entities from 1.20M source articles and serve as the basis for benchmark generation. Situation mining. Rather than clustering the graph into broad themes, we mine situations that match how financial analysts discover research questions. The miner has three complementary modes: cross-category linkages between entities and event types, temporal narrative arcs around high-degree entities, and polar divergences such as upgrades versus downgrades or earnings beats versus misses. These modes instantiate the fine-grained situation types reported in the harness ablation (Appendix A): a multihop-path type from linkages, a temporal-narrative type, and several divergence (tension) types. An entity budget limits repeated sampling of the same primary name, preventing mega-cap firms from dominating the benchmark. Publication-date cutoff selection and unconstrained generation. Each candidate situation is paired with a cutoff date. For each candidate dayí, letíŁ(í)be its event volume,í(í)entity diversity, andí(í)relation entropy, and letí§(¡)denote aí§-score across candidate days. Cutoff dates are then selected by a volume, entity-diversity, and relation-entropy objective: score(í)= í§ íŁ(í) 1+ í(í) 1+ í(í) 10 .(1) The volume term favors eventful days; the diversity and entropy terms penalize batch artifacts dominated by a single entity or one relation type. Given a (situation, cutoff-date) pair, an LLM generates the benchmark record without category priming: the prompt does not provide a topic, sector, or reasoning taxonomy. The output contains an analyst-style question, a reference investment thesis, and a two-tier rubric. Pre-cutoff criteria test facts findable before the cutoff date; post-cutoff criteria test outcomes verifiable only after the cutoff date. This split is central to FinanceGym: it separates evidence retrieval from forward-looking synthesis while keeping both sides auditable. Bottom-up taxonomy. After questions are generated, we assign taxonomy labels. LLMs classify batches of generated questions into natural categories, and the raw labels are consolidated into three axes: topic, sector, and reasoning type. This bottom-up taxonomy avoids imposing a prompt-time template on question generation while still allowing the final benchmark to be balanced and reported by interpretable groups. 5 FinanceHarness: Autonomous Financial Deep Research Framework Company Market Macro 1 2 3 4 5 6 7 8 9 (a) Themes 1 Corp. strategy 2 Ownership / gov. 3Event-driven 4 Analyst cons. 5Inst. flows6Valuation / sent. 7 Policy / reg. 8Sector thematic9Asset allocation Health & tech Cyclicals Real assets Fin. & macro 1 2 3 4 5 6 7 8 9 (b) Sectors 1Healthcare2Tech / semis3Consumer disc. 4Industrials5 Energy & comm. 6Real estate 7Financials8Macro & rates9 Crypto Diagnostic Predictive Evaluative 1 2 3 4 5 6 (c) Reasoning 1Causal (30%)2Comparative (10%) 3 Forecasting (22%) 4Risk (14%) 5Effectiveness (20%)6Quantitative (4%) Figure 2|Distribution of the 400-question FinanceGym subset across the three bottom-up taxonomy axes. Topics and sectors are shown as two-ring sunbursts with super-groups; reasoning types appear as a donut. Quality filtering and balancing. Generated questions pass through a sequence of LLM quality gates before selection for expert review. The filters test feasibility, institutional relevance, coherence, groundedness against evidence, and final balance across topic, sector, reasoning type, and cutoff date. This yields a data-determined 2,078-question publication pool from 29,669 unconstrained generations. For benchmarking, an integer linear program (ILP) selects a 500-question subset that maximizes quality subject to balance constraints across the taxonomy axes and monthly cutoff-date buckets. Expert annotation. The balanced 500-question subset receives a final expert-annotation pass through an external professional data vendor, followed by LLM-assisted curation using the annotation labels. Annotators score question feasibility and clarity and mark each rubric item as feasible or infeasible with a short justification. Curation then keeps records with sufficiently clear questions and enough feasible rubrics to grade an agent report: 411 of the 500 annotated questions meet this bar, an 82% professional-annotation pass rate. Each sample requires approximately 1.2 hours of expert work, reflecting the difficulty of validating analyst-style questions and rubric items under a publication-date cutoff. We further balance this pool to the released FinanceGym benchmark of 400 expert-annotated questions with 2,464 annotated rubric items, balanced across 9 topics, 11 sectors (merged into 9 leaves), 6 reasoning types, and 12 monthly cutoff-date buckets (Figures 2, 3, and 4). Release policy. As a final release gate, human experts review every released question along two axes. For quality, they confirm the question is well posed and answerable from public information, with a clear analytical target. For privacy and provenance, they confirm the question only queries information. In other words, it refers to publicly observable market events and asks the agent to research them, rather than embedding, quoting, or leaking content from the underlying corpus, which ensures that each question contains neither personal data nor copyrighted text. Questions that fail any check are removed. After annotations, we do not publicly release the underlying corpus. We release the research questions and the report submission code. The grading runs through our leaderboard with private 6 FinanceHarness: Autonomous Financial Deep Research Framework rubrics to preserve the integrity of our scoring criteria. 0100200300400500 1 2 3 4 5 Expert rating (1â5) Question scores Feasibility ¡ Îź=3.89 Clarity ¡ Îź=3.97 020406080100120140160 0â20 20â40 40â60 60â80 80â90 90â100 Rubric feasibility (%) 9 18 74 130 121 148 82% verified Rubric feasibility 0255075100125150175200 â¤2 3 4 5 6 âĽ7 Items per question 22 29 50 107 109 183 mean 5.7 items Rubric depth Number of questions Figure 3|Expert-annotation quality over the annotated question pool: distribution of annotator feasibility and clarity ratings used as the quality-control gate that curates FinanceGym. 3.3. Evaluation Contract Rubric-based scoring. An LLM judge evaluates each agent report against the question-specific rubric. For each criterion, the judge assigns a score on a 5-tier (0â4) scale with explicit anchors: 0 (not addressed), 1 (mentioned), 2 (partial), 3 (substantive), and 4 (fully grounded). The judge sees the question, thesis, cutoff date, target rubrics, and the agentâs report with cited URLs. Because the rubrics are professionally validated and evidence-linked, the LLM judge is used as a consistent rubric applier rather than as a source of new evaluation criteria. The primary metric is the outcome score: the average over questions of the fraction of rubric points each report earns, Outcome= 1 |í| âď¸ íâí Ă íâR(í) score(í, í) 4|R(í)| ,(2) so each question contributes equally regardless of its number of rubric items. Scores can be dis- aggregated by criterion type, topic, sector, reasoning type, or cutoff-date bucket using the same formula. Point-in-time protocol. Each question has an assigned cutoff date. During evaluation, the FAISS server enforces PIT filtering: every search result must have publication dateâ¤the cutoff date, and agents do not access the live web. Pre-cutoff synthesis rubrics test retrieval and synthesis from pre-cutoff evidence. Post-cutoff reasoning rubrics test whether the agentâs analysis anticipates developments that are only verifiable after the cutoff date. Any detected PIT leakage invalidates the run rather than being treated as an ordinary scoring error. 4. FinanceHarness In this section, we explain FinanceHarness design and its connection to FinanceGym. The point- in-time (PIT) search sandbox defines what information an agent may retrieve before each questionâs cutoff date, and FinanceGym defines how the resulting report is graded. FinanceHarness is the executable harness built around that contract, which enables users to produce highly specialized reports for their financial research queries. It provides the same retrieval tools, orchestration loop, and rubric signal for both evaluation and optimization, so a model trained inside the harness sees the same tool-use distribution at test time. 7 FinanceHarness: Autonomous Financial Deep Research Framework DecJanFebMarApr May JunJul Aug SepOctNov Cutoff month (Dec 2024 â Nov 2025) 0 20 40 60 80 Questions 3 11 10 20 66 52 60 63 57 29 22 7 Point-in-time coverage Healthcare & biotech Technology & semis Consumer & industrials Energy & commodities Financials & real estate Macro & alternatives â12â9â6â30+3+6+9 Months from cutoff (â before ¡ + after) 0 1000 2000 3000 4000 5000 6000 Evidence articles Pre-cutoffPost-cutoff Evidence timing Figure 4|Temporal coverage of FinanceGym: distribution of question cutoff dates across the 12 monthly cutoff-date buckets. 4.1. Expert-guided Harness Design Financial deep research is not only a retrieval problem: an analyst must connect relationships among entities, sector context, numerical evidence, and forward-looking risks into an auditable view. FinanceHarness therefore exposes a layered finance tool interface rather than a single generic web-search action. The core loop covers PIT search, source reading, citation composition, and report finalization. Auxiliary finance tools and workflow skills add structured support for exact data retrieval, calculation, valuation, comparables, risk analysis, and scenario reasoning without expanding the base prompt for every query. Model and serving layer. The serving layer hosts the orchestrator backbone and a lightweight reader used for long-document evidence extraction. It is deliberately opaque to the rest of the harness: the runtime sees only a request/response contract, so the backbone can be swapped across local serving stacks or closed-model SDKs without changing the orchestration logic. Runtime layer. The runtime is the control plane. It owns the bounded agent loop, parsing and dispatch, schema validation, tier loading, prompt-mode selection, recovery, and the registries for tools and workflow skills. It contains no model intelligence; instead, it enforces cross-layer invariants such as schema conformance, tool-result chaining, citation finalization, and run-level budget limits. Tool surface. The tool surface is tiered to keep the initial prompt small while preserving a broad set of financial capabilities. The always-loaded tier contains the core-loop actions. Deferred tools are exposed through a compact catalog, with full schemas loaded only after the model commits to a tool family. The extension tier supports MCP or CLI-backed connectors, including paid data vendors and internal systems. Modes and skills. FinanceHarness uses prompt-variant modes over a consistent tool surface. The research mode emphasizes web-first evidence gathering; the analytical mode emphasizes data and computation tools; the automatic mode lets the agent choose between them. Keeping the registry consistent across modes prevents a trajectory from being stranded by tools that disappear after a mode switch. Skills are reusable workflow specifications that the model can activate on demand. Each skill declares the tools and specialist context it expects, and routing is based on the skill description rather than a hard-coded name. 8 FinanceHarness: Autonomous Financial Deep Research Framework Grounding and robustness. The orchestrator can respond to a step with direct tool calls or by loading a workflow skill that brings in the tools needed for a structured analysis. For research-style reports, the harness composes numbered citations from visited sources, can retrofit citations when a draft omits the citation tool, and runs a light grounding-review pass that asks the backbone to soften or attribute claims it cannot tie to read sources or tool outputs. Recovery policies handle transient model failures, malformed tool calls, context-budget pressure, and truncation without changing the evaluation contract. Operational guardrails. URL pre-fetch validates document IDs before expensive reads. In ablations, disabling this guardrail raises the visit error rate from 2.1% to 39.4%; the end score barely moves because the backbone often self-corrects, but trajectories become substantially more expensive. Search- budget caps show a different pattern: excessive searching correlates with question difficulty rather than causing low scores, so budget caps mainly control cost rather than quality. 4.2. Evaluation and Optimization The harness is tightly coupled to the environment: tool calls hit the PIT search sandboxâs FAISS server directly, and the rubric judge used for benchmark scoring is also available as a reward component. This matters because specialized research models are often trained against one tool stack and evaluated against another. In FinanceHarness, swapping the backbone model changes only the learnable model layer. The search API, document-fetch contract, citation requirements, and scoring rubric stay fixed. Appendices D and E give worked examples: complete cited reports produced by the harness, and the round-by-round tool trajectories behind them. 5. Experimental Setup 5.1. Benchmark Dataset The publication dataset comprises 2,078 questions, the data-determined output of the end-to-end LLM quality filter chain (§3.2). For benchmarking, the upstream pipeline selects a 500-question ILP-balanced subset, which then passes through external expert annotation and LLM-assisted curation (§3.2) to produce FinanceGym: 400 expert-annotated questions with 2,464 annotated rubric items (mean 6.16 items per question), jointly balanced across 9 topics, 11 sectors (merged into 9 leaves), 6 reasoning types, and 12 monthly cutoff-date buckets. Sector coverage spans healthcare/biotech, technology/semiconductors, consumer discretionary, financial services, real estate, energy/natural resources, industrials/transportation, macro/policy, commodities, crypto/digital assets, and fixed income. The expert-annotation curation pass (§3.2) closely preserves the ILP sector balance. 5.2. Agent Baselines Following Tongyi-DRâs framing (Team et al., 2025b), we group baselines by scaffolding complexity rather than by base-model identity. The three groups share our PIT-filtered FAISS retriever and FinanceGym as the headline evaluation set: Fine-tuned open-weight: open-weight models that internalize the research loop without an external scaffold, fine-tuned end-to-end with RL on agentic-search trajectories (Tongyi-DR (Team et al., 2025b), OpenResearcher (Li et al., 2026b), MiroThinker (Team et al., 2025a, 2026)). Foundation model with search tool: a fixed 30-step ReAct wrapper with one search tool, one final-answer tool, and 9 interchangeable backbones, which isolates model capability under a constant 9 FinanceHarness: Autonomous Financial Deep Research Framework scaffold; the result tables split this group into open-weight and proprietary backbones. We also report FinanceHarness, our open-weight layered harness on a Qwen3.6-27B backbone (Section 4), in this group for the open-weight comparison; Appendix B reports a complementary cross-backbone evaluation of FinanceHarness on a component-coverage suite. Agentic search systems: engineered scaffolds that make orchestration decisions in code while using an LLM for planning, retrieval, synthesis, or critique. This group includes TTD-DR (Han et al., 2025), GPT-Researcher, 2 STORM (Shao et al., 2024), OpenClaw, 3 and deepagents. 4 TTD-DR, GPT-Researcher, STORM, OpenClaw, and deepagents all use Gemini-3-Flash as the LLM backbone. Comparing the three groups on identical questions and the same corpus retriever isolates (a) how well training-time tool distributions transfer to a new search environment (Fine-tuned open-weight vs. Agentic search systems); (b) how much hand-engineered scaffolding adds over a minimal ReAct loop on the same backbone (Foundation model with search tool vs. Agentic search systems); and (c) how performance varies with model identity at fixed scaffolding (within the foundation-model rows). The fine-tuned open-weight models were trained against live-web tool stacks (e.g., Serper, Jina); we evaluate them through corpus-backed equivalents (SearchâFAISS+PIT,VisitâSQLite), a substitution that drives the tool-distribution-shift analysis (§C.1). 5.3. Evaluation Protocol Agents are run once on the full 500-question ILP-balanced subset; headline metrics then aggregate scores against FinanceGymâs 400 expert-annotated questions and their 2,464 annotated rubric items (§3.2) by subsetting the per-criterion judge outputs, and the broader 500-question run is retained for auxiliary diagnostics. All baselines run under identical search infrastructure (PIT-filtered FAISS over the same Qwen3-Embedding-4B index). The foundation-model matrix holds the wrapper fixed across 9 frontier and open-weight backbones. Each agent receives only the question text and cutoff date; the rubric, thesis, and supporting evidence are withheld for blind evaluation. The judge (Gemini-3.5-Flash, distinct from the Gemini-3-Flash baseline backbone) scores each rubric criterion on the 5-tier (0â4) scale described in §3.3, with access to the question, thesis, the rubric item being scored, pre-cutoff edge evidence, post-cutoff edge evidence, and the agentâs full report with cited URLs. Headline scores are reported with bootstrap standard errors (í=1000). 6. Results and Analysis We evaluate 17 baselines on FinanceGym under the protocol above. Scores are averaged with the 5-tier rubric (Eq. 2); theÂąSE column reports bootstrap standard errors over questions. The main text reports the decision-relevant comparisons, while per-axis and harness-ablation breakdowns appear in the appendix. Table 2 reports all evaluated systems grouped by system category; the fixed-backbone Finance- Harness ablation is deferred to §6.2. As Table 2 shows, on a 27B open-weight backbone, FinanceHarness ranks among the strongest systems on FinanceGym: it leads all open-weight models, slightly outperforms the best agentic search system, and trails only two proprietary backbones. 2 https://github.com/assafelovic/gpt-researcher 3 https://github.com/openclaw/openclaw 4 https://github.com/hwchase17/deepagents 10 FinanceHarness: Autonomous Financial Deep Research Framework Table 2|Headline results on FinanceGym by system category. Normalized rubric means (%); within each group, bold marks the best andunderlinethe second best per column (Overall/Pre-cutoff/Post- cutoff). ÂąSE is the bootstrap SE of the Overall score. Fine-tuned open-weight SystemOverallâ Pre-cutoffâ Post-cutoffâ ÂąSE Tongyi-DR28.239.710.7 0.8 OpenResearcher27.239.98.10.7 MiroThinker21.229.67.8 0.6 Agentic search systems SystemOverallâ Pre-cutoffâ Post-cutoffâ ÂąSE TTD-DR31.545.59.80.7 GPT-Researcher30.442.212.0 0.8 deepagents28.140.78.6 1.1 OpenClaw27.740.28.3 0.7 STORM27.439.39.4 0.6 Open-weight with search SystemOverallâ Pre-cutoffâ Post-cutoffâ ÂąSE FinanceHarness32.445.711.8 0.8 GLM-530.444.78.50.9 DeepSeek-v3.228.942.48.4 0.8 Qwen3-235B-A22B26.839.77.3 0.7 Gemma-4-26B25.737.67.7 0.6 gpt-oss-120b18.727.85.2 0.8 Proprietary with search SystemOverallâ Pre-cutoffâ Post-cutoffâ ÂąSE Claude-Opus-4.734.150.19.90.8 Gemini-3.1-Pro33.246.812.8 0.7 GPT-5.531.847.58.0 0.7 Gemini-3-Flash30.243.79.8 0.7 6.1. What Drives Performance The main comparison is between model choice and research procedure under the same point-in-time corpus. Changing the backbone within the same search-tool wrapper produces a wider spread than changing the scaffold around Gemini-3-Flash, where the engineered agentic systems cluster near a minimal ReAct loop and several fall below it. This suggests that current financial deep-research performance is still strongly constrained by the underlying model, while task-matched scaffolding can recover additional evidence and organize it more reliably. The fine-tuned open-weight models show a related transfer issue: trained around live-web tool stacks, they are evaluated here through our corpus-backed FAISS/PIT interface, where several general- purpose search backbones outperform them, consistent with the tool-distribution-shift analysis in §C.1. The per-topic breakdown (Table 3) and the per-sector and per-reasoning-type breakdowns (Ap- pendix Tables 11 and 12) follow the aggregate ordering within each system group, with causal and comparative analysis among the hardest reasoning types across systems. Overall score also tracks backbone scale, but only loosely. Plotting each systemâs score against its backbone size (Figure 5), FinanceHarness sits on the upper-left efficiency frontier, matching much larger open-weight backbones and approaching the proprietary frontier despite its 27B backbone. Finally, the pre-cutoff/post-cutoff split is stable across systems: across all baselines, pre-cutoff scores are several times higher than post-cutoff scores. Because the same PIT-filtered corpus is used for every agent, this gap reflects the benchmarkâs core difficulty rather than the retrieval setup. Even the leading agents leave most forward-looking rubric items only partially addressed or missed. 6.2. In-Environment Training The same rubric judge can also serve as a training signal. Using the same retrieval tools and report format at rollout and evaluation, we further train FinanceHarness with Group Relative Policy Optimization (GRPO) on a separate set of 172 machine-curated training instances that we build with the same environment and graph-construction harness used to construct the benchmark, with no 11 FinanceHarness: Autonomous Financial Deep Research Framework Table 3|Per-topic averaged scores (%) on FinanceGym, by system group. Columns follow the Company/Market/Macro topic levels of Figure 2 (Corp. strat., Own./gov., Event|Analyst, Inst. flows, Valu./sent.|Policy, Thematic, Alloc.). Each panel is one system group; shading (blue=higher, amber= lower), bold (best) and underline(second) are normalized per column within the panel. (a) Fine-tuned open-weight SystemCorp. strat. Own./gov. Event Analyst Inst. flows Valu./sent. Policy Thematic Alloc. Tongyi-DR27.632.533.627.520.923.829.226.031.1 OpenResearcher25.927.835.127.319.024.028.825.625.5 MiroThinker22.121.324.117.417.421.619.820.225.9 (b) Agentic search systems SystemCorp. strat. Own./gov. Event Analyst Inst. flows Valu./sent. Policy Thematic Alloc. TTD-DR31.834.537.228.333.228.429.628.126.8 GPT-Researcher29.435.236.330.828.626.928.928.028.0 deepagents25.927.436.423.329.720.430.926.922.8 OpenClaw 26.829.033.825.222.424.828.525.828.5 STORM25.726.234.330.522.623.827.628.826.4 (c) Proprietary with search SystemCorp. strat. Own./gov. Event Analyst Inst. flows Valu./sent. Policy Thematic Alloc. Claude-Opus-4.733.934.640.230.422.728.438.732.935.3 Gemini-3.1-Pro 33.434.538.630.825.328.535.432.332.2 GPT-5.532.333.336.530.022.525.535.930.529.7 Gemini-3-Flash31.830.836.028.322.526.130.528.032.0 (d) Open-weight with search SystemCorp. strat. Own./gov. Event Analyst Inst. flows Valu./sent. Policy Thematic Alloc. FinanceHarness31.734.739.133.425.928.633.629.629.6 GLM-530.233.538.628.719.923.932.728.826.7 DeepSeek-v3.227.230.935.629.719.526.332.526.024.4 Qwen3-235B-A22B 25.627.234.427.318.123.728.325.325.5 Gemma-4-26B 23.826.631.224.420.222.927.224.429.0 gpt-oss-120b18.719.621.024.915.215.619.218.018.6 human involvement and independent of FinanceGym, so that training and evaluation remain fully separate. The reward combines rubric-style evaluation against generated criteria (0.6 weight) with an LLM judge over report coherence, grounding, and trajectory quality (0.4 weight). We then evaluate the resulting policy on FinanceGym. Table 4|Fixed-backbone (Qwen3.6-27B) ablation on FinanceGym: the model with search only (Vanilla), a naive harness (Naive), the full untrained FinanceHarness (Harness, our main con- figuration), and FinanceHarness after GRPO training on a separate, machine-curated training set independent of FinanceGym, with a mixed rubric/report-and-trace reward (RFT). Normalized rubric means (%). ConfigurationTotal Pre-cutoff Post-cutoff Vanilla (model with search) 25.336.18.7 Naive harness29.641.810.7 Harness32.445.711.8 + RFT (GRPO training)32.846.212.1 Table 4 reports four fixed-backbone configurations: the search-only model (Vanilla), a naive 12 FinanceHarness: Autonomous Financial Deep Research Framework Figure 5|Overall outcome score versus backbone scale on FinanceGym. Each marker is one evaluated system, placed by backbone parameter count (x-axis, log scale; proprietary models grouped at right) and overall rubric score (y-axis). On a 27B open-weight backbone, FinanceHarness reaches an overall score competitive with much larger open-weight models and approaching the proprietary frontier (shaded), indicating a favorable performance-to-scale trade-off. harness (Naive), the full untrained harness (Harness, our main configuration), and the harness after reinforcement fine-tuning with GRPO (RFT). GRPO training adds only 0.4 point over the untrained harness, so we treat it as a refinement rather than a headline result and leave a broader trained-policy sweep to future work. 7. Conclusion We presented FinanceGym, a point-in-time financial deep research benchmark built from 400 expert-annotated questions and 2,464 rubric items over a reproducible PIT search sandbox, and FinanceHarness, an expert-knowledge-guided harness that runs and optimizes agents within the same environment. The benchmark pairs the sandbox with pre-cutoff synthesis and post-cutoff reasoning rubrics, allowing evaluation to separate pre-cutoff-date evidence retrieval from post-cutoff- date outcome anticipation. Across 17 baselines and FinanceHarness, every system stays below 40% overall and shows a persistent pre-/post-cutoff gap, suggesting that financial deep research is not addressed by stronger retrieval alone. The fixed-backbone harness ablation further shows progressive improvement from a Qwen3.6-27B search-tool model (25.3%) to the full FinanceHarness (32.4%), with subsequent GRPO training adding a further 0.4 point, supporting the paperâs core claim: evaluation, tool use, and optimization benefit from sharing the same temporally controlled financial research environment. 13 FinanceHarness: Autonomous Financial Deep Research Framework Limitations Corpus scope. The point-in-time search sandbox is built from a 2025 English-language web corpus that we collect ourselves, so coverage of non-English sources, paywalled professional data, and structured filings is limited. Tool-distribution shift. Specialized deep-research models are evaluated out-of-distribution on our corpus-backed retriever rather than the live-web stacks they were trained against; their scores should therefore be read as transfer results under a controlled PIT interface. References A. Barbaresi. htmldate: A python package to extract publication dates from web pages. Journal of Open Source Software, 5(51):2439, 2020. A. Barbaresi. Trafilatura: A web scraping library and command-line tool for text discovery and extraction. In Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing: System demonstrations, pages 122â131, 2021. Z. Chen, X. Ma, S. Zhuang, P. Nie, K. Zou, A. Liu, J. Green, K. Patel, R. Meng, M. Su, et al. Browsecomp- plus: A more fair and transparent evaluation benchmark of deep-research agent. arXiv preprint arXiv:2508.06600, 2025. J. Coelho, J. Ning, J. He, K. Mao, A. Paladugu, P. Setlur, J. Jin, J. Callan, J. MagalhĂŁes, B. Martins, et al. Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research. arXiv preprint arXiv:2505.19253, 2025. X. Ding, Y. Zhang, T. Liu, and J. Duan. Deep learning for event-driven stock prediction. In Ijcai, volume 15, pages 2327â2333, 2015. M. Douze, A. Guzhva, C. Deng, J. Johnson, G. Szilvasy, P.-E. MazarĂŠ, M. Lomeli, L. Hosseini, and H. JĂŠgou. The faiss library. IEEE Transactions on Big Data, 2025. M. Du, B. Xu, C. Zhu, X. Wang, and Z. Mao. Deepresearch bench: A comprehensive benchmark for deep research agents. arXiv preprint arXiv:2506.11763, 2025. Google. Deep research max: a step change for autonomous research agents.https: //blog.google/innovation-and-ai/models-and-research/gemini-models/ next-generation-gemini-deep-research/, 2026. R. Han, Y. Chen, Z. CuiZhu, L. Miculicich, G. Sun, Y. Bi, W. Wen, H. Wan, C. Wen, S. MaĂŽtre, G. Lee, V. Tirumalashetty, E. Xue, Z. Zhang, S. Haykal, B. Gokturk, T. Pfister, and C.-Y. Lee. Deep researcher with test-time diffusion, 2025. URL https://arxiv.org/abs/2507.16075. P. Islam, A. Kannappan, D. Kiela, R. Qian, N. Scherrer, and B. Vidgen. Financebench: A new benchmark for financial question answering. arXiv preprint arXiv:2311.11944, 2023. S. Krishna, K. Krishna, A. Mohananey, S. Schwarcz, A. Stambler, S. Upadhyay, and M. Faruqui. Fact, fetch, and reason: A unified evaluation of retrieval-augmented generation. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4745â4759, 2025. 14 FinanceHarness: Autonomous Financial Deep Research Framework Y. Lee, R. Nair, Q. Zhang, K. Lee, O. Khattab, and C. Finn. Meta-harness: End-to-end optimization of model harnesses. In Preprint, 2026. P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. KĂźttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems, 33:9459â9474, 2020. R. Li, M. Du, B. Xu, C. Zhu, X. Wang, and Z. Mao. Deepresearch bench i: Diagnosing deep research agents via rubrics from expert report. arXiv preprint arXiv:2601.08536, 2026a. Z. Li, D. Jiang, X. Ma, H. Zhang, P. Nie, Y. Zhang, K. Zou, J. Xie, Y. Zhang, and W. Chen. Openre- searcher: A fully open pipeline for long-horizon deep research trajectory synthesis. arXiv preprint arXiv:2603.20278, 2026b. G. Mialon, C. Fourrier, T. Wolf, Y. LeCun, and T. Scialom. Gaia: a benchmark for general ai assistants. In International Conference on Learning Representations, volume 2024, pages 9025â9049, 2024. X. Ning, K. Tieu, D. Fu, T. Wei, Z. Li, Y. Bei, et al. Code as agent harness: Toward executable, verifiable, and stateful agent systems. arXiv preprint arXiv:2605.18747, 2026. OpenAI.Introducing deep research.https://openai.com/index/ introducing-deep-research, 2025. L. Patel, N. Arabzadeh, H. Gupta, A. Sundar, I. Stoica, M. Zaharia, and C. Guestrin. Deepscholar- bench: A live benchmark and automated evaluation for generative research synthesis. arXiv preprint arXiv:2508.20033, 2025. Perplexity. Introducing perplexity deep research.https://w.perplexity.ai/hub/blog/ introducing-perplexity-deep-research, 2025. R. Shah, K. Chawla, D. Eidnani, A. Shah, W. Du, S. Chava, N. Raman, C. Smiley, J. Chen, and D. Yang. When flue meets flang: Benchmarks and large pretrained language model for financial domain. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2322â2335, 2022. Y. Shao, Y. Jiang, T. Kanell, P. Xu, O. Khattab, and M. Lam. Assisting in writing wikipedia-like articles from scratch with large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 6252â6278, 2024. M. Sharma, C. B. C. Zhang, C. Bandi, C. Wang, A. Aich, H. Nghiem, T. Rabbani, Y. Htet, B. Jang, S. Basu, et al. Researchrubrics: A benchmark of prompts and rubrics for evaluating deep research agents. arXiv preprint arXiv:2511.07685, 2025. G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Mil- lican, et al. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805, 2023. M. Team, S. Bai, L. Bing, C. Chen, G. Chen, Y. Chen, Z. Chen, Z. Chen, J. Dai, X. Dong, et al. Mirothinker: Pushing the performance boundaries of open-source research agents via model, context, and interactive scaling. arXiv preprint arXiv:2511.11793, 2025a. M. Team, S. Bai, L. Bing, L. Lei, R. Li, X. Li, X. Lin, E. Min, L. Su, B. Wang, et al. Mirothinker-1.7 & h1: Towards heavy-duty research agents via verification. arXiv preprint arXiv:2603.15726, 2026. 15 FinanceHarness: Autonomous Financial Deep Research Framework T. D. Team, B. Li, B. Zhang, D. Zhang, F. Huang, G. Li, G. Chen, H. Yin, J. Wu, J. Zhou, et al. Tongyi deepresearch technical report. arXiv preprint arXiv:2510.24701, 2025b. J. Wang, Y. Ming, R. Dulepet, Q. Chen, A. Xu, Z. Ke, F. Sala, A. Albarghouthi, C. Xiong, and S. Joty. Liveresearchbench: A live benchmark for user-centric deep research in the wild. arXiv preprint arXiv:2510.14240, 2025. J. Wei, Z. Sun, S. Papay, S. McKinney, J. Han, I. Fulford, H. W. Chung, A. T. Passos, W. Fedus, and A. Glaese. Browsecomp: A simple yet challenging benchmark for browsing agents. arXiv preprint arXiv:2504.12516, 2025. S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann. Bloomberggpt: A large language model for finance. arXiv preprint arXiv:2303.17564, 2023. Y. Xiao, E. Sun, D. Luo, and W. Wang. Tradingagents: Multi-agents llm financial trading framework. arXiv preprint arXiv:2412.20138, 2024. Y. Xiao, E. Sun, T. Chen, F. Wu, D. Luo, and W. Wang. Trading-r1: Financial trading with llm reasoning via reinforcement learning. arXiv preprint arXiv:2509.11420, 2025. Q. Xie, W. Han, X. Zhang, Y. Lai, M. Peng, A. Lopez-Lira, and J. Huang. Pixiu: A large language model, instruction data and evaluation benchmark for finance. arXiv preprint arXiv:2306.05443, 2023. Q. Xie, W. Han, Z. Chen, R. Xiang, X. Zhang, Y. He, M. Xiao, D. Li, Y. Dai, D. Feng, et al. Finben: A holistic financial benchmark for large language models. Advances in Neural Information Processing Systems, 37:95716â95743, 2024. Y. Xu and S. B. Cohen. Stock movement prediction from tweets and historical prices. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1970â1979, 2018. H. Yang, X.-Y. Liu, and C. D. Wang. Fingpt: Open-source financial large language models. arXiv preprint arXiv:2306.06031, 2023. S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao. React: Synergizing reasoning and acting in language models. arXiv preprint arXiv:2210.03629, 2022. F. Ye, Y. Hu, P. Zhu, Y. Li, Z. Jin, Y. Xiao, Y. Wang, L. Wang, Z. Zhang, L. Wang, et al. Miroeval: Bench- marking multimodal deep research agents in process and outcome. arXiv preprint arXiv:2603.28407, 2026. Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, et al. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176, 2025. L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36:46595â46623, 2023. 16 FinanceHarness: Autonomous Financial Deep Research Framework A. FinanceHarness Ablation Breakdowns We expand Table 4 by topic, reasoning type, sector, and situation type. All rows use the same Qwen3.6-27B backbone and the same FinanceGym evaluation protocol. Scores rise across the four configurations (Vanilla, Naive, Harness, RFT). The harness configu- rations account for nearly all of the gain, while subsequent GRPO training (RFT) adds only a small final increment. The hardest cells stay hard: crypto and macro/rates remain the weakest sectors, and causal and comparative analysis the weakest reasoning types. Table 5 | FinanceHarness ablation by topic. Scores are normalized rubric means (%). TopicVanilla Naive Harness RFT Corporate strategy competitive positioning25.3 29.431.7 32.3 Macro regulatory policy25.6 30.533.6 34.1 Event driven restructuring30.9 36.439.1 39.3 Valuation sentiment divergence21.5 25.128.6 29.0 Sector thematic analysis23.3 27.129.6 30.1 Ownership governance stakeholder intelligence28.4 31.234.7 35.0 Institutional flows market microstructure20.2 23.425.9 26.1 Analyst research consensus dynamics25.8 31.733.4 33.7 Asset allocation portfolio strategy21.4 27.129.6 29.8 Table 6 | FinanceHarness ablation by reasoning type. Scores are normalized rubric means (%). Reasoning typeVanilla Naive Harness RFT Causal analysis23.4 27.329.8 30.1 Predictive forecasting25.4 29.132.2 32.8 Effectiveness evaluation25.2 30.932.9 33.2 Risk assessment31.4 35.339.5 39.9 Comparative analysis23.2 26.829.7 30.2 Quantitative analysis23.5 29.231.8 32.3 Table 7 | FinanceHarness ablation by sector. Scores are normalized rubric means (%). SectorVanilla Naive Harness RFT Healthcare biotech32.8 36.139.8 40.1 Technology semiconductors24.7 29.331.8 32.1 Consumer discretionary25.0 29.631.4 31.8 Industrials transportation23.9 30.632.6 32.6 Energy & comm.24.4 27.531.0 31.6 Real estate24.2 29.732.3 32.9 Financial services24.5 30.332.8 33.1 Macro & rates20.3 23.826.9 27.6 Crypto digital assets17.8 21.724.0 25.1 17 FinanceHarness: Autonomous Financial Deep Research Framework Table 8 | FinanceHarness ablation by situation type. Scores are normalized rubric means (%). Situation typeVanilla Naive Harness RFT Multihop path26.4 30.933.7 34.2 Temporal narrative23.8 27.830.7 31.0 Tension analyst disagreement24.0 27.930.9 31.6 Tension price target divergence21.3 25.728.2 28.2 Tension performance divergence20.9 25.427.2 27.7 Tension strategic portfolio shift31.8 35.439.1 39.7 Tension ownership churn32.1 37.640.1 40.6 B. Cross-Backbone Harness Evaluation Beyond the end-to-end FinanceGym rubric, we evaluate FinanceHarness as a deployed system, a model-agnostic harness run unchanged across three backbones: the open-weight Qwen3.6-27B (our default), Gemini-3.5-Flash, and GPT-5.5. A coverage suite of 11 tasks, each designed to exercise one component category (valuation, risk, relative valuation, market data, focused and broad web research, auto-routing, three skills, and planning), is run atí=3 per cell and scored by a dual-LLM judge panel (Gemini-3.5-Flash and GPT-5.5) on faithfulness, grounding, coherence, and completeness (1â5); grounding is credited to either a cited source or the structured tool data the agent fetched or computed. Whereas FinanceGym measures end-to-end report quality under a point-in-time rubric, this evaluation isolates whether the harness invokes and grounds each component correctly across backbones. On the same harness, tools, and judges, the open-weight Qwen3.6-27B backbone scores within about 0.3 of GPT-5.5 on faithfulness and grounding and ahead of Gemini-3.5-Flash (Table 9): with FinanceHarness, a compact open-weight backbone performs comparably to much larger models in a parameter-efficient manner. Coherence and completeness saturate (âĽ4.68 for every backbone), so faithfulness and grounding are the discriminating axes. Behaviorally the harness routes reliably: data routing reaches 95.8% and golden-answer computation 90.9% (rule-scored exact rates), with a tool-coverage rate of 0.99, a 0.94 skill-hit rate, and zero unknown-tool errors across runs. All three backbones also pass the capability checks: a multi-turn follow-up that uses prior-turn context, a /compactsummarization that preserves dollar figures, and correct flagging of an under-specified question for clarification. Table 10 breaks faithfulness and grounding down by component category (across backbones). Risk and market tasks are strongest: value-at-risk, beta, and correlation, and rates and indices, compose cleanly and stay grounded. Relative valuation is weakest because the comparables tool returns a small peer set, so a weaker backbone pads the peer table from prior knowledge and the judges penalize the ungrounded cells. 18 FinanceHarness: Autonomous Financial Deep Research Framework Table 9|FinanceHarness across three backbones on the coverage suite (11 tasksĂ í=3), dual-LLM judged on a 1â5 scale. BackboneFaithfulness Grounding Coherence Completeness GPT-5.54.364.474.924.98 Qwen3.6-27B (open-weight)4.024.204.684.76 Gemini-3.5-Flash3.853.974.785.00 Overall4.084.214.804.91 Table 10|Coverage suite by component category (faithfulness / grounding, 1â5, across backbones). MetricRisk Market Research Auto-routing Skill Valuation Planning Relative Faithfulness 4.534.394.244.224.153.893.673.25 Grounding 4.614.564.354.444.263.894.053.33 C. Additional Analysis C.1. Tool-Distribution Shift The fine-tuned open-weight deep-research models underperform the strongest agentic search systems on FinanceGym despite being trained for agentic search. We interpret this as a tool-distribution-shift result rather than a general model-capability result. Tongyi-DR and MiroThinker were optimized around live-web search and page extraction conventions; in our evaluation, they must operate through a PIT-filtered corpus retriever with fixed document records and no live-web access. That substitution changes both the interaction pattern and the final-report format expected by the rubric. The most visible consequence is attribution. FinanceGymâs score-4 anchor rewards claims that are correct, specific, and grounded in plausible source evidence. Agentic search systems such as TTD-DR and GPT-Researcher are explicitly prompted to produce source-attributed reports, while fine-tuned open-weight systems often emit cleaner answer-style prose with weaker inline attribution. This does not mean the trained models lack the underlying knowledge or search ability; it means their learned reporting interface is mismatched with a long-form, citation-sensitive financial rubric. C.2. Pre-cutoff and Post-cutoff Remain Different Failure Modes The stable gap between pre-cutoff and post-cutoff rubric scores is the most important behavioral pattern in the benchmark. Systems improve much more readily on pre-cutoff rubrics, where better search, reading, and citation discipline help recover pre-cutoff-date evidence. Post-cutoff rubrics remain compressed even for stronger systems, suggesting that financial deep research is not solved by retrieval quality alone: agents must turn historical evidence into forward-looking judgments that are only verifiable after the cutoff date. This is why the fixed-backbone FinanceHarness ablation is informative: even after training, post-cutoff performance remains well below pre-cutoff performance (about 12% versus 46%), indicating that forward-looking judgment, rather than evidence recall, is the harder part of the task. C.3. Agentic-system Trade-offs Agentic-system choice changes more than the final score. Among the agentic search systems, iterative pipelines such as TTD-DR achieve the highest scores, but they repeatedly plan, retrieve, draft, critique, 19 FinanceHarness: Autonomous Financial Deep Research Framework and revise. GPT-Researcher is slightly weaker on aggregate quality but runs a lighter, less iterative pipeline. Other agentic systems such as STORM and deepagents are useful counterpoints: they bring different composition strategies, but their default output styles are not always aligned with a rubric that rewards specific, source-grounded financial claims. We therefore treat these comparisons as evidence about orchestration fit-to-task, not as a universal ranking of agentic systems. C.4. Per-Axis Breakdowns Tables 11 and 12 give the per-sector and per-reasoning-type scores by system group, complementing the per-topic breakdown in the main text (Table 3). Table 11|Per-sector averaged scores (%) on FinanceGym, by system group. The 11 raw sectors merge into 9 leaves (Energy=energy+commodities; Macro=macro/policy+fixed income), ordered by the four coverage domains of Figure 2 (Health, Tech|Cons., Indus.|Energy, RE|Fin., Macro, Crypto). Each panel is one system group; Bold (best) andunderline(second) are normalized per column within the panel. (a) Fine-tuned open-weight SystemHealth Tech Cons. Indus. Energy RE Fin. Macro Crypto Tongyi-DR33.5 29.5 28.0 29.425.225.025.625.224.7 OpenResearcher 31.226.227.227.625.3 27.9 28.9 25.718.6 MiroThinker21.922.1 25.020.218.619.9 24.1 18.218.1 (b) Agentic search systems SystemHealth Tech Cons. Indus. Energy RE Fin. Macro Crypto TTD-DR38.2 31.531.3 30.228.930.9 36.7 26.318.6 GPT-Researcher 37.731.1 30.329.429.3 28.329.4 25.417.5 deepagents28.6 32.3 28.826.028.621.6 31.324.919.2 OpenClaw31.430.9 27.425.626.122.7 30.5 24.517.7 STORM31.027.9 28.327.126.126.8 29.4 23.719.2 (c) Proprietary with search SystemHealth Tech Cons. Indus. Energy RE Fin. Macro Crypto Claude-Opus-4.7 38.7 33.535.5 33.932.3 32.4 32.733.222.6 Gemini-3.1-Pro36.9 35.5 33.332.631.427.7 34.0 31.023.8 GPT-5.536.131.8 32.830.930.428.629.4 30.425.9 Gemini-3-Flash34.130.4 29.633.027.025.6 31.9 29.219.7 (d) Open-weight with search SystemHealth Tech Cons. Indus. Energy RE Fin. Macro Crypto FinanceHarness39.8 31.8 31.432.631.0 32.3 32.8 26.924.0 GLM-536.1 29.633.329.130.526.2 30.024.518.2 DeepSeek-v3.235.227.2 30.727.527.321.9 29.4 26.919.2 Qwen3-235B-A22B 30.725.9 26.927.324.727.428.4 25.318.4 Gemma-4-26B28.926.5 26.126.024.324.8 24.6 25.411.8 gpt-oss-120b19.919.4 17.217.120.018.5 22.1 16.715.5 20 FinanceHarness: Autonomous Financial Deep Research Framework Table 12|Per-reasoning-type averaged scores (%) on FinanceGym, by system group. Columns follow the three analytics levels of Figure 2 (Causal, Compar.|Forecast, Risk|Effect., Quant.). Each panel is one system group; Bold (best) andunderline(second) are normalized per column within the panel. (a) Fine-tuned open-weight SystemCausal Compar. Forecast Risk Effect. Quant. Tongyi-DR24.627.228.034.9 29.927.2 OpenResearcher 24.024.228.431.829.226.0 MiroThinker21.021.120.521.4 23.316.7 (b) Agentic search systems SystemCausal Compar. Forecast Risk Effect. Quant. TTD-DR27.531.031.135.934.832.4 GPT-Researcher 27.827.529.536.9 31.632.2 deepagents24.923.831.132.0 29.526.8 OpenClaw24.426.127.632.7 30.226.2 STORM25.026.129.130.5 28.425.1 (c) Proprietary with search SystemCausal Compar. Forecast Risk Effect. Quant. Claude-Opus-4.7 30.233.036.138.136.231.3 Gemini-3.1-Pro29.730.235.238.0 35.229.8 GPT-5.527.429.434.238.6 32.530.7 Gemini-3-Flash27.226.132.233.7 31.830.9 (d) Open-weight with search SystemCausal Compar. Forecast Risk Effect. Quant. FinanceHarness29.829.732.239.5 32.931.8 GLM-526.8 24.032.535.432.830.5 DeepSeek-v3.226.326.530.431.9 29.930.6 Qwen3-235B-A22B 23.723.828.031.0 28.825.6 Gemma-4-26B22.624.827.429.1 26.626.0 gpt-oss-120b17.319.918.021.9 19.914.5 21 FinanceHarness: Autonomous Financial Deep Research Framework D. FinanceHarness Case-Study Reports The boxed reports below are produced by FinanceHarness, drawn from the cross-backbone coverage evaluation (Appendix B). The six subsections span distinct tasks: an industry value-chain deep dive, a single-name discounted-cash-flow valuation, risk analytics, mega-cap relative comparables, option pricing, and fixed-income analysis. Each box gives the research question, the resulting cited report, and its sources; round-by-round tool trajectories for a representative subset are shown separately in Appendix E. These are live-web runs, separate from the point-in-time FinanceGym benchmark and not part of the headline results; the generated financial claims have not been independently verified and illustrate the harness workflow rather than provide investment or factual advice. D.1. Industry Deep Research AI Infrastructure Value Chain: Beneficiaries Beyond the GPU I . R E S E A R C H Q U E S T I O N AI compute demand is driving a historic buildout of GPUs and data-center infrastructure. As a research analyst, work outward from the AI accelerators to map the broader value chain: identify the successive layers of industries and suppliers whose demand is pulled by this buildout. For each layer you surface, name the leading public companies, explain how their business connects to AI-compute growth, and assess their financial profile and positioning (growth, margins, valuation). Conclude with a synthesized view of where along the chain the exposure looks most attractive and durable, and the key risks. I I . C I T E D R E P O RT AI Infrastructure Value Chain: Scarcity Rents Extend Well Beyond the GPU Synthesis The AI buildout is pulling demand through two interconnected chains: 1. The silicon chain: acceleratorsâ HBMâ advanced packaging/foundriesâ semiconductor equipment. 2. The physical-infrastructure chain: servers â networking/optics â electrical and cooling equipment â construction/grid interconnectionâ power generationâ data-center real estate. The most durable economics appear where technology, qualification requirements or physical constraints limit sub- stitution: NVIDIA and Broadcom in compute silicon, TSMC in advanced manufacturing and packaging, ASML in lithography, Arista in AI networking, and Eaton/Vertiv in power delivery and cooling. The better valuation asymmetry currently appears in TSMC and selected power producers, while memory, optical components and server assemblers offer stronger near-term growth but more cyclical or commoditized economics. At the other extreme, data-center REITs benefit directly from capacity scarcity but combine high capital intensity, approximately 5â6x net leverage and elevated valuation multiples. They are comparatively weak expressions of the theme unless power-secured sites can earn exceptional returns. Demand anchor Based on reported company plans cited by CNBC, Alphabet, Amazon, Microsoft and Meta together indicate roughly $734 billion of 2026 capital expenditure and finance leases. The definitions differ and not all of the spending is AI-related, but the scale establishes the demand envelope. The main question has shifted from whether infrastructure spending is rising to whether the associated revenue can generate adequate returns. Alphabetâs cash flow turned negative in Q2, while debt increased materially at both Alphabet and Amazon [9]. Alphabet illustrates the composition. Its Q2 capital expenditure was $44.9 billion, approximately 60% for servers and 40% for data centers and networking. Google Cloud revenue grew 82% and backlog reached $514 billion, suggesting that unusually high spending is still accompanied by capacity-constrained demand [8]. The physical consequence is equally significant. The IEA expects global data-center electricity consumption to more than double from 415 TWh in 2024 to approximately 945 TWh in 2030. Around 20% of planned projects could experience delays because transmission, transformers, cables and generating capacity are not being added fast enough [6]. Financial snapshot The table uses the latest available market-data snapshot as of July 28, 2026. Growth is recent year-over-year rev- 22 FinanceHarness: Autonomous Financial Deep Research Framework enue growth, margins are operating margins, and valuation is forward P/E unless otherwise indicated. These are screening metrics rather than fully normalized estimates; memory, utilities and hardware assemblers are especially cycle-sensitive. Layer and selected companiesRevenue growthOperating marginValuation Accelerators: NVIDIA / AMD / Broadcom85% / 38% / 48%66% / 14% / 49%15x / 33x / 19x Memory: Micron346%80%5x Foundry: TSMC36%60%18x Equipment: ASML / Lam / KLA21% / 24% / 12%37% / 35% / 41%27x / 32x / 37x Servers: Dell / Super Micro / HPE19% FY / 123% / 40%12% ISG / 6% / 9%17x / 8x / 11x Networking: Arista / Marvell / Cisco35% / 28% / 12%43% / 15% / 25%37x / 27x / 24x Optics: Coherent / Lumentum21% / 90%14% / 22%28x / 34x Power and cooling: Vertiv / Eaton30% / 17%16% / 16%29x / 24x Construction/grid: Quanta Services / EMCOR26% / 20%4% / 9%35x / 21x Generation: Constellation / Vistra / GE Vernova64% / 43% / 22%22% / 27% / 8%19x / 14x / 37x Data-center REITs: Equinix / Digital Realty12% / 30%24% / 26%29x / 29x EV/EBITDA Cloud buyers: Microsoft / Alphabet / Meta18% / 24% / 33%46% / 34% / 41%20x / 22x / 16x The successive layers 1. Accelerators and custom AI silicon Selected public companies: NVIDIA, AMD, Broadcom, Marvell. This is the starting point and the most profitable layer in the current financial snapshot. NVIDIA supplies accel- erators, networking products, systems and an extensive software platform. AMD is the principal listed merchant alternative in the screen, while Broadcom and Marvell provide exposure to custom silicon and networking. NVIDIAâs Q1 fiscal 2027 Data Center revenue reached $75.2 billionâapproximately 92% of total revenueâand grew 92% year over year. Its 74.9% GAAP gross margin and roughly 66% operating margin demonstrate the current scarcity rents associated with its data-center platform [1]. Broadcom provides a counterweight to NVIDIAâs merchant model. Its AI revenue reached $10.8 billion in fiscal Q2, more than doubling year over year, driven by custom accelerators, including Googleâs TPU, and networking components [11]. Broadcom is therefore exposed both to the expansion of AI clusters and to hyperscalersâ efforts to deploy internally designed accelerators. Positioning: This layer has the strongest current combination of growth, margins and technical barriers. Its weak- ness is customer concentration: a small number of hyperscalers control much of incremental spending and are simultaneously developing competing silicon. 2. High-bandwidth memory and data-center memory Selected public companies: SK Hynix, Micron, Samsung Electronics. An accelerator requires HBM, server DRAM and increasingly enterprise SSD capacity. HBM is especially demand- ing because multiple memory dies must be stacked, tested and integrated alongside the processor through ad- vanced packaging. SK Hynixâs 2026 market-outlook page describes direct exposure through HBM3E, HBM4, server DRAM and enter- prise SSDs. Citing Bank of America estimates, it places the 2026 HBM market at $54.6 billion, up 58%, while also warning that capacity additions and competition could create price pressure after 2026 [18]. Micronâs fiscal Q3 revenue rose from $9.3 billion to $41.5 billion year over year. Gross margin reached 84.6% and operating margin 80.4%, with management introducing multiyear customer agreements intended to make results more predictable [10]. Those margins should not be treated as normal through-cycle economics: Micron itself is investing at record levels in technology, products and supply, while the broader memory outlook acknowledges the possibility of future price correction [10,18]. Positioning: HBM is currently a bottleneck and may remain more differentiated than conventional memory. Nev- ertheless, the very low forward multiples reflect valid concerns about peak pricing, heavy capital expenditure and eventual supply response. 3. Foundry manufacturing and advanced packaging Leading public exposure: TSMC. AI accelerators require leading-edge logic and advanced packaging to connect compute dies with HBM. TSMCâs role in both advanced process nodes and packaging makes it more than a conventional wafer foundry. TSMC reported Q2 revenue of $40.2 billion, a 67.7% gross margin and a 60.3% operating margin [2]. Management characterized AI demand as âextremely robust,â raised expected 2026 revenue growth to slightly above 40%, and 23 FinanceHarness: Autonomous Financial Deep Research Framework increased planned capital expenditure to $60â64 billion. Approximately 70â80% is allocated to advanced processes and 10â20% to packaging, testing and related capacity [15]. The company screens at approximately 18x forward earnings despite growth and margins comparable with the highest-quality semiconductor companies in the market-data set. The valuation also reflects material risks: over- seas expansion is expected to dilute margins, the N2 ramp will carry startup costs, and a large share of current production remains geographically concentrated. Positioning: TSMC is arguably the best risk-adjusted tollbooth in the chain. Its foundry model provides expo- sure to demand from cloud providers, accelerator customers and CPU suppliers without relying on a single chip architecture. Geographic concentration is the dominant risk. 4. Wafer-fabrication and process-control equipment Selected public companies: ASML, Applied Materials, Lam Research, KLA. More AI silicon eventually requires additional lithography, deposition, etch, inspection and metrology capacity. Equipment suppliers participate with a lag: end demand first raises utilization at foundries and memory manufac- turers, which then place additional tool orders. ASML reportedâŹ32.7 billion of 2025 sales, a 52.8% gross margin and aâŹ38.8 billion year-end backlog. Q4 bookings includedâŹ7.4 billion of EUV orders. Management said stronger customer confidence in sustained AI demand had materially increased medium-term capacity plans and order intake [3]. The market-data screen shows operating margins of 35% at Lam and 41% at KLA, versus 37% at ASML. Their valuations of roughly 27â37x forward earnings reflect strong current returns and expectations that AI-related semi- conductor investment will remain elevated. Positioning: This is one of the more durable layers because advanced tools are technically complex and serve long-lived customer manufacturing programs. The trade-off is semiconductor-capital-equipment cyclicality and premium valuations. 5. Boards, servers, racks and systems integration Selected public companies: Dell, Super Micro Computer, HPE. These companies turn accelerators, CPUs, memory, network cards, power supplies and cooling components into deployable systems. AI servers generate exceptional revenue growth because accelerator content per server is extremely high, but much of that revenue is pass-through component value. Dell shipped more than $25 billion of AI-optimized servers in FY26, booked more than $64 billion of orders and en- tered FY27 with a $43 billion backlog. Its Infrastructure Solutions Group revenue grew 40%, but operating margin declined from 12.8% to 11.7%, demonstrating that volume does not automatically produce scarcity-level profitabil- ity [12]. Super Microâs market-data profile shows an 8% gross margin, a 6% operating margin and approximately 5x net debt/EBITDA. Dellâs broader reported portfolio, larger cash generation and lower net leverage in the compara- tive screen provide a more balanced profile, although its margins remain far below those of semiconductor and networking suppliers. Positioning: Strong demand exposure but weak structural economics. Buyers have alternatives, component sup- pliers capture much of the value, and working-capital requirements rise sharply during rapid growth. This is primarily a volume trade rather than a durable profit pool. 6. Ethernet switching, interconnect and optical components Selected public companies: Arista Networks, Broadcom, Marvell, Cisco, Coherent, Lumentum. Large clusters require both: ⢠Scale-up networking between accelerators within a rack or tightly coupled system. ⢠Scale-out networking connecting many racks into a larger training or inference fabric. Aristaâs announced 1.6-terabit portfolio is designed for rack-scale AI infrastructure and clusters expanding from thousands to hundreds of thousands of accelerators, including air-, liquid- and hybrid-cooled systems. The prod- ucts use Broadcom Tomahawk 6 switching silicon, illustrating how several companies can participate in the same network deployment [17]. Aristaâs approximately 64% gross margin, 43% operating margin and net-cash balance sheet distinguish it from lower-margin component suppliers. NVIDIAâs own networking revenue grew 199% to $14.8 billion, substantially faster than its 77% growth in Data Center compute revenue during the same quarter [1]. Coherent supplies optical products used in data-center com- munications; its fiscal Q3 revenue grew 21%, gross margin reached 37.7%, and management said it was expanding capacity in response to AI-data-center demand [13]. Positioning: Arista combines attractive economics with a differentiated network operating stack, but trades near 37x forward earnings. Optical suppliers offer higher operating leverage but face rapid product transitions, customer concentration and changing interconnect architectures. 24 FinanceHarness: Autonomous Financial Deep Research Framework 7. Electrical equipment and thermal management Leading public companies: Vertiv, Eaton. GPU density turns power conversion and heat removal into binding constraints. The equipment set includes switchgear, uninterruptible power supplies, power distribution and cooling systems. Vertiv is the higher-purity listed exposure in this comparison. Q1 sales grew 30%, with Americas organic growth of 44% on data-center demand. Adjusted operating margin rose 430 basis points to 20.8%, while full-year organic growth guidance increased to 29â31%. Net leverage was only about 0.2x [4]. Eaton supplies power and energy management, backup power, software, safety, design and project-management services to data-center operators. The company explicitly identifies dense GPU clusters as creating new electrical, thermal and operational stress [16]. Its 24x forward multiple is less demanding than Vertivâs 29x, though its overall business is less concentrated on data centers. Positioning: This is one of the most attractive second-order layers. Reliability requirements, manufacturing lead times and the need to redesign facilities for higher rack density support demand. Unlike a bet on one accelerator architecture, much of this spending remains necessary across competing compute platforms. The principal risk is that high valuations already discount several years of exceptional growth. 8. Data-center construction and grid interconnection Selected public companies: Quanta Services, EMCOR, Comfort Systems USA. These businesses provide exposure to transmission, substations, electrical work, mechanical systems and other large-load infrastructure. They benefit from the overlap among data-center construction, utility load growth and generation investment. Quanta reported Q1 revenue of $7.87 billion, up 26%, and record total backlog of $48.5 billion. Management cited converging utility, generation and large-load markets as the demand driver and identified data centers among potential future project areas [5]. The market-data snapshot shows higher operating margins at EMCOR and Comfort Systems than at Quanta: ap- proximately 9% and 8%, respectively, versus 4%. Quantaâs roughly 35x forward multiple is high for its current operating margin, while EMCOR screens at approximately 21x. Positioning: Backlogs provide visibility, but this remains a labor- and execution-intensive layer. Fixed-price con- tracts, skilled-labor shortages, permitting and project delays can erode the apparent backlog advantage. 9. Electricity generation and generation equipment Selected public companies: Constellation Energy, Vistra, NextEra Energy, GE Vernova, Caterpillar, Cummins. Data centers need continuous, high-quality power. The IEA expects electricity generation serving data centers to exceed 1,000 TWh by 2030. Natural gas and coal are expected to meet more than 40% of incremental demand through 2030, despite substantial renewable additions; nuclear becomes more important after 2030 [14]. Public-market exposure can therefore be sought through three broad groups: ⢠Electricity producers: Constellation, Vistra and NextEra. ⢠Generation and grid equipment: GE Vernova. ⢠Backup and distributed-generation equipment: Caterpillar and Cummins. Constellation and Vistra screen at approximately 19x and 14x forward earnings, respectivelyâbelow many equip- ment beneficiaries. Their financial exposure, however, depends on regional power markets, contract structures, operating performance and regulation. GE Vernova offers more direct equipment exposure but trades around 37x forward earnings despite an 8% operating margin. Positioning: This may offer some of the best valuation asymmetry in the physical chain, but it is not a clean or uniform exposure. Local transmission constraints mean national electricity growth will not translate evenly into company profits, and regulators can materially influence who pays for interconnection and generation investment. 10. Data-center landlords and interconnection platforms Selected public companies: Equinix, Digital Realty, Iron Mountain. These businesses monetize power-secured land, buildings and network interconnection. Equinix reported 10% Q1 revenue growth, a 51% adjusted EBITDA margin and 12% AFFO growth. Approximately 60% of its largest deals were AI-related [7]. However, Equinix and Digital Realty both trade near 29x EV/EBITDA and carry net debt around 5â6x EBITDA. They must continue investing heavily to deliver capacity: Equinix expects roughly $4.1 billion of 2026 capital ex- penditure [7]. Positioning: Demand is real, but economics depend on the spread between development returns and the cost of capital. Rising rates, power procurement, construction delays and customer concentration make this a less attrac- tive expression than supplying scarce equipment to multiple developers and operators. 25 FinanceHarness: Autonomous Financial Deep Research Framework Where the exposure is most attractive 1. Best combination of durability and valuation: TSMC TSMC combines 60% operating margins, above-40% expected 2026 revenue growth and an approximately 18x for- ward P/E. Its broad foundry and advanced-packaging role reduces dependence on which accelerator architecture ultimately gains share. Overseas expansion costs and geographic concentration prevent this from being a low-risk investment, but the operating franchise is unusually durable. 2. Highest-quality technology and ecosystem assets: ASML, NVIDIA and Broadcom ⢠ASMLâs EUV role,âŹ38.8 billion backlog and installed-base business provide exposure to multiple chipmak- ers rather than a single end product [3]. ⢠NVIDIA has the best current economics and the most integrated reported data-center platform in the group [1]. ⢠Broadcom provides exposure to custom accelerators and the networking required around them [11]. Their central risk is that a handful of large customers possess substantial buyer power and are actively funding alternatives. 3. Best physical-infrastructure beneficiaries: Eaton and Vertiv Power and cooling requirements increase with rack density across competing accelerator architectures. Eaton offers broader diversification and a lower multiple; Vertiv offers higher purity, faster growth and greater operational leverage. Neither is now undiscovered, making entry valuation important. 4. Best networking asset: Arista Aristaâs operating margin, balance sheet and AI-focused product roadmap are exceptional. Its announced 1.6T systems address both scale-up and scale-out AI fabrics [17]. The limitation is a roughly 37x forward valuation and dependence on a concentrated set of large cloud and technology customers. 5. Attractive but higher-risk value: Vistra and Constellation These companies provide electricity-generation exposure at less demanding multiples than many AI-equipment beneficiaries. The exposure is nevertheless indirect and heavily influenced by regulation, local power prices, oper- ating performance and transmission availability. 6. Tactical rather than durable: Micron, Coherent, Lumentum, Dell and Super Micro These companies can produce the strongest earnings surprises during scarcity. They also face the greatest risk of margin normalization: ⢠Memory supply eventually responds. ⢠Optical standards and architectures change quickly. ⢠Server assemblers retain relatively little of the acceleratorâs economics. ⢠Working capital and customer concentration magnify downside during pauses. 7. Least compelling at current structure: data-center REITs Equinix and Digital Realty have valuable power-secured portfolios, but high leverage, elevated EV/EBITDA mul- tiples and continuing capital requirements limit upside relative to suppliers that can sell equipment into multiple new projects. Key risks 1. Hyperscaler return-on-capital failure. A spending pause would propagate through servers, optics, semi- conductors, equipment and construction, although timing would differ by backlog and contract structure. Rising debt and negative free cash flow show that investor tolerance is not unlimited [9]. 2. Efficiency and substitution. Better models, inference optimization, lower-precision computation and cus- tom ASICs could reduce GPU requirements per unit of AI output. Broadcom may benefit from custom silicon, but merchant accelerator vendors would face pressure. 3. Customer concentration. A small number of cloud companies have bargaining power over nearly every layer, particularly custom chips, optics, servers and real estate. Broadcom identified six core custom-chip customers as drivers of AI growth [11]. 4. Capacity overshoot. HBM, optical components and server assembly are most vulnerable if suppliers extrap- olate current shortages into permanent demand. SK Hynixâs market outlook specifically identifies possible HBM price correction after 2026 as competition and capacity increase [18]. 5. Geopolitics and export controls. The chain remains heavily dependent on TSMCâs manufacturing footprint. China restrictions can also remove accelerator revenue directly: NVIDIAâs Q2 outlook assumed no Data Center compute revenue from China [1]. 26 FinanceHarness: Autonomous Financial Deep Research Framework 6. Power bottlenecks and regulation. Transmission projects can take four to eight years, and grid constraints may delay data centers even when server demand remains strong [6]. Regulators can also determine how interconnection and generation costs are distributed. 7. Technology transitions. HBM4, N2, 1.6T Ethernet and liquid cooling require coordinated execution across multiple suppliers. The relevant company sources describe continuing ramps or future availability rather than fully completed transitions [15,17,18]. 8. Valuation compression. Vertiv, Arista, Quanta and GE Vernova already trade at growth-stock multiples. Even successful earnings delivery may not protect them if the market lowers terminal-growth assumptions. Bottom line The AI infrastructure boom is not one trade but a hierarchy of profit pools. Economic value is currently highest where there is proprietary technology plus scarce capacityâaccelerators, advanced foundry and packaging, lithog- raphy and high-performance networking. The next most durable opportunity is in power distribution and cooling, where demand is less dependent on which compute architecture wins. The preferred risk-adjusted exposures are therefore TSMC first; ASML, NVIDIA and Broadcom for technology scarcity; Eaton and Vertiv for the physical bottleneck; and Arista for network scaling. Vistra and Constellation offer cheaper but more policy-sensitive exposure. Memory, optics and server assemblers are better treated as cycli- cal beneficiaries, while data-center REITs currently offer the least favorable combination of valuation, leverage and capital intensity. I I I . S O U R C E S [1] NVIDIA Announces Financial Results for First Quarter Fiscal 2027 [2] TSMC 2026 Q2 Quarterly Results - Taiwan Semiconductor ... [3] ASML reportsâŹ32.7 billion total net sales andâŹ9.6 billion ... [4] Vertiv Holdings Co. - Vertiv Reports Strong First Quarter ... [5] QUANTA SERVICES REPORTS FIRST QUARTER 2026 RESULTS [6] Executive summary â Energy and AI â Analysis - IEA [7] Equinix Reports First-Quarter Results and... :: Equinix, Inc. (EQIX) [8] 2026 Q2 Earnings Call - Alphabet Investor Relations [9] Hyperscalers face higher capex scrutiny after Alphabet report panned [10] Micron Technology, Inc. Reports Record Results for the Third ... [11] Broadcom (AVGO) earnings report Q2 2026 - CNBC [12] Dell Technologies Delivers Fourth Quarter and Full-Year ... [13] Coherent Corp. Reports Third Quarter Fiscal 2026 Results [14] Energy supply for AI â Energy and AI â Analysis - IEA [15] Q2 2026 Taiwan Semiconductor Manufacturing Co Ltd Earnings ... [16] Data centers | Digitalization | Efficiency | Eaton [17] Arista Networks - Arista Introduces Next-Generation 1.6Terabit ... [18] 2026 Market Outlook â âFocus on the HBM-Led Memory Supercycleâ 27 FinanceHarness: Autonomous Financial Deep Research Framework D.2. Valuation Compute Seam Microsoft (MSFT): DCF, WACC & Sensitivity I . R E S E A R C H Q U E S T I O N Build an auditable five-year unlevered DCF for Microsoft (MSFT) from its latest fiscal-year results. Forecast revenue, EBIT margin, taxes, D&A, capex, and changes in working capital using clearly stated assumptions; derive WACC from current market inputs, a peer-based beta, prevailing Treasury yields, and an explicit equity-risk premium; use a 2.5% perpetual-growth rate; bridge enterprise value to equity value and intrinsic value per diluted share. Include a sensitivity table spanning WACC Âą2 percentage points and terminal growth from 1.5% to 3.5%. Compare the result with the current market price without making an investment recommendation. I I . C I T E D R E P O RT Microsoft five-year unlevered DCF: $233 per diluted share Using Microsoftâs FY2025 results as the operating base, a five-year unlevered DCF produces: ⢠Enterprise value: $1.674 trillion ⢠Equity value: $1.741 trillion ⢠Intrinsic value: $233.19 per diluted share ⢠Current market price: $381.70 ⢠DCF-to-market difference: 38.9% below the current price This is a valuation result, not an investment recommendation. The modelâs largest judgment calls are the 11.60% peer-beta-based WACC and the pace at which elevated AI infrastructure spending normalizes. 1. FY2025 operating base Microsoftâs fiscal year ended June 30, 2025. Revenue was $281.724 billion and operating incomeâused here as EBITâwas $128.528 billion, producing a 45.62% EBIT margin. The reported income-tax provision was $21.795 billion on $123.627 billion of pretax income, an effective rate of 17.63%.[2] FY2025 anchorAmount Revenue$281.724B EBIT$128.528B EBIT margin45.62% Reported income-tax provision$21.795B Effective tax rate17.63% D&A and other$34.153B Capital expenditures$64.551B Operating cash flow$136.162B Reported OCF less capex$71.611B D&A increased from $22.287 billion in FY2024 to $34.153 billion in FY2025, while capital expenditures increased from $44.477 billion to $64.551 billion.[3] Those increasesâ53.2% and 45.1%, respectivelyâsupport modeling a period in which depreciation catches up with the recent infrastructure build and capex gradually normalizes. Working-capital baseline The FY2025 cash-flow statement reported the following operating-asset and liability changes:[3] ⢠Accounts receivable: $(10.581)B ⢠Inventories: +$0.309B ⢠Other current assets: $(3.044)B ⢠Accounts payable: +$0.569B ⢠Unearned revenue: +$5.438B ⢠Other current liabilities: +$5.922B The net cash effect was a $1.387 billion outflow, equivalent to 3.79% of FY2025âs $36.602 billion revenue increase. The forecast rounds this to a 4.0% investment for every dollar of incremental revenue. 2. Forecast assumptions These are explicit analyst assumptions rather than company guidance or consensus estimates. 28 FinanceHarness: Autonomous Financial Deep Research Framework AssumptionFY26EFY27EFY28EFY29EFY30E Revenue growth14.0%13.0%12.0%10.0%8.0% EBIT margin45.7%45.9%46.1%46.3%46.5% Tax rate on EBIT18.0%18.0%18.0%18.0%18.0% D&A as % of revenue13.5%14.5%15.0%15.2%15.0% Capex as % of revenue22.0%20.0%18.0%16.5%15.0% Change in NWC4% of incremental revenue in every year The revenue path represents an 11.38% five-year CAGR, decelerating from FY2025âs 15% reported growth. Mi- crosoft reported Azure annual revenue above $75 billion, up 34%, providing context for maintaining double-digit consolidated growth through FY2028E.[2] The EBIT-margin assumption provides only modest expansionâ45.62% to 46.5%âbecause scale and product mix benefits are assumed to be partly offset by rising depreciation from the infrastructure build. Capex remains above D&A through FY2029E and converges with D&A in FY2030E. 3. Five-year unlevered free-cash-flow forecast Formula: UFCF= EBITĂ(1â tax rate)+ D&Aâ capexâÎNWC Amounts are in billions of dollars. FY26EFY27EFY28EFY29EFY30E Revenue$321.165$362.917$406.467$447.114$482.883 EBIT146.773166.579187.381207.014224.540 Taxes on EBIT(26.419)(29.984)(33.729)(37.263)(40.417) NOPAT120.354136.595153.653169.751184.123 D&A43.35752.62360.97067.96172.432 Capex(70.656)(72.583)(73.164)(73.774)(72.432) Change in NWC(1.578)(1.670)(1.742)(1.626)(1.431) Unlevered FCF$91.477$114.964$139.717$162.313$182.692 The DCF calculation uses the displayed unrounded FCF schedule. 4. WACC derivation Peer-based beta The model uses the equal-weighted mean of current observed equity betas for five large cloud, enterprise-software and platform peers: PeerEquity beta Alphabet1.25 Amazon1.46 Oracle1.71 Salesforce1.18 Adobe1.43 Equal-weighted mean1.406 This is an observed peer-equity-beta approach; the individual betas are not unlevered and relevered. Cost of capital 29 FinanceHarness: Autonomous Financial Deep Research Framework InputValueBasis Risk-free rate4.68%Current 10-year Treasury yield Equity-risk premium5.00%Explicit model assumption Peer beta1.406Equal-weighted peer mean Cost of equity11.71%4.68% + 1.406Ă 5.00% Pretax cost of debt5.18%Treasury yield plus assumed 0.50% spread Tax rate17.63%FY2025 effective tax rate After-tax debt cost4.27%5.18%Ă (1â 17.63%) Debt/equity1.519%$43.151B debtá $2.840T market capitalization Equity weight98.50%Market-value weight Debt weight1.50%Market-value weight WACC11.5986%Rounded to 11.60% Microsoftâs FY2025 debt consisted of $2.999 billion of current long-term debt and $40.152 billion of long-term debt; short-term debt was zero.[4] 5. Terminal value and enterprise-to-equity bridge The terminal value uses a 2.5% perpetual-growth rate: Terminal value 2030 = $182.692íľĂ 1.025 11.5986%â 2.5% = $2.058í Valuation componentAmount PV of five explicit FCFs$484.990B PV of terminal value1,188.980B Enterprise value$1,673.970B Add: cash and short-term investments94.565B Add: equity and other investments15.405B Less: debt(43.151B) Equity value$1,740.789B Diluted shares7.465B Intrinsic value per diluted share$233.19 Microsoft reported $30.242 billion of cash, $64.323 billion of short-term investments and $43.151 billion of total debt at June 30, 2025.[4] Equity and other investments were $15.405 billion.[2] The valuation uses FY2025 weighted- average diluted shares of 7.465 billion rather than the 7.434 billion period-end basic share count.[5] The terminal value represents 71.03% of enterprise value, making the result materially sensitive to WACC and perpetual growth. 6. Sensitivity analysis Intrinsic value per diluted share: WACC \ Terminal growth1.5%2.0%2.5%3.0%3.5% 9.6%$272$285$301$319$340 10.6%$241$251$263$276$291 11.6%$216$224$233$243$255 12.6%$196$202$209$217$226 13.6%$179$184$190$196$203 The requested sensitivity range is $179â$340 per share. Even the highest value in the gridâ9.6% WACC and 3.5% perpetual growthâis 10.9% below the current $381.70 market price. Bottom line The base model values Microsoft at $233.19 per diluted share, versus a current price of $381.70. The difference indicates that the market price reflects some combination of faster or more durable cash-flow growth, a lower re- quired return, greater margin expansion, or less reinvestment than assumed here. No investment recommendation is made. 30 FinanceHarness: Autonomous Financial Deep Research Framework I I I . S O U R C E S [1] Microsoft 2025 Annual Report [2] FY25 Q4 - Press Releases - Investor Relations - Microsoft [3] FY25 Q4 - Cash Flows - Investor Relations - Microsoft [4] FY25 Q4 - Balance Sheets - Investor Relations - Microsoft [5] FY25 Q4 - Income Statements - Investor Relations - Microsoft 31 FinanceHarness: Autonomous Financial Deep Research Framework D.3. Risk Analytics Tesla (TSLA): Value-at-Risk, Beta & Correlation I . R E S E A R C H Q U E S T I O N Using daily adjusted prices for TSLA and SPY from 2023-01-03 through 2025-12-31, calculate TSLAâs 1-day 95% and 99% historical VaR and normal-parametric VaR for a $1,000,000 long position. Also calculate annualized volatility, beta to SPY, and return correlation. State the return convention, observation count, formulas, and any data gaps. I I . C I T E D R E P O RT TSLA Risk Analytics â 1-Day VaR, Volatility, Beta, and Correlation Data Window and Conventions ItemDetail Price seriesTSLA and SPY, daily adjusted close, 1255 bars each Return observationsn = 1,254 (daily simple returns from 1,255 prices) Actual window~5 years back from 2026-07-24 (the toolâs 5y preset) Return conventionSimple returns: r t = (P t â P t - 1 ) / P t - 1 Annualizationí_annual = í_dailyĂ â 252 Data gapsNo missing days reported; both series have identical bar counts (1,255). Note on date range: The price tool supports predefined windows (1mo, 3mo, 6mo, 1y, 2y, 5y, etc.) and does not accept custom start/end dates. The 5y window was used, covering approximately mid-2021 through 2026- 07-24. This is broader than the requested 2023-01-03 to 2025-12-31 range; the results therefore include additional observations outside that window. 1. Value-at-Risk (1-Day Horizon, $1,000,000 Long Position) VaR is expressed as a positive dollar loss â the maximum expected loss at the given confidence level. ConfidenceMethodVaR (%)Dollar VaR 95%Historical (empirical tail)5.8%$58,000 95%Parametric (normal)6.1%$61,000 99%Historical (empirical tail)9.5%$95,000 99%Parametric (normal)8.6%$86,000 Formulas: ⢠Historical VaR: Sort all 1,254 daily returns; take the empirical quantile at (1â íź). For 95%, the 5th percentile; for 99%, the 1st percentile. ⢠Parametric (Normal) VaR: VaR = z 1 - íźĂ í_daily, where z 1 - íź is the standard-normal quantile (1.645 for 95%, 2.326 for 99%) and í_daily = 3.76%. Interpretation: The divergence between historical and parametric VaR at the 99% level (9.5% vs. 8.6%) signals fat tails in TSLAâs return distribution â confirmed by the excess kurtosis of 3.04. The normal model underestimates tail risk at high confidence levels. The historical method, which uses the actual empirical tail, gives a more conservative 99% VaR. 2. Annualized Volatility MetricValue Daily í (sample)3.76% Annualized í = í_dailyĂ â 25259.7% This is consistent with the toolâs reported realized annualized volatility of 59.6%. 3. Beta to SPY 32 FinanceHarness: Autonomous Financial Deep Research Framework MetricValue Beta (TSLA vs. SPY)2.02 Correlation0.58 R 2 0.34 Formula: í˝ = Cov(r_TSLA, r_SPY) / Var(r_SPY), estimated via OLS regression of TSLA excess returns on SPY excess returns over the common 1,254-observation window. Interpretation: TSLA is roughly twice as sensitive to SPY market moves. An estimated 34% of TSLAâs return variance is explained by SPY (R 2 = 0.34), meaning the majority of TSLAâs risk is idiosyncratic. 4. Return Correlation PairPearson Correlation TSLAâ SPY0.58 Formula: í = Cov(r_TSLA, r_SPY) / (í_TSLAĂ í_SPY), computed over 1,254 paired daily returns. Interpretation: A moderate positive correlation â TSLA tends to move with the market but with substantial inde- pendent movement. This aligns with the beta of 2.02 and R 2 of 0.34. 5. Return Distribution Diagnostics StatisticValue Mean daily return+0.10% Median daily return+0.13% Min daily returnâ15.43% Max daily return+22.69% Skewness+0.25 (SEÂą0.07) Excess kurtosis+3.04 (fat-tailed) Coefficient of variation38.06 The positive excess kurtosis of 3.04 confirms that TSLAâs returns have significantly fatter tails than a normal distri- bution â extreme moves are far more common than the normal model predicts. This is the primary reason the 99% historical VaR (9.5%) exceeds the parametric VaR (8.6%). Bottom Line For a $1,000,000 long position in TSLA, the 1-day 95% VaR is $58,000â$61,000 and the 1-day 99% VaR is $86,000â $95,000, depending on method. TSLAâs annualized volatility of ~60% and beta of 2.02 to SPY indicate a high- volatility, high-beta stock whose tail risk is underestimated by normal-parametric models. The fat-tailed return distribution (excess kurtosis = 3.04) makes the historical VaR the more reliable measure at the 99% confidence level. 33 FinanceHarness: Autonomous Financial Deep Research Framework D.4. Relative Comparables Alphabet (GOOGL): Mega-Cap Relative Valuation I . R E S E A R C H Q U E S T I O N Perform a current relative-valuation analysis of Alphabet (GOOGL) against META, MSFT, AMZN, and AAPL. Compare trailing and forward P/E, EV/revenue, EV/EBITDA, PEG, and free-cash-flow yield alongside revenue growth, margins, and ROIC. Normalize obvious nonrecurring items, document estimate dates, and calculate the GOOGL equity values implied by peer-median forward earnings and EBITDA multiples. I I . C I T E D R E P O RT Relative Valuation Analysis: Alphabet Inc. (GOOGL) vs. Mega-Cap Peer Group Executive Synthesis As of July 28, 2026, Alphabet Inc. (GOOGL) trades at a market capitalization of $3.99 trillion ($327.00 per share). On an unadjusted trailing GAAP earnings basis, Alphabet appears discounted relative to its mega-cap technology peersâMeta Platforms (META), Microsoft (MSFT), Amazon (AMZN), and Apple (AAPL)âtrading at a trailing P/E of 16.39x compared to the peer median of 25.43x. However, this headline discount is heavily distorted by GAAP accounting under ASC 321, which requires non-operating unrealized gains on strategic equity investments (notably in Anthropic and SpaceX) to be recognized in net income [2, 3]. Over the Trailing Twelve Months (TTM) ended June 30, 2026, Alphabet recorded $157.38 billion in pre-tax non-operating equity gains [1, 2, 3, 4]. When normalizing for these non-operating investment gains, Alphabetâs core operating trailing P/E rises to ap- proximately 34.28x. On forward consensus metricsâwhich isolate recurring operating earningsâAlphabet trades at a forward P/E of 22.21x, nearly in line with the peer median of 21.71x. Valuing Alphabetâs equity by applying peer-median multiples yields an implied valuation range of $2.93 trillion to $6.17 trillion ($240.55 to $506.00 per share), depending on whether multiples are applied to TTM core metrics or forward consolidated estimates. Data Snapshots & Estimate Dates ⢠Market Price & Consensus Estimate Date: July 28, 2026. ⢠Financial Reporting Period: Trailing Twelve Months (TTM) ended June 30, 2026 (comprising Q3 2025, Q4 2025, Q1 2026, and Q2 2026), sourced directly from SEC Form 10-Q and Form 8-K filings [1, 2, 3, 4]. Trading Multiples & Valuation Overview The table below presents trailing and forward valuation multiples across the peer group. Part 1 of 2 CompanyTickerMarket Cap ($B)Enterprise Value ($B)Trailing P/E (GAAP)Trailing P/E (Norm.) Alphabet Inc.GOOGL$3,990.00$3,868.3216.39x34.28x Meta PlatformsMETA$1,510.00$1,515.5921.60x21.60x Microsoft Corp.MSFT$2,890.00$2,937.2023.17x23.17x Amazon.com Inc.AMZN$2,490.00$2,582.4527.68x27.68x Apple Inc.AAPL$4,950.00$4,966.2040.79x40.79x Peer Medianâ$2,690.00$2,759.8325.43x25.43x Part 2 of 2 CompanyForward P/EEV / RevenueEV / EBITDAPEG RatioFree Cash Flow Yield Alphabet Inc.22.21x8.72x22.46x1.26x0.57% Meta Platforms16.05x7.04x13.84x0.87x1.69% Microsoft Corp.20.09x9.23x15.93x1.18x1.28% Amazon.com Inc.23.33x3.48x16.56x1.25x0.39% Apple Inc.34.93x11.00x31.03x2.68x2.04% Peer Median21.71x8.14x16.25x1.22x1.49% Operating & Financial Fundamentals The table below summarizes revenue growth, margin profiles, balance sheet leverage, and return on capital across the target and peer set. Part 1 of 2 34 FinanceHarness: Autonomous Financial Deep Research Framework CompanyTickerRevenue TTM ($B)Revenue Growth (YoY)Gross MarginOperating Margin Alphabet Inc.GOOGL$445.8724.2%60.9%34.0% Meta PlatformsMETA$214.9633.1%81.9%40.6% Microsoft Corp.MSFT$318.2718.3%68.3%46.3% Amazon.com Inc.AMZN$742.7816.6%50.6%13.1% Apple Inc.AAPL$451.4416.6%47.9%32.3% Peer Medianâ$384.8617.45%59.45%36.45% Part 2 of 2 CompanyNet Margin (GAAP)Net Margin (Norm.)Return on Equity (ROE)ROICNet Debt / EBITDA Alphabet Inc.54.8%26.3%48.7%61.3%â0.70x Meta Platforms32.8%32.8%32.9%33.5%0.05x Microsoft Corp.39.3%39.3%34.0%28.2%0.26x Amazon.com Inc.12.2%12.2%24.3%14.8%0.59x Apple Inc.27.2%27.2%141.5%56.4%0.10x Peer Median30.00%30.00%33.45%30.85%0.18x Normalization of Nonrecurring & Non-Operating Items 1. ASC 321 Equity Security Unrealized Gains Under US GAAP (ASU 2016-01 / ASC 321), changes in the fair value of equity investments are recognized di- rectly in the Consolidated Statement of Income within "Other income (expense), net." Over the TTM ended June 30, 2026, Alphabet recorded unprecedented unrealized mark-to-market gains on its minority equity stakes in private technology companies (including Anthropic and SpaceX) [2, 3]: ⢠Q2 2026: Net gain on equity securities of $99.03 billion on pre-tax income of $138.75 billion [3]. ⢠Q1 2026: Net gain on equity securities of $36.92 billion on pre-tax income of $77.41 billion [2]. ⢠Q4 2025: Net gain on equity securities of approximately $10.70 billion [1]. ⢠Q3 2025: Net gain on equity securities of $10.73 billion on pre-tax income of $43.99 billion [4]. TTM Impact & Adjustment: Total TTM pre-tax net equity gains totaled $157.38 billion. At Alphabetâs average effective tax rate of ~19.1%, these non-operating gains boosted TTM reported Net Income by $127.32 billion ($10.38 per share). ⢠Reported TTM Net Income: $244.20 billion ($19.92 per share). ⢠Normalized TTM Operating Net Income: $116.88 billion ($9.54 per share). ⢠Unadjusted GAAP Trailing P/E: $327.00 / $19.92 = 16.39x. ⢠Normalized Trailing P/E: $327.00 / $9.54 = 34.28x. 2. Legal Settlements & Real Estate Optimization Charges In previous quarters, operating income included legal settlements ($1.40 billion in Q2 2025) and office space opti- mization charges ($607 million in Q3 2024). These items are excluded from forward earnings estimates and do not materially alter the TTM operating margin baseline of 34.0%. Implied Equity Valuation Analysis To establish a relative valuation framework, peer median multiples are applied to Alphabetâs fundamentals. 1. Valuation Implied by Peer-Median Forward P/E (21.71x) ⢠FY+1 Consensus EPS ($14.81): Implied Stock Price= $14.81Ă 21.715= $321.60 per share Implied Market Cap / Equity Value= $321.60Ă 12.20B shares= $3, 923.52 billion($3.92T) ⢠NTM / FY0 Consensus EPS ($19.84): Implied Stock Price= $19.84Ă 21.715= $430.83 per share Implied Market Cap / Equity Value= $430.83Ă 12.20B shares= $5, 256.13 billion($5.26T) 35 FinanceHarness: Autonomous Financial Deep Research Framework ⢠Using Full Peer Median Composite P/E (25.43x): Implied Stock Price= $19.84Ă 25.43= $504.53 per share Implied Market Cap / Equity Value= $504.53Ă 12.20B shares= $6, 155.27 billion($6.16T) 2. Valuation Implied by Peer-Median EV / EBITDA (16.25x) ⢠TTM EBITDA Baseline ($173.16 billion): Implied Enterprise Value= $173.16BĂ 16.245= $2, 812.98 billion Plus Net Cash Position=+$121.68 billion Implied Equity Value= $2, 812.98B+ $121.68B= $2, 934.66 billion($2.93T) Implied Stock Price= $2, 934.66B 12.20B shares = $240.55 per share ⢠Forward Projected EBITDA Composite Baseline: When applying the 16.25x peer median to consensus forward EBITDA projections (~$368B implied EV), adding net cash of $121.68 billion yields: Implied Enterprise Value= $5, 978.32 billion Implied Equity Value= $6, 100.00 billion($6.10T) Implied Stock Price= $500.00 per share Comparability Limitations & Contextual Drivers Relative valuation comparisons among mega-cap tech companies require careful interpretation due to structural differences across their business models and balance sheets: 1. Business Model & Monetization Mix: ⢠Alphabet vs. Meta: Both companies derive the majority of revenue from digital advertising (Google Search, YouTube, Meta Family of Apps). However, Google Cloud (growing over 80% YoY) and YouTube Subscriptions create an enterprise SaaS component that commands higher software multi- ples than pure-play ad models. ⢠Alphabet vs. Microsoft & Amazon: Microsoft and Amazon generate dominant cash flows from en- terprise cloud (Azure, AWS) and software licensing/e-commerce. These recurring, high-retention rev- enue streams trade at higher EV/Revenue multiples (9.23x for MSFT vs. 8.72x for GOOGL). ⢠Alphabet vs. Apple: Apple operates a hardware/consumer ecosystem with brand loyalty, driving an industry-leading ROE (141.5%) through aggressive share repurchases and low capital intensity relative to cloud infrastructure. 2. Capital Expenditure Intensity & Free Cash Flow Suppression: Alphabetâs TTM Free Cash Flow yield (0.57%) appears depressed relative to Apple (2.04%) and Meta (1.69%). This is driven by Alphabetâs mas- sive AI infrastructure buildout, with capital expenditure projected between $175 billion and $190 billion in 2026 [1, 2]. While these investments build long-term AI compute capacity, they temporarily compress cash conversion relative to TTM EBITDA. 3. Balance Sheet Strength & Net Cash: Alphabet maintains a massive net cash position of $121.68 billion ($242.47 billion cash vs. $120.79 billion debt), resulting in a net debt/EBITDA ratio of -0.70x. In contrast, Amazon carries $92.45 billion in net debt (0.59x EBITDA) and Microsoft carries $47.20B in net debt (0.26x EBITDA). Standard P/E multiples do not adjust for uninvested balance sheet liquidity, whereas EV-based multiples properly credit net cash. 4. Regulatory & Litigation Overhead: Alphabet faces unique regulatory risks, including ongoing DOJ an- titrust litigation regarding search distribution agreements and European antitrust remedies. These litigation overhangs introduce non-operational discount factors that are not present in peers like Microsoft or Apple to the same degree. Summary Findings ⢠Headline vs. Normalized Trailing P/E: Alphabetâs unadjusted trailing GAAP P/E of 16.39x reflects $157.38 billion in pre-tax non-operating equity gains [1, 2, 3, 4]. On a normalized operating basis, Alphabet trades at 34.28x trailing earnings. ⢠Forward Valuation Parity: On a forward P/E basis (22.21x), Alphabet trades closely aligned with the peer median (21.71x), reflecting robust revenue growth (24.2% YoY TTM) and operational execution in Google Cloud. 36 FinanceHarness: Autonomous Financial Deep Research Framework ⢠Implied Equity Value Range: Peer-median multiples imply a GOOGL equity value range between $2.93 trillion ($240.55 per share based on TTM EBITDA and net cash) and $5.26 trillion to $6.17 trillion ($430.83 to $506.00 per share based on forward earnings and composite EV/EBITDA multiples). I I I . S O U R C E S [1] Alphabet Announces Fourth Quarter and Fiscal Year 2025 Results [2] Alphabet Announces First Quarter 2026 Results [3] Alphabet Announces Second Quarter 2026 Results [4] GOOG Exhibit 99.1 Q3 2025 - s206.q4cdn.com 37 FinanceHarness: Autonomous Financial Deep Research Framework D.5. Derivatives Apple (AAPL): Option Pricing â Greeks & Put-Call Parity I . R E S E A R C H Q U E S T I O N Use Black-Scholes to price European AAPL call and put options with spot $225, strike $230, 90 calendar days to expiry, annual volatility 28%, continuously compounded risk-free rate 4.25%, and continuous dividend yield 0.45%. Report call and put values, delta, gamma, vega, theta, rho, and the put-call-parity check. I I . C I T E D R E P O RT Black-Scholes Values Using í= 90/365= 0.246575 years, í= $225, íž= $230, í= 28%, í= 4.25%, and í= 0.45%: MetricEuropean callEuropean put Value per share$11.1610$14.0128 Delta0.4910â0.5079 Gamma0.012730.01273 Vega, per 1 vol point$0.4451$0.4451 Theta, per calendar dayâ$0.0794 â$0.0557 Rho, per 1 percentage-point rate move$0.2449 â$0.3163 For a standard 100-share contract, the theoretical premiums are approximately $1,116.10 for the call and $1,401.28 for the put. Put-Call Parity With a continuous dividend yield: íśâ í= íí âíí â íží âíí Option-price side: 11.1610â 14.0128=â2.8518 Discounted-underlying side: 225í â0.0045í â 230í â0.0425í = 224.750â 227.602=â2.85183 The residual is approximately $0.00003 per share, so put-call parity holds to rounding. Equivalently: íś+ íží âíí â í+ íí âíí â $238.763 38 FinanceHarness: Autonomous Financial Deep Research Framework D.6. Fixed Income U.S. Treasury Bond: Duration, Convexity & Repricing I . R E S E A R C H Q U E S T I O N Analyze the U.S. Treasury 4.125% note due 2032-11-15, assuming settlement on 2026-07-27, $100 face value, semiannual coupons, and a 4.35% yield to maturity. Calculate accrued interest, clean and dirty price, Macaulay duration, modified duration, and convexity. Then estimate and exactly reprice the bond for parallel yield shocks of -50 and +50 basis points, comparing the duration-only and duration-plus-convexity approximations. I I . C I T E D R E P O RT U.S. Treasury 4.125% Note Due 2032-11-15 â Bond Analytics Bond Terms and Accrued Interest ParameterValue Coupon rate4.125% (semiannual) Maturity2032-11-15 Settlement2026-07-27 Years to maturity6.392 Face value$100 Yield to maturity4.35% Accrued interest. The last coupon was paid 2026-05-15; the next is 2026-11-15. Under the Actual/Actual day-count convention, 73 of 184 days have elapsed in the current coupon period: ⢠Semiannual coupon: $100Ă 4.125%á 2 = $2.0625 ⢠Accrued interest: (73/184)Ă $2.0625 = $0.8183 Price, Duration, and Convexity MetricValue Clean price98.7379 Accrued interest0.8183 Dirty (full) price99.5562 Current yield4.2% YTM4.35% Macaulay duration5.76 years Modified duration5.64 Convexity37.0 DV010.0557 per $100 face The bond trades at a slight discount (clean price < par) because the coupon (4.125%) is below the yield (4.35%). Yield Shock Analysis: Âą50 Basis Points The duration-only approximation isÎP/PââD_modĂÎy, and the duration-plus-convexity approximation adds the second-order term:ÎP/PââD_modĂÎy + ½à CĂÎy 2 . â50 bp shock (YTMâ 3.85%) MethodApprox. clean priceError vs. exact Exact reprice101.5681â Duration-only101.5220 â0.0458 Duration + convexity101.5680 â0.0001 39 FinanceHarness: Autonomous Financial Deep Research Framework +50 bp shock (YTMâ 4.85%) MethodApprox. clean priceError vs. exact Exact reprice95.9991â Duration-only95.9535 â0.0456 Duration + convexity95.9992+0.0001 Percentage price changes ShockExact %ÎPDuration-only %ÎPDuration+Convexity %ÎP â50 bp+2.87%+2.82%+2.87% +50 bp â2.77% â2.82%â2.77% Key Observations 1. Positive convexity benefit. The bond gains more when yields fall (+2.87%) than it loses when yields rise by the same amount (â2.77%). This asymmetry is the value of convexity: the price-yield curve is bowed toward the investor. 2. Duration-only underestimates price in both directions. By ignoring convexity, the first-order approxima- tion underprices the bond by ~$0.046 per $100 face for a 50 bp move. This error scales quadratically with the shock size. 3. Duration + convexity is nearly exact. The second-order Taylor approximation reproduces the exact reprice within $0.0001 per $100 face for aÂą50 bp parallel shift, confirming that convexity of 37.0 is well-calibrated for this bond at this yield level. 4. DV01 interpretation. A 1 bp move in yield changes the clean price by approximately $0.0557 per $100 face. For a $10 million position, that is ~$5,570 per bp. 40 FinanceHarness: Autonomous Financial Deep Research Framework E. FinanceHarness Case-Study Trajectories To keep the appendix focused we show two illustrative trajectories rather than one per report: a broad industry deep-dive and a single-name discounted-cash-flow valuation. Between them they exercise the core of the harness: web search and reading, structured equity and market-data retrieval, valuation computation (DCF/WACC, relative comps), and citation composition. Each box is the round-by-round trajectory that produced the correspondingly titled report in Appendix D, showing the agentâs tool calls with their arguments and the responses they returned. Long listings, repeated plan updates, and the calculator scratchpad are abbreviated (marked with an ellipsis) to keep the logic flow clear. E.1. Industry Deep Research AI Infrastructure Value Chain: Beneficiaries Beyond the GPU R E S E A R C H T R A J E C T O RY R1 âŚupdate_plan: [in progress] Define the AI infrastructure value-chain layers and representative public compa- nies; [pending] Research demand drivers, market structure, and company positioning by layer; [pending] Gather and compare growth, margin, and valuation profiles; [pending] Synthesize attractiveness, durability, and key risksâ Plan updated â 0/4 step(s) done. R2 âŚload_skill(industry-analysis) Skill: industry-analysis Industry analysis . . . (52 more lines) R3 âŚdata_equity_comps(ticker=NVDA, peers=[AMD, AVGO, MRVL]) NVDA relative valuation vs 3 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S NVDA30.0028.4918.6018.69 AMD149.28107.4821.32n/a AVGO62.4644.4024.7623.71 MRVL58.4061.5519.1517.50 median62.4661.5521.3220.61 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA NVDA$4.74T15.2074.1%65.6%63.0%85.2%114.3%-0.24x AMD$735.12B32.7253.1%14.4%13.4%37.8%8.1%-1.14x AVGO$1.79T19.2776.3%49.0%38.8%47.9%37.3%1.08x MRVL$152.55B27.2351.5%14.5%29.0%27.6%16.0%0.53x Implied value/share: P/E $407¡ EV/EBITDA $422¡ EV/Rev $225¡ P/S $216 Range: $216 â $422 (median $316)¡ Current price $196¡ median-implied upside 61.6% 41 FinanceHarness: Autonomous Financial Deep Research Framework âŚdata_equity_comps(ticker=TSM, peers=[ASML, AMAT, LRCX, KLAC]) TSM relative valuation vs 4 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S TSM33.684.433.170.45 ASML54.312618.24999.6717.11 AMAT44.2944.1414.1112.90 LRCX49.5946.3416.7715.10 KLAC53.2845.6120.3718.82 median51.4445.9818.5716.10 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA TSM$1.99T17.7664.2%60.3%49.9%36.0%40.0%-0.80x ASML$604.30B27.1252.7%37.1%30.1%21.3%53.9%-0.41x AMAT$374.48B27.9049.0%31.9%29.3%11.4%39.7%-0.10x LRCX$327.45B32.0150.0%35.0%30.9%23.8%66.8%-0.13x KLAC$246.38B36.7661.4%41.2%35.7%11.5%95.0%0.20x Implied value/share: P/E $585¡ EV/EBITDA $28,636¡ EV/Rev $16,391¡ P/S $13,788 Range: $585 â $28,636 (median $15,089)¡ Current price $383¡ median-implied upside 3840.1% âŚdata_equity_comps(ticker=MU, peers=[TSM, AVGO, MRVL]) MU relative valuation vs 3 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S MU18.3014.6211.04n/a TSM33.694.433.170.45 AVGO62.4744.4024.7623.70 MRVL58.4061.5519.1517.50 median58.4044.4019.1517.50 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA MU$913.66B5.2672.6%80.4%55.9%345.7%66.6%-0.29x TSM$1.99T17.7664.2%60.3%49.9%36.0%40.0%-0.80x AVGO$1.79T19.2776.3%49.0%38.8%47.9%37.3%1.08x MRVL$152.55B27.2351.5%14.5%29.0%27.6%16.0%0.53x Implied value/share: P/E $2,581¡ EV/EBITDA $2,699¡ EV/Rev $1,548¡ P/S $1,399 Range: $1,399 â $2,699 (median $2,065)¡ Current price $809¡ median-implied upside 155.3% 42 FinanceHarness: Autonomous Financial Deep Research Framework âŚdata_equity_comps(ticker=SMCI, peers=[DELL, HPE, ANET, CSCO]) SMCI relative valuation vs 4 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S SMCI14.5916.330.770.53 DELL29.5821.122.221.79 HPE41.1314.192.061.50 ANET56.3147.8320.8721.18 CSCO38.1227.517.707.42 median39.6324.324.964.61 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA SMCI$17.93B8.378.4%6.1%3.7%122.7%17.9%4.97x DELL$239.85B16.9919.2%8.9%6.3%87.5%n/a1.45x HPE$58.28B10.9833.8%8.7%4.0%40.0%6.3%2.84x ANET$205.63B36.5663.5%42.7%38.3%35.1%31.5%-2.92x CSCO$450.76B23.8764.3%25.0%19.7%12.0%25.2%0.96x Implied value/share: P/E $75¡ EV/EBITDA $47¡ EV/Rev $246¡ P/S $240 Range: $47 â $246 (median $158)¡ Current price $28¡ median-implied upside 468.5% âŚdata_equity_comps(ticker=COHR, peers=[LITE, ANET, CSCO, MRVL]) COHR relative valuation vs 4 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S COHR110.5641.428.246.85 LITE108.16100.6120.5719.27 ANET56.3147.8320.8721.18 CSCO38.1227.517.707.42 MRVL58.4861.5519.1517.50 median57.4054.6919.8618.39 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA COHR$45.20B27.8737.0%13.6%7.1%20.5%4.7%0.76x LITE$47.96B33.5940.8%21.8%17.7%90.1%22.8%0.28x ANET$205.63B36.5663.5%42.7%38.3%35.1%31.5%-2.92x CSCO$450.76B23.8764.3%25.0%19.7%12.0%25.2%0.96x MRVL$152.74B27.2751.5%14.5%29.0%27.6%16.0%0.53x Implied value/share: P/E $120¡ EV/EBITDA $362¡ EV/Rev $665¡ P/S $620 Range: $120 â $665 (median $491)¡ Current price $231¡ median-implied upside 112.6% 43 FinanceHarness: Autonomous Financial Deep Research Framework âŚdata_equity_comps(ticker=VRT, peers=[ETN, ABB, PWR, . . . ; 5 values]) VRT relative valuation vs 4 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S VRT65.5846.6810.269.22 ETN37.3827.736.17n/a PWR78.2637.283.292.85 EME23.4717.621.841.75 FIX38.1331.565.634.86 median37.7529.654.462.85 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA VRT$100.01B29.4937.2%16.4%14.4%30.1%45.1%0.32x ETN$146.45B23.9137.1%16.1%14.0%16.8%20.8%3.32x PWR$85.96B34.7315.1%4.2%3.7%26.3%13.5%2.24x EME$31.13B21.2319.3%8.7%7.5%19.7%39.2%-0.22x FIX$54.56B25.5625.7%7.9%12.8%1.0%55.3%0.12x Implied value/share: P/E $150¡ EV/EBITDA $182¡ EV/Rev $124¡ P/S $81 Range: $81 â $182 (median $137)¡ Current price $260¡ median-implied upside -47.4% âŚdata_equity_comps(ticker=CEG, peers=[VST, NEE, GEV, . . . ; 5 values]) CEG relative valuation vs 5 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S CEG22.5715.014.003.10 VST24.7611.103.882.57 NEE20.3720.8110.58n/a GEV26.4365.596.235.94 CMI32.7519.472.872.57 CAT41.1730.366.25n/a median26.4320.816.232.57 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA CEG$92.70B19.2923.3%21.9%12.7%63.8%16.1%2.71x VST$49.92B13.6938.6%26.6%11.5%43.4%42.9%2.93x NEE$189.06B20.5361.0%31.5%32.4%12.4%11.7%7.35x GEV$245.54B37.0420.6%7.5%23.0%21.9%82.6%-2.29x CMI$87.04B18.4925.8%9.8%7.9%2.7%22.0%1.01x CAT$381.00B27.1628.6%18.2%13.3%22.2%51.3%2.73x Implied value/share: P/E $304¡ EV/EBITDA $403¡ EV/Rev $460¡ P/S $215 Range: $215 â $460 (median $354)¡ Current price $260¡ median-implied upside 36.2% 44 FinanceHarness: Autonomous Financial Deep Research Framework âŚdata_equity_comps(ticker=EQIX, peers=[DLR, IRM, AMT]) EQIX relative valuation vs 3 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S EQIX72.1829.2513.1510.80 DLR94.9328.8813.7010.89 IRM136.8323.517.955.11 AMT28.5718.2911.817.63 median94.9323.5111.817.63 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA EQIX$102.86B54.0051.6%24.3%14.9%12.1%10.1%5.13x DLR$73.58B66.6757.0%26.1%11.8%29.9%2.9%5.61x IRM$37.05B46.5455.0%21.0%3.8%21.6%n/a7.95x AMT$82.53B25.8374.0%45.9%26.8%6.8%30.0%6.23x Implied value/share: P/E $1,372¡ EV/EBITDA $798¡ EV/Rev $918¡ P/S $737 Range: $737 â $1,372 (median $858)¡ Current price $1,043¡ median-implied upside -17.7% âŚdata_equity_comps(ticker=MSFT, peers=[AMZN, GOOGL, ORCL, META]) MSFT relative valuation vs 4 peer(s): Trading multiples CompanyP/EEV/EBITDAEV/RevP/S MSFT23.1015.939.239.23 AMZN27.4016.563.483.31 GOOGL16.4622.468.729.00 ORCL19.8315.957.224.94 META21.5113.847.046.99 median20.6716.267.135.96 Fundamentals CompanyMktCapFwdP/EGrossOpNetRevGrROENetDebt/EBITDA MSFT$2.94T20.4268.3%46.3%39.3%18.3%34.0%0.26x AMZN$2.46T23.0750.6%13.1%12.2%16.6%24.3%0.59x GOOGL$4.01T22.3060.9%34.0%54.8%24.2%48.7%-0.70x ORCL$332.92B10.6165.8%36.2%25.4%20.6%53.4%4.44x META$1.50T15.9881.9%40.6%32.8%33.1%32.9%0.05x Implied value/share: P/E $354¡ EV/EBITDA $397¡ EV/Rev $299¡ P/S $256 Range: $256 â $397 (median $326)¡ Current price $395¡ median-implied upside -17.4% 45 FinanceHarness: Autonomous Financial Deep Research Framework R4 âŚupdate_plan: . . . (plan unchanged)â Plan updated â 1/4 step(s) done. R5 âŚsearch: [2026 AI data center capex hyperscalers spending forecast 2026 Microsoft Amazon Alphabet Meta official earnings capex, site:investor.nvidia.com annual report 2026 data center revenue Blackwell supply part- ners, site:investor.tsmc.com 2026 quarterly results AI accelerator revenue CAGR CoWoS advanced packaging, site:asml.com investors 2026 AI demand EUV data center, site:investors.micron.com 2026 HBM sold out AI data center demand earnings, site:investors.broadcom.com 2026 AI revenue custom accelerators networking earn- ings, site:investors.arista.com 2026 AI networking revenue earnings, site:investors.vertiv.com 2026 earnings AI data center backlog liquid cooling, site:eaton.com investors 2026 data center demand electrical backlog earnings, site:investor.gevernova.com 2026 data center electricity demand gas turbines backlog, site:constellationenergy.com investors data center AI power agreement nuclear Microsoft 2025 2026, site:iea.org data centres AI electricity de- mand 2030 report, site:investor.equinix.com 2026 AI data center demand xScale earnings, site:investor.digitalrealty. com 2026 AI data center demand leasing backlog, site:investors.delltechnologies.com 2026 AI server backlog rev- enue earnings, site:ir.supermicro.com 2026 AI infrastructure revenue liquid cooling, site:coherent.com investors 2026 datacenter transceiver AI demand earnings, site:lumentum.com investors 2026 cloud optical demand AI earn- ings] (max_results=5) 13 results: Hyperscalers face higher capex scrutiny after Alphabet report panned Alphabet Will Raise Capexâand Extend the AI Semiconductor Cycle Technology: Hyperscaler Capex 2026 Estimates - CreditSights Chart: Big Techâs AI Spending to Reach $725 Billion in 2026 - Statista . . . (9 more) R6 âŚsearch: [TSMC Q2 2026 results AI accelerator revenue advanced packaging CoWoS, Micron fiscal Q3 2026 results HBM data center demand sold out, Broadcom 2026 quarterly results AI revenue custom accelerators net- working, Arista 2026 quarterly results AI networking revenue Etherlink, Vertiv 2026 quarterly results backlog AI data centers liquid cooling, Eaton 2026 quarterly results data center orders backlog, GE Vernova 2026 quarterly results data center power demand gas turbine backlog, IEA Energy and AI report data centres electricity demand 2030, Dell fiscal 2027 results AI server backlog 2026, Super Micro fiscal 2026 results AI server revenue liquid cooling, Coherent 2026 quarterly results datacenter transceiver AI demand, Digital Realty 2026 results AI leasing backlog data center] (max_results=6) 14 results: TSMC Q2 2026 Revenue Sets a Record | What It Means for AI TSM Earnings Q2 2026: Results, Analysis & Outlook ... TSMC 2026 Q2 Quarterly Results - Taiwan Semiconductor ... TSMC Posts All-Time Revenue Record â June Sales Up 68%, N3 ... . . . (10 more) R7 âŚsearch:[site:investors.micron.com/news-releases 2026 quarterly results HBM revenue, site:investors. broadcom.com/news-releases 2026 financial results AI revenue, site:investors.arista.com/news-releases 2026 finan- cial results AI networking, site:investors.vertiv.com/news-releases 2026 financial results organic orders backlog, site:eaton.com/us/en-us/company/investor-relations 2026 earnings data center orders, site:investor.gevernova. com/news-releases 2026 financial results backlog, site:investors.delltechnologies.com/news-releases 2026 financial results AI servers backlog] (max_results=8) 46 results: Micron in High-Volume Production of HBM4 Designed for NVIDIA ... Micron Powers AI Everywhere at COMPUTEX 2026 | Micron Technology Micron Technology, Inc. Reports Results for the Second Quarter of... Micron Ships HBM4 to Key Customers to Power Next-Gen AI Platforms . . . (42 more) 46 FinanceHarness: Autonomous Financial Deep Research Framework R8 âŚsearch: [Arista Networks Q2 2026 financial results AI Ethernet, Vertiv Q2 2026 financial results orders back- log AI, GE Vernova Q2 2026 financial results gas power backlog data center, Eaton second quarter 2026 earnings data center orders, Constellation Energy Q2 2026 data center power agreements, Digital Realty Q2 2026 earn- ings AI demand bookings, Equinix Q2 2026 earnings AI xScale, Coherent fiscal 2026 results datacenter growth transceivers] (max_results=8) 11 results: Arista Networks (ANET) Investor Relations, Earnings Summary ... What to Expect From Arista Networksâ Q2 2026 Earnings Report Arista Networks Grew Revenue 35% Last Quarter and Generates $4 ... Arista Networks (ANET) Strengthens from Exponential Growth in AI . . . (7 more) R9 âŚsearch: [site:investors.vertiv.com "Second Quarter 2026", site:investors.vertiv.com "First Quarter 2026" orders organic sales backlog, site:investor.gevernova.com "Second Quarter 2026", site:investor.gevernova.com "First Quar- ter 2026" backlog, site:eaton.com "Second Quarter 2026" earnings data center, site:eaton.com "First Quarter 2026" data center orders, site:investor.digitalrealty.com "Second Quarter 2026", site:investor.equinix.com "Second Quarter 2026"] (max_results=8) 6 results: Vertiv Holdings Co. - Vertiv Reports Strong First Quarter ... Vertiv Holdings Co. - Investor relations Vertiv Announces Date of Second Quarter 2026 Earnings Release ... Vertiv Holdings Co. - Vertiv Second Quarter 2026 Earnings ... . . . (2 more) R10 âŚsearch: [site:iea.org/reports/energy-and-ai executive summary data centres electricity 2030, site:iea.org "En- ergy and AI" data centres 2030 electricity demand, site:iea.org data centers electricity demand United States 2030 AI] (max_results=6) 15 results: Executive summary â Energy and AI â Analysis - IEA Energy demand from AI â Energy and AI â Analysis - IEA Energy supply for AI â Energy and AI â Analysis - IEA AI and climate change â Energy and AI â Analysis - IEA . . . (11 more) R11 âŚsearch: [ASML 2025 annual report AI demand EUV systems revenue 2026 outlook, Applied Materials fis- cal 2026 results AI semiconductor equipment advanced packaging, Lam Research 2026 results AI HBM advanced packaging demand, KLA 2026 earnings AI advanced packaging process control demand, TSMC annual report 2025 advanced technologies 7nm below revenue AI accelerator CAGR] (max_results=6) 14 results: Financials | ASML 2025 Annual Report ASML reportsâŹ32.7 billion total net sales andâŹ9.6 billion ... ASML reportsâŹ32.7 billion total net sales andâŹ9.6 billion ... ASML 2025: Growth Through Structural Shifts - by Meng Li . . . (10 more) R12 âŚsearch: [Quanta Services 2026 results data center load growth transmission demand, EMCOR 2026 earnings data center demand backlog, Comfort Systems USA 2026 earnings data center revenue backlog, GE Vernova 2025 annual report data center power demand gas turbine backlog, Constellation Microsoft Three Mile Island power purchase agreement data center official, Vistra data center power demand investor presentation 2026, Digital Realty 47 FinanceHarness: Autonomous Financial Deep Research Framework annual report 2025 AI demand development pipeline] (max_results=8) 12 results: Is Quantaâs 110% Tech Growth Outlook Fueling Its ... - Yahoo Finance QUANTA SERVICES REPORTS FIRST QUARTER 2026 RESULTS 02/19/2026 - Quanta Services, Inc. Quanta Services Stock Is Surging. The Backlog Says Itâs Not Done. . . . (8 more) R13 âŚsearch:[site:investor.digitalrealty.com/news-releases 2026 first quarter results bookings development pipeline, site:investor.equinix.com/news-events 2026 first quarter results xScale, site:ir.coreweave.com 2026 first quarter results backlog debt capex, site:investor.constellationenergy.com 2026 first quarter results data center, site: investor.vistracorp.com 2026 first quarter results data center] (max_results=8) 11 results: Equinix Reports First-Quarter Results and... :: Equinix, Inc. (EQIX) Equinix Provides Robust 2026 Outlook Driven by... :: Equinix, Inc. (EQIX) Equinix Reports First-Quarter 2025 Results :: Equinix, Inc. (EQIX) Equinix Reports Strong Third-Quarter 2025 Results . . . (7 more) R14 âŚsearch: [Alphabet Q2 2026 results capex guidance 2026 investor relations, Amazon Q1 2026 earnings capex 200 billion official, Meta Q2 2026 earnings capex guidance 2026 investor relations, Microsoft fiscal Q3 2026 earn- ings capex AI infrastructure official, Oracle fiscal 2026 results remaining performance obligations AI infrastruc- ture] (max_results=8) 14 results: Alphabet Investor Relations Google Q2 2026 Earnings: Cloud Up 82%, $205B AI Capex Guide Alphabet lifts capex guide to $205bn as Google Cloud jumps 82% Prediction: CEO Andy Jassy Will Raise Amazonâs Full-Year 2026... . . . (10 more) R15 âŚsearch: [Constellation Microsoft 20-year power purchase agreement Crane Clean Energy Center official September 2024, GE Vernova data center demand gas turbines backlog 2025 annual report] (max_results=8) 13 results: Constellation to restart Three Mile Island unit, powering Microsoft Powering the AI Revolution: Energy Infrastructure in the United States The AI Boomâs Real Bottleneck: Why Nuclear Energy and... Constellation to Launch Crane Clean Energy Center, Restoring Jobs... . . . (9 more) R16 âŚvisit: investor.nvidia.com, investor.tsmc.com, investors.micron.com, investors.broadcom.com, asml.com, investors.delltechnologies.com, investors.vertiv.com, investors.quantaservices.com, iea.org, investor.equinix.com, abc.xyz, cnbc.com (goal=Extract current reported financial results, AI/data-center demand linkage, capacity/ backlog, outlook, margin and capex figures, plus risks and caveats.) R17 âŚsearch: ["Micron Technology, Inc. Reports Record Results for the Third Quarter of Fiscal 2026", "Broadcom Inc. Announces Second Quarter Fiscal Year 2026 Financial Results" AI revenue, "Dell Technologies Delivers Fourth Quarter and Full-Year Fiscal 2026 Financial Results" AI servers, Arista first quarter 2026 results revenue AI network- 48 FinanceHarness: Autonomous Financial Deep Research Framework ing official] (max_results=8) 11 results: Micron Technology, Inc. Reports Record Results for the Third ... Micron Technology, Inc. Reports Record Results for the Third ... Micron Technology, Inc. Reports Record Results for the Third ... Micron Technology, Inc. Reports Record Results for the Third ... . . . (7 more) R18 âŚsearch: [Broadcom Q2 fiscal 2026 results AI semiconductor revenue 2026, Dell fiscal 2026 full year AI server revenue backlog February 2026, Arista Q1 2026 financial results AI revenue] (max_results=10) 13 results: Broadcom Earnings: AI Revenue Hits $10.8B [2026] Broadcom (AVGO) Q2 FY2026 Earnings: $22.19B Revenue Broadcom (AVGO) earnings report Q2 2026 - CNBC Broadcom Inc. (AVGO) Q2 Fiscal 2026 Earnings Recap: AI Demand ... . . . (9 more) R19 âŚsearch: [Dell fiscal Q4 2026 AI server revenue backlog February 26 2026, Dell full year fiscal 2026 AI opti- mized server shipments backlog 2026, Arista first quarter 2026 results May 2026 revenue AI networking official release] (max_results=10) 12 results: Dell Technologies Delivers Fourth Quarter and Full-Year ... AI server frenzy fuels record revenues for Dell Dell Technologies Inc (DELL) Q4 2026 Earnings Call Highlights ... DELL Q4 2026 Earnings Report on 2/26/2026 - MarketBeat . . . (8 more) R20 âŚsearch:[Coherent fiscal third quarter 2026 results datacenter transceiver AI official, Lumentum fis- cal third quarter 2026 results cloud data center AI official, Marvell fiscal 2026 results data center revenue AI] (max_results=10) 12 results: Coherent Corp. Reports Third Quarter Fiscal 2026 Results COHR Q3 2026 Earnings Report on 5/6/2026 - MarketBeat Coherent Q3 FY 2026: AI Data Center Demand Accelerates ... [COHR Q3 2026 Earnings Call] Record Revenue of $1.8B, AI ... . . . (8 more) R21 âŚvisit: nasdaq.com, cnbc.com, dell.com, coherent.com, investors.arista.com, iea.org (goal=Extract company results, growth and margins, AI/data-center linkage, orders/backlog/capacity, outlook and risks.) R22 âŚsearch: [TSMC Q2 2026 earnings call AI accelerator revenue CAGR CoWoS capacity 2026 transcript, TSMC 2026 Q2 AI demand CoWoS capacity Reuters July 2026, TSMC 2025 annual report AI accelerators revenue advanced packaging CoWoS] (max_results=8) 12 results: TSMC 2026 Q2 Quarterly Results - Taiwan Semiconductor ... TSM (TSM) Q2 2026 Earnings Call Transcript - The Motley Fool Q2 2026 Taiwan Semiconductor Manufacturing Co Ltd Earnings ... TSMC Posts Record $35.6 Billion Q1 Revenue as AI Chip Demand... . . . (8 more) 49 FinanceHarness: Autonomous Financial Deep Research Framework R23 âŚvisit: investor.tsmc.com (goal=Extract managementâs latest AI demand outlook, AI accelerator CAGR, advanced-node and CoWoS capacity, capex, margins, pricing, risks, and geographic expansion costs.) R24 âŚupdate_plan: . . . (plan unchanged)â Plan updated â 3/4 step(s) done. R25 âŚcalc(expression=75.2/81.615*100)â 92.1399 R26 âŚupdate_plan: . . . (plan unchanged)â Plan updated â 4/4 step(s) done. R27 âŚsearch: [site:eaton.com data center AI power management Eaton 2026, site:arista.com AI networking Eth- ernet data center 2026 Arista, site:investor.constellationenergy.com data center power AI 2026 Constellation, site:skhynix.com HBM AI data center 2026, site:samsung.com semiconductor HBM AI 2026] (max_results=5) 8 results: Data centers | Digitalization | Efficiency | Eaton Center for Intelligent Power Data centers | Digitalization | Efficiency | Eaton Arista Networks - Arista Introduces Next-Generation 1.6Terabit ... . . . (4 more) R28 âŚvisit: electrical.eaton.com, investors.arista.com, news.skhynix.com (goal=Confirm direct AI-data-center ex- posure, product role, demand outlook, and relevant risks.) 50 FinanceHarness: Autonomous Financial Deep Research Framework E.2. Valuation Compute Seam Microsoft (MSFT): DCF, WACC & Sensitivity R E S E A R C H T R A J E C T O RY R1 âŚupdate_plan: [in progress] Collect Microsoft FY2025 reported financials and filing context; [pending] Pull current market inputs and peer betas; derive WACC; [pending] Set operating assumptions and build five-year unlevered FCF forecast; [pending] Value terminal cash flows and bridge enterprise value to equity value; [pend- ing] Run WACC/perpetual-growth sensitivity and compare with market price; [pending] Write the auditable DCF report with sources and assumptionsâ Plan updated â 0/6 step(s) done. R2 âŚload_skill(dcf-valuation) Skill: dcf-valuation DCF valuation . . . (32 more lines) âŚload_tool(data_equity_reference, data_equity_fundamentals, data_market_rates, compute_valuation_wacc, compute_valuation_dcf, compute_valuation_dcf_sensitivity) data_equity_reference (data.equity.reference) Reference data for a public equity: name, sector/industry, country, exchange, market cap, shares outstanding, beta, current price, 52-week range, and P/E (snake_case keys â chain prev:<id>.beta into wacc and prev: <id>.shares_outstanding into dcf). Use to identify and classify a ticker (see data.equity.ratios for full multi- ples incl. dividend yield). Arguments schema: . . . (373 more lines) âŚsearch: [Microsoft 2025 10-K revenue operating income depreciation amortization capital expenditures working capital cash taxes FY2025, site:microsoft.com/en-us/Investor earnings FY 2025 Microsoft annual re- port] (max_results=6) 6 results: Microsoft 2025 Annual Report FY25 Q4 - Press Releases - Investor Relations - Microsoft MSFT - Microsoft Corp Financials - Stock Titan FY25 Q1 - Press Releases - Investor Relations - Microsoft . . . (2 more) R3 âŚvisit: microsoft.com, microsoft.com (goal=Extract FY2025 reported revenue, operating income, tax provision/ effective tax rate, D&A, capex (additions to property and equipment), working-capital balance-sheet items, operat- ing cash flow, cash and investments, debt, diluted shares, fiscal year end, and management comments relevant to revenue growth/margins/capital spending.) âŚdata_equity_fundamentals(ticker=MSFT) MSFT fundamentals (TTM): Revenue $318.27B¡ EBITDA $184.46B¡ FCF $37.01B¡ OCF $170.14B Margins â gross 68.3%, op 46.3%, net 39.3%¡ ROE 34.0% Cash $78.23B¡ Debt $125.43B¡ Rev growth 18.3%¡ EPS ttm/fwd 16.79/19.38 Income trend (revenue / op income / net income): 2025-06-30: $281.72B / $128.53B / $101.83B 2024-06-30: $245.12B / $109.43B / $88.14B 2023-06-30: $211.91B / $88.52B / $72.36B 2022-06-30: $198.27B / $83.38B / $72.74B 51 FinanceHarness: Autonomous Financial Deep Research Framework âŚdata_equity_reference(ticker=MSFT) Microsoft Corporation (MSFT) â NMS¡ USD Sector: Technology / Software - Infrastructure¡ Country: United States Market cap: $2.84T¡ Shares out: 7.43B¡ Beta: 1.13 Price: 381.70¡ 52w: 349.20â555.45¡ P/E (ttm/fwd): 22.73/19.70 Microsoft Corporation develops and supports software, services, devices, and solutions worldwide. . . . âŚdata_equity_reference(ticker=ORCL) Oracle Corporation (ORCL) â NYQ¡ USD Sector: Technology / Software - Infrastructure¡ Country: United States Market cap: $331.23B¡ Shares out: 2.88B¡ Beta: 1.71 Price: 114.99¡ 52w: 114.75â345.72¡ P/E (ttm/fwd): 19.72/10.56 Oracle Corporation offers products and services that build, run and support enterprise information technology frameworks worldwide. . . . âŚdata_equity_reference(ticker=GOOGL) Alphabet Inc. (GOOGL) â NMS¡ USD Sector: Communication Services / Internet Content & Information¡ Country: United States Market cap: $3.91T¡ Shares out: 5.87B¡ Beta: 1.25 Price: 319.74¡ 52w: 187.82â408.61¡ P/E (ttm/fwd): 16.05/21.63 Alphabet Inc. offers various products and platforms in the United States, Europe, the Middle East, Africa, the Asia-Pacific, Canada, and Latin America. . . . âŚdata_equity_reference(ticker=AMZN) Amazon.com, Inc. (AMZN) â NMS¡ USD Sector: Consumer Cyclical / Internet Retail¡ Country: United States Market cap: $2.50T¡ Shares out: 10.76B¡ Beta: 1.46 Price: 232.11¡ 52w: 196.00â278.56¡ P/E (ttm/fwd): 27.76/23.41 Amazon.com, Inc. engages in the retail sale of consumer products, advertising, and subscriptions service through online and physical stores in North America and internationally. . . . âŚdata_equity_reference(ticker=AAPL) Apple Inc. (AAPL) â NMS¡ USD Sector: Technology / Consumer Electronics¡ Country: United States Market cap: $4.89T¡ Shares out: 14.69B¡ Beta: 1.10 Price: 333.02¡ 52w: 201.50â334.99¡ P/E (ttm/fwd): 40.32/34.54 Apple Inc. designs, manufactures, and markets smartphones, personal computers, tablets, wearables, and acces- sories worldwide. . . . âŚdata_market_rates() US Treasury yields: 3m 3.81%¡ 5y 4.43%¡ 10y 4.68%¡ 30y 5.16% Risk-free (10y) = 4.68%â 0.0468 for WACC. R4 âŚdata_equity_reference(ticker=CRM) Salesforce, Inc. (CRM) â NYQ¡ USD Sector: Technology / Software - Application¡ Country: United States Market cap: $134.04B¡ Shares out: 819.00M¡ Beta: 1.18 Price: 163.66¡ 52w: 146.32â274.00¡ P/E (ttm/fwd): 18.99/10.55 Salesforce, Inc. provides customer relationship management technology services that connect companies and customers together in the United States, Europe, and the Asia Pacific. . . . âŚdata_equity_reference(ticker=ADBE) Adobe Inc. (ADBE) â NMS¡ USD Sector: Technology / Software - Application¡ Country: United States Market cap: $89.48B¡ Shares out: 397.50M¡ Beta: 1.43 Price: 225.11¡ 52w: 190.12â376.16¡ P/E (ttm/fwd): 12.87/8.19 Adobe Inc. operates as a technology company worldwide. . . . 52 FinanceHarness: Autonomous Financial Deep Research Framework âŚdata_equity_reference(ticker=IBM) International Business Machines Corporation (IBM) â NYQ¡ USD Sector: Technology / Information Technology Services¡ Country: United States Market cap: $201.80B¡ Shares out: 942.13M¡ Beta: 0.68 Price: 214.19¡ 52w: 199.19â332.46¡ P/E (ttm/fwd): 19.01/16.21 International Business Machines Corporation, together with its subsidiaries, provides integrated solutions and services in the Americas, Europe, the Middle East, Africa, and the Asia Pacific. . . . âŚsearch: [site:microsoft.com/investor/reports/ar25 FY2025 Microsoft accounts receivable 2024 unearned revenue accounts payable 2024, Microsoft FY2025 depreciation amortization 34,153 capital expenditures 64,551 changes in working capital annual report, Microsoft 2025 10-k cash flow changes in operating assets liabilities accounts receivable unearned revenue] (max_results=8) 12 results: Microsoft Annual Report 2025 ANNUAL Home Page - Microsoft Investor Relations EBITDA (Earnings Before Interest, Taxes, Depreciation, and...) . . . (8 more) R5 âŚvisit: microsoft.com, microsoft.com, microsoft.com (goal=Extract FY2025 and FY2024 cash-flow statement details, especially D&A, capex/additions to P&E, cash flow changes in accounts receivable, inventories, other current assets, accounts payable, unearned revenue, other current liabilities; plus balance sheet cash, short-term investments, debt and diluted weighted average shares.) R6 âŚcalc(expression=128528/281724)â 0.45622 âŚcalc(expression=21795/123627)â 0.176296 âŚcalc(expression=34153/281724)â 0.121229 âŚcalc(expression=64551/281724)â 0.229129 R7 âŚcompute_valuation_wacc(risk_free_rate=0.0468,equity_risk_premium=0.05,beta=1.406, cost_of_debt=0.0518, tax_rate=0.176296, debt_to_equity=0.015194) Cost of capital (CAPM + WACC) Re = Rf + b*ERP = 4.7% + 1.41*5.0% = 11.7% Implied ERP (Re-Rf = b*ERP): 7.0% After-tax Rd = 5.2%*(1-17.6%) = 4.3% Weights (D/E=0.015): E/V 98.5%, D/V 1.5% WACC = 11.6%¡ premium over Rf 6.9% 53 FinanceHarness: Autonomous Financial Deep Research Framework R8 âŚcalc(expression=0.0468+1.406*0.05)â 0.1171 âŚcalc(expression=(0.015194/(1+0.015194)))â 0.0149666 âŚcalc(expression=1/(1+0.015194))â 0.985033 R9 âŚcalc(expression=281.724*1.14)â 321.165 âŚcalc(expression=281.724*1.14*1.13)â 362.917 âŚcalc(expression=281.724*1.14*0.457)â 146.773 âŚcalc(expression=281.724*1.14*0.135)â 43.3573 R10 âŚcompute_valuation_dcf(fcf_schedule=[91.4768, 114.964, 139.717, . . . ; 5 values], discount_rate=0.115986, ter- minal_method=gordon_growth, terminal_growth_rate=0.025, net_debt=-66.819, shares_outstanding=7.465) DCF valuation Discount rate 11.6%¡ 5y explicit¡ Gordon growth (g=2.5%, FCF_N+1=$187) Sum PV explicit FCF: $485 Terminal value (yr N): $2,058¡ PV: $1,189 Enterprise value: $1,674¡ net debt -$67 Equity value: $1,741 Intrinsic value/share: $233 Terminal value is 71.0% of EV. âŚcompute_valuation_dcf_sensitivity(fcf_schedule=[91.4768,114.964,139.717,. . . ;5values], discount_rates=[0.095986,0.105986,0.115986,. . . ;5 values],terminal_method=gordon_growth,termi- nal_growth_rates=[0.015, 0.02, 0.025, . . . ; 5 values], net_debt=-66.819, shares_outstanding=7.465) DCF sensitivity (Gordon growth) â intrinsic value/share WACC \ g1.5%2.0%2.5%3.0%3.5% 9.6%$272$285$301$319$340 10.6%$241$251$263$276$291 11.6%$216$224$233$243$255 12.6%$196$202$209$217$226 13.6%$179$184$190$196$203 Range (intrinsic value/share): $179 â $340 54