Paper deep dive
Vibe-FDTR: An agent-oriented framework for reproducible frequency-domain thermoreflectance data analysis
Fuwei Yang, Weiheng Li, Bai Song
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 88%
Last extracted: 8/3/2026, 9:51:13 AM
Summary
The paper introduces Vibe-FDTR, an agent-oriented framework that utilizes Large Language Model (LLM) agents to perform reproducible frequency-domain thermoreflectance (FDTR) data analysis. The framework integrates a configuration-driven code package with procedural agent skills to translate natural language requests into executable analysis steps. Evaluated on synthetic and real-data benchmarks, Vibe-FDTR achieves near-perfect success rates (100% and 98.9%) while significantly reducing computational cost and execution time compared to baseline agent-only or code-only variants. It also includes an expert mode for autonomous experimental planning and uncertainty analysis.
Entities (8)
Relation Signals (6)
Vibe-FDTR â uses â LLM
confidence 95% · Vibe-FDTR... enables large language model (LLM) agents to perform reliable and reproducible FDTR analyses
Vibe-FDTR â outperforms â Code-agent
confidence 90% · Vibe-FDTR also reduces computational cost by 87.7% relative to the Code-agent variant
Vibe-FDTR â outperforms â Agent-only
confidence 90% · drops further to 38.6% and 0% when the domain package is also omitted (Agent-only)
FDTR â usedfor â Thermal property measurement
confidence 90% · FDTR is a laser pump-probe technique widely used to measure thermal properties at the micro- and nanoscale
Vibe-FDTR â evaluatedon â Gold-coated graphite
confidence 85% · real-data multi-step tasks based on measurements of gold-coated graphite samples
Vibe-FDTR â supports â Expert mode
confidence 85% · an optional expert mode supports experimental planning via autonomous sensitivity and uncertainty evaluations
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Frequency-domain thermoreflectance (FDTR) is a laser pump-probe technique widely used to measure thermal properties at the micro- and nanoscale; however, it relies on a complex data analysis procedure that demands substantial domain expertise and is susceptible to subtle human errors. Here, we present Vibe-FDTR, an agent-oriented framework that enables large language model (LLM) agents to perform reliable and reproducible FDTR analyses directly from natural language requests. This framework couples a configuration-driven FDTR code package, which enforces physical and parametric consistency, with procedural agent skills that translate user intentions into organized and verifiable analysis steps. We evaluate Vibe-FDTR using a controlled benchmark with two levels: synthetic single-step tasks and real-data multi-step tasks based on measurements of gold-coated graphite samples. Across the two levels, agents using Vibe-FDTR achieve success rates of 100% and 98.9%, respectively. In sharp contrast, ablating skills (Code-agent) reduces performance to 91.4% and 36.7%, which drops further to 38.6% and 0% when the domain package is also omitted (Agent-only). Beyond success rate, Vibe-FDTR also reduces computational cost by 87.7% relative to the Code-agent variant and cuts execution time by more than 60%. Finally, an optional expert mode supports experimental planning via autonomous sensitivity and uncertainty evaluations, and formulates physically grounded recommendations for underspecified tasks. These results demonstrate that encapsulating domain code and expert knowledge into agent skills offers a promising route toward low-barrier, autonomous, and trustworthy thermal metrology.
Tags
Links
- Source: https://arxiv.org/abs/2607.28200v1
- Canonical: https://arxiv.org/abs/2607.28200v1
Trouble viewing inline? Open PDF directly â
Full Text
62,559 characters extracted from source content.
Expand or collapse full text
Vibe-FDTR: An agent-oriented framework for reproducible frequency-domain thermoreflectance data analysis Fuwei Yang 1,2,3* Weiheng Li 1,4* Bai Song 1,4â 1 National Key Laboratory of Advanced Micro and Nano Manufacture Technology, Peking University, Beijing 100871, China 2 Center for Nano and Micro Mechanics, Tsinghua University, Beijing 100084, China 3 Department of Engineering Mechanics, Tsinghua University, Beijing 100084, China 4 School of Mechanics and Engineering Science, Peking University, Beijing 100871, China * These authors contributed equally to this work. â songbai@pku.edu.cn Abstract Frequency-domain thermoreflectance (FDTR) is a laser pump-probe technique widely used to measure thermal properties at the micro- and nanoscale; however, it relies on a complex data analysis procedure that demands substantial domain expertise and is susceptible to subtle human errors. Here, we present Vibe-FDTR, an agent-oriented framework that enables large language model (LLM) agents to perform reliable and reproducible FDTR analyses directly from natural language re- quests. This framework couples a configuration-driven FDTR code package, which enforces physical and parametric consistency, with procedural agent skills that trans- late user intentions into organized and verifiable analysis steps. We evaluate Vibe- FDTR using a controlled benchmark with two levels: synthetic single-step tasks and real-data multi-step tasks based on measurements of gold-coated graphite samples. Across the two levels, agents using Vibe-FDTR achieve success rates of 100% and 98.9%, respectively. In sharp contrast, ablating skills (Code-agent) reduces perfor- mance to 91.4% and 36.7%, which drops further to 38.6% and 0% when the domain package is also omitted (Agent-only). Beyond success rate, Vibe-FDTR also reduces computational cost by 87.7% relative to the Code-agent variant and cuts execution time by more than 60%. Finally, an optional expert mode supports experimen- tal planning via autonomous sensitivity and uncertainty evaluations, and formu- lates physically grounded recommendations for underspecified tasks. These results demonstrate that encapsulating domain code and expert knowledge into agent skills offers a promising route toward low-barrier, autonomous, and trustworthy thermal metrology. Keywords: Thermal measurement, frequency-domain thermoreflectance, data anal- ysis, LLM agent, agent-oriented framework 1 arXiv:2607.28200v1 [physics.app-ph] 30 Jul 2026 1 Introduction Accurate measurement of thermal properties at the micro- and nanoscale is essential for understanding heat transport and guiding the thermal design of modern electronic and energy systems 1,2,3,4,5 . To probe thermal transport in samples of diverse sizes, struc- tures, and dimensions, a variety of experimental approaches have been developed, in- cluding the 3Ï method 6,7 , scanning thermal microscopy (SThM) 8,9 , thermal bridge 10,11,12 , opto-thermal Raman spectroscopy 13,14 , and all sorts of laser pump-probe techniques 15,16 . Among these approaches, time-domain thermoreflectance (TDTR) and frequency-domain thermoreflectance (FDTR) are widely recognized for their reliability, versatility, and rel- ative ease of use 17,18 . By using a modulated pump laser to generate a localized thermal response and a probe laser to monitor the corresponding surface thermoreflectance varia- tion, TDTR and FDTR can resolve heat transport over fast time scales and small length scales with high sensitivity, enabling the precise determination of thermal conductivity, specific heat, and interfacial thermal conductance. Furthermore, these techniques are inherently non-contact and eliminate the need for integrating electrical heaters or ther- mometers onto the sample 7,11 . In light of these advantages, TDTR and FDTR have been extensively employed to characterize both bulk materials and thin films, with thermal conductivity values spanning five orders of magnitude from ultralow (âŒ0.01 W m â1 K â1 ) to ultrahigh (>1000 W m â1 K â1 ) 19,20,21,22,23,24 . In addition, TDTR and FDTR have also been expanded to buried interfaces and microfabricated structures, and combined with in situ experiments involving external fields or mechanical manipulations 25,26,27,28,29 . Despite the widespread application of TDTR and FDTR, a complex and highly de- manding data-processing workflow is required to accurately extract the target thermal properties. Unlike the thermal bridge method where thermal conductivity can be readily derived via Fourierâs law from the applied heat flux and measured temperature differ- ence 7,11 , thermoreflectance measurements rely on an indirect model fitting procedure 30 . This nonlinear inverse problem is inherently limited by parameter sensitivities and corre- lations, therefore a single signal or fitting window often provides insufficient identifiability for extracting multiple thermal properties. To obtain reliable results, complementary in- formation from different signals may be required, such as those measured with different laser spot sizes or pump-probe offset distances 31,32,28 . In practice, a complete thermore- flectance data analysis is rarely a single fitting operation, but rather a sequence of coupled decisions involving data organization, sensitivity evaluation, fitting-parameter selection, and uncertainty assessment. Such a workflow is tedious and inefficient when performed manually, especially for large datasets or iterative analyses, and is susceptible to subtle human errors. Minor mistakes may still lead to apparently successful code execution, but ultimately produce physically invalid results. 2 Very recently, the synergistic development of various large language models (LLMs) and harnesses has enabled artificial intelligence (AI) to not only serve as isolated prediction and acceleration tools, but also act as agents that proactively execute workflows. Leading open-source LLMs such as DeepSeek-V4 33 have focused increasingly on agentic capabili- ties. By enabling AI to interpret user intent and translate it into concrete actionsâsuch as configuration writing, tool calling, and data managementâthis progress has the potential to significantly improve the efficiency of scientific research. In fact, such a paradigm shift has already been preliminarily demonstrated in a variety of settings, including chemistry agents equipped with domain tools and LLM-guided experimental planning 34 , closed-loop materials synthesis 35 , automated atomic force microscopy 36 , and AI-driven scanning probe microscopy 37 . Of central importance is the establishment of agentic frameworks that use procedural guidance to translate high-level scientific intent into traceable tool invocations and verifiable outputs. In thermoreflectance metrology, however, previous efforts with AI have mainly focused on solving data-driven inverse problems, including neural networks for inferring thermal parameters from TDTR signals 38 , reconstructing depth-dependent thermal conductivity from FDTR responses 39 , and evaluating interfacial bond quality from FDTR phase maps 40 . To the best of our knowledge, no agent-oriented framework has been reported to reliably and reproducibly perform TDTR/FDTR analysis with nat- ural language requests, which can minimize the need for expert knowledge and largely reduce operational burden associated with thermoreflectance metrology. Here, we take FDTR as an example and present Vibe-FDTR, an agent-oriented frame- work inspired by recent vibe-coding practices that translate natural language requests into executable code through LLM-based agents 41,42,43 . Vibe-FDTR adapts this idea to thermal metrology by combining domain-specific procedural skills and configuration- driven FDTR code package. The framework is designed primarily to reliably execute well-specified post-measurement tasks, while an optional expert mode adds an additional skill layer for FDTR experimental planning and for analysing underspecified problems. To quantitatively evaluate agent performance, we employ a controlled two-level benchmark comprising both single-operation tasks on synthetic data and multi-step tasks on real data obtained from gold-coated graphite samples. Through repeated runs, the high success rates, low computational cost, and short execution time of Vibe-FDTR are systemati- cally quantified and highlighted. For comparison, two reference scenarios are considered, one with the skills ablated (Code-agent) and the other with the domain package further omitted (Agent-only). Finally, we demonstrate the expert mode as an intelligent copi- lot to autonomously perform sensitivity and uncertainty analyses, and provide physically grounded recommendations regarding a series of partially specified tasks. 3 2 Principles of FDTR measurement and data analysis 2.1 FDTR experimental setup We first briefly introduce the FDTR platform and basic measurement principle using our experimental setup as a representative example 24,23,28 . As shown in Fig. 1(a), the plat- form employs a two-laser pump-probe configuration, in which a 405 nm continuous-wave (CW) pump laser is intensity-modulated by the reference output of a lock-in amplifier and focused onto the metal-coated sample surface to generate a periodic temperature field. A 532 nm CW probe laser is aligned to the heated region and monitors the resulting temper- ature oscillation through the temperature-dependent reflectance of the metal transducer (typically gold). Both the pump and probe beams feature a Gaussian cross-sectional inten- sity profile. The reflected probe beam is converted into an electric current via a balanced photodetector to improve the signal-to-noise ratio. Subsequently, the lock-in amplifier is used to resolve the amplitude and phase signals relative to the pump modulation. In the most commonly used frequency-sweep (f-sweep) mode, the modulation frequency is var- ied to obtain the corresponding thermal response, while beam-offset measurements scan the lateral pump-probe separation at fixed f, which is superior for probing in-plane heat spreading and for characterizing the spot size. The schematics of the two measurement modes are demonstrated in the insets of Fig. 1(b) and (c). Sample positioning is achieved using a piezoelectric stage, and temperature-dependent measurements are performed by integrating a heating/cooling stage (INSTEC HCP421V in this work) on top of it. 2.2 Thermal model for fitting After obtaining the amplitude and phase signals, thermal properties of the sample are extracted by comparing the measured response with a multilayer Fourier heat diffusion model 21 . With an appropriately small pump power, the thermoreflectance signal of the metal transducer remains linear with respect to the surface temperature oscillation. The absorbed pump power is commonly approximated as a modulated Gaussian heat source at the transducer surface. The current FDTR code package supports isotropic and trans- versely isotropic thermal models, and more complex models can be readily incorporated in future implementations when needed. For a transversely isotropic layer, the heat conduction equation is given by Îș i â 2 T âr 2 + 1 r âT âr + Îș o â 2 T âz 2 = C âT ât ,(1) where Îș i and Îș o are the in-plane and out-of-plane thermal conductivities, respectively, and C is the volumetric heat capacity. After transformation into the frequency domain and Hankel space, the temperature and heat flux at the two surfaces of each layer can be 4 related through a transfer matrix as Î b Q b ! =M Î t Q t ! ,(2) where Î and Q denote the transformed temperature and heat flux, respectively; and the subscripts t and b refer to the top and bottom surfaces of the layer, respectively. For a homogeneous layer of thickness d, the matrix is M = " cosh(ηd)â 1 Îș o η sinh(ηd) âÎș o η sinh(ηd)cosh(ηd) # ,(3) η 2 = Îș i ÎČ 2 + 2ÏifC Îș o ,(4) where ÎČ is the Hankel transform variable and f is the modulation frequency. An interface with thermal conductance G is described by M int = " 1 1/G 01 # .(5) The complete multilayer response is obtained by multiplying the transfer matrices of all layers and interfaces. Denoting the total transfer matrix as M tot = " A M B M C M D M # ,(6) and applying the bottom boundary condition, the complex surface temperature response can be calculated and averaged over the probe-beam intensity profile as H(f,x 0 ) = Z 2 Ïw 2 1 exp â 2 ((xâ x 0 ) 2 + y 2 ) w 2 1 Ă P 0 2Ï Z â 0 ÎČJ 0 ÎČ p x 2 + y 2 â D M C M exp â ÎČ 2 w 2 0 8 dÎČ dxdy. (7) Here, P 0 is the absorbed pump power, w 0 and w 1 are the pump and probe beam radii, respectively, J 0 is the zeroth-order Bessel function, and x 0 is the lateral pump-probe offset. The modeled complex response H(f,x 0 ) provides the calculated FDTR amplitude and phase. Thermal properties are then obtained by adjusting selected model parameters, such as thermal conductivity, heat capacity, or interfacial thermal conductance, until the modeled response best matches the measured frequency-sweep and/or beam-offset data. Independently calibrated quantities, such as transducer thickness and laser spot size, are usually fixed during fitting. 5 2.3 Sensitivity and uncertainty analysis Because FDTR extracts thermal properties through nonlinear fitting, sensitivity analysis is employed to assess whether a given set of parameters can be reliably fitted from the selected signal and data range. For an observable y, such as the measured amplitude or phase, the normalized sensitivity to a parameter x is defined as 22 S x = â lny â lnx .(8) This quantity describes the relative change in the calculated signal caused by a small variation in the parameter of interest. The uncertainty of the fitted parameters can be estimated by Monte Carlo simulations or a full-error propagation formula 44 . Here, we adopt the latter method. Briefly, the fitting process can be written as the minimization of the residual between the measured and modeled FDTR responses: R = M p X i=1 [y i â f (α i ,X U ,X C )] 2 .(9) Here, α i is the scanned variable such as the modulation frequency or pumpâprobe offset, y i is the measured signal at α i , X U denotes the fitted unknown parameters, and X C denotes the fixed input parameters. Let F and Y be the M p Ă 1 vectors formed by f (α i ,X U ,X C ) and y i , respectively. The Jacobian matrices with respect to the fitted and fixed parameters are J U = âF âX U ,J C = âF âX C .(10) The varianceâcovariance matrix of the fitted parameters is then given by Var(X U ) = (J â U J U ) â1 J â U h Var(Y ) + J C Var(X C )J â C i J U (J â U J U ) â1 , (11) where the superscriptâ denotes the complex conjugate transpose. Here, Var(Y ) represents the variance of the experimental signal, and Var(X C ) represents the uncertainties of the fixed input parameters. The contribution from Var(Y ) is often negligible. The square roots of the diagonal elements of Var(X U ) yield the uncertainties of the fitted parameters. We note that such uncertainty analysis is local in the parameter space, and does not capture optimization-induced errors associated with convergence to a local minimum. 6 3 Vibe-FDTR framework We first introduce the architectural design of the Vibe-FDTR framework. As shown in Fig. 3, the core of Vibe-FDTR consists of a code layer and a procedural guidance layer, with the former enforcing physical and parametric consistency, while the latter governing how user requests are translated into organized and verifiable analyses. Together, the two layers make agent executions predictable and auditable. Furthermore, an optional expert layer extends this core architecture when users explicitly request support for experimental design, dealing with underspecified problems, or failure diagnosis. In such cases, the expert skill provides additional structured scientific guidance for AI agents based on predefined expert experience. 3.1 Code layer The code layer provides the execution foundation of Vibe-FDTR. Conventional FDTR analysis often relies on separate scripts for different processing and fitting operations, re- quiring manual transfer of parameters and intermediate results. Vibe-FDTR instead inte- grates data discovery, material-property lookup, model configuration, numerical analysis, and standardized output within a unified package (Fig. 3). All modules share consistent parameter definitions, unit conventions, and data structures, allowing successive analysis steps to operate on a common basis. A clear interface also separates agent-directed work- flow construction from the implementation of the thermal model and numerical solvers. A central design choice of the code layer is to use configuration files as persistent, machine-readable specifications. These files preserve the model definition and analysis settings so that each calculation can be inspected, reproduced, or modified without re- assembling information from dispersed files or relying on manual records. Multi-step anal- ysis is described by a separate pipeline specification that defines the execution order and transfers fitted quantities between successive steps. Before numerical execution, the code layer validates the configuration files and rejects incomplete, incompatible, or physically invalid inputs. This code-level validation provides the hard constraints of Vibe-FDTR, setting the stage for the procedural rules encoded in the guidance layer. 3.2 Procedural guidance layer The procedural guidance layer translates user requests into executable FDTR operations through agent skills. Even when a user provides a clear analysis request, the agent still needs to map that request to the correct code entry point, construct the required configu- ration, select the appropriate analysis module, and pass intermediate files between steps. The role of this layer is therefore to provide task-level guidance for code use, so that the agent does not need to infer the analysis logic via the underlying source code. 7 Vibe-FDTR uses a progressive-disclosure structure to limit the information exposed to the agent at each stage. The main skill defines the overall workflow, parameter- naming conventions, supported analysis functions, and standard execution order, in- cluding frequency-sweep fitting, beam-offset fitting, spot-size fitting, sensitivity analysis, uncertainty propagation, and iterative fitting. More specialized skills are loaded only when required. Configuration generation is handled by a dedicated skill that converts the sample structure, material properties, fixed inputs, fitting targets, parameter bounds, signal channels, and data ranges into a machine-readable configuration file. Sensitivity and uncertainty analyses are handled by separate skills that extend an existing fitting configuration with the additional information needed for reliability evaluation. For multi- step analysis involving parameter transfer, an iterative-pipeline skill specifies the fitting sequence, update rules, convergence criteria, and required artifacts. 3.3 Expert layer Built on top of the code layer and the procedural guidance layer, the expert layer is implemented as an optional upper-level skill which is invoked only by explicit user re- quests. Routine tasks therefore remain direct, while more ambiguous requests can be examined within a broader FDTR reasoning procedure. With this skill, the agent can better interpret the overall FDTR data-processing workflow and incorporate relevant prior experience, therefore facilitating less experienced users with complicated tasks. As shown in Fig. 4, the expert skill first distinguishes between design-only and data- backed requests. For design-only tasks, missing physical quantities are assigned explicit physically reasonable assumptions, followed by parameter-role assignment to separate fitting parameters from constrained inputs. The skill then develops candidate analysis schemes and evaluates them through sensitivity and uncertainty calculations performed by the underlying task skills. The results are reviewed and the scheme is revised with additional calculations when needed before a structured recommendation is produced. For data-backed tasks, the workflow begins with data-quality screening to identify corrupted points, spikes, or inconsistent repeats. It then applies the same pre-fit design process to select a justifiable fitting scheme. After execution, the fitted results are reviewed from several aspects including residual quality and physical plausibility, with unsatisfac- tory results triggering the revision of the assumptions or fitting schemes. Throughout the workflow, each decision stage is informed by expert experience which is extracted through AI-mediated elicitation of human FDTR experience from representative cases. 8 4 Benchmark design 4.1 Benchmark tasks To quantitatively evaluate Vibe-FDTR for the reliable execution of well-defined FDTR data analysis tasks, we constructed a two-level benchmark with seven Level-1 (L1) tasks and nine Level-2 (L2) tasks. These tasks cover commonly used analysis methods, including both f-sweep and beam-offset fitting, as well as sensitivity and uncertainty calculations. The L1 tasks are simple single-step operations based on synthetic FDTR data generated from the forward model described in Section 2, which are intended to test the reliability of an AI agent in the most straight-forward scenario. In terms of materials selection, we consider representative isotropic and anisotropic glasses and crystals, including fused silica (SiO 2 ), silicon (Si), graphite, and hexagonal boron nitride (hBN). For L2 tasks, we employ real FDTR experimental data from a gold-coated graphite (Au/graphite) sample. More complex workflows are designed, involving combined data preprocessing, analysis, multiple fitting steps and batch processing of different data groups. The L2 tasks closely reflect the end-to-end workflow typically followed by researchers in real-world settings. The key inputs for all these tasks are summarized in Table 1, with the ground truth independently determined via manual inspection. The Au/graphite specimen was prepared from commercially available highly oriented pyrolytic graphite (HOPG, ZYA grade, NT-MDT). Before metal deposition, the HOPG was mechanically cleaved to expose a fresh basal-plane surface. The graphite specimen and a fused-silica reference sample were mounted on the same silicon carrier wafer and coated in a single electron-beam evaporation process. The metal transducer consisted of a Ti adhesion layer and an Au layer. The Au thickness was obtained by measuring patterned step edges on the fused-silica reference using atomic force microscopy (AFM, Cypher ES, Asylum Research). The temperature-dependent electrical conductivity of the Au film was measured using a standard four-probe method, and its thermal conductivity was calculated from the Wiedemann-Franz law. In Fig. 1(b) and (c), we demonstrate the measured beam-offset and f-sweep phase signals at three representative temperatures. Beyond the well-specified L1 and L2 tasks, a set of partially specified requests without quantitative ground truth further challenges the optional expert mode to autonomously plan sensitivity and uncertainty analyses and to condense the resulting evidence into physically grounded recommendations. These underspecified expert-mode (E) tasks are summarized in Table 2. Spanning from routine fitting requests to open-ended experi- mental design, this layered benchmark allows a comprehensive examination of where the Vibe-FDTR framework excels and where its current limits lie. 9 4.2 Evaluation protocol We build a containerized environment to execute and record automated FDTR analyses using the Vibe-FDTR framework. To isolate the respective contributions of the domain code and the procedural skills, each task is additionally executed under two reference configurations, namely a Code-agent variant with the skill documents ablated and an Agent-only setting where the domain package is also withheld. As shown in Fig. 5, this environment is based on OpenCode 45 , an open-source harness which we connect to the DeepSeek-V4-Pro API. Each run takes place in an isolated Docker container. At the beginning of a test, the evaluation program populates the container with the required inputs and injects the task prompt; the agent subsequently carries out the request without human intervention. For the Code-agent and Agent-only settings, where no skill documents are available, we additionally inform the agent of the runtime details. Upon completion, the output files, total wall-clock time, and API call cost are collected for assessment. All tests are conducted on a dual-socket EPYC 9654 workstation with 192 cores and ample memory. Fifty containers run in parallel, compressing the full suite into approx- imately one hour and thereby minimizing temporal fluctuations in API service perfor- mance. Monitoring of the computational resources confirms that CPU and memory usage on the host remains low throughout, indicating that execution is not bottlenecked by the underlying hardware. A run is deemed successful only when it satisfies all of the following criteria: completion within the time limit, generation of the required outputs, and fitted results falling within predefined numerical tolerances. For the L1 synthetic tasks, the ground truth used for the forward data generation serves as the reference, with the time limit set to 600 s; for the L2 real-data tasks, the target values are instead the analysis results obtained by human experts using the same FDTR backend, and the time limit is extended to 900 s. Given the stochastic nature of agent behavior, each task was independently repeated 10 times to sample as broad a range of outcomes as possible and to obtain statistically meaningful results. For each task, the success rate, token cost, and wall-clock time are recorded. 5 Results and discussion 5.1 Reliability in well-defined FDTR tasks We first evaluate success rates on the L1 and L2 benchmark tests. Figure 6 shows the number of successful runs out of 10 repetitions for each task, which reflects both numer- ical correctness and repeated-run reliability. Across both benchmark levels, Vibe-FDTR consistently achieves the highest success rates in all tasks, while the Code-agent and Agent-only references become progressively less reliable as task complexity increases. 10 For the L1 single-step tasks, Vibe-FDTR completes all 70 test runs successfullyâa success rate of 100%, as shown in Fig. 6(a). Code-agent achieves a success rate of 91.4%, showing that access to the domain code was sufficient for most explicitly specified fitting tasks. Its remaining failures are concentrated in sensitivity and uncertainty analyses, where correct execution requires more complicated configuration files and interface use. Agent-only achieves a success rate of only 38.6%, with successful runs limited mainly to simpler fitting tasks. These results indicate that the code layer of Vibe-FDTR provides the primary foundation for reliable single-step analysis, while structured procedural guidance further guarantees the success of tasks with more demanding requirements. The distinction becomes much more prominent for the L2 real-data workflows. Vibe- FDTR maintains a rather high success rate of 98.9%, with only one failed run out of 90, whereas Code-agent only succeeds in 36.7% of total runs and Agent-only produces no successful runs at all. Looking into the details of Fig. 6(b), Code-agent performs com- paratively better on the more direct data-processing and fitting tasks, but its reliability decreases sharply for composite, iterative, and batch workflows that requires uncertainty propagation, parameter transfer, or multiple analysis stages. The contrast between L1 and L2 shows that the procedural guidance layer of Vibe-FDTR becomes increasingly im- portant when a complete FDTR workflow must preserve consistency across data handling, configuration updates, and successive calculations. In addition, the above results also in- dicate that finishing the complex FDTR workflows from scratch remains challenging for current frontier open-source large language models. 5.2 Failure mode analysis To better understand the large performance gap between Vibe-FDTR and the two refer- ence variants, we examined the failure modes in the L1 and L2 benchmarks. For L1 tasks, Code-agent failures arise mainly from errors in the configuration setup and command-line interface (CLI) usage, including bypassing the existing CLI tools and consequently mis- using the underlying code functions in L1-F01, passing parameters in an invalid format in L1-S01, and mixing the parameter naming for isotropic and anisotropic materials in L1- U01. Although minor, these operational mistakes lead to wrong outputs and can largely be eliminated by the explicit configuration guidance as implemented in Vibe-FDTR. For the Agent-only configuration, the agent has to rely on the general knowledge of the model to interpret each task and develop dedicated code, leading to two main types of failure. First, substantial reasoning and execution resources are spent on model derivation and debugging. For example, in L1-U01, the agent spends much of its reasoning context to derive Jacobian matrices and uncertainty-propagation formulas, while in the L1-F03 and F04 tasks, repeated code modification and looping eventually triggers debugging time- outs. Second, the agent meets the challenges of concepts understanding, such as thermal conductivity and thermal resistance (L1-S01). 11 Moving to L2 tasks, we first look at the only failed run of Vibe-FDTR occurred in L2-C01, where the agent omits the requested frequency range, leading to a 1.2% devia- tion of the fitted result. This isolated configuration error does not indicate a systematic limitation. For Code-agent, the relatively low success rates already observed in the L1 sensitivity and uncertainty tasks carries over to L2 workflows containing these analysis steps. Further failures arise in iterative multi-step fitting and batch processing, where the agent often bypasses the intended configuration workflow, uses inappropriate initial values or bounds, or fails to maintain consistent parameter transfer across stages and tempera- tures. By contrast, Code-agent achieves relatively high success rates in L2-D01 and D02 tasks because they mainly extends the simple L1 fitting tasks with clearly specified data- selection and preprocessing steps, without introducing additional analysis dependencies. Agent-only completes no L2 task, as most runs exhausts their time budgets planning the multi-step tasks, exploring data directories, and deriving the underlying physical model before fitting could even begin. Reviewing these failure modes, the core value of Vibe-FDTR becomes evident: the framework encapsulates invaluable human-expert experience into skills together with the physical model and domain code, enabling domain-specific tasks to be executed efficiently and stably, and substantially saving the trial-and-error costs of writing relevant programs from scratch. 5.3 Cost and efficiency We then analyze the token cost and time consumption during the L1 and L2 benchmark tasks. As shown in Fig. 7(a), Vibe-FDTR requires only $0.0098 per L1 run and $0.0133 per L2 run, representing an 87.7% reduction relative to Code-agent at both benchmark levels. Agent-only also shows a lower cost than Code-agent, however, this should not be interpreted as greater efficiency since many of its runs simply fail or terminate before completing the full analysis. The much higher cost for Code-agent can be attributed mainly to the parallel inspection of the source code and repeated exploration of available interfaces, which substantially increase token consumption. In Fig. 7(b), we further demonstrate the average time consumption in the benchmark tests. Vibe-FDTR completes the L1 and L2 tasks with average runtimes of 88 and 106 s, respectively. These values are 64.9% and 69.0% lower than those for Code-agent, and 78.1% and 85.8% lower than those for Agent-only. In contrast to the cost, time consump- tion of Code-agent is lower than Agent-only due to the parallel code inspection process. Agent-only typically follows a more sequential process of constructing, testing, and debug- ging the code and the analysis workflow, resulting in the longest execution time. Because unsuccessful runs are stopped at preset time limits, the reported Agent-only runtime provides a conservative estimate of this overhead. 12 To gain a intuitive feeling of the efficiency advantage of Vibe-FDTR, we take L2- B02, the most complex batch-processing task as an example. Vibe-FDTR completes this task within 2 minutes on average, which involves a full iterative fitting workflow at room temperature (RT, 295 K), then extending the fitting to eight additional temperature datasets. Even with ready-made analysis codes, the same sequence would normally require repeated manual efforts and much longer hands-on processing time. The above results clearly demonstrate the advantages of Vibe-FDTR not only in its reliability and accuracy but also in efficiency and cost savings, which holds great potential as a copilot for well-defined FDTR tasks. 5.4 Expert mode demonstration Unlike the L1 and L2 benchmarks, the optional expert mode is evaluated qualitatively with a focus on scientifically reasonable decision process and final recommendation. As shown in Table 2, the E tasks involve eight design-only and four data-backed cases, cov- ering a wide range of situations and challenges encountered by researchers in practice. Generally, Vibe-FDTR in the expert mode transforms open-ended user intent into ex- plicit analysis assumptions and executable evaluation schemes. The recommendations are broadly consistent with expert experiences. In particular, Vibe-FDTR demonstrates its potential as an intelligent copilot in recommending complementary measurement modes, identifying poorly separable parameters, and recognizing abnormal data or unreliable fitting conditions. To gain further insights, we look into the execution trace of the agent. Here, the E-7 task is used as an example and the main agent decisions and outputs are detailed in Fig. 8. First, the agent frames the task as a design-only problem and addresses the miss- ing information of the thin film by adopting the properties of Si, while assuming a fixed thermal boundary conductance for G 1 . Next, the agent organizes the design space ac- cording to the competing serial thermal resistances, d 2 /Îș 2 and 1/G 3 , and performs batch sensitivity calculations over broad ranges of film thickness and thermal conductivity. The agent discovers that the sensitivity curves of Îș 2 and G 3 closely follow each other, and further validates the inseparability of these two parameters via representative uncertainty analysis. Moreover, the agent also tries alternative fitting schemes, including beam-offset fitting and separate single-parameter fits before concluding that a single FDTR measure- ment could not reliably recover both quantities simultaneously. Finally, the agent gives a structured recommendation including clear statement on the assumptions and risks. This execution aligns closely with the intended expert-mode workflow in Fig. 4. The execution traces nevertheless reveal limitations in the scientific decisions made during this process. In E-7, the original user prompt does not specify which interface should be independently constrained. The agent jumps to fix G 1 and focuses on the sep- 13 arability of Îș 2 and G 3 , which is counterintuitive for human experts, although the final recommendation acknowledges the risk of an uncalibrated G 1 . Similarly, in several data- backed cases, the backtracking process is also incomplete. For instance, E-10 retains both interfacial thermal conductances in an already coupled anisotropic fitting; E-11 recognizes the low-frequency inconsistency but does not redefine the fitting window and repeat the analysis using the more reliable high-frequency range; and E-12 improves the fitting by releasing another Au property without systematically testing whether the supplied trans- ducer thickness is correct. These cases show that the expert mode is not yet capable of fully substituting for human expertise. We also note that the performance of the expert mode strongly depends on the underlying large language model of the agent. 6 Conclusion In summary, we have developed Vibe-FDTR, an agent-oriented framework that enables state-of-the-art LLM agents to perform reliable and reproducible FDTR analyses directly from natural language requests. To quantitatively assess its performance, we design a controlled two-level benchmark comprising single-step tasks (L1) with synthetic data and multi-step tasks (L2) based on real measurements of gold-coated graphite samples, to- gether with a repeated-run evaluation protocol. Agents using Vibe-FDTR achieve success rates of 100% and 98.9% on the L1 and L2 tasks, respectively, compared to 91.4% and 36.7% for the Code-agent variant and 38.6% and 0% for the Agent-only setting. The per- run cost and runtime are also reduced by 87.7% and over 60%, respectively, relative to the Code-agent variant. These controlled comparisons show that the domain code pack- age provides the physical foundation, while explicit procedural guidance via agent skills becomes increasingly important for maintaining reliability and efficiency as task com- plexity grows. Beyond executing well-specified instructions, the optional expert mode autonomously plans and performs sensitivity and uncertainty analyses and formulates physically grounded recommendations for underspecified tasks, although its scientific de- cisions still fall short of those of human experts and remain sensitive to the underlying language model. By encapsulating validated domain code and expert procedural knowl- edge into agent skills, Vibe-FDTR lowers the barrier to FDTR analysis for both new practitioners and specialists from other fields. This framework can be readily extended to other thermoreflectance techniques such as TDTR, and integrated with instrument control and data acquisition, offering a concrete route toward fully autonomous thermoreflectance metrology. 14 Data and code availability The agent inputs and traces will be deposited at this repository. The Vibe-FDTR source, skills and benchmark runner will be available on GitHub. Acknowledgments This work was supported by the Science Fund for Creative Research Groups from the Na- tional Natural Science Foundation of China (No. 52521007), the Scientific Research Inno- vation Capability Support Project for Young Faculty (ZYGXQNJSKYCXNLZCXM-E1) from the Ministry of Education of China, the National Key R&D Project from the Ministry of Science and Technology of China (No. 2022YFA1203100 and No. 2024YFA1207900), and the High-performance Computing Platform of Peking University. We thank Shuang- dui Wu for her help with the figures. B.S. acknowledges support from the New Cornerstone Science Foundation through the XPLORER PRIZE. Author declarations Conflict of interest. The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Author contributions. F.Y.: Conceptualization, Methodology, Software, Writing â Original Draft, Writing â Review & Editing. W.L.: Software, Validation, Investigation, Data Curation, Writing â Original Draft, Writing â Review & Editing. B.S.: Conceptu- alization, Supervision, Writing â Review & Editing, Funding acquisition. AI-assisted writing disclosure. During the preparation of this work, the authors used OpenCode and DeepSeek API service in order to assist coding and language polishing. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article. 15 Tables Table 1. Overview of L1 and L2 benchmark tasks. L1 denotes single-operation tests on synthetic datasets, and L2 denotes integrated workflows on measured Au/graphite data. IDOperationSampleInput constraintMain targets L1-F01 Frequency fitAu/fused silica Fixed Îș SiO 2 and spot sizeÎș Au , G L1-F02 Frequency fitAu/SiFixed Îș Au and spot sized Au , Îș Si , G L1-F03 Frequency fitAu/graphiteFixed Îș i,Gr and spot sizeÎș o,Gr , G L1-F04 Frequency fitAu/hBN/SiFixed Îș o,hBN and hBN/Si GÎș i,hBN , Au/hBN G L1-O01 Beam-offset fitAu/graphiteFixed Îș o,Gr and modulation frequencyÎș i,Gr , G L1-S01 Sensitivity analysisAu/fused silica Specified parameters and spot sizesSensitivity to Îș SiO 2 , G, and d Au L1-U01 Uncertainty propagationAu/graphite/Si Prescribed fixed-parameter uncertaintiesUncertainties of Îș i,Gr and Îș o,Gr L2-D01 Averaging + frequency fitAu/graphiteRT; one repeated point excludedÎș o,Gr , G L2-D02 Averaging + offset fitAu/graphite60 ⊠C; selected offset windowÎș i,Gr , G L2-C01 Sensitivity + frequency fitAu/graphiteRT; selected frequency rangeSensitivity, Îș o,Gr , G L2-C02 Offset fit + uncertaintyAu/graphiteRT; selected offset windowÎș i,Gr , G, uncertainty L2-C03 Sensitivity + fit + uncertainty Au/graphiteRTSensitivity, Îș o,Gr , G, uncertainty L2-I01 Spotfit + frequency fitAu/graphiteRT; highest-frequency amplitudeSpot size, Îș o,Gr , G L2-I02 Iterative fit + uncertaintyAu/graphiteRT; spotfit, offset fit, and frequency fitSpot size, Îș i,Gr , Îș o,Gr , G, uncertainty L2-B01 Temperature-batch fittingAu/graphiteTemperature-dependent dataÎș i,Gr (T), G(T) L2-B02 Temperature-batch iterationAu/graphiteRT spotfit followed by temperature-dependent iteration Spot size, Îș i,Gr (T), Îș o,Gr (T), G(T) 16 Table 2. Summary of expert-mode benchmark tasks and the final outputs by the agent using Vibe-FDTR. Here, only key information is selected for display convenience. ID User intentionVibe-FDTR recommendations E-1 Design FDTR measurement for Au/diamond. Target: Îș diamond . Assume a nominal 3 ÎŒm spot size and heat capacities from the database. Recommend f-sweep fitting over 5 kHz-20 MHz for Îș diamond and G Au/diamond . May use high-frequency offset amplitude for spot calibra- tion and offset phase fit as a cross-check. Justify both fitted quantities with about 8% uncertainty. E-2 Design FDTR measurement for Au/graphite. Targets: Îș i,Gr and Îș o,Gr . Assume the spot size is unknown and use heat capacities from the database. Recommend a sequential fitting workflow with 50 MHz offset amplitude for spot fitting, 5 MHz offset phase for Îș i,Gr , and 100 kHz- 50 MHz f-sweep phase for Îș o,Gr and G Au/Gr . Justify by sensitivity and uncertainty analysis. E-3 Design FDTR measurement for Au/graphene/Si. Targets: Îș i,Gr and Îș o,Gr . Assume graphene heat capacity from graphite-based data and fix G Gr/Si at 5Ă 10 7 W m â2 K â1 . Recommend beam-offset fitting for Îș i,Gr and f-sweep fitting for G Au/Gr . Reject simultaneous fitting of Îș o,Gr and G Au/Gr because film and interfacial thermal conductance are degener- ate. E-4 Recommend Au thickness for FDTR measurement of sapphire. Consider different possible G. Target: Îș sapphire . Assume a 3 ÎŒm spot size and heat capacities from the database. Rec- ommend 70-80 nm thick Au and fit Îș sapphire and G by f-sweep fitting over 50 kHzâ10 MHz. Justify by batch sensitivity runs with different d Au and G, and uncertainties at 80 nm (8.4% and 6.2% for Îș sapphire and G, respectively). E-5 Recommend spot size and fre- quency range for f-sweep measure- ment on Au/graphite. Targets: Îș i,Gr and Îș o,Gr . Assume graphite heat capacity from the database. Recommend a 5- 8 ÎŒm spot and a 200 kHzâ20 MHz f-sweep window. Fit Îș i,Gr , Îș o,Gr , and G Au/Gr simultaneously. Justify by sensitivity analysis and uncertainties at different spot size. Warn about the local minima risk for three- parameter fitting. E-6 Recommend the reliable fitting range for G of Au/fused sil- ica. Target: G from 10 6 to 10 9 W m â2 K â1 . Assume a 3 ÎŒm spot size and thermal properties from the database. Recommend reliable fitting for G †3 Ă 10 7 W m â2 K â1 , with an optimum near 10 7 W m â2 K â1 ; reject Gâ„ 10 8 W m â2 K â1 as poorly identifiable from f-sweep phase. Justify with sensitivity and uncertainty results for G from 10 6 to 10 9 W m â2 K â1 . E-7 Explore separable fitting condi- tions for Au/thin film/SiC. Con- sider a typical thin-film heat ca- pacity. Targets: film Îș and the two G. Assume an 80 nm Au transducer, thin-film heat capacity of 2.5 MJ m â3 K â1 , G Au/film = 200 MW m â2 K â1 , and a 3 ÎŒm spot size. Identify poor separability of film Îș and G film/SiC in a single mea- surement. Use d film G/Îș film to decide whether to fit film Îș or interface G while fixing the other. Justify by sensitivity and uncertainty results for different film thickness, film Îș, and G film/SiC . E-8 Explore identifiable co-fitting sets for Au/fused silica f-sweep phase data. Targets: properties of Au and SiO 2 , spot-size, and G. Assume about 48 nm Au, 3 ÎŒm spot size, G Au/SiO 2 â 50 MW m â2 K â1 , and two-parameter fitting as the practical limit. Recommend viable pa- rameter pairs and reject strongly correlated Au-property and substrate- property sets by sensitivity and uncertainty results for different co-fitting sets involving d Au , Îș Au , C Au , spot size, Îș SiO 2 , C SiO 2 , and G. E-9 Execute f-sweep fitting for Au/graphite using the given dataset. Targets: Îș o,Gr and G. (Bad data with phase spike present, which is the same dataset used in L2-D01) Detect and exclude the point11 data file with phase spikes and fit 10 clean f-sweep repeats. Report Îș o,Gr = 5.9 W m â1 K â1 and G = 9.4Ă 10 7 W m â2 K â1 . E-10 Choose proper fitting methods for Au/hBN/Si using the given data. Targets: Îș i,hBN and Îș o,hBN . Choose iterfit with offset phase for Îș i,hBN and f-sweep phase for Îș o,hBN , G Au/hBN , and G hBN/Si after several fitting methods comparison. Re- port Îș i,hBN = 396 W m â1 K â1 , Îș o,hBN = 3.7 W m â1 K â1 , and both G values near 5.0Ă 10 7 W m â2 K â1 . E-11 Execute f-sweep fitting for Au/graphite using the given dataset. Targets: Îș o,Gr and G. (One outlier data present, together with data problems in low-frequency ranges) Use prompt-specified properties and spot size. Exclude one anoma- lous repeat. Report Îș o,Gr = 16.1 W m â1 K â1 and G = 3.8 Ă 10 7 W m â2 K â1 . Identify and warn about the low-frequency data shift which gives a high fitting residual. E-12 Execute f-sweep fitting for Au/fused silica. Targets: Îș Au and G. (Incorrect Au thickness provided) Use the prompt-stated 50 nm Au thickness and report Îș Au = 203 W m â1 K â1 and G Au/SiO 2 = 9.5Ă 10 7 W m â2 K â1 . Identify the poor fit problem and try co-fitting the heat capacity of Au to im- prove. Warn that the fitted Au heat capacity deviates from the provided value. 17 Figures Probe 532 nm Pump 405 nm SignalReference PBS Dichroic mirror BS Camera XYZ piezo λ/2 λ/4 Balanced photodetector Signal in Waveform out (f = 50 kHz - 50 MHz) f Lock-in amplifer Objective Sample Transducer Layer 1 Substrate [ C 0 Îș 0 d 0 ] [ C 2 Îș i2 Îș o2 d 2 ] [ C 4 Îș i4 Îș o4 ] G 1 G 3 Offset distance (ÎŒm) -15010-10-5515 -10 0 -20 -30 -40 -50 -60 -70 T = 363 K T = 295 K T = 213 K Frequency (Hz) 10 5 10 6 10 7 -30 -20 -10 0 Phase (Deg) -40 Phase (Deg) (a)(b) (c) Concentric pump-probe Sample Sample Pump Probe Beam offset f-sweep T = 363 K T = 295 K T = 213 K Figure 1. Basic principles of the FDTR technique. (a) Schematic of FDTR experimental setup. (b) Representative beam-offset phase signals at f = 1 MHz measured on an Au/graphite sample at three different temperatures. (c) Representative f-sweep signals measured on the same sample. The insets show the schematics of the two measurement modes. 18 I have temperature-dependent FDTR data for the Au/graphite sample stored in the folder ... Please perform the following iterative fitting workflow for room-temperature data. Following the FDTR routing rules, I'l first scan the data directory. Fit converged. Plots saved to ... Both in-plane (Îș r) and out-of-plane (Îș z) thermal conductivities decrease with increasing temperature, while the Au/graphite TBC increases with temperature. Vibe-FDTR Manual data selection Manual records Manually run multiple scripts Parameter editing Inspection and possible rerun Traditional manual FDTR analysis Temperature (K) 400200300 Thermal conductivity (Wm -1 K -1 ) 10 1 10 2 10 3 Out-of-plane Îș o In-plane Îș i Figure 2. Schematic showing the concept of Vibe-FDTR in comparison with traditional manual analysis process. 19 User intentions Analysis execution Sensitivi ty Unce rtaintyIterative fit Single-st ep fit Pr ocedural guidance layer Code layer Task routing Scan data Input protocol Data naming convention Standardized output Analysis guide Fit guideSensitivi ty guide Unce rtainty guide Iterative fit pipeline Full parameter list Configuration guide Materi al lookup Configuration generation Valid ity check Figure 3. Core architecture of Vibe-FDTR which is composed of a procedural guidance layer and a code layer. 20 Fit execution Fit result review Traceable expert recommendation Data-backed request Design-only request Assumption specification Data assessment Feasibilit y review Fit st rategy Measurement ranges Parameter ranges Candidate schemes Parameter-rol e assignment Scheme design Analysis execution Sensitivi tyUnce rtainty Design- only Figure 4. Optional expert mode workflow of Vibe-FDTR. 21 Figure 5. Design of the evaluation environment for benchmark tasks. 22 L1-F01L2-B01 L2-C01 L2-C02 L2-C03L2-D01 L2-D02 L2-I01 L2-I02 L2-B02 L1-U01L1-F02 L1-F03 L1-F04L1-O01 L1-S01 Vibe-FDTR Code-agent Agent-only Success counts out of 10 repeated runs per task 2 4 6 8 10 2 4 6 8 10 Figure 6. Number of successful runs of L1 and L2 benchmark tasks for Vibe-FDTR and the two reference variants. Each task is repeatedly run for 10 times. 23 0.00 0.02 0.04 0.06 0.08 0.10 Average cost (USD) Vibe-FDTR Code-agent Agent-only L1L2 0 100 200 300 400 500 600 700 Average time (s) Figure 7. Average time and cost per run for L1 and L2 benchmark tasks using Vibe-FDTR and the two reference variants. 24 Roles assignment Assumptions made Target: film conductivity Îș2; film/SiC conductance G3 Calibrated/fixed: Au transducer properties; spot size ... Task framing User query Mode Goal Under what film-thickness and conductivity conditions can FDTR separately extract film conductivity and interface conductance? When should the interface conductance be fitted versus fixed? Design-only expert task; no measurement data provided. For Au / thin-film / SiC, assess identifiability of Îș2 and G3 for planned FDTR measurements. Review and additional calculation Alternative scheme: offset fit Alternative scheme: single-parameter fit Uncertainty for offset fit: Îș2 = 105.9%, G3 = 82.0%. Offsetfit does not rescue simultaneous extraction. Film-dominated: fit Îș2 only â 15.1%. Interface-dominated: fit G3 only â 11.8%. Use Si as a concrete library material; vary d2 and Îș2. For the focused study, fix assumed G1 = 200 MW m -2 K -1 and use nominal G3 = 50 MW m -2 K -1 . Parameter roles and assumptions Practical guidance Key risk Assumptions stated Direct recommendation Assumed thin film properties Focused study fixes assumed G1; nominal G3 used for sensitivity... Îș2 and G3 cannot be simultaneously extracted from one FDTR sweep. Fix one parameter to fit the other. d2G3/Îș2 > 5 â fit Îș2, fix G3 d2G3/Îș2 < 0.05 â fit G3, fix Îș2 0.05 < d2G3/Îș2 < 5 â fit either one, not both G1 was initially unknown and had high sensitivity. If not independently calibrated, it becomes a third correlated serial-resistance parameter. Final recommendation Batch design Physics basis Inference Representative co-fit uncertainty checks Key results Scheme design and sensitivity analysis Serial thermal resistance: R film = d2/Îș2 and Rint = 1/G3. Use d2 G3/Îș2 = R film/Rint to organize the design space. Map sensitivity across the d2-Îș2 design space. d2 = 10-1000 nm; Îș2 = 1-500 W m -1 K -1 . 7 thicknesses Ă 8 conductivities = 56 cases. Îș2 and G3 sensitivity curves are highly correlated (~0.99-1.0): inseparable in a single frequency sweep. Sensitivity alone does not determine whether two parameters can be jointly recovered. Uncertainty is used to quantify co-fit identifiability. Joint targets: Îș2 and G3. At d2 G3/Îș2 = 1.0: Îș2 uncertainty = 61.6%; G3 uncertainty = 34.7% ... Even in the best co-fit case, one parameter retains 35-60% uncertainty; the two parameters cannot be jointly extracted. Uncertainty analysis Figure 8. Representative execution trace under the optional expert mode of Vibe-FDTR. The E-7 task is demonstrated as an example, and the agent outputs are selected and reorganized for clearance. 25 References [1] D. G. Cahill, W. K. Ford, K. E. Goodson, et al. âNanoscale Thermal Transportâ. In: Journal of Applied Physics 93.2 (Jan. 2003), p. 793â818. [2] Arden L. Moore and Li Shi. âEmerging Challenges and Materials for Thermal Man- agement of Electronicsâ. In: Materials Today 17.4 (May 2014), p. 163â174. (Visited on 01/16/2022). [3] Bai Song, Anthony Fiorino, Edgar Meyhofer, and Pramod Reddy. âNear-field ra- diative thermal transport: From theory to experimentâ. In: AIP Advances 5.5 (Apr. 2015), p. 053503. doi: 10.1063/1.4919048. url: https://doi.org/10.1063/1. 4919048 (visited on 07/28/2026). [4] Dongliang Zhao, Xin Qian, Xiaokun Gu, Saad Ayub Jajja, and Ronggui Yang. âMeasurement Techniques for Thermal Conductivity and Interfacial Thermal Con- ductance of Bulk and Thin Film Materialsâ. In: Journal of Electronic Packaging 138.040802 (Oct. 2016). (Visited on 06/25/2026). [5] Jie Chen, Xiangfan Xu, Jun Zhou, and Baowen Li. âInterfacial Thermal Resis- tance: Past, Present, and Futureâ. In: Reviews of Modern Physics 94.2 (Apr. 2022), p. 025002. (Visited on 04/25/2022). [6] David G. Cahill. âThermal Conductivity Measurement from 30 to 750 K: The 3Ï Methodâ. In: Review of Scientific Instruments 61.2 (Feb. 1990), p. 802â808. (Vis- ited on 06/10/2022). [7] Chris Dames. âMEASURING THE THERMAL CONDUCTIVITY OF THIN FILMS: 3 OMEGA AND RELATED ELECTROTHERMAL METHODSâ. In: Annual Re- view of Heat Transfer 16.1 (2013), p. 7â49. (Visited on 06/25/2026). [8] A. Majumdar. âScanning Thermal Microscopyâ. In: Annual Review of Materials Science 29.1 (Aug. 1999), p. 505â585. (Visited on 03/08/2022). [9] David G. Cahill, Kenneth Goodson, and Arunava Majumdar. âThermometry and Thermal Transport in Micro/Nanoscale Solid-State Devices and Structuresâ. In: Journal of Heat Transfer 124.2 (Dec. 2001), p. 223â241. (Visited on 06/25/2026). [10] Sunmi Shin and Renkun Chen. âThermal Transport Measurements of Nanostruc- tures Using Suspended Micro-Devicesâ. In: Nanoscale Energy Transport: Emerging Phenomena, Methods and Applications. IOP Publishing, Mar. 2020. (Visited on 06/25/2026). [11] Haiyu He, Yuxi Wang, Zhiyao Jiang, and Bai Song. âBig MEMS for Thermal Mea- surementâ. In: ASME Journal of Heat and Mass Transfer (Sept. 2024), p. 1â28. (Visited on 11/17/2024). 26 [12] Haiyu He, Yuxi Wang, and Bai Song. âBig-MEMS for high thermal conductivity measurementâ. en. In: Review of Scientific Instruments 96.2 (Feb. 2025). doi: 10. 1063/5.0245017. url: https://pubs.aip.org/aip/rsi/article/96/2/024902/ 3333957/Big-MEMS-for-high-thermal-conductivity-measurement (visited on 09/15/2025). [13] Michel Kazan. âRaman Spectroscopy for Thermal Transport Characterization: Prin- ciples, Techniques, and Applicationsâ. In: Journal of Applied Physics 139.8 (Feb. 2026), p. 081101. (Visited on 06/25/2026). [14] Shen Xu, Aoran Fan, Haidong Wang, Xing Zhang, and Xinwei Wang. âRaman- Based Nanoscale Thermal Transport Characterization: A Critical Reviewâ. In: In- ternational Journal of Heat and Mass Transfer 154 (June 2020), p. 119751. (Visited on 06/25/2026). [15] Puqing Jiang, Xin Qian, and Ronggui Yang. âTutorial: Time-domain Thermore- flectance (TDTR) for Thermal Property Characterization of Bulk and Thin Film Materialsâ. In: Journal of Applied Physics 124.16 (Oct. 2018), p. 161103. [16] Dylan J. Kirsch, Joshua Martin, Ronald Warzoha, Mark McLean, Donald Windover, and Ichiro Takeuchi. âAn Instrumentation Guide to Measuring Thermal Conductiv- ity Using Frequency Domain Thermoreflectance (FDTR)â. In: Review of Scientific Instruments 95.10 (Oct. 2024), p. 103006. (Visited on 11/04/2024). [17] David G. Cahill, Paul V. Braun, Gang Chen, et al. âNanoscale Thermal Transport. I. 2003-2012â. In: Applied Physics Reviews 1.1 (Mar. 2014). [18] David H. Olson, Jeffrey L. Braun, and Patrick E. Hopkins. âSpatially Resolved Thermoreflectance Techniques for Thermal Conductivity Measurements from the Nanoscale to the Mesoscaleâ. In: Journal of Applied Physics 126.15 (2019). [19] Ke Chen, Bai Song, Navaneetha K. Ravichandran, et al. âUltrahigh Thermal Con- ductivity in Isotope-Enriched Cubic Boron Nitrideâ. In: Science 367.6477 (Jan. 2020), p. 555â559. [20] Suixuan Li, Chuanjin Su, Zihao Qin, et al. âMetallicΞ-Phase Tantalum Nitride Has a Thermal Conductivity Triple That of Copperâ. In: Science 391.6786 (Feb. 2026), p. 707â711. (Visited on 06/25/2026). [21] Aaron J. Schmidt, Ramez Cheaito, and Matteo Chiesa. âA Frequency-Domain Ther- moreflectance Method for the Characterization of Thermal Propertiesâ. In: Review of Scientific Instruments 80.9 (Sept. 2009), p. 094901. (Visited on 06/08/2022). [22] Fei Tian, Bai Song, Xi Chen, et al. âUnusual High Thermal Conductivity in Boron Arsenide Bulk Crystalsâ. In: Science 361.6402 (Aug. 2018), p. 582â585. (Visited on 11/27/2022). 27 [23] Yuxi Wang, Nianjie Liang, Xingxing Zhang, et al. âThermal Transport in a 2D Amorphous Materialâ. In: Physical Review X 15.3 (Sept. 2025), p. 031077. (Visited on 01/12/2026). [24] Fuwei Yang, Wenjiang Zhou, Tian Gu, et al. âThermal Transport in Rhombohedral Boron Nitrideâ. In: Physical Review B 112.11 (Sept. 2025), p. 115422. (Visited on 09/21/2025). [25] Kiumars Aryana, John A. Tomko, Ran Gao, et al. âObservation of Solid-State Bidirectional Thermal Conductivity Switching in Antiferroelectric Lead Zirconate (PbZrO3)â. In: Nature Communications 13.1 (Mar. 2022), p. 1573. (Visited on 06/25/2026). [26] Brandi L. Wooten, Ryo Iguchi, Ping Tang, et al. âElectric FieldâDependent Phonon Spectrum and Heat Conduction in Ferroelectricsâ. In: Science Advances 9.5 (Feb. 2023), eadd7194. (Visited on 06/25/2026). [27] Ronald J. Warzoha, Adam A. Wilson, Brian F. Donovan, et al. âMeasurements of Thermal Resistance Across Buried Interfaces with Frequency-Domain Thermore- flectance and Microscale Confinementâ. In: ACS Applied Materials & Interfaces 16.31 (Aug. 2024), p. 41633â41641. (Visited on 10/03/2024). [28] Fuwei Yang, Wenjiang Zhou, Zhibin Zhang, et al. âUltrahigh Thermal Conduc- tance across Superlubric Interfaces in Twisted Graphiteâ. In: Physical Review Letters 134.14 (Apr. 2025), p. 146302. (Visited on 04/14/2025). [29] Yufeng Zhang, Yanzheng Du, Xiao Wan, et al. âAnomalous Enhancement of Ther- mal Conduction across Twisted van Der Waals Heterointerfacesâ. In: Proceedings of the National Academy of Sciences 123.9 (Mar. 2026), e2531049123. (Visited on 03/02/2026). [30] David G. Cahill. âAnalysis of Heat Flow in Layered Structures for Time-Domain Thermoreflectanceâ. In: Review of Scientific Instruments 75.12 (2004), p. 5119â 5122. [31] P. Jiang, X. Qian, and R. Yang. âTime-Domain Thermoreflectance (TDTR) Mea- surements of Anisotropic Thermal Conductivity Using a Variable Spot Size Ap- proachâ. In: Review of Scientific Instruments 88.7 (July 2017), p. 074901. [32] Puqing Jiang, Dihui Wang, Zeyu Xiang, Ronggui Yang, and Heng Ban. âA New Spatial-Domain Thermoreflectance Method to Measure a Broad Range of Anisotropic in-Plane Thermal Conductivityâ. In: International Journal of Heat and Mass Trans- fer 191 (Aug. 2022), p. 122849. (Visited on 10/24/2022). [33] DeepSeek-AI, Anyi Xu, Bangcai Lin, et al. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence. 2026. arXiv: 2606.19348 [cs.CL]. url: https: //arxiv.org/abs/2606.19348. 28 [34] Chris Lu, Cong Lu, Robert Tjarko Lange, et al. âTowards End-to-End Automa- tion of AI Researchâ. In: Nature 651.8107 (Mar. 2026), p. 914â919. (Visited on 03/30/2026). [35] Richard B. Canty, Jeffrey A. Bennett, Keith A. Brown, et al. âScience Acceleration and Accessibility with Self-Driving Labsâ. In: Nature Communications 16.1 (Apr. 2025), p. 3856. (Visited on 06/25/2026). [36] Indrajeet Mandal, Jitendra Soni, Mohd Zaki, et al. âEvaluating Large Language Model Agents for Automation of Atomic Force Microscopyâ. In: Nature Communi- cations 16.1 (Oct. 2025), p. 9104. (Visited on 05/20/2026). [37] SangHeon Lee. âJ 3 SPM AI: An Integrated Open-Source Platform for AI-assisted Image Analysis and Image-Guided Workflows in Scanning Probe Microscopyâ. In: Micron 204 (May 2026), p. 104017. (Visited on 06/25/2026). [38] Yu Pang, Puqing Jiang, and Ronggui Yang. âMachine Learning-Based Data Pro- cessing Technique for Time-Domain Thermoreflectance (TDTR) Measurementsâ. In: Journal of Applied Physics 130.8 (Aug. 2021), p. 084901. (Visited on 05/09/2025). [39] Yasuaki Ikeda, Yuki Akura, Masaki Shimofuri, Amit Banerjee, Toshiyuki Tsuchiya, and Jun Hirotani. âEstimating Depth-Directional Thermal Conductivity Profiles Using Neural Network with Dropout in Frequency-Domain Thermoreflectanceâ. In: Journal of Applied Physics 137.5 (Feb. 2025), p. 055106. (Visited on 05/09/2025). [40] Amun Jarzembski, Siddharth Nair, Wyatt Hodges, et al. âWide-Field Bond Quality Evaluation Using Frequency Domain Thermoreflectance with Deep Neural Network Feature Reconstructionâ. In: Advanced Materials Interfaces 2401039.n/a (2025). (Visited on 05/15/2025). [41] Yuyao Ge, Lingrui Mei, Zenghao Duan, et al. A Survey of Vibe Coding with Large Language Models. Oct. 2025. (Visited on 06/25/2026). [42] Partha Pratim Ray. A Review on Vibe Coding: Fundamentals, State-of-the-art, Challenges and Future Directions. May 2025. (Visited on 06/25/2026). [43] Advait Sarkar and Ian Drosos. Vibe Coding: Programming through Conversation with Artificial Intelligence. Oct. 2025. arXiv: 2506.23253 [cs]. (Visited on 06/25/2026). [44] Jia Yang, Elbara Ziade, and Aaron J. Schmidt. âUncertainty Analysis of Ther- moreflectance Measurementsâ. In: Review of Scientific Instruments 87.1 (Jan. 2016), p. 014901. doi: 10.1063/1.4939671. [45] SST. OpenCode. Version 0.1.x. Open-source AI coding agent for the terminal. Ac- cessed: 2026-07-08. 2025. url: https://github.com/anomalyco/opencode. 29