Paper deep dive
Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
Yusen Tan, Yixuan Chen, Zheng Fang, Pan Liu, Yifan Li, Qinyu Guo, Zhedong Lin, Yuqiang Li, Xiangxiang Zeng, Tong Wang, Jun Xia
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/17/2026, 3:59:24 AM
Summary
The paper introduces UltraIR, a foundation model for infrared (IR) spectroscopy with over 100 million parameters that utilizes simulation-to-real transfer learning. Pretrained on approximately 60 million simulated IR spectra using spectral reconstruction, molecular fingerprint alignment, and functional-group prediction, UltraIR is adapted to downstream tasks such as functional-group prediction, molecular structure elucidation, physicochemical property prediction, mixture analysis, bacterial classification, and soil property prediction. It outperforms conventional machine learning and task-specific deep learning baselines, demonstrating strong performance with limited labeled data and zero-shot inference across different spectrometers and laboratories.
Entities (21)
Relation Signals (17)
UltraIR → pretrainedon → Simulated IR Spectra
confidence 95% · UltraIR is pretrained on approximately 60 million simulated IR spectra
UltraIR → usesmethod → Simulation-to-Real Transfer Learning
confidence 95% · UltraIR... enables simulation-to-real transfer learning for chemical sensing and analysis
UltraIR → hascomponent → Transformer
confidence 90% · The encoder combines... a patch-based Transformer.
UltraIR → outperforms → FCGFormer
confidence 90% · UltraIR consistently and substantially outperformed both task-specific deep-learning models... FCGFormer
UltraIR → outperforms → IRAnalysis
confidence 90% · UltraIR consistently and substantially outperformed both task-specific deep-learning models... IRAnalysis
UltraIR → outperforms → XGBoost
confidence 90% · UltraIR consistently and substantially outperformed... conventional machine-learning baselines... XGBoost
UltraIR → performstask → Microplastics Classification
confidence 90% · Across... microplastics classification... UltraIR outperforms...
UltraIR → performstask → Soil Property Prediction
confidence 90% · Across... soil property prediction, UltraIR outperforms...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation is labor-intensive, relies on prior knowledge and reference spectra, and is difficult to scale, whereas most machine-learning methods are tailored to individual tasks or datasets, require large labeled training sets, and transfer poorly across analytical objectives and experimental datasets. Here we introduce UltraIR, a foundation model for IR spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. UltraIR is pretrained on approximately 60 million simulated IR spectra using spectral reconstruction, molecular fingerprint similarity alignment, and functional-group prediction, then adapted to downstream objectives with task-specific labels or targets. Across functional-group prediction, molecular structure elucidation, physicochemical property prediction, mixture-component identification and quantification, bacterial classification, medicinal-herb geographic origin traceability and constituent quantification, microplastics classification, and soil property prediction, UltraIR outperforms conventional machine-learning and task-specific deep-learning baselines. It performs strongly with limited labeled experimental spectra and in zero-shot inference for the same analytical task across Fourier-transform infrared spectrometers and laboratories, providing a route to adaptable, data-efficient chemical sensing from complex real-world samples.
Tags
Links
- Source: https://arxiv.org/abs/2608.13341v2
- Canonical: https://arxiv.org/abs/2608.13341v2
Trouble viewing inline? Open PDF directly →
Full Text
156,523 characters extracted from source content.
Expand or collapse full text
1]Information Hub, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, China 2]College of Computer Science and Technology, Jilin University, Changchun 130012, China 3]School of Computer Science, University of Auckland, Auckland 1142, New Zealand 4]AI for Chemistry Center, Shanghai Artificial Intelligence Laboratory, Shanghai 200232, China 5]College of Computer Science and Electronic Engineering, Hunan University, Changsha 410082, China 6]State Key Laboratory of Chemo and Biosensing, College of Chemistry and Chemical Engineering, Hunan University, Changsha 410082, China 7]Department of Computer Science and Engineering, The Hong Kong University of Science and Technology, Hong Kong SAR, China Zeng (xzeng@foxmail.com); Tong Wang (wangtong@hnu.edu.cn); Jun Xia (junxia@hkust-gz.edu.cn) Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples Yusen Tan Yixuan Chen Zheng Fang Pan Liu Yifan Li Qinyu Guo Zhedong Lin Yuqiang Li Xiangxiang Zeng† Tong Wang† Jun Xia† Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [ Abstract Infrared (IR) spectroscopy is widely used for chemical sensing, from molecular characterization to the analysis of complex samples, but extracting reliable chemical information from spectra remains challenging. Conventional interpretation based on expert peak assignment, rule-based analysis, and library matching is labor-intensive and difficult to scale across analytical objectives and datasets while maintaining consistently high accuracy, and its reliance on prior knowledge and reference spectra makes it less reliable for unfamiliar compounds and complex samples. Machine-learning methods have enabled more automated and scalable chemical inference from IR spectra, but most approaches remain tailored to individual tasks or datasets, require large training sets of labeled spectra, and show limited transferability across analytical objectives and experimental datasets. Together, these limitations and the scarcity of labeled experimental IR spectra motivate simulation-to-real transfer learning to enable data-efficient chemical inference under limited supervision. Here we introduce UltraIR, a foundation model for IR spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. UltraIR learns a shared spectral representation by pretraining on approximately 60 million simulated IR spectra using three complementary pretraining objectives for spectral reconstruction, molecular fingerprint similarity alignment, and functional-group prediction. The pretrained encoder is then adapted to each downstream objective using task-specific labels or targets paired with the corresponding IR spectra, typically with only a limited number of labeled spectra available. Across benchmark evaluations of functional-group prediction, molecular structure elucidation, physicochemical property prediction, and mixture-component identification and quantification, as well as real-world chemical sensing applications involving bacterial classification, medicinal-herb geographic origin traceability and constituent quantification using two newly generated in-house datasets, microplastics classification, and soil property prediction, UltraIR outperforms conventional machine-learning and task-specific deep-learning baselines. UltraIR demonstrates the practical value of simulation-to-real transfer learning through strong performance in two challenging real-world chemical sensing scenarios: downstream adaptation using limited labeled experimental spectra and zero-shot inference for the same analytical task across Fourier-transform infrared spectrometers and laboratories. More broadly, this approach provides a route to adaptable, data-efficient chemical sensing systems for extracting reliable chemical information from IR spectra of complex real-world samples. 1 Introduction Infrared (IR) spectroscopy is an analytical technique that probes molecular vibrations through their interaction with infrared radiation to generate chemically informative spectral fingerprints [39, 16, 23]. Its rapid, label-free, and often non-destructive measurements support widespread use in chemical sensing and analysis, ranging from molecular characterization to complex-sample analysis in environmental monitoring, clinical analysis, and materials characterization [11, 20, 36, 46]. However, the ability to acquire IR spectra at scale has not been matched by an equally scalable ability to convert them into reliable chemical information [14, 5, 15]. Conventional workflows still rely heavily on expert peak assignment, rule-based interpretation, and library matching. These approaches depend on prior expert knowledge, predefined rules, or reference spectra, which can limit reliable interpretation when suitable reference information is unavailable, as may occur for unfamiliar compounds and complex samples [14, 19]. Such dependencies also make conventional workflows difficult to scale across analytical objectives and datasets while maintaining consistent accuracy [38, 31, 39]. A central challenge for next-generation IR chemical sensing is therefore to translate spectra into reliable chemical information at scale across analytical objectives and datasets while maintaining high accuracy. Machine learning has begun to address the challenge of automated IR analysis at scale by learning predictive relationships directly from spectral data [24, 35, 50]. Recent studies have shown that IR spectra support functional-group recognition, molecular structure elucidation, library-based molecular retrieval, and mixture-component identification [12, 34, 1, 45, 22, 41]. Together, these advances demonstrate that IR spectra contain molecular information that can support diverse analytical tasks. However, most approaches are tailored to individual tasks or datasets and require large task-specific collections of labeled spectra, limiting the development of shared spectral representations that transfer across analytical objectives and experimental datasets. Despite enabling more rapid and scalable analysis than expert peak assignment, rule-based interpretation, and library matching, these models often achieve only modest accuracy gains over strong conventional machine-learning baselines, and these gains may disappear when few labeled spectra are available for training. These limitations motivate a general representation-learning approach that enables scalable, reliable, transferable, and data-efficient IR analysis. Foundation models provide such a general representation-learning approach, as demonstrated by large language models (LLMs), whose large-scale pretraining supports adaptation across diverse downstream tasks [9, 6]. Similar large-scale representation-learning paradigms have emerged across sensing domains, including medical imaging [49, 29], continuous glucose monitoring [28], neural-activity modeling [42], remote sensing [44], and robotic tactile sensing [25, 40]. However, foundation-model approaches remain underexplored in IR spectroscopy, with their development constrained by the scarcity of large, chemically diverse collections of labeled experimental spectra [34, 3]. Simulated IR spectra offer a scalable, complementary source of pretraining data that can broaden molecular coverage and support shared representation learning [30, 2, 52, 51]. To date, however, simulated IR spectra have mainly been used within individual chemical sensing and analysis tasks, with reported performance gains demonstrating their value for data-driven IR analysis [1]. Together, these observations motivate a simulation-to-real representation-learning paradigm that combines large-scale pretraining on simulated spectra with supervised adaptation using task-specific labeled experimental spectra, offering a promising route toward transferable and data-efficient IR foundation models. Specifically, pretraining can capture recurring relationships among IR spectral patterns, molecular structures and chemical properties across a broad chemical space, while downstream supervision adapts the shared representation to each chemical sensing and analysis objective. Here we introduce UltraIR, a foundation model for IR spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. UltraIR uses a two-stage learning framework in which a shared spectral encoder is first pretrained on approximately 60 million simulated IR spectra. For each downstream objective, the pretrained encoder is paired with a newly initialized task-specific module, and all components are jointly optimized using labeled experimental IR spectra. The encoder combines a derivative-aware multi-channel input, hierarchical convolutional modules, and a patch-based Transformer. These components capture local line shapes and peak shifts, extract multiscale vibrational signatures, and model longer-range spectral dependencies, respectively. Pretraining integrates three complementary objectives. Wavelet-domain spectral reconstruction encourages preservation of broad spectral envelopes and fine absorption features. Molecular fingerprint similarity alignment organizes the latent space according to graded structural relationships among molecules, while multi-label functional-group prediction anchors the representation to interpretable chemical motifs. Together, this design couples broad molecular exposure from simulated spectra with supervised adaptation using task-specific labeled experimental IR spectra, often available only in limited numbers, to provide a shared representation backbone for molecular interpretation, mixture analysis, and chemical sensing in complex samples. We evaluate UltraIR across molecular and mixture benchmarks, as well as real-world chemical sensing applications. The benchmarks cover functional-group prediction, molecular structure elucidation, physicochemical property prediction, and mixture-component identification and quantification. The real-world applications include bacterial classification, medicinal-herb geographic origin traceability and constituent quantification using two newly generated in-house datasets, microplastics classification, and soil property prediction. Across these evaluations, UltraIR outperforms conventional machine-learning and task-specific deep-learning baselines. The practical value of UltraIR’s simulation-to-real transfer learning is demonstrated under two demanding conditions encountered in real-world chemical sensing: downstream adaptation with limited labeled experimental spectra and zero-shot inference for the same analytical task across Fourier-transform infrared (FTIR) spectrometers and laboratories. More broadly, this learning framework provides a basis for adaptable and data-efficient chemical sensing systems that extract reliable chemical information from IR spectra of complex real-world samples. 2 Results UltraIR overview UltraIR is a foundation model for IR spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. The framework integrates two complementary resources: approximately 60 million simulated IR spectra for large-scale pretraining and a diverse collection of downstream datasets for supervised adaptation and evaluation across molecular-level benchmarks, mixture analysis, and real-world chemical sensing applications (Fig. 1a). The standardized pretraining dataset comprises molecule-associated IR spectra from three sources: approximately 1.5 million publicly released simulated spectra from IRtoMol, the multimodal spectroscopy dataset, and QM9S [1, 2, 52]; approximately 7.5 million newly generated spectra from molecular-dynamics simulations; and approximately 51 million spectra generated using a machine-learning-based spectral predictor [30]. Before pretraining, simulated spectra corresponding to molecules present in the downstream molecular datasets were excluded from the pretraining pool to prevent molecule-level overlap. Together, these sources provide the scale and molecular diversity required for large-scale pretraining. For supervised adaptation and evaluation, we assembled approximately 120,000 experimental IR spectra together with the simulated USPTO molecular benchmark. Molecular-level benchmarks use spectra from the National Institute of Standards and Technology (NIST) Chemistry WebBook [26], the Spectral Database for Organic Compounds (SDBS) [32], and USPTO [51] for functional-group prediction, molecular structure elucidation, and physicochemical property prediction. Mixture analysis uses NIST, SDBS, and USPTO spectra for targeted component detection and targeted fractional contribution estimation, the DeepMIR liquid-mixture dataset for additional evaluation of targeted component detection [41], and FTIR mixture spectra for mixture-level component quantification [4]. Real-world chemical sensing applications include genus-level classification of bacterial isolates from green snow [37], medicinal-herb geographic origin traceability and constituent quantification, microplastics classification using environmentally sourced spectra [46], and soil property prediction using the Open Soil Spectral Library [36]. Together, these datasets span chemical sensing and analysis from molecules and mixtures to biological and botanical samples and complex environmental samples. UltraIR follows a two-stage simulation-to-real learning framework (Fig. 1a). During pretraining, a shared encoder learns spectral representations from simulated IR spectra through three complementary objectives. Wavelet-domain spectral reconstruction encourages the encoder to retain coarse spectral envelopes and fine absorption features; molecular fingerprint similarity alignment encourages the latent space to reflect graded structural similarity among molecules; and multi-label functional-group prediction links the learned representations to interpretable chemical motifs. Together, these objectives integrate spectral morphology, molecular structural relationships, and functional-group information into a shared representation. For downstream adaptation, the pretrained encoder initializes the shared representation backbone, while a newly initialized task-specific head is added for each downstream objective. The encoder and head are then jointly optimized using the corresponding supervised data. This framework combines broad exposure to chemical diversity through pretraining on simulated spectra with task-specific supervision, enabling transfer of the shared representation across analytical objectives. Figure 1: Overview of the UltraIR simulation-to-real framework, encoder architecture, and molecular- and mixture-level analytical tasks. a, UltraIR data resources and the two-stage simulation-to-real learning framework. The pretraining dataset contains approximately 60 million simulated IR spectra assembled from publicly released spectrum–molecule pairs and spectra predicted or simulated for PubChem-derived molecules using machine-learning models and molecular dynamics. The shared encoder is pretrained with three chemically motivated objectives: wavelet-domain spectral reconstruction, molecular fingerprint similarity alignment and multi-label functional-group prediction. For downstream adaptation, the pretrained encoder initializes the shared representation backbone and is jointly optimized with a newly initialized task-specific head using supervised downstream data, including approximately 120,000 experimental IR spectra from gas-, liquid-, and solid-phase samples. b, UltraIR encoder architecture. The original spectrum is concatenated with first- and second-derivative channels generated by the derivative encoding module and processed through hierarchical residual spectral blocks. The resulting features are concatenated with a [CLS] token, combined with positional encodings, and processed by a Transformer encoder to produce the final [CLS] representation. Colors distinguish UltraIR components and separate trainable from non-trainable operations; symbols denote concatenation, summation, and element-wise multiplication. c, Molecular- and mixture-level analytical tasks used to evaluate UltraIR, comprising functional-group prediction, molecular structure elucidation, physicochemical property prediction, and mixture-component identification and quantification. Mixture analysis includes targeted component detection, targeted fractional contribution estimation, and mixture-level component quantification. Downstream task landscape To evaluate whether the shared spectral representation transfers across analytical objectives and datasets, we assessed UltraIR across a downstream task landscape spanning molecular- and mixture-level benchmarks and real-world chemical sensing applications. The molecular- and mixture-level benchmarks link IR spectra to functional-group annotations, molecular structures, physicochemical properties, and mixture composition (Fig. 1c), whereas the real-world applications evaluate chemical inference from complex biological, botanical, and environmental samples. Functional-group prediction evaluates whether UltraIR can identify functional groups from their characteristic vibrational features. Molecular structure elucidation evaluates whether UltraIR can identify the correct molecular structure from spectral information under a specified molecular-formula constraint. Physicochemical property prediction examines whether the shared spectral representation encodes continuous molecular attributes beyond discrete structural annotations. Mixture analysis assesses component detection and quantification through three tasks. In targeted component detection, the model receives a single-compound reference spectrum and a mixture spectrum and predicts whether the corresponding compound is present in the mixture. In targeted fractional contribution estimation, the same pair of spectra is used to estimate the target compound’s fractional contribution. In mixture-level component quantification, the model receives only a mixture spectrum and predicts the proportions of multiple components. Real-world chemical sensing applications span four sample domains and multiple analytical objectives. Genus-level bacterial classification evaluates whether biological spectral fingerprints can distinguish fast-growing bacterial isolates collected from green snow. Medicinal-herb characterization combines geographic origin traceability with chemical constituent quantification, requiring both categorical and continuous inference from botanical spectra. Microplastics classification uses environmental IR spectra to distinguish polymer types. Soil property prediction evaluates whether IR spectra support the prediction of continuous soil physicochemical properties. The following sections benchmark UltraIR against established conventional machine-learning and task-specific deep-learning methods appropriate to each task. UltraIR enables molecular structural interpretation from infrared spectra We first evaluated UltraIR on functional-group prediction and molecular structure elucidation using the public experimental NIST and SDBS datasets and the simulated USPTO benchmark [26, 32, 51]. Functional-group prediction. For functional-group prediction, UltraIR was benchmarked against task-specific deep-learning models, including FCGFormer and IRAnalysis [12, 34], and conventional machine-learning baselines, including XGBoost, random forest (RF), k-nearest neighbors (KNN), and a logistic classifier implemented using scikit-learn [10, 8, 33]. Performance was evaluated using three standard and complementary metrics: Micro-F1, Macro-F1, and exact match ratio (EMR). Micro-F1 and Macro-F1 are commonly used label-level metrics for multi-label classification. Micro-F1 aggregates true positives, false positives, and false negatives across all functional-group labels, providing an overall measure of label-wise prediction performance, whereas Macro-F1 first computes the F1 score for each functional group and then averages across groups, giving equal weight to frequent and infrequent functional groups. In contrast, EMR provides a molecule-level assessment of complete functional-group assignment: a prediction is counted as correct only when the full set of functional groups is recovered without missing labels or false assignments. We assessed model behavior from three perspectives: overall performance on each dataset, performance variation as a function of labeled training-data size, and performance for individual functional groups. Across the three datasets and all three metrics, UltraIR consistently and substantially outperformed both task-specific deep-learning models and conventional machine-learning baselines (Fig. 2a). On SDBS and USPTO, FCGFormer and IRAnalysis either only marginally outperformed XGBoost or performed comparably to it. On NIST, both models slightly underperformed XGBoost across all three metrics. These results demonstrate that task-specific deep learning alone does not confer a decisive advantage over strong conventional baselines for functional-group prediction. By contrast, UltraIR established a clear and consistent performance advantage over both classes of methods, with the most pronounced gains in EMR on NIST and USPTO. These gains are chemically meaningful rather than merely statistical. Functional-group prediction converts an IR spectrum into an interpretable molecular annotation that chemists can use to assess molecular composition, validate spectral interpretations, and evaluate plausible structural hypotheses. A partially correct label-wise prediction may still miss a key group or introduce a false assignment, thereby altering the chemical interpretation of the molecule even when Micro-F1 or Macro-F1 remains high. Higher EMR therefore shows that UltraIR more frequently recovers the complete functional-group set, rather than merely improving isolated label detections. We next examined whether UltraIR’s performance advantage persisted as the amount of labeled training data varied. Using EMR as the evaluation metric, UltraIR consistently outperformed all competing methods at every training proportion on NIST and USPTO (Fig. 2b). On SDBS, UltraIR performed comparably to the strongest competing method at the two lowest training proportions, 0.1% and 1%, and clearly outperformed it at all higher proportions. By contrast, the two task-specific deep-learning models, particularly FCGFormer, underperformed the strongest conventional baseline at the two lowest training proportions and were markedly inferior to UltraIR on NIST and USPTO. These results support the central simulation-to-real premise of UltraIR: large-scale simulated spectra provide transferable chemical information that enables reliable inference across a broad range of labeled-data availability, particularly when downstream supervision is scarce. Finally, we examined whether the aggregate performance advantage of UltraIR extended across individual functional groups rather than being driven by a small number of frequent labels. Per-group F1-score rankings placed UltraIR consistently among the top-ranked methods across NIST, SDBS, and USPTO (Fig. 2c). This pattern covered 17 functional groups spanning hydrocarbons and unsaturated motifs, oxygen- and nitrogen-containing groups, halogenated groups, and carbonyl-containing derivatives. The aggregate gains therefore reflect broad performance across chemically diverse functional groups rather than gains driven by a narrow subset of common labels. This breadth supports the use of UltraIR for functional interpretation across diverse molecular chemistry. Molecular structure elucidation. For molecular structure elucidation, UltraIR was benchmarked against representative IR structure-elucidation models, including IRtoMol, AISE, PBSA, and DLIR [1, 3, 45, 47]. Given an IR spectrum and its molecular formula, each model generates a ranked list of molecular candidates represented as Simplified Molecular Input Line Entry System (SMILES) strings. Performance was evaluated using top-1, top-5, and top-10 accuracy, together with fingerprint-based Tanimoto similarity. The top-k metrics assess exact structure recovery by determining whether the reference molecule appears among the k highest-ranked candidates. Fingerprint-based Tanimoto similarity provides a complementary assessment of structural agreement between the top-ranked candidate and the reference molecule, with higher values indicating closer structural similarity. We examined performance from three perspectives: overall exact-recovery accuracy across datasets, structural similarity of top-ranked candidates, and the overlap and characteristics of molecules correctly solved by different methods. Across NIST, SDBS, and USPTO, UltraIR achieved the strongest overall performance among the compared methods for molecular structure elucidation (Fig. 2d). Its top-1, top-5, and top-10 accuracies show that the learned spectral representation can prioritize the reference structure within a molecular-formula-constrained candidate set. In contrast, the other methods exhibited greater dataset-to-dataset variation in both overall performance and relative ranking. Figure 2: UltraIR enables molecular structural interpretation from infrared spectra. a, Overall functional-group prediction performance on NIST, SDBS, and USPTO, evaluated using Micro-F1, Macro-F1, and exact match ratio (EMR). EMR counts a prediction as correct only when the complete functional-group set is recovered. b, Training-data scaling analysis for functional-group prediction, showing EMR as the fraction of labeled training spectra varies from 0.1% to 100%. c, Functional-group-wise F1-score rankings across 17 chemically diverse labels on NIST, SDBS, and USPTO, with lower ranks indicating stronger relative performance for each functional group. d, Molecular structure elucidation on NIST, SDBS, and USPTO, evaluated using top-1, top-5, and top-10 accuracy. A prediction is counted as correct when at least one of the top k candidates matches the ground-truth molecular structure after canonicalization. e, Paired comparison of fingerprint-based Tanimoto similarity between each ground-truth molecule and the corresponding top-1 candidates generated by UltraIR and IRtoMol on the NIST test set. Lines connect predictions for the same molecule. f, Correct-case overlap among UltraIR, IRtoMol, AISE, PBSA, and DLIR on the NIST test set. The 15 most frequent correct-case subsets are shown together with the fraction of molecules for which all methods were incorrect. Blue and red denote subsets including and excluding UltraIR, respectively, whereas gray denotes cases missed by all methods. g, Two representative top-1 structure-elucidation examples from the NIST test set, comparing molecular structures predicted by UltraIR and competing methods. We further assessed the structural similarity of the top-1 candidates on NIST using fingerprint-based Tanimoto similarity. For each spectrum, we compared the Tanimoto similarity between the ground-truth molecule and the top-1 candidates generated by UltraIR and IRtoMol, the strongest competing method on NIST in the exact-recovery evaluation. UltraIR produced candidates with higher, equal, and lower similarity than IRtoMol for 44%, 41%, and 15% of spectra, respectively (Fig. 2e). Thus, UltraIR matched or exceeded IRtoMol in structural similarity for 85% of the evaluated spectra, and higher-similarity candidates outnumbered lower-similarity candidates by 29 percentage points. Together with the exact-recovery results, this comparison shows that UltraIR not only more often ranks the ground-truth molecule among its leading candidates, but also more often produces a top-ranked candidate that is structurally closer to the ground-truth molecule. We next analyzed the overlap of correctly solved molecules across UltraIR, IRtoMol, AISE, PBSA, and DLIR on the NIST dataset, focusing on the 15 most frequent correct-case subsets (Fig. 2f). The largest subset comprised molecules that were not solved by any of the five methods, accounting for 44.0% of the test set. Among the displayed correct-case subsets, those including UltraIR accounted for 52.0% of test samples, whereas subsets excluding UltraIR accounted for only 2.9%. Notably, all five methods correctly solved 20.2% of the test set, while UltraIR alone correctly solved a further 9.0%. This pattern shows that UltraIR both recovers a substantial set of molecules shared with existing methods and contributes a sizable set of correct predictions not obtained by the competing structure-elucidation approaches. Representative NIST test cases further illustrate the chemical basis of this advantage (Fig. 2g). In the first case, the ground-truth molecule comprised an oxygen-rich fused polycyclic scaffold containing a methoxy substituent, cyclic ethers, and two ring-associated carbonyl groups. UltraIR recovered the exact top-1 structure with a Tanimoto similarity of 1.00. AISE and PBSA retained most of the fused scaffold but introduced an incorrect double bond within an oxygen-containing ring, each yielding a similarity of 0.66. IRtoMol introduced the same bond-order error and replaced a ring carbon with an oxygen atom, resulting in a similarity of 0.53, whereas DLIR generated a substantially different and less oxygenated polycyclic framework with a similarity of 0.20. In the second case, the ground-truth molecule contained a nitrogen-rich heteroaromatic core bearing carboxamide and formamide functionalities and two benzyl substituents. UltraIR again recovered the exact structure with a similarity of 1.00. By contrast, the competing candidates altered the heteroaromatic core and the connectivity of the carbonyl-containing and aromatic substituents, with DLIR additionally introducing a chlorine atom absent from the ground truth. IRtoMol, AISE, PBSA, and DLIR achieved similarities of only 0.34, 0.34, 0.30, and 0.21, respectively. Collectively, these results establish UltraIR as a strong approach for functional and structural interpretation of IR spectra, spanning functional-group prediction and molecular structure elucidation. UltraIR enables physicochemical property prediction from infrared spectra Beyond molecular structural interpretation, we evaluated whether UltraIR could predict continuous molecular physicochemical properties from IR spectra. Physicochemical property prediction was evaluated on NIST, SDBS, and USPTO for 11 structure-derived molecular descriptors and scores: synthetic accessibility (SA) score, logP P, topological polar surface area (TPSA), hydrogen-bond donors, hydrogen-bond acceptors, rotatable bonds, the fraction of sp3sp^3-hybridized carbon atoms (Fraction Csp3), quantitative estimate of drug-likeness (QED), aromatic rings, aliphatic rings, and Bertz complexity (BertzCT). UltraIR was compared with conventional regression baselines, including XGBoost regression, support-vector regression (SVR), KNN regression, and partial least-squares regression (PLSR) [10, 33, 43]. Performance was evaluated using normalized mean absolute error (normalized MAE), normalized root mean squared error (normalized RMSE), and the coefficient of determination R2R^2. Normalized MAE and normalized RMSE quantify scale-adjusted prediction error, whereas R2R^2 measures the proportion of variance explained by the model. We further analyzed property-wise R2R^2 rankings and BertzCT complexity-thresholded prediction error. Across all three datasets, UltraIR achieved the lowest normalized MAE and normalized RMSE, together with the highest R2R^2 among the evaluated methods (Fig. 3a). XGBoost was the strongest conventional baseline, whereas SVR, KNN, and PLSR showed larger normalized errors or lower R2R^2 values in several settings. Property-wise R2R^2 rankings for the ten properties included in the ranking analysis placed UltraIR first in all 30 property–dataset combinations (Fig. 3b). This consistent ranking shows that UltraIR’s advantage extends across the ten ranked properties rather than being driven by a single target. Figure 3: UltraIR enables physicochemical property prediction from infrared spectra. a, Overall physicochemical property prediction on NIST, SDBS, and USPTO, evaluated using normalized MAE, normalized RMSE, and R2R^2. b, Property-wise R2R^2 rankings for 10 of the 11 physicochemical properties across NIST, SDBS, and USPTO. The remaining property, BertzCT, is examined separately in c. Lower ranks indicate stronger relative performance for each property. c, Mean relative error for BertzCT prediction across cumulative molecular-complexity thresholds on NIST, SDBS, and USPTO. At each threshold, performance is evaluated for molecules with true BertzCT values greater than or equal to that threshold. We next examined BertzCT prediction across cumulative molecular-complexity thresholds. At each threshold, mean relative error was calculated for molecules whose true BertzCT value met or exceeded that threshold. UltraIR achieved the lowest mean relative error at every threshold across NIST, SDBS, and USPTO, whereas conventional regression baselines showed larger errors and more pronounced threshold-dependent variation, particularly for high-complexity subsets (Fig. 3c). These results indicate that UltraIR maintains accurate continuous-property prediction across a broad range of molecular complexity. Together, these results extend the molecular-level interpretation capability of UltraIR beyond discrete functional-group labels and molecular structure elucidation. The learned IR representations capture continuous physicochemical attributes across diverse molecular properties and complexity regimes. UltraIR enables mixture-component identification and quantification from infrared spectra We evaluated targeted component detection and fractional contribution estimation across NIST, SDBS, and USPTO. Targeted component detection under increasing mixture complexity was further evaluated using the DeepMIR liquid-mixture dataset [41], whereas mixture-level component quantification was assessed using FTIR mixture datasets [4]. Targeted component detection and fractional contribution estimation. For pairwise mixture analysis, we evaluated UltraIR on targeted component detection and targeted fractional contribution estimation across NIST, SDBS, and USPTO. In both tasks, the input consisted of a reference single-compound spectrum and a mixture spectrum. For targeted component detection, UltraIR was benchmarked against DeepMIR [41], reverse match, and the hit quality index (HQI). Performance was evaluated using accuracy, Macro-F1, and the area under the receiver operating characteristic curve (ROC-AUC). For targeted fractional contribution estimation, UltraIR was compared with a DeepMIR-derived regression baseline and conventional regression baselines, including XGBoost regression, support-vector regression (SVR), KNN regression, and partial least-squares regression (PLSR). Performance was evaluated using mean absolute error (MAE), root mean squared error (RMSE), and R2R^2. UltraIR achieved strong targeted component-detection performance across all three datasets (Fig. 4a). On NIST, UltraIR attained higher accuracy and Macro-F1 than DeepMIR while achieving comparable ROC-AUC. On SDBS and USPTO, UltraIR and DeepMIR performed comparably across all three metrics. Across all datasets, UltraIR consistently outperformed the direct spectral-matching baselines: reverse match and HQI. Figure 4: UltraIR enables mixture-component identification and quantification from infrared spectra. a, Targeted component detection on NIST, SDBS, and USPTO, evaluated using accuracy, Macro-F1, and ROC-AUC. UltraIR is compared with DeepMIR, reverse match, and hit quality index (HQI). b, Targeted component detection in binary, ternary, and quaternary mixtures, comparing UltraIR and DeepMIR using accuracy and Macro-F1. c, Targeted fractional contribution estimation on NIST, SDBS, and USPTO, evaluated using mean absolute error (MAE), root mean squared error (RMSE), and R2R^2. d, Parity plots for targeted fractional contribution estimation, comparing predicted and true target-component contributions for UltraIR and competing methods across NIST, SDBS, and USPTO. The identity line denotes ideal agreement, and dashed lines show fitted relationships. e, Normalized residual distributions for mixture-level multi-component quantification from mixture spectra. f, Component-wise R2R^2 values for mixture-level quantification of acrylonitrile, adiponitrile, propionitrile, and glycerol. We next examined whether this performance persisted as mixture complexity increased. UltraIR and DeepMIR performed comparably for binary and ternary mixtures, whereas UltraIR established a clearer advantage in both accuracy and Macro-F1 for quaternary mixtures (Fig. 4b). This pattern indicates that UltraIR retains stronger target-component detection as additional components increase spectral overlap and make the target contribution more difficult to isolate. UltraIR also achieved the strongest performance in targeted fractional contribution estimation across all three datasets, substantially and significantly outperforming the strong DeepMIR-derived regression baseline. It attained the lowest MAE and RMSE, together with the highest R2R^2 among all evaluated methods (Fig. 4c). Its R2R^2 values reached 0.956 on NIST, 0.986 on SDBS, and 0.996 on USPTO (Fig. 4d). The parity plots showed that UltraIR predictions closely followed the identity line, whereas competing methods exhibited larger deviations and weaker calibration (Fig. 4d). These results establish that UltraIR supports both qualitative identification of a target component and quantitative estimation of its contribution to a mixture spectrum. Mixture-level component quantification. We finally evaluated mixture-level component quantification, in which the model receives only a mixture spectrum and predicts the proportions of multiple components. This task followed a simulation-to-experiment transfer protocol: models were first trained on a simulated FTIR mixture dataset and then fine-tuned and evaluated only on experimentally acquired FTIR mixture spectra [4]. UltraIR was compared with the mixture-analysis models AIMWSP and ML-FTIR [48, 4], as well as conventional regression baselines including XGBoost regression, SVR, KNN regression, and PLSR. Performance was assessed using normalized residual distributions and component-wise R2R^2. UltraIR produced residuals most tightly concentrated around zero, indicating the most accurate and stable mixture-level concentration estimates among the evaluated methods (Fig. 4e). Component-wise analysis further confirmed this advantage. UltraIR achieved the highest R2R^2 for each of the four quantified components: acrylonitrile, adiponitrile, propionitrile, and glycerol (Fig. 4f). Several competing methods showed pronounced component-dependent degradation. PLSR yielded negative R2R^2 values across all four components, and KNN also produced a negative R2R^2 for propionitrile, indicating performance below that of a mean-value predictor in these settings. Together, these results show that UltraIR extends molecular-level IR analysis from individual compounds to chemical mixtures. It enables reference-guided component detection and fractional contribution estimation, as well as multi-component quantification from a single mixture spectrum. UltraIR thus supports functional-group prediction, molecular structure elucidation, physicochemical property prediction, and mixture analysis. UltraIR supports biological and botanical infrared sensing Having established UltraIR across molecular-level tasks, we next examined whether it could transfer to biological and botanical IR sensing tasks involving experimentally acquired spectra from complex real-world samples. We considered two applications with distinct sample types and analytical objectives: genus-level bacterial classification and medicinal-herb characterization, encompassing geographic origin traceability and chemical constituent quantification. Bacterial classification. For bacterial classification (Fig. 5a), UltraIR was evaluated on FTIR spectra of fast-growing bacterial isolates from green snow [37]. The task was formulated as genus-level classification and benchmarked against conventional machine-learning baselines, including XGBoost, RF, KNN, and a logistic classifier. Performance was evaluated using accuracy, Macro-F1, and the Matthews correlation coefficient (MCC), which provides a robust measure under class imbalance. UltraIR achieved the highest accuracy, Macro-F1, and MCC among all compared methods (Fig. 5c). Genus-wise F1 analysis further showed that UltraIR consistently outperformed the conventional baselines across all nine evaluated genera, whereas baseline performance varied substantially among genera (Fig. 5d). Thus, UltraIR’s aggregate advantage reflected broad genus-level discrimination rather than performance gains confined to a small subset of genera. Medicinal herb characterization. We next evaluated UltraIR on two newly generated in-house experimental datasets of Jinyinhua (Lonicerae Japonicae Flos) and Shanyinhua (Lonicerae Flos), comprising newly acquired IR spectra for both geographic origin traceability and chemical constituent quantification (Fig. 5b). The origin-traceability datasets comprised 120 Jinyinhua spectra from Shandong, Henan, Hebei, and Sichuan and 150 Shanyinhua spectra from Hunan, Hubei, Sichuan, Henan, and Guangdong, with 30 spectra per origin. Liquid chromatography–mass spectrometry (LC-MS)-derived relative abundances of the target constituents, expressed in arbitrary units (a.u.), served as regression labels for 60 Jinyinhua and 75 Shanyinhua samples. Details of the in-house datasets, including dataset composition, LC-MS sample preparation, and data acquisition, are provided in the Appendix. We retained a matched UltraIR model without simulated pretraining as an ablation for these tasks. For geographic origin traceability, UltraIR and the no-pretraining ablation were compared with XGBoost, RF, KNN, and a logistic classifier using accuracy, Macro-F1, and MCC. For chemical constituent quantification, they were compared with XGBoost regression, SVR, KNN regression, and PLSR using normalized residual distributions and constituent-wise R2R^2. For geographic origin traceability, UltraIR achieved the highest mean accuracy, Macro-F1 and MCC on both Jinyinhua and Shanyinhua, outperforming all conventional baselines (Fig. 5e). By contrast, the no-pretraining ablation underperformed UltraIR across all three metrics on both Jinyinhua and Shanyinhua and also fell below XGBoost, the strongest conventional baseline, in every comparison. A Uniform Manifold Approximation and Projection (UMAP) of embeddings extracted by UltraIR provided a complementary qualitative view of the learned spectral organization (Fig. 5f). Spectra from the same geographic origin largely co-localized within each herb species, while Jinyinhua and Shanyinhua occupied largely distinct regions of the embedding space rather than being broadly intermixed. This organization emerged even though herb species identity was not used as a supervised label in the origin-traceability task and is consistent with UltraIR encoding spectral variation associated with both medicinal-herb identity and geographic origin. We next moved from categorical origin assignment to quantitative estimation of chemical constituents. For both Jinyinhua and Shanyinhua, UltraIR produced normalized residual distributions that were more tightly concentrated around zero than those of the no-pretraining ablation and the conventional regression baselines, with fewer large residuals (Fig. 5g,i). The no-pretraining ablation nevertheless produced a compact interquartile residual distribution. Its interquartile range was narrower than that of KNN regression, the strongest conventional baseline by this criterion, on Jinyinhua and comparable to that of KNN regression on Shanyinhua. However, the ablation also produced more large-residual outliers. Because residual centering and dispersion do not establish how much constituent-specific variation a model explains, we next examined constituent-wise R2R^2. UltraIR achieved the highest R2R^2 for every evaluated constituent in both Jinyinhua and Shanyinhua, with particularly large gains over the no-pretraining ablation (Fig. 5h,j). Several conventional regression baselines showed marked constituent-dependent degradation and, in some cases, negative R2R^2 values. The no-pretraining ablation also showed weak constituent-level generalization despite its residual distributions being centered close to zero. On Jinyinhua, its R2R^2 values were negative for all six constituents and exceeded only those of PLSR, the weakest method on this dataset. On Shanyinhua, it produced the lowest R2R^2 among all methods for three of the four constituents. The consistent contrast between UltraIR and the matched ablation provides direct evidence that simulated pretraining was critical under limited labeled supervision. Figure 5: UltraIR supports biological and botanical infrared sensing. a, Schematic of genus-level bacterial classification from infrared spectra. b, Schematic of medicinal-herb characterization using in-house experimental datasets, comprising geographic origin traceability and chemical constituent quantification for Jinyinhua and Shanyinhua. c, Overall genus-level bacterial classification, evaluated using accuracy, Macro-F1, and the Matthews correlation coefficient (MCC). d, Genus-wise F1 scores across the nine evaluated bacterial genera. e, Geographic origin traceability of Jinyinhua and Shanyinhua, evaluated using accuracy, Macro-F1, and MCC. UltraIR is compared with its matched no-pretraining ablation and conventional classifiers. f, Uniform Manifold Approximation and Projection (UMAP) of medicinal-herb spectral embeddings extracted by UltraIR. Points are colored by geographic origin, and marker shapes indicate herb species. g, Normalized residual distributions for chemical constituent quantification in Jinyinhua, comparing UltraIR, its matched no-pretraining ablation, and conventional regression baselines. h, Constituent-wise R2R^2 values for the six quantified Jinyinhua constituents across the same methods. i, Normalized residual distributions for chemical constituent quantification in Shanyinhua, comparing UltraIR, its matched no-pretraining ablation, and conventional regression baselines. j, Constituent-wise R2R^2 values for the four quantified Shanyinhua constituents across the same methods. Together with the bacterial classification results, these findings show that UltraIR supports IR analysis of complex biological and botanical samples across genus-level identification, medicinal-herb geographic origin traceability, and chemical constituent quantification. UltraIR supports environmental infrared sensing Having established UltraIR’s performance in biological and botanical IR sensing, we finally evaluated the model for environmental IR sensing using experimentally acquired spectra from complex samples. We considered two representative environmental applications: microplastics classification and soil physicochemical property prediction. Microplastics classification. For microplastics classification (Fig. 6a), UltraIR was evaluated on IR spectra of environmentally sourced microplastics [46]. UltraIR was compared with task-specific deep-learning baselines, including Softmax and DB-CNN-CBAM [46, 18], and conventional machine-learning baselines, including XGBoost, RF, KNN, and a logistic classifier. Performance was evaluated using accuracy, Macro-F1, and MCC. We further analyzed the true-class margin, defined as the correct-class logit minus the highest logit among the non-true classes. Larger positive values indicate stronger separation of the correct class from its nearest competitor. Across accuracy, Macro-F1, and MCC, UltraIR achieved the strongest performance among all evaluated methods (Fig. 6b). Its advantage over both task-specific deep-learning and conventional machine-learning baselines demonstrates effective transfer from molecular benchmarks to polymer identification in complex environmental spectra. We next examined whether this aggregate performance advantage was accompanied by larger true-class margins. In the pairwise margin plots, many spectra lay above the diagonal, indicating that UltraIR assigned a larger true-class margin than Softmax or DB-CNN-CBAM, including numerous cases in which both methods correctly predicted the polymer class (Fig. 6c). UltraIR also corrected substantially more baseline errors than it introduced. Relative to Softmax, UltraIR corrected 472 baseline errors while introducing 112 errors. Relative to DB-CNN-CBAM, it corrected 657 baseline errors while introducing 97 errors. Thus, UltraIR not only corrected more misclassified spectra, but also achieved stronger separation between the correct polymer class and its nearest alternatives. Finally, class-wise F1 analysis showed that UltraIR maintained consistently high performance across all 18 polymer categories, including common commodity polymers and more specialized materials (Fig. 6d). This broad class-wise advantage shows that the overall improvement was not driven by a narrow subset of readily distinguishable polymers, but extended across the microplastics classification label space. Soil property prediction. For soil property prediction (Fig. 6e), we evaluated whether UltraIR could estimate continuous physicochemical properties from soil IR spectra using the Open Soil Spectral Library (OSSL) [36]. UltraIR was compared with soil spectral-modeling baselines, including GADF-Swin and LSTM-CNN [17, 27], as well as conventional regression baselines, including XGBoost regression, SVR, KNN regression, and PLSR. Performance was evaluated using normalized MAE and RMSE, together with property-wise R2R^2, across ten soil properties: total carbon, total nitrogen, total sulfur, clay content, pH (H2O)(H_2O), cation-exchange capacity (CEC), and exchangeable Ca, Mg, K, and Na. Figure 6: UltraIR supports environmental infrared sensing across microplastics and soil analysis. a, Schematic of microplastics classification from infrared spectra of environmentally sourced microplastics. b, Overall microplastics classification, evaluated using accuracy, Macro-F1, and the Matthews correlation coefficient (MCC). c, Pairwise true-class margin comparisons for UltraIR versus Softmax and UltraIR versus DB-CNN-CBAM. The true-class margin is defined as the correct-class logit minus the highest logit among the non-true classes, with points above the diagonal indicating a larger margin for UltraIR. Corrected and degraded samples denote cases classified correctly by UltraIR but incorrectly by the comparator, and vice versa, respectively. d, Class-wise F1 scores across 18 polymer categories. e, Schematic of soil physicochemical property prediction from infrared spectra. f, Overall soil physicochemical property prediction, evaluated using normalized mean absolute error (normalized MAE), normalized root mean squared error (normalized RMSE), and R2R^2. g, Training-data scaling analysis for soil physicochemical property prediction, reporting normalized RMSE and R2R^2 as the labeled training fraction varies from 0.1% to 100%. h, Parity plots for soil property prediction, comparing predicted and measured values for UltraIR, GADF-Swin, and LSTM-CNN across ten properties: total carbon, total nitrogen, total sulfur, clay content, pH (H2O)(H_2O), cation-exchange capacity (CEC), and exchangeable Ca, Mg, K, and Na. The identity line denotes ideal agreement. i, Cross-instrument and cross-laboratory soil property prediction. Bars show property-wise R2R^2 for soil texture and acidity and for soil chemical properties. Across the three aggregate regression metrics, UltraIR achieved the strongest overall performance in soil property prediction (Fig. 6f). It attained the lowest normalized MAE and normalized RMSE and the highest R2R^2 among all evaluated methods. We further assessed performance under reduced labeled-data settings. As the training-set fraction decreased from 100% to 0.1%, normalized RMSE increased and R2R^2 decreased across all methods (Fig. 6g). At the 1% training fraction, UltraIR achieved lower normalized RMSE than XGBoost regression and KNN regression and attained the highest R2R^2 among all evaluated methods, although its normalized RMSE was slightly higher than that of PLSR. At this fraction, UltraIR substantially outperformed the task-specific neural baselines GADF-Swin and LSTM-CNN, with LSTM-CNN showing particularly severe performance degradation. As the labeled training fraction increased to 10%, 50%, and 100%, both task-specific models achieved progressively lower normalized RMSE and higher R2R^2, but remained inferior to UltraIR. This strong performance with only 1% of the labeled training data demonstrates the data efficiency of UltraIR for soil spectroscopy, which is practically important because soil reference measurements are expensive to acquire and are often unevenly available across soil properties and sampling regions. Property-wise analysis further confirmed the consistency of this advantage. UltraIR achieved the highest R2R^2 for all ten soil properties (Fig. 6h). Its predictions closely followed the identity line for total carbon, total sulfur, clay content, exchangeable calcium, and pH, for which it achieved particularly strong performance. Total nitrogen remained the most challenging property across the evaluated methods, yet UltraIR still achieved the highest R2R^2. These results demonstrate broad predictive capability across soil attributes with distinct chemical origins, concentration ranges, and spectral signatures. Together with the microplastics results, these findings show that UltraIR supports environmental IR sensing across both discrete material classification and continuous quantitative estimation. Its accuracy, robustness under limited labeled data, and consistent property-wise performance support reliable IR-based analysis of complex environmental samples. Soil property prediction as a case study of cross-instrument and cross-laboratory generalization To examine UltraIR’s capacity to generalize across FTIR instruments and laboratories, we used soil property prediction as a case study. UltraIR was compared with the same soil spectral-modeling baselines: GADF-Swin, LSTM-CNN, XGBoost regression, SVR, KNN regression, and PLSR. Relative to the preceding experiment, this comparison used different subsets of soil IR spectra from OSSL [36] and a partially different set of soil properties. Performance was evaluated using property-wise R2R^2 for clay content, silt content, sand content, pH (H2O)(H_2O), CEC, and exchangeable Ca, Mg, K, and Na. All models were trained using spectra from the Kellogg Soil Survey Laboratory (KSSL) subset of OSSL and evaluated using zero-shot target-domain inference on the ICRAF–ISRIC subset of OSSL. The source and target subsets differed in laboratory, FTIR instrument, and sample preparation procedures. Detailed metadata and label distributions for both subsets are provided in Appendix Tables D and E. UltraIR achieved the highest R2R^2 for all soil texture and acidity properties and for most soil chemical properties (Fig. 6i). UltraIR outperformed all competing methods for CEC and exchangeable Mg, K, and Na, while performing comparably to GADF-Swin and better than the remaining methods for exchangeable Ca. UltraIR was also the only method to maintain positive R2R^2 across all evaluated properties. Performance differences were most pronounced for the chemical properties, for which several baselines yielded negative R2R^2, particularly for exchangeable Na. This case study demonstrates that UltraIR achieved more stable and generally stronger performance during zero-shot inference across instruments and laboratories, which represents an important practical challenge in real-world chemical sensing. 3 Discussion IR spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from complex spectra at scale remains difficult across analytical objectives and datasets. UltraIR addresses this challenge by enabling simulation-to-real transfer learning through a foundation model for IR spectroscopy with more than 100 million parameters. The model learns a shared spectral representation by pretraining on approximately 60 million simulated IR spectra using complementary objectives for spectral reconstruction, molecular fingerprint similarity alignment, and functional-group prediction. The pretrained encoder is then adapted to each downstream analytical objective using task-specific supervision. This design separates large-scale spectral representation learning from downstream adaptation, allowing the same pretrained encoder to support diverse forms of IR chemical sensing and analysis. A central advance of UltraIR lies in the breadth of chemical inference supported by its shared pretrained encoder. The evaluated tasks span molecular structural interpretation, physicochemical property prediction, mixture-component identification and quantification, and real-world chemical sensing in biological and botanical samples and complex environmental samples, with two newly generated in-house experimental datasets used for botanical sensing. These tasks differ in their inputs, output spaces, and analytical objectives, encompassing multi-label classification, molecular generation, regression, reference-guided mixture analysis, and multi-component quantification. Across these evaluations, UltraIR outperformed conventional machine-learning and task-specific deep-learning baselines. This consistent performance across tasks and datasets indicates that the learned representation can transfer after supervised adaptation rather than remaining tied to a single label space or dataset. UltraIR therefore provides a practical basis for reusing a shared pretrained encoder with a task-specific output module for each analytical objective. The practical value of UltraIR’s simulation-to-real transfer learning was particularly evident in two challenging real-world chemical sensing scenarios: downstream adaptation with limited labeled experimental spectra and zero-shot target-domain inference for the same analytical task across FTIR instruments and laboratories. In training-data scaling experiments, UltraIR retained stronger performance than task-specific neural baselines as labeled supervision decreased. The medicinal-herb ablation further showed that simulated pretraining benefited complex sample analysis under limited labeled supervision. A complementary soil property prediction case study showed that UltraIR maintained more stable and generally stronger performance than competing methods during zero-shot target-domain inference across instruments and laboratories. Simulated pretraining provides broad chemical exposure, whereas task-specific labels align the shared representation with individual analytical objectives, together supporting data-efficient learning and generalization across measurement conditions. Although UltraIR demonstrated strong performance across diverse infrared chemical analysis tasks, several limitations remain and point to directions for further improvement. UltraIR still requires supervised adaptation and a task-specific output module for each analytical objective. The present study therefore demonstrates broad transferability across analytical objectives after adaptation, rather than zero-shot or universal inference across tasks. A simulation-to-real domain gap also remains because simulated IR spectra cannot fully reproduce experimental variation. This variation can arise from instrument response, spectral resolution, measurement geometry, sample state, preparation procedures, temperature, concentration, scattering, complex matrices, and baseline artifacts. Improving simulation fidelity will require explicit modeling of these factors. Future adaptation strategies could exploit unlabeled experimental IR spectra to reduce the need for labeled experimental spectra. Overall, UltraIR establishes a scalable framework for simulation-to-real transfer learning in foundation modeling for IR spectroscopy. By coupling pretraining on chemically diverse simulated spectra with supervised downstream adaptation, UltraIR supports chemical sensing and analysis across analytical objectives and datasets, spanning molecules and mixtures, biological and botanical samples, and complex environmental samples. This framework provides a route beyond collections of independently trained task-specific models toward adaptable and data-efficient IR-based chemical sensing systems built on reusable spectral representations. 4 Methods Overview of UltraIR UltraIR is a foundation model for IR spectroscopy with more than 100 million parameters that enables simulation-to-real transfer learning for chemical sensing and analysis from molecules to complex samples. The framework comprises a shared spectral encoder and task-specific output modules. During pretraining, the encoder is optimized on approximately 60 million simulated IR spectra to learn a shared, chemically informative spectral representation. For downstream adaptation, the pretrained encoder is paired with a newly initialized task-specific output module, and both components are jointly optimized using labeled spectra for the corresponding task. The encoder architecture and pretrained parameters provide a common initialization across downstream objectives, whereas the output module, output dimensionality, and supervised training objective are defined separately for each task. In UltraIR, the shared encoder is initialized from the pretrained checkpoint. In designated ablation analyses, we additionally evaluated a matched no-pretraining control in which the same encoder architecture and task-specific output modules were initialized randomly and trained using the same downstream data and optimization procedure. Formally, given a preprocessed IR spectrum x, the shared encoder produces a latent representation =fθ(),z=f_θ(x), (1) where fθf_θ denotes the shared encoder with parameters θ. For UltraIR, θ is initialized from the pretrained parameters θpre _pre. For the no-pretraining control, θ is instead initialized randomly. A task-specific output module hϕ(t)h_φ^(t) then transforms the shared representation, together with any auxiliary input required by task t, into the corresponding analytical output: ^(t)=hϕ(t)(,(t)), y^(t)=h_φ^(t) (z,u^(t) ), (2) where (t)u^(t) denotes optional task-specific inputs, such as a molecular formula or a reference single-compound spectrum. The downstream outputs include functional-group annotations, formula-conditioned molecular structure candidates, molecular physicochemical property estimates, mixture-component detection and quantification outputs, bacterial genus labels, medicinal-herb geographic origin labels and constituent relative abundances, polymer-type labels for microplastics, and soil physicochemical property estimates. Spectral preprocessing and data augmentation Before being supplied to the model, all IR spectra were converted to a common spectral range, sampling grid, and intensity scale. The same deterministic preprocessing pipeline was used for large-scale pretraining on simulated spectra and downstream adaptation on experimental spectra. Stochastic augmentation was stage-specific. A broader set of augmentation operations was used during pretraining, whereas only weak, validation-selected augmentations were considered during downstream adaptation. Standard spectral preprocessing pipeline. Each spectrum was first cropped to the mid-IR range of 400 cm−1 to 4000 cm−1400\,cm^-14000\,cm^-1. Spectra that did not cover the complete range were zero-padded at the missing boundaries. The resulting spectrum was resampled by piecewise-linear interpolation onto a canonical 3,600-point grid and then min–max normalized to the interval [0,1][0,1]. This canonical representation was subsequently mapped to the fixed input length required by each model implementation. Pretraining augmentation. During pretraining, four stochastic augmentation operations were applied to each spectrum. Additive Gaussian noise was sampled at a random signal-to-noise ratio to represent electronic measurement noise. A random shift along the wavenumber axis represented calibration uncertainty. An additive intensity offset represented baseline drift. Finally, a random subset of spectral positions was masked by setting the corresponding intensities to zero. This masking objective encouraged the encoder to use information from the broader spectral context rather than relying exclusively on isolated local features. Downstream adaptation augmentation. For downstream adaptation, stochastic augmentation was not imposed as a universal part of the training pipeline. Instead, weak augmentations, including additive noise, small wavenumber shifts, and baseline offsets, were enabled only when selected using the validation split for the corresponding model–task setting. Random spectral masking was not used during downstream adaptation because most downstream objectives require access to the complete experimental spectral profile. Hybrid convolutional–Transformer architecture As illustrated in Fig. 1b, UltraIR uses a shared hybrid convolutional–Transformer encoder to map an IR spectrum to a latent spectral representation z for downstream chemical analysis. The encoder comprises three principal modules: a derivative-aware multi-channel input module, a hierarchical convolutional backbone with multi-scale feature fusion, and a patch-based Transformer encoder. Together, these modules combine derivative-aware spectral feature extraction, multi-scale local feature aggregation, and long-range spectral-context modeling. Derivative-aware multi-channel input module. Given the preprocessed IR spectrum ∈ℝ1×Lx ^1× L, where L denotes the number of spectral data points, this module augments the original absorbance signal with derivative-based spectral representations to enhance sensitivity to spectral line shape features. Specifically, x is first smoothed via a Savitzky–Golay (SG) filter to yield ~ x. Two learnable derivative convolutions then compute the first- and second-order spectral derivatives 1d_1 and 2d_2, where the corresponding kernels 1w_1 and 2w_2 are initialized as first- and second-order central-difference operators, respectively, and remain trainable throughout optimization. Each derivative signal is subsequently normalized using layer normalization, and the three signals are concatenated to form the three-channel representation cat=[,1,2]∈ℝ3×Lx_cat=[x,\,d_1,\,d_2] ^3× L. To adaptively reweight each spectral channel, catx_cat is passed through a gated input fusion module that computes a position-wise, input-dependent channel gate: =σ(2gGELU(1gcat))∈ℝ3×L,g=σ\! (W_2^g\,GELU\! (W_1^g\,x_cat ) ) ^3× L, (3) where 1gW_1^g and 2gW_2^g are 1×11×1 convolutional projections, GELU denotes the Gaussian error linear unit activation, and σ(⋅)σ(·) denotes the sigmoid function. The fused multi-channel representation is then obtained via element-wise modulation: multi=cat⊙∈ℝ3×L,x_multi=x_cat ^3× L, (4) where ⊙ denotes element-wise multiplication. Convolutional spectral encoder. The fused representation multi∈ℝ3×Lx_multi ^3× L is fed into a convolutional encoder that hierarchically extracts local spectral features across multiple spectral scales. Prior to feature extraction, multix_multi is rescaled by a learnable per-channel scale vector ∈ℝ3 γ ^3, broadcast along the spectral dimension and initialized as [1.0, 0.5, 0.5]⊤[1.0,\,0.5,\,0.5] to encode the inductive bias that the original absorbance signal should initially dominate over the two derivative channels, while allowing the network to adjust this balance during training, yielding ^multi∈ℝ3×L x_multi ^3× L. The encoder then begins with a convolutional stem that applies a wide-kernel convolution followed by batch normalization and a GELU activation to ^multi x_multi, mapping it to an initial feature map 0∈ℝC0×LH_0 ^C_0× L, where C0C_0 denotes the base channel width. Four successive residual blocks then progressively encode the representation. The first three blocks each apply a strided convolution with stride 2, yielding intermediate feature maps 1∈ℝC1×L/2F_1 ^C_1× L/2, 2∈ℝC×L/4F_2 ^C× L/4, and 3∈ℝC×L′F_3 ^C× L , where C1=2C0C_1=2C_0, C=4C0C=4C_0, and L′=L/8L =L/8. The fourth block operates at the same spectral resolution, producing 4∈ℝC×L′F_4 ^C× L . Each residual block comprises two convolutional layers, each followed by batch normalization, with a GELU activation applied after the first convolution–batch normalization (Conv-BN) layer and again after the residual addition, together with a squeeze-and-excitation (SE) module for adaptive channel-wise recalibration. Formally, for block i∈1,2,3,4i∈\1,2,3,4\ with input i−1F_i-1 (where 0=0F_0=H_0), the residual branch is defined as: fi()=BN(i(2)∗GELU(BN(i(1)∗))),f_i(F)=BN\! (W_i^(2) \! (BN\! (W_i^(1) ) ) ), (5) where i(1)W_i^(1) and i(2)W_i^(2) are the convolutional filters of the two Conv-BN layers within block i. The block output is then obtained by adding the residual branch to the shortcut and applying a final activation: ~i=GELU(fi(i−1)+i(i−1)), F_i=GELU\! (f_i(F_i-1)+P_i(F_i-1) ), (6) where i(⋅)P_i(·) is a strided pointwise projection for the first three blocks and the identity mapping for the fourth. The SE module then computes a channel-wise attention vector: i=σ(i,2seReLU(i,1seGAP(~i))),s_i=σ\! (W_i,2^se\,ReLU\! (W_i,1^se\,GAP( F_i) ) ), (7) where ReLU(⋅)ReLU(·) denotes the rectified linear unit activation, GAP(⋅)GAP(·) denotes global average pooling, and i,1seW_i,1^se and i,2seW_i,2^se are pointwise linear projections. The recalibrated block output is then i=~i⊙iF_i= F_i _i, where is_i is broadcast along the spectral dimension. To preserve fine-grained spectral details from shallower layers, multi-scale feature fusion is performed by resampling 2F_2 and 1F_1 to resolution L′L via nearest-neighbor interpolation ℛ(⋅)R(·), with 1F_1 additionally projected to C channels via a pointwise convolution ϕ(⋅)φ(·). The resampled features are concatenated with 4F_4 along the channel dimension to form =[4,ℛ(2),ϕ(ℛ(1))]∈ℝ3C×L′.Y= [F_4,\;R(F_2),\;φ\! (R(F_1) ) ] ^3C× L . (8) A pointwise channel-mixing multilayer perceptron (MLP) ψ(⋅)ψ(·) is then applied independently at each spectral position to reduce the channel dimension from 3C3C to C, producing the unified feature map: =ψ()∈ℝC×L′.F=ψ(Y) ^C× L . (9) Patch-based Transformer encoder. The convolutional feature map ∈ℝC×L′F ^C× L is tokenized by a strided convolutional patch embedding layer with kernel size P and stride P/2P/2, which projects each overlapping local segment into a d-dimensional token embedding and produces a sequence of N patch tokens ∈ℝN×dE ^N× d. A learnable [CLS] token cls∈ℝde_cls ^d is prepended, and a learnable positional embedding pos∈ℝ(N+1)×dE_pos ^(N+1)× d is added to encode sequential positional information, yielding the initial token sequence: (0)=[cls;]+pos∈ℝ(N+1)×d.T^(0)= [e_cls;\,E ]+E_pos ^(N+1)× d. (10) The sequence is then processed by NlayersN_layers stacked Transformer encoder layers with pre-layer normalization, where each layer applies a multi-head self-attention sub-layer (H heads, per-head dimension dh=d/Hd_h=d/H) followed by a two-layer feed-forward network with GELU activation and hidden dimension 4d4d, each preceded by layer normalization and wrapped with a residual connection. The final hidden state corresponding to the [CLS] token is extracted from the last Transformer layer as the global latent spectral representation: =[0](Nlayers)∈ℝd.z=T^(N_layers)_[0] ^d. (11) Task-specific downstream architectures The hybrid convolutional–Transformer encoder is shared across downstream tasks. Most classification and regression tasks apply the shared spectral encoder to a single input spectrum and pass the resulting representation to a task-specific MLP head. Two task families require additional task-specific modules. Formula-conditioned molecular structure elucidation takes a molecular formula as an additional input and generates SMILES strings. Pairwise mixture analysis takes a reference single-compound spectrum together with a mixture spectrum to determine whether the reference compound is present and to estimate its fractional contribution. By contrast, mixture-level component quantification from a mixture spectrum alone follows the standard single-spectrum regression pipeline. The following sections describe the additional modules used for structure generation and pairwise mixture analysis. Molecular structure elucidation. Molecular structure elucidation is formulated as formula-conditioned autoregressive SMILES generation. Given an IR spectrum and its associated molecular formula, the model generates a ranked list of candidate SMILES strings by connecting the pretrained encoder to a molecular formula encoder, a Perceiver-style IR token resampler, a spectral vocabulary prompt module, and an autoregressive SMILES decoder. The molecular formula string is tokenized at the character level and encoded by a formula encoder. The encoder consists of a token embedding layer with sinusoidal positional encodings and dropout, followed by NfN_f pre-norm Transformer encoder layers, each with multi-head self-attention and a GELU feed-forward sublayer, and a final layer normalization. It produces formula hidden states f∈ℝMf×dH_f ^M_f× d, where MfM_f denotes the number of formula tokens. On the spectral side, the backbone extracts both the global [CLS] feature ∈ℝdz ^d and the full patch token sequence patch∈ℝN×dT_patch ^N× d from its final Transformer layer. A projection head comprising layer normalization, a linear layer, and GELU activation maps z to a single IR memory token IR∈ℝdm_IR ^d. The patch tokens are further compressed by a Perceiver-style resampler into K latent vectors. The resampler maintains K learnable query embeddings and refines them over NrN_r cross-attention blocks. Unlike standard post-norm Perceivers, each block applies pre-norm layer normalization separately to the queries and the key–value source before cross-attention. A pre-norm GELU feed-forward sublayer follows, with residual connections throughout. Denoting the latent state at layer l as (l)∈ℝK×dX^(l) ^K× d, the update rule per block is: (l) ^(l) =(l−1)+CrossAttn(LNq((l−1)),LNkv(patch)), =X^(l-1)+CrossAttn\! (LN_q\! (X^(l-1) ),\;LN_kv\! (T_patch ) ), (12) (l) ^(l) ←(l)+FFN((l)), ^(l)+FFN\! (X^(l) ), (13) with (0)=lat∈ℝK×dX^(0)=Q_lat ^K× d the learnable query matrix, yielding lat=(Nr)∈ℝK×dM_lat=X^(N_r) ^K× d. A soft vocabulary prompt token is constructed from z to bias the decoder toward chemically plausible token sequences. A task-specific MLP head maps z to logits over the SMILES vocabulary of size V. A temperature-scaled softmax then converts these logits into a probability distribution, which is used to compute a soft embedding as a weighted sum over the decoder embedding matrix ∈ℝV×dE ^V× d. The soft embedding is then projected and scaled by a learnable scalar gate α to yield the prompt token: =softmax(MLPcls()τ)∈ℝV,prompt=α⋅fproj(⊤)∈ℝd,p=softmax\! ( MLP_cls(z)τ ) ^V, _prompt=α· f_proj\! (p E ) ^d, (14) where MLPclsMLP_cls denotes the task-specific vocabulary-classification head, τ is a fixed temperature hyperparameter, and fprojf_proj denotes a three-stage projection comprising layer normalization, a linear transformation, and GELU activation. The complete decoder memory is formed by concatenating all spectral tokens and the formula hidden states: =[IR;lat;prompt;f]∈ℝ(K+2+Mf)×d.M= [m_IR;\;M_lat;\;m_prompt;\;H_f ] ^(K+2+M_f)× d. (15) The SMILES decoder consists of NsN_s pre-norm Transformer layers, each applying causal self-attention under a triangular mask, cross-attention to M, and a GELU feed-forward sublayer, followed by a final layer normalization and a linear projection to vocabulary logits. During training, teacher forcing is applied, and the primary loss is token-level cross-entropy ℒLML_LM. Two contrastive losses further shape the representation. An IR–formula contrastive loss ℒIR-FL_IR -F aligns z with the masked mean-pooled formula encoding of fH_f, whereas an IR–target contrastive loss ℒIR-TL_IR -T aligns z with the masked mean-pooled decoder hidden states at the final Transformer layer. Both use symmetric information noise-contrastive estimation (InfoNCE) with two-layer projection heads that map the inputs into a shared dcd_c-dimensional contrastive space. The training objective is: ℒstruct=ℒLM+λIFℒIR-F+λITℒIR-T.L_struct=L_LM+ _IF\,L_IR -F+ _IT\,L_IR -T. (16) At inference, beam search with length-normalized scoring generates multiple ranked candidate SMILES sequences. A formula-consistency reranking step then promotes candidates whose SMILES match the exact atomic composition of the provided molecular formula and demotes chemically invalid structures. Targeted component detection and contribution estimation. The pairwise mixture-analysis tasks, targeted component detection and targeted fractional contribution estimation, are unified under a common spectral comparison framework. Given a reference single-compound spectrum and a mixture spectrum, the model predicts whether the reference compound is present in the mixture and estimates its fractional contribution. The two tasks share the same pairwise representation backbone and differ only in their final prediction head. Each input pair (ref,mix)∈ℝ1×L×ℝ1×L(x_ref,\,x_mix) ^1× L×R^1× L is encoded by three parallel branches that extract complementary representations of the spectral pair. The first branch uses a weight-shared encoder to process each spectrum independently, producing a global [CLS] feature and a patch token sequence for each spectrum. The patch tokens are projected to dpd_p dimensions and aggregated by attentive pooling, combining an attention-weighted sum with an element-wise maximum: pool=12(∑n=1Nαnn+maxnn),αn=exp(sn)∑n′exp(sn′),p_pool= 12\! ( _n=1^N _n\,t_n+ *max_n\,t_n ), _n= (s_n) _n (s_n ), (17) where nt_n are the projected patch tokens and sns_n is an attention score produced by layer normalization followed by a linear projection. The second branch operates directly on the raw input signals. The reference spectrum, mixture spectrum, and their absolute pointwise difference are stacked into a three-channel tensor [ref,mix,|mix−ref|]∈ℝ3×L[x_ref,\,x_mix,\,|x_mix-x_ref|] ^3× L. Successive Conv-BN-GELU blocks with progressively doubling channel widths process this tensor and reduce its spectral resolution by a fixed factor. The resulting feature sequence is augmented with sinusoidal positional encodings, processed by a shallow Transformer encoder, attentively pooled, and projected to yield the joint feature joint∈ℝdph_joint ^d_p. The third branch provides complementary frequency-domain features. A real-input fast Fourier transform (FFT) is applied to each spectrum, yielding complex coefficients r^f r_f and m^f m_f for the reference and mixture spectra at frequency bin f. Eight frequency-domain channels are then constructed: the log-amplitude spectra log(1+|r^f|) (1+| r_f|) and log(1+|m^f|) (1+| m_f|), their sum and absolute difference, the normalized cross-magnitude term κf=|r^fm^f∗|/(|r^f||m^f|+ε) _f=| r_f m_f^*|/(| r_f|\,| m_f|+ ), the normalized cross-spectral phase ϕf=∠(r^fm^f∗)/π _f= ( r_f m_f^*)/π, the log-amplitude ratio ρf=log(1+(|m^f|+ε)/(|r^f|+ε)) _f= \! (1+(| m_f|+ )/(| r_f|+ ) ), and the log-amplitude product πf=log(1+|r^f||m^f|) _f= \! (1+| r_f|\,| m_f| ), where (⋅)∗(·)^* denotes complex conjugation, ∠(⋅) (·) denotes the complex argument, and ε is a small constant for numerical stability. These channels are concatenated to form the frequency-domain tensor freq∈ℝ8×FX_freq ^8× F, which is processed by three convolutional sub-branches with short, medium, and long receptive fields. The resulting feature maps are concatenated, compressed by strided convolutions, globally average-pooled, and projected to yield the frequency representation freq∈ℝdph_freq ^d_p. The pair representation pairr_pair is assembled by concatenating five dpd_p-dimensional feature vectors with a ten-dimensional scalar statistics vector statss_stats. The five feature vectors are the projected [CLS] features refc_ref and mixc_mix, their absolute difference |ref−mix||c_ref-c_mix|, the absolute difference of the pooled token features |ref−mix||p_ref-p_mix|, and the frequency feature freqh_freq. The statistics vector statss_stats collects the cosine similarity between the [CLS] features, the cosine similarity between the pooled-token features, the cosine similarity between jointh_joint and freqh_freq, the mean squared activation energy of the joint Transformer tokens, the cosine similarity of the log-amplitude spectra, the mean and maximum absolute log-amplitude difference, the mean and maximum normalized cross-magnitude term, and the mean absolute normalized cross-spectral phase: pair=[ref;mix;|ref−mix|;|ref−mix|;freq;stats].r_pair= [c_ref;\;c_mix;\;|c_ref-c_mix|;\;|p_ref-p_mix|;\;h_freq;\;s_stats ]. (18) The pair representation pairr_pair is then passed to a task-specific prediction head. For targeted component detection, a binary classification head maps pairr_pair to a scalar logit and is trained with binary cross-entropy loss. For targeted fractional contribution estimation, a regression head maps pairr_pair to a scalar contribution estimate a^∈[0,1] a∈[0,1] through a sigmoid output layer and is trained with mean squared error against the normalized target contribution. Pretraining In the large-scale simulation-based pretraining stage, each preprocessed IR spectrum x is first stochastically augmented to obtain an encoder input ~ x. The augmented spectrum is passed through the hybrid convolutional–Transformer encoder fθf_θ, producing the latent representation =fθ(~)z=f_θ( x). Three complementary objectives are then applied to outputs derived from z: wavelet-domain spectral reconstruction, molecular fingerprint similarity alignment, and multi-label functional-group prediction. Together, these objectives encourage the learned representation to capture spectral morphology, molecular structural relationships, and interpretable functional-group signatures. The total pretraining loss is their weighted combination: ℒpre=λreconℒrecon+λcontrastℒcontrast+λfgℒfg,L_pre= _recon\,L_recon+ _contrast\,L_contrast+ _fg\,L_fg, (19) where λrecon _recon, λcontrast _contrast, and λfg _fg are scalar loss weights. The encoder and pretraining heads are optimized jointly by backpropagating ℒpreL_pre using AdamW. Each objective is described below. Wavelet-domain spectral reconstruction. To encourage the encoder to preserve fine-grained spectral structure, wavelet-domain spectral reconstruction is used as one of the pretraining objectives. A dedicated wavelet reconstruction head maps z through a two-stage bottleneck, each stage comprising layer normalization, a linear transformation, a GELU activation, and dropout, yielding a compressed representation ∈ℝdbb ^d_b. This representation is decoded into the approximation coefficients and J=4J=4 levels of detail coefficients of a discrete wavelet transform (DWT) using a Daubechies-4 wavelet. Specifically, the predicted approximation coefficients ^∈ℝL0 a ^L_0 and the j-th-level detail coefficients ^j∈ℝLj d_j ^L_j are given by: ^=eα0a,^j=eαjd,j,j=1,…,J, a=e _0\,W_a\,b, d_j=e _j\,W_d,j\,b, j=1,…,J, (20) where aW_a and each d,jW_d,j are linear projections to the respective band lengths, and α0,α1,…,αJ _0, _1,…, _J are learnable log-scale parameters that explicitly accommodate the heterogeneous energy scales across wavelet subbands. The predicted coefficients are then passed through the inverse DWT (IDWT) to recover the full-resolution spectral estimate ^∈ℝL x ^L. The reconstruction objective combines a wavelet-coefficient-domain loss and a signal-domain loss. Denoting by a and jj=1J\d_j\_j=1^J the target DWT coefficients of the unaugmented spectrum x, the coefficient loss applies the smooth-ℓ1 _1 loss ℓH(⋅,⋅) _H(·,·) independently to each wavelet band and averages uniformly across all J+1J+1 bands: ℒcoeff=1J+1[ℓH(^,)+∑j=1JℓH(^j,j)].L_coeff= 1J+1 [ _H\! ( a,\,a )+ _j=1^J _H\! ( d_j,\,d_j ) ]. (21) The signal-domain loss is computed between x and the ground-truth x using a spatially weighted smooth-ℓ1 _1 loss. Leveraging the binary mask m introduced during data augmentation, masked positions are assigned an elevated weight mw>1m_w>1 to concentrate reconstruction capacity on the corrupted spectral regions. The per-position weight is wl=1+(mw−1)mlw_l=1+(m_w-1)\,m_l, and the signal-domain loss is computed as the weighted average: ℒsignal=∑l=1LwlℓH(x^l,xl)∑l=1Lwl.L_signal= _l=1^Lw_l\; _H\! ( x_l,\,x_l ) _l=1^Lw_l. (22) The two reconstruction terms are combined as a weighted sum: ℒrecon=λcoeffℒcoeff+λsignalℒsignal.L_recon= _coeff\,L_coeff+ _signal\,L_signal. (23) By jointly supervising coarse-scale spectral envelopes through the approximation band and fine-scale absorption line shapes through the detail bands, this pretraining objective encourages the encoder to learn representations that capture both global spectral morphology and local absorption features. Molecular fingerprint similarity alignment. To encourage the latent space to reflect graded structural relationships among molecules, UltraIR is additionally trained with a soft Tanimoto contrastive objective. This objective operates on a dedicated projection of the encoder output: a two-layer projection head maps z to a dpd_p-dimensional embedding ∈ℝdpp ^d_p via layer normalization, a linear projection, a GELU activation, dropout, and a final linear projection. Rather than defining hard positive and negative pairs, this objective constructs a continuous structural similarity target from precomputed binary molecular fingerprints. For a batch of B samples with binarized fingerprints i∈0,1Mf_i∈\0,1\^M, the pairwise Tanimoto similarity is: Tij=i⊤j‖i‖1+‖j‖1−i⊤j+ε.T_ij= f_i f_j\|f_i\|_1+\|f_j\|_1-f_i f_j+ . (24) A soft target distribution pijp_ij over off-diagonal batch members is obtained by applying a softmax with temperature τt _t to each row of T after masking the diagonal. Student similarity logits are computed as scaled inner products between ℓ2 _2-normalized embeddings ¯i=i/‖i‖2 p_i=p_i/\|p_i\|_2, giving sij=¯i⊤¯j/τs_ij= p_i p_j/ _s. The contrastive loss is the soft cross-entropy between the student distribution and the Tanimoto-derived target: ℒcontrast=−1B∑i∑j≠ipijlogexp(sij)∑k≠iexp(sik).L_contrast=- 1B _i _j≠ ip_ij (s_ij) _k≠ i (s_ik). (25) This objective encourages spectra of structurally related compounds to occupy nearby regions of the latent space in proportion to their fingerprint overlap. Multi-label functional-group prediction. A multilayer classification head maps the latent representation z to logits over the functional-group classes. It consists of layer normalization followed by three linear projections with progressively reducing dimensionality, interleaved with GELU activations and dropout. The model is trained to predict the presence or absence of each functional group as an independent binary classification task using binary cross-entropy with logits. This objective provides explicit chemical supervision that anchors the latent space to interpretable structural motifs. Downstream adaptation For downstream adaptation, the pretrained checkpoint was used to initialize only the shared spectral encoder components present in both the pretraining and downstream models. These components included the derivative-aware multi-channel input module, the convolutional spectral encoder, and the patch-based Transformer encoder. Pretraining-specific heads, including the reconstruction head, contrastive projection head, and pretraining functional-group prediction head, were not transferred to downstream tasks. Consequently, even when a downstream task shared a prediction type or label space with a pretraining objective, its task-specific output module was newly initialized rather than inherited from the pretraining checkpoint. Other components introduced exclusively for downstream prediction were also initialized randomly. Downstream task heads Classification head. All classification tasks use task-specific instances of the same MLP head architecture. The head consists of layer normalization followed by three linear projections with progressively reducing dimensionality, interleaved with GELU activations and a dropout layer. The final projection maps the hidden representation to C output logits, where C denotes the number of classes for each task. Single-spectrum classification tasks receive the global spectral representation z as input, whereas targeted component detection receives the dedicated pair representation pairr_pair. For functional-group prediction, each output logit corresponds to an independent binary classification target, and the head is trained with binary cross-entropy loss. Bacterial classification, medicinal-herb geographic origin traceability, and microplastics classification are formulated as single-label multiclass classification tasks and trained with cross-entropy loss, with label smoothing applied for medicinal-herb geographic origin traceability. Targeted component detection produces a scalar logit and is trained with binary cross-entropy loss. Regression head. For regression tasks, including physicochemical property prediction, targeted fractional contribution estimation, mixture-level component quantification, medicinal-herb constituent quantification, and soil property prediction, a task-specific three-layer MLP regression head processes the input representation. This head follows the same architecture as the classification head but uses task-specific output layers. Single-spectrum regression tasks receive the global spectral representation ∈ℝdz ^d as input, whereas targeted fractional contribution estimation receives the dedicated pair representation pairr_pair. To stabilize training across target dimensions with heterogeneous numerical scales, regression targets were transformed using target-wise scaling. For single-target regression tasks, including targeted fractional contribution estimation, the final output is passed through a sigmoid activation and optimized using mean squared error loss. For multi-target regression tasks, including physicochemical property prediction, mixture-level component quantification, medicinal-herb constituent quantification, and soil property prediction, the regression head adopts a linear output layer and is optimized using a smooth-ℓ1 _1 regression loss. 5 Data and code availability The code is available at https://github.com/AIMS-Lab-HKUSTGZ/UltraIR. The checkpoints of UltraIR are available at https://huggingface.co/yusentan/UltraIR. The newly generated UltraIR molecular-dynamics simulated infrared dataset used for UltraIR pretraining can be accessed at https://huggingface.co/yusentan/UltraIR. The simulated infrared dataset from IRtoMol [1] used for UltraIR pretraining can be accessed through https://zenodo.org/records/7928396. The simulated infrared dataset from the multimodal spectroscopy dataset [2] used for UltraIR pretraining can be accessed through https://zenodo.org/records/14770232. The simulated infrared dataset from QM9S [52] used for UltraIR pretraining can be accessed through https://figshare.com/articles/dataset/QM9S_dataset/24235333. The real infrared dataset from the NIST Chemistry WebBook [26] used for molecular-level interpretation and mixture analysis benchmarks can be accessed through https://webbook.nist.gov/chemistry/. The real infrared dataset from SDBS [32] used for molecular-level interpretation and mixture analysis benchmarks can be accessed through https://sdbs.db.aist.go.jp/. The simulated infrared dataset from USPTO [51] used for molecular-level interpretation and mixture analysis benchmarks can be accessed through https://zenodo.org/records/16417648. The external liquid-mixture benchmark from DeepMIR [41] used for mixture analysis benchmarks can be accessed through https://github.com/LinTan-CSU/DeepMIR. The experimental FTIR mixture dataset [4] used for mixture analysis benchmarks can be accessed through https://doi.org/10.5281/zenodo.5498197. The green-snow bacterial FTIR dataset [37] used for bacterial classification can be accessed through https://zenodo.org/records/4297950. The two newly generated in-house experimental datasets of Jinyinhua and Shanyinhua used for medicinal-herb characterization can be accessed at https://huggingface.co/yusentan/UltraIR. The environmentally sourced microplastics infrared dataset [46] used for microplastics classification can be accessed through https://drive.google.com/drive/folders/11MofhjEchgZelWPcHUvIMRPNEQPaLfUO?usp=sharing. The OSSL dataset [36] used for soil property prediction can be accessed through https://docs.soilspectroscopy.org/db-access.html. 6 Acknowledgments This research was supported by the National Natural Science Foundation of China Project (No. 623B2086), CCF-GHFund (No. OF 2026005), the CIPS-SMP-Zhipu Large Model Fund, Ant Group, Shanghai Artificial Intelligence Laboratory, and TeleAI of China Telecom. 7 Author contributions Conceptualization, Y.T. and J.X.; methodology, Y.T., Y.C., and J.X.; software, Y.T.; data curation, Y.T., Y.C., Z.F., P.L., and Y.F.L.; formal analysis, Y.T., Y.C., and Z.F.; investigation, Y.T., Y.C., Z.F., P.L., and Y.F.L.; validation, Y.T., Y.C., Z.F., P.L., and Y.F.L.; visualization, Y.T., Y.C., Q.G., and Z.L.; writing – original draft, Y.T.; writing – review & editing, all authors; resources, J.X., Y.Q.L., and T.W.; supervision, J.X., X.Z., and T.W.; project administration, J.X.; funding acquisition, J.X. 8 Competing interests The authors declare no competing interests. References [1] M. Alberts, T. Laino, and A. C. Vaucher (2024) Leveraging infrared spectroscopy for automated structure elucidation. Communications Chemistry 7 (1), p. 268. Cited by: §A, Table C, §1, §1, §2, §2, §5. [2] M. Alberts, O. Schilter, F. Zipoli, N. Hartrampf, and T. Laino (2024) Unraveling molecular structure: A multimodal spectroscopic dataset for chemistry. In Advances in Neural Information Processing Systems, Vol. 37, p. 125780–125808. Cited by: §A, §1, §2, §5. [3] M. Alberts, F. Zipoli, and T. Laino (2025) Setting new benchmarks in AI-driven infrared structure elucidation. Digital Discovery 4 (7), p. 1936–1943. Cited by: Table C, §1, §2. [4] A. Angulo, L. Yang, E. S. Aydil, and M. A. Modestino (2022) Machine learning enhanced spectroscopic analysis: towards autonomous chemical mixture characterization for rapid process optimization. Digital Discovery 1 (1), p. 35–44. Cited by: §A, Table C, Table C, §2, §2, §2, §5. [5] M. J. Baker, J. Trevisan, P. Bassan, R. Bhargava, H. J. Butler, K. M. Dorling, P. R. Fielden, S. W. Fogarty, N. J. Fullwood, K. A. Heys, et al. (2014) Using Fourier transform IR spectroscopy to analyze biological materials. Nature Protocols 9 (8), p. 1771–1791. Cited by: §1. [6] R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al. (2021) On the opportunities and risks of foundation models. Cited by: §1. [7] S. Boothroyd, P. K. Behara, O. C. Madin, D. F. Hahn, H. Jang, V. Gapsys, J. R. Wagner, J. T. Horton, D. L. Dotson, M. W. Thompson, J. Maat, T. Gokey, L. Wang, D. J. Cole, M. K. Gilson, J. D. Chodera, C. I. Bayly, M. R. Shirts, and D. L. Mobley (2023) Development and Benchmarking of Open Force Field 2.0.0: The Sage Small Molecule Force Field. Journal of Chemical Theory and Computation 19 (11), p. 3251–3275. Cited by: §A. [8] L. Breiman (2001) Random forests. Machine Learning 45 (1), p. 5–32. Cited by: §2. [9] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. (2020) Language models are few-shot learners. In Advances in Neural Information Processing Systems, Vol. 33, p. 1877–1901. Cited by: §1. [10] T. Chen and C. Guestrin (2016) XGBoost: A Scalable Tree Boosting System. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, p. 785–794. Cited by: §2, §2. [11] S. De Bruyne, M. M. Speeckaert, and J. R. Delanghe (2018) Applications of mid-infrared spectroscopy in the clinical laboratory setting. Critical Reviews in Clinical Laboratory Sciences 55 (1), p. 1–20. Cited by: §1. [12] V. H. M. Doan, C. D. Ly, S. Mondal, T. T. Truong, T. D. Nguyen, J. Choi, B. Lee, and J. Oh (2024) Fcg-Former: identification of functional groups in FTIR spectra using enhanced transformer-based model. Analytical Chemistry 96 (30), p. 12358–12369. Cited by: Table C, §1, §2. [13] P. Eastman, J. Swails, J. D. Chodera, R. T. McGibbon, Y. Zhao, K. A. Beauchamp, L. Wang, A. C. Simmonett, M. P. Harrigan, C. D. Stern, et al. (2017) OpenMM 7: Rapid development of high performance algorithms for molecular dynamics. PLOS Computational Biology 13 (7), p. e1005659. Cited by: §A. [14] T. Eissa, L. Voronina, M. Huber, F. Fleischmann, and M. Žigman (2024) The perils of molecular interpretations from vibrational spectra of complex samples. Angewandte Chemie International Edition 63 (50), p. e202411596. Cited by: §1. [15] R. Gautam, S. Vanga, F. Ariese, and S. Umapathy (2015) Review of multidimensional data processing approaches for Raman and infrared spectroscopy. EPJ Techniques and Instrumentation 2 (1), p. 8. Cited by: §1. [16] P. R. Griffiths (1983) Fourier transform infrared spectrometry. Science 222 (4621), p. 297–302. Cited by: §1. [17] W. Guo, S. Gao, Y. Ding, and D. Dong (2025) Simultaneous prediction of multiple soil components using Mid-Infrared Spectroscopy and the GADF-Swin Transformer model. Computers and Electronics in Agriculture 237, p. 110507. Cited by: Table C, §2. [18] M. He, J. Tong, X. Li, X. Han, Y. Qin, R. Fang, Z. Chen, and M. Gao (2025) Research on hybrid microplastic recognition method based on dual-branch convolutional neural network combined with attention mechanism. Microchemical Journal 218, p. 115131. Cited by: Table C, §2. [19] K. J. Houthuijs, G. Berden, U. F. Engelke, V. Gautam, D. S. Wishart, R. A. Wevers, J. Martens, and J. Oomens (2023) An in silico infrared spectral library of molecular ions for metabolite identification. Analytical Chemistry 95 (23), p. 8998–9005. Cited by: §1. [20] M. Huber, K. V. Kepesidis, L. Voronina, M. Božić, M. Trubetskov, N. Harbeck, F. Krausz, and M. Žigman (2021) Stability of person-specific blood-based infrared molecular fingerprints opens up prospects for health monitoring. Nature Communications 12 (1), p. 1511. Cited by: §1. [21] Y. Jin, J. Wang, F. Xu, X. Ji, Z. Gao, L. Zhang, G. Ke, R. Zhu, and W. E (2026) NMR-Solver: automated structure elucidation via large-scale spectral matching and physics-guided fragment optimization. Nature Communications 17 (1), p. 4740. Cited by: §A. [22] G. C. Kanakala, B. Sridharan, and U. D. Priyakumar (2024) Spectra to structure: contrastive learning framework for library ranking and generating molecular structures for infrared spectra. Digital Discovery 3 (12), p. 2417–2423. Cited by: §1. [23] S. G. Kazarian and K. A. Chan (2013) ATR-FTIR spectroscopic imaging: recent advances and applications to biological systems. Analyst 138 (7), p. 1940–1951. Cited by: §1. [24] J. L. Lansford and D. G. Vlachos (2020) Infrared spectroscopy data- and physics-driven machine learning for characterizing surface microstructure of complex materials. Nature Communications 11 (1), p. 1513. Cited by: §1. [25] S. Li, T. Wu, J. Xu, Y. Huang, Z. Zhang, H. Zhao, Q. Xu, Z. Wang, L. Ye, Y. Yang, et al. (2026) Biomimetic multimodal tactile sensing enables human-like robotic perception. Nature Sensors 1 (1), p. 52–62. Cited by: §1. [26] P. J. Linstrom and W. G. Mallard (2001) The NIST Chemistry WebBook: A chemical data resource on the internet. Journal of Chemical & Engineering Data 46 (5), p. 1059–1063. Cited by: §A, Table C, Table C, Table C, Table C, Table C, §2, §2, §5. [27] Y. Liu, L. Shen, X. Zhu, Y. Xie, and S. He (2024) Spectral data-driven prediction of soil properties using LSTM-CNN-attention model. Applied Sciences 14 (24), p. 11687. Cited by: Table C, §2. [28] G. Lutsker, G. Sapir, S. Shilo, J. Merino, A. Godneva, J. R. Greenfield, D. Samocha-Bonet, R. Dhir, F. Gude, S. Mannor, et al. (2026) A foundation model for continuous glucose monitoring data. Nature 650 (8103), p. 978–986. Cited by: §1. [29] D. Ma, J. Pang, M. B. Gotway, and J. Liang (2025) A fully open AI foundation model applied to chest radiography. Nature 643 (8071), p. 488–498. Cited by: §1. [30] C. McGill, M. Forsuelo, Y. Guan, and W. H. Green (2021) Predicting infrared spectra with message passing neural networks. Journal of Chemical Information and Modeling 61 (6), p. 2594–2609. Cited by: §A, §A, §1, §2. [31] L. M. Miller and J. P. Coates (2025) Interpretation of Infrared Spectra: A Practical and Systematic Approach. In Encyclopedia of Analytical Chemistry, p. 1–24. Cited by: §1. [32] National Institute of Advanced Industrial Science and Technology (2026) SDBS: Spectral Database for Organic Compounds. External Links: Link Cited by: §A, Table C, Table C, Table C, Table C, Table C, §2, §2, §5. [33] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, et al. (2011) Scikit-learn: Machine learning in Python. Journal of Machine Learning Research 12, p. 2825–2830. Cited by: §2, §2. [34] D. Punjabi, Y. Huang, L. Holzhauer, P. Tremouilhac, P. Friederich, N. Jung, and S. Bräse (2025) Infrared spectrum analysis of organic molecules with neural networks using standard reference data sets in combination with real-world data. Journal of Cheminformatics 17 (1), p. 24. Cited by: Table C, §1, §1, §2. [35] Z. Ren, Z. Zhang, J. Wei, B. Dong, and C. Lee (2022) Wavelength-multiplexed hook nanoantennas for machine learning enabled mid-infrared spectroscopy. Nature Communications 13 (1), p. 3859. Cited by: §1. [36] J. L. Safanelli, T. Hengl, L. L. Parente, R. Minarik, D. E. Bloom, K. Todd-Brown, A. Gholizadeh, W. d. S. Mendes, and J. Sanderman (2025) Open Soil Spectral Library (OSSL): Building reproducible soil calibration models through open development and community engagement. PLOS ONE 20 (1), p. e0296545. Cited by: §A, Table C, Table D, Table D, Table E, Table E, §1, §2, §2, §2, §5. [37] M. Smirnova, A. Kohler, and V. Shapaval (2020) FTIR Dataset. Zenodo. External Links: Document, Link Cited by: §A, Table C, §2, §2, §5. [38] B. C. Smith (2018) Infrared spectral interpretation: a systematic approach. CRC Press. Cited by: §1. [39] B. H. Stuart (2004) Infrared spectroscopy: fundamentals and applications. John Wiley & Sons. Cited by: §1. [40] H. Sun, X. Gao, W. Niu, L. Li, Y. Song, L. Liu, H. Zhou, H. Liu, Z. L. Wang, X. Pu, et al. (2026) A spike–language dual framework bridges fast perception and deep reasoning in artificial tactile somatosensory systems. Nature Sensors, p. 1–12. Cited by: §1. [41] L. Tan, Y. Wang, H. Zhang, J. Sun, Q. Yang, X. Yang, Z. Zhang, and H. Lu (2025) DeepMIR: A Hybrid Convolutional Neural Network-Transformer Framework for Accurate Identification of Target Components from Mid-Infrared Spectra of Mixtures. Analytical Chemistry 97 (50), p. 27706–27715. Cited by: §A, Table C, Table C, Table C, §1, §2, §2, §2, §5. [42] E. Y. Wang, P. G. Fahey, Z. Ding, S. Papadopoulos, K. Ponder, M. A. Weis, A. Chang, T. Muhammad, S. Patel, Z. Ding, et al. (2025) Foundation model of neural activity predicts response to new stimulus types. Nature 640 (8058), p. 470–477. Cited by: §1. [43] S. Wold, M. Sjöström, and L. Eriksson (2001) PLS-regression: a basic tool of chemometrics. Chemometrics and Intelligent Laboratory Systems 58 (2), p. 109–130. Cited by: §2. [44] K. Wu, Y. Zhang, L. Ru, B. Dang, J. Lao, L. Yu, J. Luo, Z. Zhu, Y. Sun, J. Zhang, et al. (2025) A semantic-enhanced multi-modal remote sensing foundation model for Earth observation. Nature Machine Intelligence 7 (8), p. 1235–1249. Cited by: §1. [45] W. Wu, A. Leonardis, J. Jiao, J. Jiang, and L. Chen (2025) Transformer-based models for predicting molecular structures from infrared spectra using patch-based self-attention. The Journal of Physical Chemistry A 129 (8), p. 2077–2085. Cited by: Table C, §1, §2. [46] J. Xie, A. Gowen, and J. Xu (2026) Open-set convolutional neural network for infrared spectral classification of environmentally sourced microplastics. npj Emerging Contaminants 2 (1), p. 6. Cited by: §A, Table C, Table C, §1, §2, §2, §5. [47] C. Zhang and Y. Ha (2026) Toward Complete Molecular Structure Prediction from Infrared Spectroscopy Using Deep Learning. Journal of Chemical Information and Modeling 66 (1), p. 100–109. Cited by: Table C, §2. [48] J. Zhou, Z. Zhang, B. Dong, Z. Ren, W. Liu, and C. Lee (2022) Midinfrared spectroscopic analysis of aqueous mixtures using artificial-intelligence-enhanced metamaterial waveguide sensing platform. ACS Nano 17 (1), p. 711–724. Cited by: Table C, §2. [49] Y. Zhou, M. A. Chia, S. K. Wagner, M. S. Ayhan, D. J. Williamson, R. R. Struyven, T. Liu, M. Xu, M. G. Lozano, P. Woodward-Court, et al. (2023) A foundation model for generalizable disease detection from retinal images. Nature 622 (7981), p. 156–163. Cited by: §1. [50] J. Zhu, S. Ji, Z. Ren, W. Wu, Z. Zhang, Z. Ni, L. Liu, Z. Zhang, A. Song, and C. Lee (2023) Triboelectric-induced ion mobility for artificial intelligence-enhanced mid-infrared gas spectroscopy. Nature Communications 14 (1), p. 2524. Cited by: §1. [51] F. Zipoli, M. Alberts, and T. Laino (2025) IR-NMR multimodal computational spectra dataset for 177K patent-extracted organic molecules. Scientific Data 12 (1), p. 1375. Cited by: §A, Table C, Table C, Table C, Table C, Table C, §1, §2, §2, §5. [52] Z. Zou, Y. Zhang, L. Liang, M. Wei, J. Leng, J. Jiang, Y. Luo, and W. Hu (2023) A deep learning model for predicting selected organic molecular spectra. Nature Computational Science 3 (11), p. 957–964. Cited by: §A, §1, §2, §5. A Supplementary experimental details Pretraining hyperparameters The principal hyperparameters used for large-scale UltraIR pretraining are summarized in Table A. Table A: Pre-training hyperparameters of UltraIR. Hyperparameter Value Optimization Learning rate 1×10−41× 10^-4 Batch size 128 Epochs 5 Warmup ratio 5% Transformer encoder Hidden dimension d 1024 Patch length P 16 Number of attention heads H 16 Number of Transformer layers NlayersN_layers 8 Dropout 0.05 Pre-training heads Fingerprint embedding dimension dpd_p 512 Number of wavelet decomposition levels J 4 Wavelet reconstruction bottleneck dimension dbd_b 256 Pre-training objectives Wavelet-domain spectral reconstruction loss weight λrecon _recon 50.0 Molecular fingerprint similarity alignment loss weight λcontrast _contrast 0.2 Multi-label functional-group prediction loss weight λfg _fg 1.25 Wavelet-coefficient-domain loss weight λcoeff _coeff 0.2 Signal-domain loss weight λsignal _signal 0.8 Masked-position weighting factor mwm_w 3.0 Fingerprint student temperature τs _s 0.1 Fingerprint target temperature τt _t 0.05 Pretraining dataset construction The UltraIR pretraining dataset comprised 60,000,000 simulated infrared (IR) spectra assembled from three complementary sources: publicly released simulated spectra, spectra generated using molecular dynamics, and spectra predicted using a machine-learning model. After removing molecules overlapping with the National Institute of Standards and Technology (NIST) Chemistry WebBook, Spectral Database for Organic Compounds (SDBS), and USPTO datasets, the publicly released sources contributed 630,395 spectra from IRtoMol, 743,955 from the multimodal spectroscopy dataset, and 127,898 from QM9S, yielding 1,502,248 spectra in total [1, 2, 52]. From structures in the PubChem-derived molecular collection distributed with the SimNMR-PubChem database [21] that were not represented in NIST, SDBS, or USPTO, we generated 7,534,293 spectra through molecular-dynamics simulation and predicted the remaining 50,963,459 spectra were predicted using Chemprop-IR [30]. For the molecular-dynamics component, three-dimensional molecular conformers were generated from Simplified Molecular Input Line Entry System (SMILES) representations and subjected to geometry relaxation before simulation. Each molecule was simulated without periodic boundary conditions at 300 K using a Langevin middle integrator in OpenMM, with the Open Force Field Sage 2.2.1 and fixed atom-centered partial charges [13, 7]. Molecular dipole trajectories were calculated from the fixed partial charges and instantaneous atomic positions. After temporal mean removal and Hann windowing, the three Cartesian dipole components were transformed independently using a fast Fourier transform. The resulting power spectra were summed, weighted by wavenumber, interpolated onto the common wavenumber grid, and normalized to a maximum intensity of one. For the machine-learning-generated component, IR spectra were predicted using the two officially released pretrained Chemprop-IR checkpoints [30]. The two models were applied independently, and their predicted intensities were combined using an equally weighted arithmetic mean to obtain the final spectrum for each molecular structure. Functional-group labels and molecular fingerprints used during pretraining were generated from molecular structures using RDKit. Functional-group labels were encoded as 17-dimensional multi-hot vectors, with each class assigned when the molecular graph matched its corresponding SMILES Arbitrary Target Specification (SMARTS) definition. Multiple classes could therefore be assigned to the same molecule. The functional-group classes, output order, and operational SMARTS definitions are provided in Table B. For molecular fingerprint similarity alignment, each structure was represented by a 2,048-bit Morgan fingerprint generated with a radius of 2. Table B: Functional-group classes and SMARTS definitions used for pretraining and downstream functional-group prediction. The listed order corresponds to the 17 dimensions of the multi-label target used in both stages. A class was assigned when the corresponding SMARTS pattern matched the molecular graph, and each molecule could receive multiple labels. Index Functional-group class SMARTS definition 1 Alkane [CX4;H3,H2] 2 Methyl [CH3] 3 Alkene [CX3]=[CX3] 4 Alkyne [CX2]#[CX2] 5 Alcohols [#6][OX2H] 6 Amines [NX3;H2,H1,H0;!$(N[CX3](=O))] 7 Nitriles [NX1]#[CX2] 8 Aromatics [$([cX3](:*):*),$([cX2+](:*):*)] 9 Alkyl halides [#6][F,Cl,Br,I] 10 Esters [#6][CX3](=O)[OX2H0][#6] 11 Ketones [#6][CX3](=O)[#6] 12 Aldehydes [CX3H1](=O)[#6] 13 Carboxylic acids [CX3](=O)[OX2H1] 14 Ether [OD2]([#6;!$(C=O)])([#6;!$(C=O)]) 15 Acyl halides [CX3](=[OX1])[F,Cl,Br,I] 16 Amides [NX3][CX3](=[OX1])[#6] 17 Nitro [$([N+](=O)[O-]),$([NX3](=O)=O)][#6] Downstream datasets and benchmark summary The eight downstream tasks were evaluated using ten datasets, with the Jinyinhua and Shanyinhua datasets each contributing separate subsets for geographic origin traceability and chemical constituent quantification. For the molecular-level benchmarks, we used NIST, SDBS, and USPTO. The NIST Chemistry WebBook, maintained as NIST Standard Reference Database 69, is a public resource that compiles chemical and spectroscopic data, including IR spectra [26]. We retained 23,286 NIST IR spectra after filtering. SDBS is a public database maintained by the National Institute of Advanced Industrial Science and Technology and contains experimentally measured Fourier-transform infrared (FTIR) spectra with corresponding molecular records for organic compounds [32]. We retained 18,253 SDBS IR spectra after filtering. The USPTO dataset is a computational multimodal spectroscopy resource constructed from patent-extracted organic molecules and contains 177,461 IR spectra generated using long-timescale molecular-dynamics simulations and machine-learning-assisted dipole-moment prediction [51]. For functional-group prediction, labels for NIST, SDBS, and USPTO were generated using the same 17 classes, class order, and SMARTS definitions as in pretraining (Table B). For physicochemical property prediction, all 11 targets were calculated from molecular structures using RDKit. The synthetic accessibility score (SAScore) was obtained using the RDKit SA_Score implementation, and the remaining targets were calculated using their corresponding RDKit descriptor functions. Mixture-level and application-facing evaluations used several additional data sources. For external evaluation of targeted component detection, we used 360 spectra from the external liquid-mixture benchmark described in DeepMIR, comprising binary, ternary and quaternary mixtures [41]. Mixture-level component quantification was evaluated using 74 experimentally acquired spectra from a four-component FTIR mixture dataset comprising acrylonitrile, adiponitrile, propionitrile and glycerol [4]. For this task, models were first trained on a simulated four-component mixture dataset containing 2,400 synthetic mixtures, and subsequently fine-tuned on the experimentally acquired spectra for evaluation [4]. Bacterial genus classification used 795 FTIR spectra from 45 fast-growing bacterial isolates collected from Antarctic green snow and spanning nine genera [37]. For medicinal-herb geographic origin traceability and constituent quantification, the dataset composition, LC-MS-derived relative-abundance regression targets, together with the corresponding analytical procedures, are detailed in the following subsection. The microplastics benchmark contained 32,965 IR spectra assembled from laboratory and open-source collections to represent diverse environmentally relevant polymer classes [46]. After filtering for complete wavenumber coverage and availability of the labels required for soil property prediction, we retained 46,099 mid-IR spectra from the Kellogg Soil Survey Laboratory (KSSL) collection within the Open Soil Spectral Library (OSSL) [36]. For cross-instrument and cross-laboratory evaluation, we independently filtered the datasets using the labels required for the cross-laboratory tasks and enforced complete wavenumber coverage. This procedure retained 55,974 KSSL spectra for model training and 3,753 ICRAF–ISRIC spectra for zero-shot target-domain inference. Table C summarizes the datasets and comparison methods used for each downstream task. Table C: Summary of downstream benchmarks and comparison methods. The table summarizes the datasets and competing methods evaluated for each downstream task. Task Dataset Compared methods Functional-group prediction NIST [26], SDBS [32], USPTO [51] FCGFormer [12], IRAnalysis [34], XGBoost, Random Forest, KNN, Logistic classifier Molecular structure elucidation NIST [26], SDBS [32], USPTO [51] IRtoMol [1], AISE [3], PBSA [45], DLIR [47] Physicochemical property prediction NIST [26], SDBS [32], USPTO [51] XGBoost regression, SVR, KNN regression, PLSR Targeted component detection NIST [26], SDBS [32], USPTO [51] , DeepMIR liquid-mixture dataset [41] DeepMIR [41], reverse match, HQI Targeted fractional contribution estimation NIST [26], SDBS [32], USPTO [51] DeepMIR [41], XGBoost regression, SVR, KNN regression, PLSR Mixture-level component quantification Experimental FTIR mixture dataset [4] AIMWSP [48], ML-FTIR [4], XGBoost regression, SVR, KNN regression, PLSR Bacterial genus classification Green-snow bacterial FTIR spectra [37] XGBoost, Random Forest, KNN, Logistic classifier Medicinal-herb origin traceability Jinyinhua, Shanyinhua XGBoost, Random Forest, KNN, Logistic classifier Medicinal-herb constituent quantification Jinyinhua, Shanyinhua XGBoost regression, SVR, KNN regression, PLSR Microplastics classification Environmentally sourced microplastics IR dataset [46] Softmax [46], DB-CNN-CBAM [18], XGBoost, Random Forest, KNN, Logistic classifier Soil property prediction OSSL [36] GADF-Swin [17], LSTM-CNN [27], XGBoost regression, SVR, KNN regression, PLSR In-house medicinal-herb datasets and LC-MS-derived regression targets In addition to the public and external benchmarks summarized above, we assembled two in-house medicinal-herb datasets: Jinyinhua (Lonicerae Japonicae Flos) and Shanyinhua (Lonicerae Flos). For geographic origin traceability, the Jinyinhua dataset contains 120 IR spectra from Shandong, Henan, Hebei, and Sichuan, and the Shanyinhua dataset contains 150 IR spectra from Hunan, Hubei, Sichuan, Henan, and Guangdong; each origin is represented by 30 spectra. For chemical constituent quantification, LC-MS-derived target relative abundances measured for 60 Jinyinhua and 75 Shanyinhua samples served as regression labels and were expressed in arbitrary units (a.u.). The Jinyinhua targets were 4-methylcoumarin, milrinone, geniposide, phenylacetaldehyde, cis-cinnamic acid, and lumazine; the Shanyinhua targets were coniferyl aldehyde, quercetin, kaempferol, and luteolin. For LC-MS analysis, 30 mg of powdered material was extracted with 1.5 mL of 70% aqueous methanol. The extract was vortexed, ultrasonicated at 30 ∘C for 30 min, centrifuged at 12,000 r.p.m. for 15 min, and passed through a 0.22 μ membrane filter. Chromatographic separation was performed on a Kinetex F5 column (2.6 μ , 100 m × 2.1 m; Phenomenex) at 40 ∘C. The mobile phases were 0.1% formic acid in water (A) and 0.1% formic acid in acetonitrile (B), with a flow rate of 0.2 mL min-1 and an injection volume of 4 μ . The percentage of A was programmed as 100% at 0 min, 99% at 2 min, 95% at 3 min, 90% at 6 min, 85% at 14 min, 82% at 15 min, 80% at 18 min, 70% at 20 min, 60% at 23 min, 22% at 31 min, 10% at 33 min, 0% at 36–44 min, and 100% at 44.1–50 min. Full-scan time-of-flight mass spectrometry (TOF-MS) and information-dependent acquisition (IDA) tandem mass spectrometry (MS/MS) spectra were acquired in positive-ion mode over m/zm/z 100–1,000 and m/zm/z 50–1,000, respectively. The source temperature was 500 ∘C, the ion-spray voltage was +5,500 V, ion-source gases 1 and 2 were each 50 psi, curtain gas was 30 psi, declustering potential was +80 V, and collision energy was 35 V. Dataset metadata and label distributions for cross-instrument and cross-laboratory soil-property prediction For cross-instrument and cross-laboratory soil-property prediction, the source- and target-domain dataset metadata and corresponding label distributions across the nine soil properties are summarized in Tables D and E, respectively. Table D: Dataset metadata for cross-instrument and cross-laboratory soil-property prediction. Spectra from the Kellogg Soil Survey Laboratory (KSSL) and ICRAF–ISRIC, together with their associated metadata, were obtained from the Open Soil Spectral Library (OSSL) [36]. Characteristic Kellogg Soil Survey Laboratory database Source training domain ICRAF–ISRIC Soil Spectral Library Target test domain Institution USDA National Soil Survey Center World Agroforestry Centre / ISRIC Sample origin United States Not recorded in OSSL FTIR system Bruker Vertex 70 with HTS-XT Bruker Tensor 27 with HTS-XT Sample preparation Ground to <80<80 mesh Ground to <0.1<0.1 m Table E: Label distributions for cross-instrument and cross-laboratory soil-property prediction. The Kellogg Soil Survey Laboratory (KSSL) and ICRAF–ISRIC spectra and associated soil-property labels were obtained from the Open Soil Spectral Library (OSSL) [36]. Values are reported as median [minimum, maximum]. Property Unit Kellogg Soil Survey Laboratory database Source training domain ICRAF–ISRIC Soil Spectral Library Target test domain Clay Content % 20.88 [0, 96.14] 30.10 [0, 96.80] Silt Content % 39.10 [0, 94.50] 24.90 [0.20, 100.00] Sand Content % 32.60 [0.10, 100.00] 33.05 [0, 99.50] pH (H2O) – 6.37 [2.29, 10.70] 5.80 [3.00, 10.50] CEC cmolckg−1cmol_c\,kg^-1 15.62 [0, 584.59] 11.50 [0, 189.60] Exchangeable Ca cmolckg−1cmol_c\,kg^-1 11.92 [0, 410.41] 3.80 [0, 168.20] Exchangeable Mg cmolckg−1cmol_c\,kg^-1 2.86 [0, 172.64] 1.10 [0, 68.00] Exchangeable K cmolckg−1cmol_c\,kg^-1 0.364 [0, 32.33] 0.20 [0, 9.80] Exchangeable Na cmolckg−1cmol_c\,kg^-1 0 [0, 868.36] 0.10 [0, 31.60] Cross-validation and evaluation protocol All downstream evaluations requiring supervised model training were conducted using five-fold cross-validation. For each such dataset, samples were partitioned into five non-overlapping folds, with each fold used once as the test set. In each cross-validation run, one fold, corresponding to approximately 20% of the dataset, was held out for testing. The remaining samples were divided into training and validation subsets to approximate an overall 70:10:20 train–validation–test ratio. When this ratio could not be achieved exactly because of discrete sample counts, the partition with the smallest deviation from the target ratio was used. For the molecule-level NIST, SDBS, and USPTO benchmarks, scaffold-based partitioning was used to minimize structural overlap among the training, validation, and test sets. Identical partitions were used for UltraIR, the matched no-pretraining ablation where included, and all comparison methods, ensuring directly comparable evaluations. For the reduced-training-data experiments, the validation and test sets remained fixed within each cross-validation run. Only the original training subset was randomly subsampled to the specified fractions of labeled training data. These fractions therefore refer to the available training subset rather than the complete dataset. At each fraction, identical subsampled training sets were used for all compared methods. Model configurations, training duration, early-stopping criteria, checkpoint selection, and the optional use of downstream spectral augmentation were determined exclusively using the training and validation subsets within each cross-validation run. Weak physically motivated augmentations, when enabled, were applied only to training samples. Except for the DeepMIR liquid-mixture dataset, downstream performance was aggregated across the five held-out test folds. For DeepMIR, checkpoints trained on the corresponding NIST mixture-analysis tasks with matched target labels were applied directly for inference. No additional model training or five-fold cross-validation was performed on the DeepMIR dataset. Evaluation metrics Statistical significance testing. All significance annotations shown in the figures were calculated independently for each dataset–metric combination. Methods were ranked according to their mean performance across the five held-out test folds, using the appropriate direction for each metric, and only the top-ranked and second-ranked methods were compared. Their fold-level results were paired by the corresponding test fold and analyzed using a two-sided paired t-test implemented with scipy.stats.ttest_rel. The null hypothesis was that the mean paired difference between the two methods across the five folds was zero. Significance was annotated as **** for P<0.0001P<0.0001, *** for 0.0001≤P<0.0010.0001≤ P<0.001, ** for 0.001≤P<0.010.001≤ P<0.01, * for 0.01≤P<0.050.01≤ P<0.05, and n.s. for P≥0.05P≥ 0.05. Classification metrics. Single-label classification tasks, including bacterial classification, medicinal-herb geographic origin traceability, and microplastics classification, were evaluated using accuracy, Macro-F1, and the multiclass Matthews correlation coefficient (MCC). Targeted component detection was evaluated using accuracy, Macro-F1, and the area under the receiver operating characteristic curve (ROC-AUC), calculated from the continuous component-presence scores before thresholding. Macro-F1 was calculated as the unweighted mean of the one-versus-rest F1 scores across classes, thereby assigning equal weight to each class. For the class-wise radar analyses of bacterial and microplastics classification, the value reported for each bacterial genus or polymer class was its corresponding one-versus-rest F1 score. These class-specific F1 scores were calculated independently for each class and were distinct from the Macro-F1 values reported for overall task performance. Functional-group prediction was formulated as multi-label classification and evaluated using Micro-F1, Macro-F1, and exact match ratio (EMR). Micro-F1 was calculated by aggregating true positives, false positives, and false negatives across all samples and functional-group labels, whereas Macro-F1 was the unweighted mean of the label-specific F1 scores. EMR required the complete predicted functional-group vector to match the ground-truth vector: EMR=1N∑i=1N(i=^i),EMR= 1N _i=1^NI (y_i= y_i ), (26) where N is the number of evaluated spectra, iy_i and ^i y_i are the ground-truth and predicted binary label vectors, respectively, and (⋅)I(·) is the indicator function. A prediction containing either a missed functional group or an additional false-positive group was therefore counted as incorrect. Molecular structure elucidation metrics. Formula-conditioned molecular structure elucidation was evaluated using top-k accuracy for k∈1,5,10k∈\1,5,10\. Let sis_i denote the ground-truth SMILES string for sample i, and let i(k)P_i^(k) denote the first k generated candidate SMILES strings in their original ranked order. Before structure matching, the ground-truth SMILES and every generated candidate were parsed using RDKit and converted to canonical isomeric SMILES. Denoting this canonicalization operation by (⋅)C(·), top-k accuracy was defined as Top-k=1N∑i=1N[∃s^∈i(k)such that(s^)=(si)].Top -k= 1N _i=1^NI [∃\, s _i^(k)\ such that\ C( s)=C(s_i) ]. (27) A sample was counted as correct when at least one of the first k generated candidates represented the same canonical molecular structure as the ground truth. Candidates that could not be parsed into valid molecular structures remained in their generated rank positions but could not contribute a correct match. Stereochemical annotations were retained during canonicalization where present in the corresponding SMILES strings. Structural similarity between a generated candidate and the ground-truth molecule was assessed using the Tanimoto similarity of their binary molecular fingerprints. For fingerprint vectors f and g, the Tanimoto similarity was defined as T(,)=∥1+∥1−.T(f,g)= f Tg _1+ _1-f Tg. (28) The score ranges from 00 to 11, with larger values indicating greater fingerprint overlap. Regression metrics. Regression tasks were evaluated using mean absolute error (MAE), root mean squared error (RMSE), and the coefficient of determination R2R^2. For single-target regression tasks, including targeted fractional contribution estimation, MAE, RMSE, and R2R^2 were calculated directly in the original target space. For multi-target regression tasks, including physicochemical property prediction, mixture-level component quantification, medicinal-herb constituent quantification, and soil property prediction, normalized MAE and normalized RMSE were used to account for differences in numerical ranges among targets. These metrics were calculated after target-wise normalization and averaged across targets such that each target contributed equally. Overall R2R^2 was calculated as the arithmetic mean of target-specific R2R^2 values across targets. For normalized residual analyses, residuals were calculated independently for each target and normalized by the corresponding target range. Let yijy_ij and y^ij y_ij denote the ground-truth and predicted values, respectively, for target j of sample i. The normalized residual was defined as rij′=y^ij−yijyj,max−yj,min,r _ij= y_ij-y_ijy_j, -y_j, , (29) where yj,maxy_j, and yj,miny_j, denote the maximum and minimum values of target j, respectively. To visualize multi-target regression residuals, the normalized residuals from all samples and targets were pooled such that each sample–target pair contributed equally to the resulting distribution. BertzCT complexity-thresholded analysis. Prediction errors across molecular complexity were assessed using cumulative BertzCT thresholds. For threshold τ, the evaluated subset was defined as τ=i:Bi≥τ,S_τ= \i B_i≥τ \, (30) where BiB_i is the ground-truth BertzCT value of molecule i. The mean relative error (MRE) was then calculated as MRE(τ)=1|τ|∑i∈τ|B^i−Bi||Bi|+ϵ,MRE(τ)= 1|S_τ| _i _τ | B_i-B_i ||B_i|+ε, (31) where ϵε is a small constant. True-class margin analysis. For microplastics classification, prediction confidence was further examined using the true-class margin. Given the logit sics_ic assigned by a model to class c for sample i, with yiy_i denoting the ground-truth class label of sample i, the margin was defined as mi=siyi−maxc≠yisic.m_i=s_iy_i- _c≠ y_is_ic. (32) A positive margin indicates that the ground-truth class received the highest logit, whereas a negative margin indicates that at least one competing class received a higher score. Larger positive values indicate greater separation between the true class and its closest competing class. B More results Extended ablation analysis of UltraIR Table F: Extended ablation analysis of simulated pretraining across representative downstream tasks. UltraIR is compared with the matched no-pretraining ablation across four downstream tasks. Values are means ± standard deviations across five test folds. Bold values indicate the better result for each metric. Arrows indicate the preferred direction. a, Functional-group prediction Results on the NIST dataset Model Macro-F1 ↑ Micro-F1 ↑ EMR ↑ UltraIR 0.919±0.0040.919_± 0.004 0.943±0.0020.943_± 0.002 0.772±0.0070.772_± 0.007 w/o pretraining 0.901±0.0050.901_± 0.005 0.932±0.0020.932_± 0.002 0.719±0.0090.719_± 0.009 b, Molecular structure elucidation Results on the NIST dataset Model Top-1 accuracy ↑ Top-5 accuracy ↑ Top-10 accuracy ↑ UltraIR 0.522±0.0050.522_± 0.005 0.575±0.0070.575_± 0.007 0.576±0.0070.576_± 0.007 w/o pretraining 0.477±0.0030.477_± 0.003 0.563±0.0060.563_± 0.006 0.567±0.0070.567_± 0.007 c, Mixture-level component quantification Model Normalized MAE (×10−3× 10^-3) ↓ Normalized RMSE (×10−3× 10^-3) ↓ R2R^2 ↑ UltraIR 2.036±0.3662.036_± 0.366 2.621±0.5022.621_± 0.502 0.782±0.0470.782_± 0.047 w/o pretraining 2.288±0.2252.288_± 0.225 2.919±0.3702.919_± 0.370 0.714±0.0290.714_± 0.029 d, Soil property prediction Model Normalized MAE ↓ Normalized RMSE ↓ R2R^2 ↑ UltraIR 0.086±0.0020.086_± 0.002 0.325±0.1470.325_± 0.147 0.921±0.0260.921_± 0.026 w/o pretraining 0.104±0.0020.104_± 0.002 0.360±0.1380.360_± 0.138 0.901±0.0250.901_± 0.025 Additional functional-group prediction results Figure A: Functional-group prediction across molecular-complexity strata and precision–recall trade-offs. a, Exact match ratio (EMR) for functional-group prediction stratified by the number of functional groups in each molecule on NIST, SDBS, and USPTO. b, EMR for functional-group prediction stratified by the number of heavy atoms in each molecule on NIST, SDBS, and USPTO. c, Precision–recall curves for functional-group prediction on NIST, SDBS, and USPTO. Table G: Per-functional-group classification performance on the NIST benchmark. F1 scores are reported as means ± standard deviations across five folds. The best result for each functional group is highlighted in bold. Functional group Positive rate (%) Logistic KNN RF XGBoost IRAnalysis FCGFormer UltraIR Alkane 79.7679.76 0.919±0.0050.919_± 0.005 0.931±0.0080.931_± 0.008 0.952±0.0030.952_± 0.003 0.966±0.0030.966_± 0.003 0.960±0.0030.960_± 0.003 0.962±0.0020.962_± 0.002 0.978±0.0010.978_± 0.001 Methyl 63.0863.08 0.785±0.0090.785_± 0.009 0.865±0.0070.865_± 0.007 0.877±0.0100.877_± 0.010 0.920±0.0050.920_± 0.005 0.910±0.0060.910_± 0.006 0.911±0.0060.911_± 0.006 0.950±0.0020.950_± 0.002 Alkene 13.1613.16 0.449±0.0240.449_± 0.024 0.684±0.0240.684_± 0.024 0.689±0.0270.689_± 0.027 0.783±0.0120.783_± 0.012 0.761±0.0110.761_± 0.011 0.731±0.0080.731_± 0.008 0.869±0.0150.869_± 0.015 Alkyne 2.042.04 0.544±0.0470.544_± 0.047 0.821±0.0310.821_± 0.031 0.747±0.0570.747_± 0.057 0.874±0.0230.874_± 0.023 0.874±0.0150.874_± 0.015 0.884±0.0350.884_± 0.035 0.949±0.0220.949_± 0.022 Alcohols 23.8423.84 0.682±0.0370.682_± 0.037 0.827±0.0080.827_± 0.008 0.850±0.0090.850_± 0.009 0.887±0.0030.887_± 0.003 0.889±0.0080.889_± 0.008 0.889±0.0100.889_± 0.010 0.928±0.0040.928_± 0.004 Amines 22.7422.74 0.621±0.0440.621_± 0.044 0.816±0.0050.816_± 0.005 0.812±0.0170.812_± 0.017 0.865±0.0110.865_± 0.011 0.849±0.0050.849_± 0.005 0.850±0.0050.850_± 0.005 0.918±0.0050.918_± 0.005 Nitriles 3.823.82 0.291±0.0480.291_± 0.048 0.528±0.0570.528_± 0.057 0.466±0.0480.466_± 0.048 0.795±0.0250.795_± 0.025 0.659±0.0150.659_± 0.015 0.586±0.0330.586_± 0.033 0.931±0.0160.931_± 0.016 Aromatics 58.8458.84 0.872±0.0160.872_± 0.016 0.924±0.0080.924_± 0.008 0.920±0.0110.920_± 0.011 0.963±0.0040.963_± 0.004 0.953±0.0090.953_± 0.009 0.956±0.0070.956_± 0.007 0.982±0.0050.982_± 0.005 Alkyl halides 25.4125.41 0.542±0.0340.542_± 0.034 0.764±0.0090.764_± 0.009 0.750±0.0100.750_± 0.010 0.810±0.0100.810_± 0.010 0.800±0.0040.800_± 0.004 0.791±0.0130.791_± 0.013 0.889±0.0100.889_± 0.010 Esters 11.7811.78 0.732±0.0210.732_± 0.021 0.884±0.0140.884_± 0.014 0.832±0.0290.832_± 0.029 0.907±0.0130.907_± 0.013 0.912±0.0130.912_± 0.013 0.902±0.0190.902_± 0.019 0.945±0.0100.945_± 0.010 Ketones 8.968.96 0.419±0.0450.419_± 0.045 0.756±0.0230.756_± 0.023 0.677±0.0260.677_± 0.026 0.783±0.0290.783_± 0.029 0.801±0.0110.801_± 0.011 0.785±0.0220.785_± 0.022 0.894±0.0140.894_± 0.014 Aldehydes 1.971.97 0.552±0.0500.552_± 0.050 0.789±0.0340.789_± 0.034 0.827±0.0250.827_± 0.025 0.873±0.0210.873_± 0.021 0.866±0.0350.866_± 0.035 0.891±0.0220.891_± 0.022 0.944±0.0160.944_± 0.016 Carboxylic acids 6.856.85 0.706±0.0170.706_± 0.017 0.856±0.0140.856_± 0.014 0.854±0.0170.854_± 0.017 0.899±0.0160.899_± 0.016 0.881±0.0090.881_± 0.009 0.889±0.0120.889_± 0.012 0.931±0.0150.931_± 0.015 Ether 13.8613.86 0.580±0.0240.580_± 0.024 0.784±0.0230.784_± 0.023 0.717±0.0220.717_± 0.022 0.816±0.0220.816_± 0.022 0.818±0.0240.818_± 0.024 0.818±0.0200.818_± 0.020 0.905±0.0140.905_± 0.014 Acyl halides 0.920.92 0.559±0.0330.559_± 0.033 0.827±0.0680.827_± 0.068 0.781±0.0460.781_± 0.046 0.843±0.0490.843_± 0.049 0.864±0.0370.864_± 0.037 0.861±0.0500.861_± 0.050 0.897±0.0620.897_± 0.062 Amides 6.196.19 0.443±0.0400.443_± 0.040 0.680±0.0260.680_± 0.026 0.478±0.0520.478_± 0.052 0.685±0.0210.685_± 0.021 0.712±0.0100.712_± 0.010 0.678±0.0160.678_± 0.016 0.791±0.0140.791_± 0.014 Nitro 5.635.63 0.761±0.0180.761_± 0.018 0.845±0.0260.845_± 0.026 0.811±0.0310.811_± 0.031 0.889±0.0270.889_± 0.027 0.884±0.0170.884_± 0.017 0.867±0.0190.867_± 0.019 0.919±0.0060.919_± 0.006 Table H: Per-functional-group classification performance on the SDBS benchmark. F1 scores are reported as means ± standard deviations across five folds. The best result for each functional group is highlighted in bold. Functional group Positive rate (%) Logistic KNN RF XGBoost IRAnalysis FCGFormer UltraIR Alkane 74.1974.19 0.858±0.0150.858_± 0.015 0.875±0.0150.875_± 0.015 0.866±0.0160.866_± 0.016 0.915±0.0110.915_± 0.011 0.909±0.0120.909_± 0.012 0.904±0.0090.904_± 0.009 0.941±0.0080.941_± 0.008 Methyl 55.7855.78 0.695±0.0390.695_± 0.039 0.767±0.0110.767_± 0.011 0.776±0.0110.776_± 0.011 0.852±0.0110.852_± 0.011 0.833±0.0130.833_± 0.013 0.815±0.0050.815_± 0.005 0.881±0.0080.881_± 0.008 Alkene 12.0712.07 0.401±0.0420.401_± 0.042 0.470±0.0730.470_± 0.073 0.301±0.0850.301_± 0.085 0.565±0.0320.565_± 0.032 0.614±0.0680.614_± 0.068 0.550±0.0590.550_± 0.059 0.699±0.0500.699_± 0.050 Alkyne 1.071.07 0.415±0.0890.415_± 0.089 0.536±0.1040.536_± 0.104 0.102±0.0710.102_± 0.071 0.525±0.0670.525_± 0.067 0.662±0.1080.662_± 0.108 0.576±0.0900.576_± 0.090 0.726±0.0610.726_± 0.061 Alcohols 29.7229.72 0.631±0.0600.631_± 0.060 0.800±0.0260.800_± 0.026 0.773±0.0410.773_± 0.041 0.844±0.0270.844_± 0.027 0.848±0.0180.848_± 0.018 0.827±0.0220.827_± 0.022 0.893±0.0180.893_± 0.018 Amines 27.3527.35 0.587±0.1050.587_± 0.105 0.753±0.0080.753_± 0.008 0.703±0.0260.703_± 0.026 0.783±0.0150.783_± 0.015 0.797±0.0190.797_± 0.019 0.785±0.0140.785_± 0.014 0.872±0.0170.872_± 0.017 Nitriles 2.982.98 0.586±0.0390.586_± 0.039 0.278±0.0920.278_± 0.092 0.152±0.0770.152_± 0.077 0.800±0.0400.800_± 0.040 0.752±0.0620.752_± 0.062 0.747±0.0490.747_± 0.049 0.876±0.0250.876_± 0.025 Aromatics 60.4760.47 0.831±0.0170.831_± 0.017 0.865±0.0090.865_± 0.009 0.855±0.0500.855_± 0.050 0.935±0.0140.935_± 0.014 0.935±0.0060.935_± 0.006 0.929±0.0040.929_± 0.004 0.968±0.0040.968_± 0.004 Alkyl halides 18.6218.62 0.363±0.0350.363_± 0.035 0.429±0.0550.429_± 0.055 0.296±0.0450.296_± 0.045 0.497±0.0300.497_± 0.030 0.557±0.0280.557_± 0.028 0.487±0.0370.487_± 0.037 0.607±0.0350.607_± 0.035 Esters 11.9511.95 0.735±0.0090.735_± 0.009 0.799±0.0170.799_± 0.017 0.719±0.0240.719_± 0.024 0.836±0.0170.836_± 0.017 0.852±0.0300.852_± 0.030 0.825±0.0260.825_± 0.026 0.884±0.0180.884_± 0.018 Ketones 9.249.24 0.384±0.0510.384_± 0.051 0.545±0.0400.545_± 0.040 0.226±0.1060.226_± 0.106 0.569±0.0360.569_± 0.036 0.638±0.0300.638_± 0.030 0.603±0.0420.603_± 0.042 0.720±0.0310.720_± 0.031 Aldehydes 2.012.01 0.251±0.0600.251_± 0.060 0.407±0.0750.407_± 0.075 0.175±0.0740.175_± 0.074 0.446±0.0820.446_± 0.082 0.579±0.0920.579_± 0.092 0.479±0.0860.479_± 0.086 0.647±0.0730.647_± 0.073 Carboxylic acids 12.1712.17 0.677±0.0170.677_± 0.017 0.791±0.0430.791_± 0.043 0.785±0.0390.785_± 0.039 0.844±0.0360.844_± 0.036 0.846±0.0370.846_± 0.037 0.827±0.0350.827_± 0.035 0.882±0.0300.882_± 0.030 Ether 14.6314.63 0.518±0.0330.518_± 0.033 0.615±0.0310.615_± 0.031 0.368±0.0920.368_± 0.092 0.673±0.0300.673_± 0.030 0.709±0.0160.709_± 0.016 0.667±0.0250.667_± 0.025 0.784±0.0210.784_± 0.021 Acyl halides 0.990.99 0.606±0.1080.606_± 0.108 0.728±0.1260.728_± 0.126 0.560±0.1570.560_± 0.157 0.692±0.1190.692_± 0.119 0.740±0.1530.740_± 0.153 0.704±0.0980.704_± 0.098 0.719±0.1200.719_± 0.120 Amides 7.957.95 0.468±0.0450.468_± 0.045 0.612±0.0360.612_± 0.036 0.279±0.1110.279_± 0.111 0.583±0.0550.583_± 0.055 0.672±0.0220.672_± 0.022 0.618±0.0220.618_± 0.022 0.746±0.0300.746_± 0.030 Nitro 6.176.17 0.685±0.0760.685_± 0.076 0.699±0.0660.699_± 0.066 0.606±0.0670.606_± 0.067 0.794±0.0310.794_± 0.031 0.824±0.0250.824_± 0.025 0.752±0.0510.752_± 0.051 0.871±0.0400.871_± 0.040 Table I: Per-functional-group classification performance on the USPTO benchmark. F1 scores are reported as means ± standard deviations across five folds. The best result for each functional group is highlighted in bold. Functional group Positive rate (%) Logistic KNN RF XGBoost IRAnalysis FCGFormer UltraIR Alkane 93.1593.15 0.989±0.0010.989_± 0.001 0.964±0.0010.964_± 0.001 0.965±0.0010.965_± 0.001 0.994±0.0000.994_± 0.000 0.993±0.0010.993_± 0.001 0.997±0.0000.997_± 0.000 0.999±0.0000.999_± 0.000 Methyl 73.1073.10 0.914±0.0040.914_± 0.004 0.839±0.0040.839_± 0.004 0.851±0.0040.851_± 0.004 0.964±0.0010.964_± 0.001 0.964±0.0020.964_± 0.002 0.967±0.0010.967_± 0.001 0.981±0.0010.981_± 0.001 Alkene 11.3411.34 0.160±0.0400.160_± 0.040 0.204±0.0060.204_± 0.006 0.000±0.0000.000_± 0.000 0.240±0.0120.240_± 0.012 0.308±0.0180.308_± 0.018 0.375±0.0260.375_± 0.026 0.603±0.0170.603_± 0.017 Alkyne 1.911.91 0.301±0.0220.301_± 0.022 0.165±0.0140.165_± 0.014 0.000±0.0000.000_± 0.000 0.653±0.0350.653_± 0.035 0.644±0.0400.644_± 0.040 0.806±0.0200.806_± 0.020 0.929±0.0130.929_± 0.013 Alcohols 26.2226.22 0.636±0.0180.636_± 0.018 0.467±0.0120.467_± 0.012 0.634±0.0090.634_± 0.009 0.840±0.0050.840_± 0.005 0.849±0.0030.849_± 0.003 0.905±0.0040.905_± 0.004 0.953±0.0020.953_± 0.002 Amines 44.4844.48 0.682±0.0220.682_± 0.022 0.613±0.0030.613_± 0.003 0.659±0.0080.659_± 0.008 0.806±0.0040.806_± 0.004 0.817±0.0020.817_± 0.002 0.857±0.0030.857_± 0.003 0.917±0.0030.917_± 0.003 Nitriles 6.696.69 0.426±0.0310.426_± 0.031 0.321±0.0150.321_± 0.015 0.001±0.0010.001_± 0.001 0.813±0.0120.813_± 0.012 0.804±0.0140.804_± 0.014 0.841±0.0090.841_± 0.009 0.945±0.0060.945_± 0.006 Aromatics 91.4291.42 0.976±0.0040.976_± 0.004 0.954±0.0070.954_± 0.007 0.955±0.0130.955_± 0.013 0.983±0.0050.983_± 0.005 0.983±0.0040.983_± 0.004 0.987±0.0020.987_± 0.002 0.993±0.0010.993_± 0.001 Alkyl halides 46.5046.50 0.653±0.0170.653_± 0.017 0.600±0.0020.600_± 0.002 0.661±0.0060.661_± 0.006 0.731±0.0050.731_± 0.005 0.722±0.0040.722_± 0.004 0.740±0.0030.740_± 0.003 0.809±0.0040.809_± 0.004 Esters 18.2318.23 0.650±0.0250.650_± 0.025 0.551±0.0240.551_± 0.024 0.312±0.0120.312_± 0.012 0.790±0.0160.790_± 0.016 0.813±0.0150.813_± 0.015 0.854±0.0100.854_± 0.010 0.925±0.0060.925_± 0.006 Ketones 8.228.22 0.319±0.0410.319_± 0.041 0.268±0.0120.268_± 0.012 0.002±0.0010.002_± 0.001 0.433±0.0430.433_± 0.043 0.459±0.0330.459_± 0.033 0.550±0.0320.550_± 0.032 0.692±0.0270.692_± 0.027 Aldehydes 2.672.67 0.653±0.0370.653_± 0.037 0.426±0.0250.426_± 0.025 0.117±0.0200.117_± 0.020 0.956±0.0090.956_± 0.009 0.968±0.0070.968_± 0.007 0.967±0.0060.967_± 0.006 0.988±0.0030.988_± 0.003 Carboxylic acids 9.739.73 0.806±0.0140.806_± 0.014 0.524±0.0190.524_± 0.019 0.739±0.0210.739_± 0.021 0.947±0.0050.947_± 0.005 0.943±0.0030.943_± 0.003 0.964±0.0020.964_± 0.002 0.985±0.0010.985_± 0.001 Ether 37.0537.05 0.656±0.0230.656_± 0.023 0.571±0.0150.571_± 0.015 0.627±0.0130.627_± 0.013 0.785±0.0080.785_± 0.008 0.797±0.0070.797_± 0.007 0.837±0.0040.837_± 0.004 0.905±0.0030.905_± 0.003 Acyl halides 0.590.59 0.132±0.0300.132_± 0.030 0.277±0.0440.277_± 0.044 0.000±0.0000.000_± 0.000 0.205±0.0290.205_± 0.029 0.436±0.0490.436_± 0.049 0.549±0.0310.549_± 0.031 0.719±0.0260.719_± 0.026 Amides 25.6125.61 0.657±0.0090.657_± 0.009 0.620±0.0030.620_± 0.003 0.558±0.0150.558_± 0.015 0.764±0.0040.764_± 0.004 0.796±0.0030.796_± 0.003 0.829±0.0030.829_± 0.003 0.899±0.0030.899_± 0.003 Nitro 5.275.27 0.346±0.0360.346_± 0.036 0.190±0.0170.190_± 0.017 0.000±0.0000.000_± 0.000 0.518±0.0230.518_± 0.023 0.527±0.0170.527_± 0.017 0.649±0.0110.649_± 0.011 0.814±0.0140.814_± 0.014 Additional molecular structure elucidation results Table J: Ablation of IR and molecular-formula conditioning for molecular structure elucidation on the NIST benchmark. Tanimoto@k denotes the maximum Morgan Tanimoto similarity between the ground-truth structure and the first k candidates. Results are reported as means ± standard deviations across five folds. Model Top-1 Top-5 Top-10 Acc (%) ↑ Tanimoto ↑ Acc (%) ↑ Tanimoto ↑ Acc (%) ↑ Tanimoto ↑ UltraIR w/o IR 9.71±0.219.71_± 0.21 0.293±0.0020.293_± 0.002 25.54±0.4825.54_± 0.48 0.476±0.0030.476_± 0.003 33.17±0.4333.17_± 0.43 0.548±0.0030.548_± 0.003 UltraIR w/o Formula 38.94±0.2638.94_± 0.26 0.594±0.0030.594_± 0.003 46.66±0.5646.66_± 0.56 0.676±0.0050.676_± 0.005 49.45±0.4649.45_± 0.46 0.702±0.0050.702_± 0.005 UltraIR 52.20±0.5452.20_± 0.54 0.682±0.0040.682_± 0.004 57.48±0.7157.48_± 0.71 0.742±0.0030.742_± 0.003 57.57±0.7457.57_± 0.74 0.748±0.0030.748_± 0.003 Figure B: Additional molecular structure elucidation examples. Six representative examples from the NIST test set comparing the ground-truth molecular structures with candidates generated by UltraIR, UltraIR without pretraining, and the competing structure-elucidation models IRtoMol, AISE, PBSA, and DLIR. Fingerprint-based Tanimoto similarities between each generated candidate and the corresponding ground-truth structure are reported. Green check marks indicate exact structure recovery. Additional physicochemical property prediction results Figure C: Property-wise physicochemical prediction performance of UltraIR across NIST, SDBS, and USPTO. Predicted-versus-true parity plots are shown for synthetic accessibility (SA) score, logP P, topological polar surface area (TPSA), hydrogen-bond donors, hydrogen-bond acceptors, rotatable bonds, the fraction of sp3sp^3-hybridized carbon atoms (Fraction Csp3), quantitative estimate of drug-likeness (QED), aromatic rings, and aliphatic rings. The identity line denotes ideal agreement. Additional bacterial classification results Figure D: Confusion matrices for genus-level bacterial classification. Confusion matrices are shown for UltraIR and XGBoost across the nine evaluated bacterial genera. Rows denote actual genera and columns denote predicted genera, with each cell indicating the percentage of samples assigned to the corresponding predicted genus. Additional medicinal-herb geographic origin traceability results Figure E: Confusion matrices for medicinal-herb geographic origin traceability. a, Confusion matrices for Jinyinhua geographic origin traceability, comparing UltraIR and XGBoost across Shandong, Henan, Hebei, and Sichuan. b, Confusion matrices for Shanyinhua geographic origin traceability, comparing UltraIR and XGBoost across Hunan, Hubei, Sichuan, Henan, and Guangdong. Rows denote actual origins and columns denote predicted origins, with each cell indicating the percentage of samples assigned to the corresponding predicted origin. Additional microplastics classification results Figure F: Confusion matrices for microplastics classification. Confusion matrices are shown for UltraIR and Softmax across the 18 evaluated polymer classes. Rows denote actual plastic types and columns denote predicted plastic types, with each cell indicating the percentage of samples assigned to the corresponding predicted class. Additional cross-instrument and cross-laboratory generalization results Figure G: Cross-instrument and cross-laboratory generalization for soil property prediction. All models were trained on the Kellogg Soil Survey Laboratory (KSSL) subset of OSSL and evaluated by zero-shot inference on the ICRAF–ISRIC subset. a, Prediction performance for soil texture and acidity properties, including clay content, silt content, sand content, and pH (H2O)(H_2O), evaluated using Normalized MAE, Normalized RMSE, and R2R^2. b, Prediction performance for soil chemical properties, including cation exchange capacity (CEC) and exchangeable Ca, Mg, K, and Na, evaluated using the same metrics.