Paper deep dive
Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization
Quanjia Xiao, Weimin Ouyang, Zonglin Yang, Tianhao Wu, Qingguo Zhou, Runze Mao, Zhi X. Chen
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 7:36:16 AM
Summary
This paper proposes a full-stack domain-enhanced Large Language Model (LLM) workflow tailored for combustion science to address hallucinations and lack of physical consistency in general-purpose models. The methodology includes constructing a large-scale hybrid corpus, performing incremental pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning with verifiable rewards (RLVR). The authors introduce FlameBench, a standardized benchmark for evaluating complex reasoning in combustion science, and demonstrate that their approach significantly outperforms state-of-the-art general models and retrieval-augmented generation (RAG) methods.
Entities (10)
Relation Signals (8)
FlameBench â evaluates â Combustion Science Reasoning
confidence 95% ¡ FlameBench, a standardized evaluation benchmark specifically designed for complex reasoning tasks in combustion science.
Full-Stack Domain Enhancement Workflow â includes â SFT
confidence 95% ¡ ...supervised fine-tuning (SFT), and reinforcement learning with verifiable rewards (RLVR)...
Full-Stack Domain Enhancement Workflow â includes â RLVR
confidence 95% ¡ ...reinforcement learning with verifiable rewards (RLVR)...
Full-Stack Domain Enhancement Workflow â includes â CPT
confidence 95% ¡ The proposed framework jointly improves domain knowledge acquisition... encompassing incremental pre-training...
Qwen-8B â servesasbasefor â Proposed Model
confidence 95% ¡ We adopt Qwen-8B (Qwen Team et al., 2025) as the base model...
Proposed Model â outperforms â RAG
confidence 92% ¡ Experimental results demonstrate that the model developed in this work significantly outperforms... traditional retrieval-augmented generation methods
RLVR â basedon â GRPO
confidence 90% ¡ We introduce reinforcement learning with verifiable rewards (RLVR) based on the GRPO framework
Proposed Model â outperforms â
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) in the direction of task adaptation and capability enhancement for professional fields demonstrate significant application potential. Nevertheless, for complex physical systems such as combustion science, general-purpose LLMs often generate severe hallucinations due to insufficient domain knowledge and the inability to adhere to physical conservation laws. To address this issue, we propose the first full-stack domain-enhanced LLM workflow tailored for the field of combustion science, which integrates automated domain corpus construction, incremental pre-training, instruction fine-tuning, and verifiable reward-based reinforcement learning. This workflow ensures that the model truly internalizes physical laws rather than merely learning textual statistical patterns. We also release FlameBench, a standardized evaluation benchmark specifically designed for complex reasoning tasks in combustion science. Experimental results demonstrate that the model developed in this work significantly outperforms state-of-the-art general-purpose closed-source models and traditional retrieval-augmented generation methods on combustion science reasoning tasks. This work lays a solid technical and resource foundation for the subsequent development of domain-specific scientific research agents with reliable scientific reasoning capabilities.
Tags
Links
- Source: https://arxiv.org/abs/2603.19268v1
- Canonical: https://arxiv.org/abs/2603.19268v1
Trouble viewing inline? Open PDF directly â
Full Text
32,184 characters extracted from source content.
Expand or collapse full text
Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization Quanjia Xiao 1 Weimin Ouyang 2 Zonglin Yang 1 Tianhao Wu 2 Qingguo Zhou 3 Runze Mao â 1 2 Zhi X. Chen â 1 2 Abstract Large language models (LLMs) in the direction of task adaptation and capability enhancement for professional fields demonstrate significant appli- cation potential. Nevertheless, for complex physi- cal systems such as combustion science, general- purpose LLMs often generate severe hallucina- tions due to insufficient domain knowledge and the inability to adhere to physical conservation laws. To address this issue, we propose the first full-stack domain-enhanced LLM workflow tai- lored for the field of combustion science, which integrates automated domain corpus construction, incremental pre-training, instruction fine-tuning, and verifiable reward-based reinforcement learn- ing. This workflow ensures that the model truly in- ternalizes physical laws rather than merely learn- ing textual statistical patterns. We also release FlameBench, a standardized evaluation bench- mark specifically designed for complex reason- ing tasks in combustion science. Experimental results demonstrate that the model developed in this work significantly outperforms state-of-the- art general-purpose closed-source models and tra- ditional retrieval-augmented generation methods on combustion science reasoning tasks. This work lays a solid technical and resource foundation for the subsequent development of domain-specific scientific research agents with reliable scientific reasoning capabilities. 1. Introduction AI for Science (AI4S) has demonstrated remarkable success in advancing model construction and problem-solving for specialized fields, exemplified by advances in protein struc- â Equal contribution corresponding authors. 1 Peking Univer- sity 2 AI for Science Institute, Beijing 3 DP Technology. Corre- spondence to: Runze Mao<maorz@pku.edu.cn>, Zhi X. Chen <chenzhi@pku.edu.cn>. Preprint. March 23, 2026. ture prediction (Jumper et al., 2021; Lin et al., 2023) and materials property optimization (Merchant et al., 2023). A key enabler of this progress is the emergence of Large Lan- guage Models (LLMs), which provide strong capabilities in natural language understanding, knowledge representation, and multi-step reasoning, enabling applications such as sci- entific literature analysis, hypothesis generation (Kumbhar et al., 2025), and experimental planning (Boiko et al., 2023). Despite these successes, most existing LLMs are trained predominantly on general-domain corpora (Taylor et al., 2022), resulting in limited coverage of specialized scientific knowledge. This limitation is particularly pronounced in engineering sciences, where problem-solving is governed by strict physical constraints and complex multi-physics interactions (Raissi et al., 2019). Combustion science, a foundational discipline in energy systems and aerospace engineering, studies the coupled dy- namics of chemical reactions, fluid transport, and energy conversion during fuel oxidation (Wu et al., 2025). From a modeling perspective, combustion problems exhibit sev- eral properties that pose substantial challenges for general- purpose LLMs. First, effective reasoning requires the inte- gration of heterogeneous domain knowledge spanning chem- ical kinetics, fluid mechanics, and thermodynamics (Ouyang et al., 2024). Second, combustion processes are inherently spatiotemporal and nonlinear, involving cross-scale interac- tions from microscopic reaction pathways to macroscopic flow structures. Third, valid reasoning is constrained by fun- damental physical laws, including mass, momentum, and energy conservation, violations of which render predictions physically meaningless (Baez et al., 2024). Consequently, the primary challenge in combustion-related tasks lies not in static state prediction, but in process-level reasoning under explicit physical constraints. Current LLMs face two major bottlenecks when applied to combustion science. The first is a domain knowledge deficit: combustion-specific concepts, equations, and mod- eling paradigms are sparsely represented in general training corpora, limiting the modelâs ability to interpret special- ized symbolic systems such as reaction mechanisms and turbulent combustion models (M. Bran et al., 2024; Zhao 1 arXiv:2603.19268v1 [cs.CL] 27 Feb 2026 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization et al., 2025). The second is the lack of physicochemical consistency in reasoning. Without explicit awareness of con- servation laws and kinetic constraints, LLMs are prone to generating physically implausible outputs when confronted with tightly coupled multi-parameter scenarios, such as reac- tion pathways that violate chemical kinetics or efficiency es- timates that contradict energy conservation. While retrieval- augmented generation (RAG) can partially mitigate knowl- edge sparsity by injecting external documents at inference time (Lewis et al., 2020; Zhong et al., 2025), it does not address inconsistencies arising during multi-step reasoning. Existing domain adaptation approaches, which typically fo- cus on a single training stage, therefore remain insufficient for robust deployment in combustion-related applications. To address these challenges, we propose a full-stack domain- enhanced LLM workflow tailored for combustion sci- ence. The proposed framework jointly improves domain knowledge acquisition and physically consistent reasoning through dedicated corpus construction, multi-stage model optimization, and standardized evaluation. Our contribu- tions are summarized as follows: â˘Construction of a large-scale, combustion-specific cor- pus derived from hundreds of thousands of relevant scientific publications, providing high-quality training data for domain-specific language models. ⢠Development of a full-stack domain adaptation work- flow that integrates data generation, multi-stage op- timization, and objective evaluationâencompassing incremental pre-training, supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and the creation of a benchmark dataset (FlameBench). This workflow simultaneously addresses critical knowl- edge gaps and establishes a standardized basis for eval- uating domain-specific model performance. 2. Related Work 2.1. Domain-Specific LLMs in AI4S Large Language Models have recently been adapted to a range of scientific domains under the AI for Science (AI4S) paradigm. In bioinformatics, the AlphaFold series demon- strates that incorporating domain structure can enable accu- rate protein structure prediction (Jumper et al., 2021). In the biomedical domain, models such as BioGPT leverage large-scale medical corpora to support tasks including ques- tion answering and document summarization (Peng et al., 2023). Similar trends have emerged in materials science and molecular design, where domain-specialized approaches are trained on curated literature and experimental data to sup- port composition design and property prediction (G Ě omez- Bombarelli et al., 2018). These efforts collectively suggest that domain-specific data and inductive biases are essential for extending the capabilities of LLMs to scientific problem settings. However, despite the central role of combustion science in energy and aerospace applications, the develop- ment of domain-specific foundation models for combustion remains largely unexplored. 2.2. Domain Adaptation Methods for LLMs A variety of strategies have been proposed to adapt general- purpose LLMs to specialized domains. Continued or incre- mental pre-training on domain-specific corpora has been shown to improve vocabulary coverage and factual knowl- edge acquisition (Gururangan et al., 2020). Supervised fine-tuning (SFT) further aligns models with downstream tasks through instruction-style datasets (Wei et al., 2021). Retrieval-augmented generation (RAG) augments model outputs with external knowledge sources, mitigating fac- tual hallucinations in knowledge-intensive settings. In ad- dition, post-training alignment methods, including Direct Preference Optimization (DPO) (Rafailov et al., 2023) and Reinforcement Learning from Human Feedback (RLHF) (Ouyang et al., 2022), aim to improve output consistency and preference alignment. While effective in general do- mains, these approaches exhibit limitations in combustion- related applications: incremental pre-training alone does not enforce physical consistency, SFT provides limited supervi- sion for tightly coupled dynamical reasoning, RAG strug- gles with integrated multi-physics reasoning, and RLHF is constrained by the scalability of expert feedback in highly specialized scientific domains. 2.3. AI Applications in Combustion Science Machine learning has long been extensively applied to combustion-related research (Wu et al., 2025), and it has yielded remarkable results particularly in the aspects of pre- dictive modeling for combustion parameters and reduced- order characterization of complex physical processes. In recent years, with the advancement of large language models (LLMs), the academic community has begun to explore their novel applications in combustion scienceâfor instance, con- structing domain-adaptive solutions via retrieval-augmented generation (RAG) (Sharma & Raman, 2024). Such explo- rations hold certain value, yet they remain overall in the proof-of-concept stage: suffering from a narrow knowledge coverage and weak scene generalization ability, these ap- proaches not only cannot avoid factual hallucinations but also fail to support practical engineering deployment. 2 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization 3. Methodology 3.1. Overview We propose a full-stack domain-enhanced workflow that adapts general-purpose LLMs to combustion science by jointly addressing domain knowledge acquisition and physi- cally consistent reasoning. The workflow consists of four components: (i) construction of a large-scale combustion- specific corpus, (i) multi-stage model adaptation, (i) re- inforcement learning, and (iv) standardized evaluation us- ing a domain-specific benchmark. Starting from a general base model, the framework progressively injects domain knowledge and enforces physical consistency through in- cremental pre-training, supervised fine-tuning, and rein- forcement learning, followed by systematic evaluation on FLAMEBENCH. 3.2. Corpus Construction 3.2.1. DATA SOURCES To balance domain specialization and general reasoning capability, we construct a mixed corpus of approximately 30B tokens. The corpus includes 5B tokens of combustion- related data and 25B tokens of general pre-training data. The domain corpus is composed of English and Chinese aca- demic publications in combustion science, as well as curated scientific encyclopedic resources covering physics, chem- istry, and engineering fundamentals. The general corpus is sampled from open-source mid-training datasets (Team Olmo et al., 2025; Soldaini et al., 2024), ensuring reten- tion of general language understanding and commonsense reasoning. 25.7% 19.8% 16.1% 15.9% 11.0% 6.6% 4.8% PDFs and Web pages (25.9%) Combustion (20.0%) Math (16.2%) Code and Python (16.0%) QA (11.1%) Thinking (6.7%) Instruction (4.1%) Figure 1. Dataset token distribution by category. 3.2.2. AUTOMATED PROCESSING PIPELINE Raw documents are transformed into structured training data through an automated pipeline. PDF documents are first parsed and converted into Markdown format(Wang et al., 2024; Wei et al., 2025) . Multi-level deduplication is then performed at both exact and approximate levels, using hashing for exact matching and MinHash-based simi- larity detection for approximate matching (Jennings et al., 2023). Quality control is performed through a combination of rule-based filtering and model-based evaluation, where low-quality fragments are identified using perplexity and relevance scoring and subsequently repaired or removed. This process yields a high-quality structured corpus suitable for large-scale model training. Figure 2. Data processing pipeline for the combustion-specific pre-training corpus. 3.2.3. POST-TRAINING DATA CONSTRUCTION Based on the curated corpus, we construct datasets for su- pervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR). The SFT dataset includes 800K general instruction-following samples (Team Olmo et al., 2025) and 12K combustion-specific chain-of-thought ex- amples spanning knowledge queries, formula derivation, experimental analysis, and literature summarization. The RLVR dataset consists of 7K complex reasoning instances targeting multi-parameter coupling and process-level deduc- tion in combustion scenarios. 3 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization 3.3. Multi-Stage Model Adaptation 3.3.1. CONTINUE PRE-TRAINING We perform Continue pre-training(CPT) on the mixed cor- pus to inject combustion-specific knowledge while preserv- ing general language capabilities(Ibrahim et al., 2023; Que et al., 2024). Training is conducted for one epoch using a conservative learning rate schedule to avoid catastrophic forgetting. This stage enables the model to acquire domain terminology, core concepts, and fundamental physical rela- tionships relevant to combustion science. Figure 3. The construction pipeline of FlameBench. 3.3.2. SUPERVISED FINE-TUNING Following CPT, we apply supervised fine-tuning in two phases. The model is first aligned with general instruction- following behavior using a large-scale open-source instruc- tion dataset. It is then adapted to combustion-specific tasks through fine-tuning on domain-specific chain-of-thought data. Optimization is performed using cross-entropy loss with a cosine learning rate schedule. 3.3.3. REINFORCEMENT LEARNING To further improve reasoning reliability under physical con- straints, we introduce reinforcement learning with verifiable rewards (RLVR) based on the GRPO framework(DeepSeek- AI, 2025). A binary reward function is defined to explic- itly penalize violations of domain knowledge and physical consistency. We regularize policy updates using a KL diver- gence constraint to maintain training stability. This stage enables the model to generate physically plausible and logi- cally consistent solutions for complex combustion reasoning tasks. 3.4. FlameBench: A Domain-Specific Evaluation Benchmark To evaluate domain knowledge retention and constrained reasoning, we introduce FlameBench, a benchmark tailored to combustion science. The benchmark is constructed from high-information-density fragments extracted from peer- reviewed literature, dissertations, and domain-specific code repositories. Questions are generated and verified through an automated pipeline with dual validation, followed by expert refinement. The final benchmark consists of 436 high-quality questions covering eight combustion subfields, with each question grounded in a unique source reference to ensure reproducibility and verifiability. Fig. 3 illustrates the construction workflow of FlameBench. 4. Experiments and Result Analysis 4.1. Experimental Setup 4.1.1. BASE MODEL AND HARDWARE ENVIRONMENT We adopt Qwen-8B (Qwen Team et al., 2025) as the base model, chosen for its favorable trade-off between general- language capability and training efficiency. Continued pre- training (CPT), supervised fine-tuning (SFT), and evaluation were conducted on a distributed cluster equipped with 80 Huawei Ascend 910B accelerators. The reinforcement learn- ing with verifiable rewards (RLVR) stage was executed on a separate node with 8 NVIDIA A800 GPUs. 4.1.2. TRAINING STAGES AND CONTROL GROUPS The primary experimental variable is the training stage. To quantify the incremental contribution of each component in our pipeline, we consider the following model variants: â˘Baseline: The original Qwen-8B model without do- main adaptation. â˘CPT: Qwen-8B is further trained on a 30B-token hy- brid corpus (5B combustion-domain tokens and 25B general-domain tokens). â˘SFT-General: The CPT model fine-tuned on 800K general-purpose instruction-following samples. â˘SFT-Combustion:The SFT-General model fur- ther fine-tuned on 12K combustion-specific chain-of- thought (CoT) samples, targeting domain-specific rea- soning tasks , and serving as the cold-start model for subsequent RLVR optimization. ⢠RLVR-Opt: The SFT-Combustion model further trained with RLVR on 7K complex samples. 4 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization â˘RAG-Methods: As illustrated in Figure 4, the RAG system is implemented using FAISS for indexing, BGE- M3 (Chen et al., 2024) for embedding generation, and LangChain (LangChain Team, 2022) for orchestration. Figure 4. Architecture of the Closed-loop RAG Assessment Sys- tem. 4.1.3. EVALUATION PROTOCOL We evaluate all models on FlameBench, a domain-specific benchmark designed to assess reasoning in combustion sci- ence. Performance is measured using multiple-choice ac- curacy, which directly reflects the modelâs ability to apply fundamental principles such as chemical kinetics, transport phenomena, and multi-physics coupling. All evaluation items undergo a three-stage validation processâdifficulty filtering, correctness verification, and expert calibrationâto ensure coverage, reliability, and reproducibility. 4.2. Results and Analysis 4.2.1. PERFORMANCE ACROSS TRAINING STAGES Table 1 reports the performance of different training stages on FlameBench. We observe a clear monotonic improve- ment as the training pipeline progresses, indicating that each stage contributes complementary benefits while maintaining practical inference efficiency. Table 1. Performance Comparison of Multi-stage Training Model GroupAccuracy (%) Qwen3-8B-Base26.8 CPT33.3 SFT-General33.5 SFT-Combustion35.1 RLVR-Opt43.8 Table 2. Performance Comparison of Different Models ModelAccuracy (%) GPT-515.60 GLM-432.64 Gemini Pro32.10 DeepSeek-R128.37 Average Score27.18 Impact of Continued Pre-training Relative to the base- line (Qwen3-8B-Base, accuracy = 26.8%), Continued Pre- training (CPT) yields a substantial improvement of 6.5 per- centage points (accuracy = 33.3%), confirming that large- scale hybrid-corpus pre-training effectively injects founda- tional combustion knowledge. Notably, CPT outperforms the average accuracy of RAG-based methods (26.24%) by 7.06 percentage points, and even surpasses the best- performing RAG model (RAG + GLM-4, 32.09%). This result demonstrates that internalized domain knowledge pro- vides a stronger inductive bias for combustion reasoning than external retrieval, which often suffers from irrelevant context injection and limited multi-step reasoning capabil- ity. Effects of Supervised Fine-tuningGeneral-purpose SFT (SFT-General, 33.5%) only brings marginal improvement over CPT, indicating that it primarily enhances instruc- tion adherence rather than domain reasoning. In contrast, domain-specific SFT (SFT-Combustion, 35.1%) further boosts accuracy by 1.8 percentage points, as it aligns the model with combustion-specific problem-solving patterns such as chemical kinetics derivation and thermodynam- ics constraint application. Specifically, SFT-Combustion Table 3. Performance of RAG with Different Models MethodAccuracy (%) RAG + GPT-516.52 RAG + GLM-432.09 RAG + Gemini Pro27.80 RAG + DeepSeek-R128.54 Average Score26.24 5 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization Figure 5. Loss and learning rate curves during the Continue Pre-training and SFT stages. outperforms closed-source models including DeepSeek-R1 (28.37%) and Gemini Pro (32.10%) on FlameBench, and approaches the performance of GLM-4 (32.64%) in the turbulent combustion subfield. The modest gain also re- flects the high complexity of FlameBench, where standard SFT alone is insufficient to address multi-physics coupling challenges. Benefits of RLVR Optimization Reinforcement Learn- ing from Verifiable Rewards (RLVR) substantially improves upon the SFT-Combustion baseline, increasing accuracy from 35.1% to 43.8%. As shown in Figure 6, the vanilla Qwen-8B RL baseline exhibits a rapid reduction in response lengthâfrom approximately 1,500 tokens to 100â200 to- kensâreflecting a tendency toward shortcut-driven pol- icy collapse during optimization. In contrast, the RLVR model initialized from SFT-Combustion maintains a stable response length of around 2,000 tokens throughout training. This length stability, together with consistently higher val- idation accuracy and mean rewards under a low-entropy policy regime, is indicative of convergence toward a more deterministic and stable optimization regime, rather than reliance on exploitative or reward-hacking behaviors. While sustained output length alone does not guarantee improved reasoning fidelity, these results are consistent with more systematic and physically grounded reasoning processes. Overall, our findings suggest that RLVR, when fortified with domain-specific priors, stabilizes policy optimization and reduces physics-inconsistent generations in combustion science tasks. 4.3. Comparison with RAG Methods A direct quantitative comparison between RLVR-Opt and RAG baselines (Table 3) highlights three distinct advantages of our end-to-end training paradigm. First, knowledge inter- nalization enables RLVR-Opt to outperform the best RAG model (RAG + GLM-4, 32.09%) by 11.71 percentage points, avoiding the context interference and performance ceiling of retrieval-based systems. Second, domain-specific reasoning 050100150200 Steps 0 1000 2000 Value Response Length 050100150200 Steps 0.0 0.1 0.2 0.3 0.4 Value Val Accuracy Qwen-8B RL SFT-Combustion RL (a) Response Length and Validation Accuracy 050100150200 Steps 1 2 3 Value Policy Entropy 050100150200 Steps 0.0 0.2 0.4 Value Mean Reward Qwen-8B RL SFT-Combustion RL (b) Mean Reward and Policy Entropy Figure 6.Comparison of training metrics between SFT- Combustion RL and Qwen-8B RL. emerges from multi-stage optimization, allowing the model to handle cross-disciplinary coupling (e.g., fluid mechanics + chemical reactions) more reliably than RAG, which often struggles with integrating heterogeneous knowledge frag- ments. Third, inference efficiency is drastically improved: RLVR-Opt eliminates the overhead of document retrieval and embedding matching, reducing inference latency for interactive scientific analysis and real-time simulation. 4.4. Experimental Conclusion Overall, the results demonstrate that the proposed CPTâSFTâ RLVR framework substantially improves LLM performance in combustion science. Hybrid-corpus CPT is necessary for effective knowledge injection, while RLVR constitutes the key step for surpassing the limitations of both SFT and RAG by explicitly optimizing scientific reasoning. The resulting 6 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization model achieves state-of-the-art performance on FlameBench while maintaining efficient end-to-end inference. 5. Conclusion This work addresses the limitations of general-purpose large language models (LLMs) in combustion science, particu- larly the lack of domain-specific knowledge. We propose a full-stack domain-enhanced workflow that integrates three core components: construction of a dedicated combustion corpus, multi-stage model training (CPT, SFT, and RLVR), and a standardized evaluation benchmark (FlameBench). Experimental results demonstrate that the proposed work- flow systematically improves both domain knowledge reten- tion and reasoning capabilities. Specifically, the model opti- mized through CPTâSFTâRLVR achieves the highest per- formance on FlameBench, outperforming general-purpose LLMs and retrieval-augmented baselines. Beyond combustion science, this workflow provides a gen- eralizable paradigm for enhancing LLMs in engineering domains characterized by strict physical laws and multidis- ciplinary coupling. Future work will explore extending this approach to broader AI-for-Science applications and more complex experimental scenarios, aiming to further integrate LLMs into scientific discovery workflows. Impact Statement This paper presents work whose goal is to advance the field of Machine Learning. There are many potential societal consequences of our work, none which we feel must be specifically highlighted here. References Baez, A., Zhang, W., Ma, Z., Das, S., Nguyen, L. M., and Daniel, L. Guaranteeing conservation laws with projec- tion in physics-informed neural networks. arXiv preprint arXiv:2410.17445, 2024. URLhttps://arxiv.or g/abs/2410.17445. Boiko, D. A., MacKnight, R., Kline, B., and Gomes, G. Au- tonomous chemical research with large language models. Nature, 624(7992):570â578, 2023. doi: 10.1038/s415 86-023-06792-0. URLhttps://w.nature.com /articles/s41586-023-06792-0. Chen, J. et al.BGE M3-Embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216, 2024. URLhttps://arxiv.or g/abs/2402.03216. DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning. arXiv preprint arXiv:2501.12948, 2025. G Ě omez-Bombarelli, R., Duvenaud, D., S Ě anchez-Lengeling, B., Sheberla, D., Aguilera-Iparraguirre, J., Hirzel, T. D., Adams, R. P., and Aspuru-Guzik, A. Automatic chemical design using a data-driven continuous representation of molecules. ACS Central Science, 4(2):268â276, 2018. doi: 10.1021/acscentsci.7b00572. Gururangan, S. et al. Donât stop pretraining: Adapt language models to domains and tasks. In ACL, 2020. Ibrahim, A. et al. Simple and scalable strategies to con- tinually pre-train large language models. arXiv preprint arXiv:2312.06946, 2023. URLhttps://arxiv.or g/abs/2312.06946. Jennings, J. et al. Curating trillion-token datasets: Intro- ducing NVIDIA NeMo data curator. NVIDIA Developer Blog, Aug 2023. URLhttps://developer.nvid ia.com/blog/curating-trillion-token-d atasets-introducing-nemo-data-curator /. Jumper, J., Evans, R., Pritzel, A., Green, T., Figurnov, M., Ronneberger, O., Tunyasuvunakool, K., Bates, R., Ë Z Ě Äądek, A., Potapenko, A., Bridgland, A., et al. Highly accurate protein structure prediction with AlphaFold. Nature, 596 (7873):583â589, 2021. doi: 10.1038/s41586-021-03819 -2. URLhttps://w.nature.com/article s/s41586-021-03819-2. Kumbhar, S., Mishra, V., Coutinho, K., Handa, D., Iquebal, A., and Baral, C. Hypothesis generation for materials discovery and design using goal-driven and constraint- guided llm agents, 2025. URLhttps://arxiv.or g/abs/2501.13299. LangChain Team. LangChain: Building applications with LLMs through composability, 2022.URLhttps: //github.com/langchain-ai/langchain. Accessed: 2025. Lewis, P. et al.Retrieval-augmented generation for knowledge-intensive nlp tasks. In NeurIPS, 2020. Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., Smetanin, N., Verkuil, R., Kabeli, O., Shmueli, Y., dos Santos Costa, A., et al. Evolutionary-scale predic- tion of atomic-level protein structure with a language model. Science, 379(6637):1123â1130, 2023.doi: 10.1126/science.ade2574. URLhttps://w.scie nce.org/doi/10.1126/science.ade2574. M. Bran, A., Cox, S., Schilter, O., Baldassari, C., White, A. D., and Schwaller, P. Augmenting large language 7 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization models with chemistry tools. Nature Machine Intelli- gence, 6(5):525â535, 05 2024. ISSN 2522-5839. doi: 10.1038/s42256-024-00832-8. URLhttps://doi. org/10.1038/s42256-024-00832-8. Merchant, A., Batzner, S., Schoenholz, S. S., Aykol, M., Cheon, G., and Cubuk, E. D. Scaling deep learning for materials discovery. Nature, 624(7990):80â85, 2023. doi: 10.1038/s41586- 023- 06735-9. URLhttps: //w.nature.com/articles/s41586-023 -06735-9. Ouyang, L. et al. Training language models to follow in- structions with human feedback. In NeurIPS, 2022. Ouyang, S., Zhang, Z., Yan, B., Liu, X., Choi, Y., Han, J., and Qin, L. Structured chemistry reasoning with large lan- guage models. Conference on Machine Learning, 2024. Peng, Y., Wang, Z., Zhang, C., Li, H., Liu, J., and Xu, J. BioGPT: Generative pre-trained transformer for biomed- ical text generation and mining. Bioinformatics, 39(5): btad325, 2023. doi: 10.1093/bioinformatics/btad325. Que, H. et al. D-cpt law: Domain-specific continual pre- training scaling law for large language models. arXiv preprint arXiv:2406.01375, 2024. URLhttps://ar xiv.org/abs/2406.01375. Qwen Team et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. URLhttps://arxiv.or g/abs/2505.09388. Rafailov, R. et al. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, 2023. Raissi, M., Perdikaris, P., and Karniadakis, G. E. Physics- informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. Journal of Computational Physics, 378:686â707, 2019. doi: 10.1016/j.jcp.2018.10. 045. URLhttps://w.sciencedirect.com/ science/article/pii/S0021999118307125. Sharma, V. and Raman, V. A reliable knowledge process- ing framework for combustion science using foundation models. Energy and AI, 16:100365, 2024. Soldaini, L. et al. Dolma: An open corpus of three trillion tokens for language model pretraining research. arXiv preprint arXiv:2402.00159, 2024. URLhttps://ar xiv.org/abs/2402.00159. Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R. Galactica: A large language model for science, 2022. URL https://arxiv.org/abs/2211.09085. Team Olmo et al. Olmo 3. arXiv preprint arXiv:2512.13961, 2025. URLhttps://arxiv.org/abs/2512.1 3961. Wang, B. et al. Mineru: An open-source solution for precise document content extraction. arXiv preprint arXiv:2409.18839, 2024. URLhttps://arxiv.or g/abs/2409.18839. Wei, H. et al. Deepseek-ocr: Contexts optical compression. arXiv preprint arXiv:2510.18234, 2025. URLhttps: //arxiv.org/abs/2510.18234. Wei, J. et al. Finetuned language models are zero-shot learn- ers. In International Conference on Machine Learning, 2021. Wu, J., Wang, X., Zhang, G., Liu, J., Li, X., Zhang, Y., Zhang, H., Lyu, J., Wang, B., and Wu, Y. Physics- informed machine learning for combustion: A review, 2025. URLhttps://arxiv.org/abs/2509.0 3347. Zhao, Z., Ma, D., Chen, L., Sun, L., Li, Z., Xia, Y., Chen, B., Xu, H., Zhu, Z., Zhu, S., Fan, S., Shen, G., Yu, K., and Chen, X. Developing chemdfm as a large language foundation model for chemistry. Cell Reports Physical Science, 6(4), 04 2025. ISSN 2666-3864. doi: 10.1016/j. xcrp.2025.102523. URLhttps://doi.org/10.1 016/j.xcrp.2025.102523. Zhong, X., Jin, B., Ouyang, S., Shen, Y., Jin, Q., Fang, Y., Lu, Z., and Han, J. Benchmarking retrieval-augmented generation for chemistry. 2025. URLhttps://arxi v.org/abs/2505.07671. 8 Full-Stack Domain Enhancement for Combustion LLMs: Construction and Optimization A. Appendix: Detailed Training Configurations A.1. CPT and SFT Stages Table 4 presents the hyperparameters for CPT, SFT-General, and SFT-Combustion. All stages were conducted using the LLaMA-Factory framework with DeepSpeed ZeRO-3. HyperparameterCPTSFT-GeneralSFT-Combustion Learning Rate2.0e-55.0e-52.0e-5 LR SchedulerWSDCosineCosine Max Length16,38416,38420,000 Batch Size (Total)256256128 Precisionbf16bf16bf16 Warmup Ratio0.010.050.03 Table 4. Hyperparameters for CPT and SFT Stages. A.2. Reinforcement Learning (GRPO) The final stage employed the Group Relative Policy Optimization (GRPO) algorithm using theVerlframework. We utilized vLLM as the rollout engine to efficiently sample multiple responses for reward normalization. ParameterValue AlgorithmGRPO Actor Learning Rate 2.0Ă 10 â6 Global Training Batch Size128 Max Prompt Length1,024 Max Response Length8,192 KL Coefficient (β)0.005 KL Loss Typelowvarkl Number of Samples (n)8 Rollout EnginevLLM PrecisionMixed (bf16/fp32) Table 5. Hyperparameters for the GRPO reinforcement learning stage. 9