Paper deep dive
A Locally Deployable Tool-Grounded LLM Multi-agent Framework for Automating Methane Emission Analysis and Reporting
Yang Yan, Zifan Zhou, Xuan Wang, Erum Hassan, Bilguunzaya Mijiddorj, Jie Cao, Bin Li, Binbin Weng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/20/2026, 3:45:02 AM
Summary
This paper presents a locally deployable, tool-grounded Large Language Model (LLM) multi-agent framework designed to automate methane emission analysis and reporting. The system integrates field measurements, meteorological data, and deterministic sensor-processing routines (specifically Gaussian plume inversion) to estimate emission rates and locate sources. Deployed across wastewater treatment facilities, landfills, and oil and gas sites, the framework achieved high accuracy in workflow routing (92.0%), emission estimation (85.0%), and report generation (95.0%), significantly reducing processing time and manual effort while maintaining data privacy through local execution.
Entities (10)
Relation Signals (10)
LLM Multi-agent Framework → utilizes → Gaussian Plume Model
confidence 98% · The framework uses LLM agents as workflow coordinators that link... deterministic sensor-processing routines, Gaussian plume inversion...
LLM Multi-agent Framework → deploys → Llama-3.1-8B
confidence 95% · Llama 3.1 8B was used as the default model for general interaction and task routing.
LLM Multi-agent Framework → deploys → Qwen 30B
confidence 95% · Qwen 30B was used for more complex text reasoning tasks.
LLM Multi-agent Framework → uses → AIMNet
confidence 95% · In this work, we apply this tool-grounded design to methane field monitoring by developing a locally deployable LLM multi-agent framework using the AIMNet platform.
LLM Multi-agent Framework → testedat → Oil and Gas Sites
confidence 92% · Extensive field deployments across diverse real-world environments (e.g., ... and oil and gas sites)
LLM Multi-agent Framework → testedat → Wastewater Treatment Facilities
confidence 92% · Extensive field deployments across diverse real-world environments (e.g., wastewater treatment facilities...)
LLM Multi-agent Framework → testedat → Landfills
confidence 92% · Extensive field deployments across diverse real-world environments (e.g., ... landfills ...)
→ →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Methane field monitoring requires the integration of sampling design, meteorological interpretation, sensor processing, plume analysis, visualization, and reporting, but these steps are often distributed across separate expert-driven workflows. We developed a locally deployable, tool-grounded large language model (LLM) multi-agent framework for our low-cost methane sensing and field-monitoring campaigns. The framework uses LLM agents as workflow coordinators that link field measurements, meteorological data, deterministic sensor-processing routines, Gaussian plume inversion, and report generation, rather than directly estimating methane concentrations or emissions. Extensive field deployments across diverse real-world environments (e.g., wastewater treatment facilities, landfills, and oil and gas sites) demonstrate that our framework can achieve 92.0\% accuracy in workflow routing and parameter extraction, 85.0\% success in emission-rate estimation and plume prediction, and 95.0\% success in generating editable reports under practical operating conditions. Compared with manual and general-purpose LLM-assisted workflows, it reduced workflow time from hours-level to minutes-level, lowered manual coordination and prompt-engineering requirements, and retained traceable plume-based outputs. In addition, most processing can be performed locally, reducing exposure of sensitive facility and field data to cloud services. These results indicate that tool-grounded LLM coordination can reduce the time, labor, usability, and data-security barriers of methane field monitoring.
Tags
Links
- Source: https://arxiv.org/abs/2608.18473v1
- Canonical: https://arxiv.org/abs/2608.18473v1
Trouble viewing inline? Open PDF directly →
Full Text
77,572 characters extracted from source content.
Expand or collapse full text
A Locally Deployable Tool-Grounded LLM Multi-agent Framework for Automating Methane Emission Analysis and Reporting Yang Yan Zifan Zhou Xuan Wang Erum Hassan Bilguunzaya Mijiddorj Jie Cao Bin Li Binbin Weng Abstract Methane field monitoring requires the integration of sampling design, meteorological interpretation, sensor processing, plume analysis, visualization, and reporting, but these steps are often distributed across separate expert-driven workflows. We developed a locally deployable, tool-grounded large language model (LLM) multi-agent framework for our low-cost methane sensing and field-monitoring campaigns. The framework uses LLM agents as workflow coordinators that link field measurements, meteorological data, deterministic sensor-processing routines, Gaussian plume inversion, and report generation, rather than directly estimating methane concentrations or emissions. Extensive field deployments across diverse real-world environments (e.g., wastewater treatment facilities, landfills, and oil and gas sites) demonstrate that our framework can achieve 92.0% accuracy in workflow routing and parameter extraction, 85.0% success in emission-rate estimation and plume prediction, and 95.0% success in generating editable reports under practical operating conditions. Compared with manual and general-purpose LLM-assisted workflows, it reduced workflow time from hours-level to minutes-level, lowered manual coordination and prompt-engineering requirements, and retained traceable plume-based outputs. In addition, most processing can be performed locally, reducing exposure of sensitive facility and field data to cloud services. These results indicate that tool-grounded LLM coordination can reduce the time, labor, usability, and data-security barriers of methane field monitoring. keywordsMethane monitoring, Large language models, Multi-agent systems, Source localization, Emission quantification, Environmental sensing †affiliation: School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, 73019 USA†affiliation: Department of Electrical Engineering, Pennsylvania State University, University Park, PA, 16801 USA†affiliation: Department of Electrical Engineering, Pennsylvania State University, University Park, PA, 16801 USA†affiliation: School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, 73019 USA†affiliation: School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, 73019 USA†affiliation: School of Computer Science, University of Oklahoma, Norman, OK, 73019 USA†affiliation: Department of Electrical Engineering, Pennsylvania State University, University Park, PA, 16801 USA†affiliation: School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, 73019 USA†email: binbinweng@ou.edu†abbreviations: IR,NMR,UV 1 Introduction Methane (CH4) is a short-lived greenhouse gas with a strong warming effect and reducing methane emissions is therefore an important part of efforts to limit near-term climate warming 29; 31; 23. Beyond its climate impact, methane leakage also raises safety concerns because the gas is highly flammable and can accumulate to hazardous levels in confined or poorly ventilated spaces 14. Anthropogenic methane emissions arise largely from fossil fuel production and agriculture, with landfills and wastewater treatment facilities also contributing important local sources 31; 23. At these sites, emissions can be intermittent and vary considerably across relatively short distances, which makes them difficult to characterize using sparse or infrequent measurements 7; 3; 10; 32; 36. More detailed measurements at individual facilities are therefore needed to determine where methane is released and to estimate the associated emission rate 15; 10; 6. As shown in Figure 1, current methane monitoring mainly relies on satellite observations, airborne surveys, and ground-based measurements 24; 15; 33; 34. Satellite and airborne observations provide broad spatial coverage, whereas ground-based and mobile measurements are needed when the objective is to examine individual sites more closely, locate emission sources, or conduct repeated surveys 22; 10; 5; 27; 12; 44; 47. Vehicle-based methane surveys can further provide rapid spatial coverage for leak detection, source localization, and emission-rate estimation.37 To support these field measurements, our previous studies developed AIMNet, a low-cost nondispersive infrared methane sensing platform that has been used for fixed-site monitoring as well as vehicle-based and portable surveys 48; 41; 42; 20; 39. However, substantial manual effort is still required before and after field measurements because the measurement plan must be prepared in advance and the collected data still need to be processed afterward. Figure 1: Conventional multi-expert methane monitoring workflows. The figure was generated with AI assistance based on author-provided content and subsequently reviewed and refined by the authors. In practice, field planning and post-measurement analysis are often handled using separate tools and procedures. For example, locating a suspected methane source first requires the operator to determine where measurements should be collected based on the site and weather conditions. After the survey, the measurement data may need to be organized, aligned with meteorological records, converted to consistent units, and corrected for background concentration before source localization or emission-rate estimation can be performed. Previous studies have developed methods for individual stages of this process, including monitoring design, source attribution, and emission quantification 27; 12; 5. Combining these stages into a complete workflow can still require users to transfer data between tools and maintain consistent analysis settings. This becomes particularly inconvenient when field objectives are given as simple natural-language requests rather than model-ready inputs. A framework that can interpret these requests and coordinate the required tools could therefore reduce the manual effort between field measurement and final analysis. Large language models (LLMs) have demonstrated strong capabilities in interpreting natural-language instructions 35; 8; 1; 28. More recent agentic frameworks extend these capabilities by allowing LLMs to coordinate external tools and specialized workflow components 40; 17; 13; 18. Tool-grounded multi-agent systems provide a way to separate these functions by assigning each agent a specific task and a predefined set of tools or data sources. The LLM can then interpret the request and coordinate the workflow while quantitative analysis is performed by the corresponding processing routines and physics-based models. Recent studies have begun to apply LLMs and LLM agents to environmental data analysis, decision support, and geospatial applications46; 26; 11; 9. However, reliable tool execution and domain-specific quantitative reasoning remain important challenges in such applications. In addition, field sensing requires the coordination of measured data, meteorological information, and physical analysis models rather than the interpretation of existing digital data alone. Integrating these elements into an agentic workflow remains an important step toward practical environmental monitoring. In this work, we apply this tool-grounded design to methane field monitoring by developing a locally deployable LLM multi-agent framework using the AIMNet platform. As shown in Figure 2, the framework begins with a natural-language request and first identifies the type of monitoring task. It can then support measurement planning or process collected methane data for further analysis. For source-analysis tasks, deterministic plume-modeling tools are used to reconstruct the methane plume and estimate the likely leak location and emission rate, and the resulting analysis is organized into a standardized field report. To evaluate the framework under realistic conditions, we deployed it in wastewater, landfill, and oil and gas environments and assessed its performance across the complete monitoring workflow. Unlike conventional methane monitoring workflows that require analysts to connect these steps manually, the proposed framework coordinates them within a single local system while retaining established processing routines and physical models for quantitative analysis. Field evaluation shows that this integrated workflow substantially reduces manual effort and processing time while maintaining reliable quantitative performance. 2 Overall Multi-agent System Description An LLM agent is a software component built around a large language model and assigned a specific role within a workflow. Its behavior is defined by task instructions together with the tools and structured information available to it, allowing the agent to receive a task and return results in a form that can be used by other parts of the system. A multi-agent system extends this structure by coordinating several agents, with each agent handling a different part of the overall task and passing intermediate results to other agents when needed. By separating a complex workflow into well-defined stages, the system allows each agent to focus on a limited responsibility while intermediate outputs can be checked before they are used in subsequent steps 40; 25; 19; 17. Figure 2: System architecture and workflow of the tool-grounded LLM multi-agent framework for methane monitoring. Figure 2 shows how this multi-agent structure is applied to methane monitoring. The framework is organized into four layers that move from user input and task routing through mission planning and methane data analysis to decision support and report generation. Within this structure, each agent is assigned a limited role and exchanges structured information with other agents or deterministic tools as the task progresses. The LLM agents use the user request to determine the required task and parameters, then coordinate the corresponding workflow and pass the resulting information to the next stage. Quantitative methane analysis is not performed directly by the LLM. Instead, sensor-processing routines and physics-based plume models are used to generate the numerical results. This separation allows the LLM to coordinate the overall workflow without replacing the established calculations used for methane analysis, keeping the quantitative results traceable and physically grounded. Local deployment was adopted because methane monitoring data may contain sensitive facility and location information21. Reliable Internet access may also be unavailable during field campaigns, making local execution advantageous for field deployment. In this study, the framework was deployed on a workstation running 64-bit Microsoft Windows 11. The workstation was equipped with an Intel Core i9-13900KF CPU with 24 cores and 32 logical processors, 64 GB of RAM, an NVIDIA GeForce RTX 4090 GPU with 24 GB of VRAM, and approximately 3 TB of SSD storage. The local models were served through Ollama 30. Llama 3.1 8B was used as the default model for general interaction and task routing. Qwen 30B was used for more complex text reasoning tasks. Qwen-VL 8B was used for image interpretation. 16; 30; 43; 4. User prompts, sensor measurements, and intermediate analytical results therefore remained within the local computing environment. External weather and mapping services were queried only when required. The framework can also support other locally deployed open-weight models, including the Llama and Qwen families 16; 43, or commercial LLM services when cloud-based inference is preferred. 2.1 Data Input Layer As shown in Fig. 2, the input layer serves as the entry point between the user and the downstream task-specific agents. Users can describe a monitoring task in natural language or upload a CSV file containing field measurements and their associated location and time information. These inputs can be used for general inquiries, measurement planning, or analysis of previously collected methane data. The Request Parser and Task Classifier converts the user input into a structured task specification that identifies what the user is asking, where and when the task applies, and what information is needed to complete it. The Agent Router then uses this specification to direct the request to the appropriate downstream agent and analytical tools. Figure 3: Example conversion of a natural-language request into structured parameters for methane leak localization. As illustrated in Fig. 3, planning and analysis requests are converted into a machine-readable format with standardized fields and units. General questions are routed directly to the selected LLM, whereas task-oriented requests may invoke Google Maps for geographic information, Open-Meteo for meteorological data, and the Message Queuing Telemetry Transport (MQTT)-based AIMNet database for methane measurements 48. 2.2 Mission Planning Layer The mission planning layer is designed to support efficient methane emission monitoring before field deployment. In conventional field measurements, operators usually carry the sensing device around the target facility and repeatedly search the downwind area to locate potential emission sources. Although this strategy is practical, it can be time consuming because the effective search region is strongly influenced by the target location, site layout, wind direction, wind speed, and short-term meteorological fluctuations2. To address this limitation, the mission planning agent extracts the planned measurement time, target location, and facility type from the user request. It retrieves the meteorological conditions for the site and uses the current wind speed and direction to estimate the downwind region where methane enhancement is most likely to be detected. The agent also considers forecast wind conditions over the following three hours, which approximates the upper duration of a typical field survey in this study, and adjusts the search region according to the expected wind trend. By accounting for both the current conditions and their expected changes during the survey, the resulting plan identifies the most favorable area for methane measurements and reduces unnecessary searching in the field. Based on the estimated monitoring region, the path planning agent generates an executable route for field deployment. The agent combines the mission planning output with GPS route information, road accessibility, and site geometry to identify efficient driving or walking paths. It further highlights high-probability detection areas where repeated measurements are recommended. 2.3 Data Analysis Layer For analysis-oriented tasks, the Data Analysis Agent retrieves methane measurements either from deployed sensing nodes through MQTT or from user-uploaded datasets collected using portable or vehicle-based sensing platforms. The input datasets may contain methane concentrations, GPS coordinates, timestamps, meteorological variables, and source information. The agent then coordinates data preprocessing and invokes the task-specific analytical routine. For GPS-referenced mobile measurements, the preprocessing routine checks the coordinate and timestamp fields and identifies the methane concentration field before analysis. Measurement units are converted to a consistent basis, and the background methane concentration is estimated from the observed data when no value is provided. Column names may differ among sensing instruments and processing workflows, so the methane field is selected based on the available metadata and expected naming conventions. Distribution-based checks are used when the field remains ambiguous. If the user specifies a methane column, that field is used after validation; otherwise, the most likely field is selected automatically. The numerical distribution and available metadata are subsequently evaluated to distinguish ppm- and ppb-scale measurements. When the median methane value is below 100, the measurements are interpreted as ppm-scale values unless the column name or metadata indicate ppb-scale concentrations or enhancements. If no background concentration is provided, the background methane level is estimated from the lower portion of the observed concentration distribution. Meteorological parameters are obtained using a priority-based procedure. User-provided wind speed, wind direction, and atmospheric stability class are used when available. If these parameters are incomplete but valid GPS and time information are provided, the Weather Agent retrieves wind speed, wind direction, temperature, humidity, and related variables from the selected meteorological source. Atmospheric stability is then assigned using a deterministic rule-based classifier based on wind speed, time of day, and other available weather conditions. If the required meteorological information remains unavailable, predefined fallback parameters are applied and the resulting analysis is assigned lower validity. For known-source plume reconstruction, missing meteorological parameters can alternatively be inferred by fitting candidate plume conditions to the observed methane measurements. In this study, a steady-state Gaussian plume model was used for source localization, emission-rate estimation, and plume reconstruction. The model was selected because it is widely used and computationally efficient for demonstrating the end-to-end analysis workflow. Other dispersion or inversion methods can also be incorporated into the same analysis layer when needed. The steady-state Gaussian plume concentration is expressed as C(x,y,z)=Q2πuσyσzexp(−y22σy2)[exp(−(z−H)22σz2)+exp(−(z+H)22σz2)],C(x,y,z)= Q2π u _y _z (- y^22 _y^2 ) [ (- (z-H)^22 _z^2 )+ (- (z+H)^22 _z^2 ) ], (1) where C(x,y,z)C(x,y,z) is the methane concentration enhancement at the receptor location. Q is the emission rate, u is the wind speed, and H is the effective source height. The variables x, y, and z represent the downwind, crosswind, and vertical distances from the source. The parameters σy _y and σz _z represent the lateral and vertical dispersion coefficients. They are determined from the downwind distance and the assigned atmospheric stability class. In this study, C represents methane enhancement above the estimated background concentration rather than the absolute methane concentration. For GPS-referenced mobile measurements, the receptor height is denoted as zrz_r, and the effective source height may be specified by the user or approximated based on the facility configuration. The processed measurements are then used to jointly estimate the source location and emission rate. For a candidate source location s, the predicted methane concentration at the i-th measurement point is represented as c^i(,Q)=b+Qki(), c_i(s,Q)=b+Qk_i(s), (2) where b is the estimated background methane concentration, Q is the methane emission rate, and ki()k_i(s) is the unit-emission plume response at the i-th receptor for the candidate source location s. The plume response depends on the relative downwind and crosswind positions of the receptor, wind speed, wind direction, atmospheric stability class, receptor height, and effective source height. The term Qki()Qk_i(s) therefore represents the modeled methane enhancement above the background concentration. For each candidate source location, the optimal non-negative emission rate is estimated by weighted least squares: Q∗()=argmin∑i=1NQ≥0wi[ci−b−Qki()]2,Q^*(s)= _Q≥ 0 _i=1^Nw_i [c_i-b-Qk_i(s) ]^2, (3) where cic_i is the observed methane concentration at the i-th measurement point, N is the number of measurements, and wiw_i is a data-dependent weight. Methane-enhanced observations are assigned greater weights to emphasize plume-affected measurements, while background-level observations are retained as negative evidence. This treatment penalizes candidate sources that would predict strong methane enhancements at locations where no corresponding enhancement was observed. Because the Gaussian plume response is linear with respect to Q, the optimal emission rate for a given candidate source can be obtained as Q∗()=max[0,∑i=1Nwiki()(ci−b)∑i=1Nwiki2()].Q^*(s)= [0, _i=1^Nw_ik_i(s)(c_i-b) _i=1^Nw_ik_i^2(s) ]. (4) The fit associated with each candidate source is evaluated using the weighted root-mean-square error, WRMSE()=∑i=1Nwi[ci−b−Q∗()ki()]2∑i=1Nwi,WRMSE(s)= _i=1^Nw_i [c_i-b-Q^*(s)k_i(s) ]^2 _i=1^Nw_i, (5) and the final source location is selected as ∗=argminWRMSE().s^*= _sWRMSE(s). (6) Source localization is conducted using a two-stage inversion procedure. First, a coarse spatial grid search is performed around the mobile measurement path to identify candidate source regions. At each grid point, the unit-emission plume response is calculated, the corresponding optimal emission rate is estimated, and the weighted residual error is recorded. Second, the most promising candidate regions are refined using numerical optimization, allowing the source location and emission rate to vary continuously beyond the initial grid resolution. This procedure reduces computational requirements while retaining continuous source-location estimates. To characterize localization uncertainty, the spatial error surface is converted into a relative source-likelihood distribution. Candidate locations with lower weighted residual errors are assigned greater relative likelihood, and the normalized distribution is used to determine the reported source region and confidence score. These quantities describe model-based spatial uncertainty under the assumptions of the Gaussian plume formulation. The resulting output includes the predicted source coordinate, model-based emission-rate estimate, estimated background concentration, weighted error metrics, source-likelihood information, and plume visualization. For surveys containing multiple spatially separated methane enhancements, the analysis routine first calculates the methane enhancement at each GPS location relative to the estimated background concentration. Elevated observations are identified using an adaptive threshold based on the background level and the observed concentration distribution and are then clustered in projected spatial coordinates. If multiple clusters are separated by more than a predefined distance and each contains a sufficient number of elevated observations, the dataset is treated as a multiple-source case. Source localization is then performed independently for each cluster. The final output includes an overview of the detected methane-enhancement regions together with cluster-level source estimates, plume visualizations, and diagnostic results. If no robust spatially separated clusters are identified, the standard single-source localization workflow is used. The same analysis layer also supports plume reconstruction when the source coordinate is known but meteorological information is incomplete. In this case, the source location is fixed, and candidate wind directions, wind speeds, and atmospheric stability classes are evaluated. For each candidate meteorological configuration, the Gaussian plume response is calculated at the observed measurement locations, and the corresponding emission rate is estimated using weighted least squares. The configuration that minimizes the weighted RMSE between the modeled and observed methane concentrations is selected as the best-fitting plume condition. The inferred wind direction, wind speed, and atmospheric stability class are reported as model-derived fitting parameters rather than direct meteorological measurements. These parameters represent the conditions that best reproduce the observed methane distribution under the Gaussian plume assumptions and do not replace independent weather observations. 2.4 Decision Support and Report Generation Layer After the localization, emission estimation, and plume reconstruction workflows are completed, the final outputs are processed by the decision support and report generation layer. This layer is designed to evaluate the reliability of the analytical results, recommend follow-up sampling when needed, and convert intermediate outputs into structured decision support reports. Instead of generating isolated figures or text summaries, the layer integrates historical user inputs, task specifications, processed sensing data, plume prediction results, validity information, generated figures, and previous interaction context to assemble a complete report according to the user’s request. A key component of this layer is the Validity Agent, which evaluates whether the localization or plume reconstruction result is sufficiently constrained by the available data. The validity score is computed from several diagnostic factors, including the RMSE relative to methane enhancement, GPS spatial spread, number of elevated methane points, distance from the predicted source to the sampled route, emission rate magnitude, and source likelihood confidence. The output is a high, medium, or low validity label with explanatory flags. Low validity is assigned when the model residual is high, GPS coverage is too compact, too few elevated points are available, or the predicted source lies far from the sampled route. This validity assessment does not recompute the leak location; instead, it provides an interpretable reliability check for the existing analytical result. Based on the validity assessment, the framework provides follow-up sampling recommendations to support adaptive field decision making. When the validity is high, the system may recommend an optional confirmation transect to verify the predicted plume structure. When the validity is medium or low, additional sampling is recommended before a final field decision is made. The default recommendation is a cross plume loop around the candidate source, including an upwind background point, two cross plume edge points, and a downwind centerline point. For multi-source cases, the system recommends local verification routes for each detected cluster rather than one global route. In this way, the recommendation module supports practical field planning without claiming fully autonomous route optimization. The Report Generator Agent then converts the structured results and generated figures into an AIMNet Methane Report. The report can include Geographic Information System (GIS)-based monitoring maps, methane plume visualizations, estimated source locations, emission rate predictions, uncertainty information, validity labels, follow-up sampling suggestions, conclusions, and recommended field actions. Recent results are deduplicated by data source and location so that repeated runs do not dominate the report. Multi-source results are summarized as a single multi-source case, while cluster level outputs are included as supporting evidence when needed. Each figure is paired with a numerical summary and a short visual interpretation. The Report Generator Agent also supports context-aware figure retrieval and layout refinement. Previously generated plume maps, GIS plots, route planning figures, and analysis tables can be recalled and incorporated into the final report. If the user is not satisfied with the font size, label placement, figure arrangement, caption style, or other formatting details, the feedback can be returned to the agent for iterative revision. User preferred report formats, such as section order, figure style, and conclusion structure, can also be saved and reused in subsequent monitoring tasks to ensure consistent formatting and reduce repetitive manual editing. When a vision language model is available, it is used only to describe visual context in the generated figures, such as plume footprint, nearby roads, buildings, vegetation, or open land. Numerical methane conclusions, including source location, emission rate, uncertainty, and validity level, are derived from the structured analysis results rather than from the vision model. This design prevents visual interpretation from overriding quantitative plume inversion results while still allowing the report to provide useful spatial context. In addition to visualization, the agent summarizes the generated data and interprets the analytical results to produce an overall conclusion and recommendation section. This text component remains adjustable by the user, allowing the report to combine automated generation with user directed refinement. 3 Applications and Discussion To evaluate the practical performance of the proposed framework, the developed functions were tested in real-world methane monitoring applications. The application section examines how the system supports key tasks in practice, including planning, data retrieval, methane data analysis, leak localization, emission rate estimation, and report generation. Through field-oriented use cases, the study assesses whether these integrated functions can effectively assist users under realistic operational conditions. Figure 4: Route-planning outputs generated by the AIMNet agent for three methane monitoring scenarios Figure 4 presents representative route planning results generated by the proposed framework for methane monitoring under different site conditions. Three field locations were selected for the route planning experiments: a wastewater treatment facility in Norman, Oklahoma, a landfill in Oklahoma City, Oklahoma, and an oil and gas facility in El Reno, Oklahoma. The experiments were conducted at different times to examine the applicability of the proposed framework under different environmental conditions and site layout characteristics. The field measurements followed our previously established methane monitoring setup, consisting of a vehicle-based platform equipped with the LI-COR LI-7810, LI-COR LI-7700, and AIMNet sensing device to collect methane concentration data along the routes generated by the proposed framework.20 The purpose of this setup was to evaluate whether the planned routes could guide the monitoring platform to areas where methane emission signals from potential leak points could be captured. The top row shows three example planning scenarios over satellite basemaps. In each case, the system identifies candidate monitoring regions and generates directional search sectors associated with different methane enhancement levels. The color bar, ranging from 0 to 12 ppm, represents the relative methane concentration level considered during the planning process. The highlighted polygons indicate suggested monitoring coverage regions, while the directional sectors represent candidate downwind search orientations for field deployment. The magenta dashed rectangles denote selected target areas used to constrain local route generation and focus subsequent measurements. The green regions represent the currently estimated potential emission areas, whereas the blue regions indicate future or extended search areas predicted by the planning module. The bottom row illustrates the corresponding route planning outputs for field execution. The blue lines indicate planned driving or traversal paths, while the colored markers represent measurement points or sampled methane observations collected along the route. These examples demonstrate that the proposed framework can adapt route design to different site geometries, including compact industrial facilities, linear roadside corridors, and open field or undeveloped areas. The enlarged insets further show how the planned routes can be refined for local inspection around suspected source regions. Figure 5: Reverse Prediction of Leak Points from Field Measurement Results at a Wastewater Facility Overall, the figure demonstrates that the route planning module integrates site layout, target region selection, wind related search orientation, and methane related spatial cues to generate operationally feasible monitoring paths for field measurements. In the first and third field tests, the major methane enhancement peaks, shown as red points, were located within the overlapping regions of the current and predicted potential emission areas. In the second field test, although part of the observed methane enhancement region was not fully covered by the planned route, the framework still identified the primary emission prone area. This discrepancy may be attributed to local turbulence, temporary wind direction changes, or other short-term meteorological variations during field deployment. To evaluate the leak source prediction capability of the proposed agent system, a vehicle-based methane survey was conducted around the wastewater facility. Figure 5 presents the automatically generated result from an input CSV file containing raw sensor measurements and GPS data. The proposed agent system removes abnormal signals, preprocesses the methane measurements, applies the machine learning model, and generates a GIS-based concentration map. Blue points indicate methane readings near the processed baseline, whereas measurements exceeding the background concentration by at least 1 ppm are selected for reverse prediction of the potential leak area. The resulting 80% confidence region successfully covered the actual emission points, which were later confirmed through follow-up field detection. Figure 6: Gaussian plume reconstruction under different meteorological conditions using manual and automated input workflows To validate the emission data analysis and prediction function, the proposed framework was tested using data collected from a real leak monitoring mission. During a field task conducted in collaboration with the Caddo Nation in Oklahoma, methane measurements were manually collected using the LI-COR LI-7810 to investigate abandoned wells with potentially high leakage risk. The selected test location had already been identified as a high methane emission point based on prior field observations, making it a suitable case for evaluating whether the generated plume report was consistent with real field measurements. Figure 6 presents the quick report generated by the agent. The left panel shows the workflow in which sensor parameters and weather information were manually entered by the user, whereas the right panel shows the automated workflow in which the system extracts sensor readings from the CSV file based on the timestamps and the provided GPS coordinates. In the automated workflow, the system also retrieves the corresponding weather conditions and estimates the atmospheric stability class for plume prediction. In addition to the input measurement location, multiple manual measurements were collected at other positions within and around the predicted Gaussian plume region for comparison. The observed methane concentration patterns from these manual field measurements were consistent with the plume distribution generated by the agent, indicating that the predicted plume range corresponded well with the area where methane enhancements were detected in the field. Both workflows organize the sensor readings, wind direction, wind speed, sampling labels, measurement positions, and predicted plume distribution into a standardized report format. To evaluate the repeatability and accuracy of the overall AIMNet agent system, multiple tests were conducted across different facilities under three methane monitoring scenarios, including wastewater facilities, landfills, and two oil and gas facilities. These tests were performed under different time and weather conditions to evaluate the robustness of the system. The evaluation was divided into several workflow level and output quality tasks to assess whether the agent could correctly route user requests, interpret input data, select appropriate environmental parameters, perform methane analysis, and generate usable reports. As shown in Table 1, the proposed agent demonstrated reliable decision making capability across different workflow and data interpretation tasks. For workflow routing and parameter extraction, the agent achieved a success rate of 92.0% over 50 prompts, indicating that most user instructions were correctly mapped to the corresponding workflow and that the required parameters were successfully extracted. The remaining errors mainly occurred when a CSV file was uploaded without sufficient user instructions. In such cases, the agent occasionally selected an inappropriate workflow and generated an incorrect graph type. The methane column and unit selection task achieved a success rate of 90.0%, showing that the agent was generally able to identify the correct methane concentration column and handle ppm/ppb unit conversion. The failed case occurred when the uploaded dataset contained a short data sequence, unusually high concentration values, and missing unit metadata. Under this condition, the agent misinterpreted ppm level values as ppb level measurements. Table 1: Agent decision accuracy across workflow and data interpretation tasks. Agent Decision Task Test Cases Success Criterion Success Rate Failure Description Workflow routing and parameter extraction 50 prompts Correct workflow selected and all required parameters extracted 92.0% Errors mainly occurred when prompts contained incomplete field descriptions for graphing. Methane column and unit selection 10 CSV files Correct methane column selected and ppm/ppb unit handled correctly 90.0% One case failed due to unclear column naming and missing unit for high methane concentration Weather source decision 10 cases Correctly selected user provided wind, weather API, inferred wind, or fallback 100% None Single source vs multi-source decision 8 CSV files Correctly classified the survey as single source or multi-source 87.5% Errors were mainly caused by overlapping methane enhancement regions from nearby potential sources. Validity decision 10 cases Validity label matched the actual error category 80.0% Misclassification occurred when multiple error types appeared simultaneously in the same input case. The weather source decision task reached 100%, suggesting that the agent correctly selected among user provided wind information, weather API retrieval, inferred wind settings, and fallback configuration. The single source versus multi-source decision task achieved 87.5%. The remaining errors were mainly caused by overlapping methane enhancement regions from nearby potential sources, which made it difficult to distinguish whether the observed methane plume was produced by one source or multiple sources. The validity decision task achieved a success rate of 80.0%, which was the lowest score among the workflow level decision tasks. This lower performance was mainly due to cases in which multiple error types appeared simultaneously in the same input file or user request. Under these conditions, the agent occasionally assigned the wrong validity label because the input contained more than one possible failure category. Table 2: Methane analysis performance of AIMNet after workflow execution. Agent Decision Task Test Cases Success Criterion Success Rate Failure Description Missing wind plume reconstruction 30 cases Correctly used the weather API when time information was available, or applied the default wind setting when detailed time information was missing. 100% None Potential measurement area prediction 10 cases Correctly cover the potential methane detection area 90% Only occurs when wind and environmental conditions are unstable Leak source localization 15 cases Correctly localize the Highest reading point, and Predicted source error <5<5 m 73.3% Errors mainly occur under unstable or weak wind conditions, which reduce the reliability of plume-based source localization. Emission Rate and Emission plume prediction 20 cases Leak rate estimates and plume visualizations were physically meaningful and consistent with manual calculation. 85% Incorrect stability class was selected for calculation-based on the weather conditions. Report Generation 20 cases Generated an editable report with meaningful content and structure aligned with the user’s request 95% Failures mainly occurred when the revised report contained corrupted text or misplaced elements Table 2 further evaluates the methane analysis performance after workflow execution. The Supporting Information includes representative examples that illustrate the expected output and success criterion for each task. The missing-wind plume reconstruction task achieved 100%, demonstrating that the agent could correctly use the weather API when time information was available or apply a default wind setting when detailed timing information was missing. Potential measurement area prediction achieved 90%, indicating that the agent could generally identify the likely methane detection area. The incorrect predictions mainly occurred under unstable environmental conditions, such as rainy weather or high temperature sunny noon periods, where wind direction and atmospheric conditions were less stable. Leak source localization achieved 73.3%, which was the lowest score among the analysis tasks. This result is reasonable because source localization is highly sensitive to wind stability, weak wind conditions, environmental uncertainty, and the spatial coverage of the measurement points. In particular, unstable or low wind conditions can reduce plume consistency, and insufficient test points may further limit the reliability of source localization. The emission rate and plume prediction task achieved 85%, showing that the generated leak rate estimates and plume visualizations were generally physically meaningful and consistent with manual calculation. The remaining errors mainly occurred when the agent selected an incorrect atmospheric stability class, which resulted in less meaningful plume calculation results in a small number of cases. Report generation achieved a success rate of 95%, indicating that the agent could produce editable reports aligned with the user’s request in most cases. The remaining cases produced report files with localized formatting or readability issues, such as corrupted text segments, partially unreadable content, or misplaced elements during revision. Representative examples are provided in the Supporting Information. Figure 7 provides a detailed evaluation of the core methane analysis functions, including source localization, emission rate estimation, and plume prediction. This imbalance mainly reflects differences in monitoring priorities, facility availability, and field campaign logistics. Oil and gas facilities were sampled more frequently because they are a major target for methane detection and are often more numerous and spatially clustered, allowing multiple cases to be collected during one field campaign. In comparison, accessible landfill and wastewater facilities were more limited and subject to site permission, scheduling, and suitable meteorological conditions. Therefore, the landfill and wastewater results should be interpreted as representative case studies rather than comprehensive facility-type evaluations. The results show that the landfill, the oil, and the gas facility cases achieved better overall performance. In the oil and gas facility cases, only one predicted source was outside the expected prediction range, while the mean and median source errors remained relatively low and within an acceptable range for field detection. Figure 7: Quantitative validation of leak source localization, emission rate estimation, and plume reconstruction across field scenarios. In contrast, the wastewater facility cases showed lower localization performance, with a success rate of only 50%. Two cases produced much larger source errors, which increased both the mean and median source error values. This reduced performance was mainly caused by the complex built environment around the wastewater facility, where nearby buildings, facility structures, and other obstacles can disturb local wind fields and affect methane plume transport. 45 As a result, the Gaussian plume-based localization method is more reliable in open or less obstructed environments, such as landfill and some oil and gas facility scenarios, but it can become less accurate in areas with multiple buildings, trees, or other physical obstructions. These results indicate that the proposed agent is effective for predicting methane emission sources when the plume transport is not strongly blocked or distorted by surrounding structures. However, in complex environments with dense buildings or vegetation, human review and additional field measurements are still needed to verify the predicted source location. Similar challenges may also occur in future oil and gas field applications if the monitoring area contains trees, equipment clusters, or other obstacles. This limitation could be addressed by integrating more advanced gas dispersion models, such as Gaussian puff models, Lagrangian particle dispersion models, or CFD assisted wind field models, to improve plume prediction and source localization under complex field conditions. Figure 8: External LLM-based and expert evaluation of the locally deployed AIMNet agent across different report generation tasks The external evaluation results in Figure 8 show that the generated reports received consistently high scores from Gemini Judge, ChatGPT 5.5 Judge, and an expert judge. The average scores for language quality, visualization, and report format were 9.3, 9.2, and 8.4, respectively, resulting in an overall average score of 9.0. Among the three criteria, language quality and visualization received the highest scores, suggesting that the agent could generate clear technical descriptions and meaningful figures for methane analysis reports. The report format score was slightly lower, mainly because report layout and formatting are more sensitive to figure placement, table alignment, and revision consistency. The expert judge provided the highest overall score of 9.3, which further supports the usability of the generated reports for practical methane monitoring analysis. Figure 9 compares the estimated time required by the manual workflow, general purpose LLMs (ChatGPT), and the AIMNet agent. In the manual workflow, the task was assumed to be completed by a two to three person expert team through discussion, manual data interpretation, calculation, visualization, and report preparation. For the general purpose LLM workflow, an expert user was still required to provide domain knowledge, interpret the model outputs, and manually refine the analysis results. In contrast, the AIMNet agent was operated by a student user with limited prior experience in methane emission analysis and visualization, demonstrating the potential of the agent to reduce the required level of domain expertise. Figure 9: Estimated time comparison among the manual workflow, general purpose LLMs (ChatGPT), and the AIMNet agent For field test planning, the manual workflow required approximately 30 minutes, while general purpose LLMs required 5 to 10 minutes and the AIMNet agent completed the task in less than 1 minute. For data collection and processing, the manual workflow required 30 minutes to 2 hours, compared with 30 minutes to 1 hour for general purpose LLMs and less than 5 minutes for the AIMNet agent. Similar improvements were observed in emission rate estimation, plume prediction, visualization, and report generation. In particular, visualization and report generation were reduced from 1 to 2 hours in the manual workflow to less than 15 minutes using the AIMNet agent. These results indicate that the proposed agent substantially reduces the time required for methane monitoring workflows while maintaining meaningful analytical and reporting quality. More importantly, the AIMNet agent reduces the dependence on expert level manual operation by integrating workflow routing, data processing, plume-based calculation, visualization, and report generation into a single automated framework. 4 Conclusion The central finding of this study is that LLMs can provide practical value in methane monitoring when they are used to coordinate, rather than replace, deterministic environmental analysis. By translating natural-language requests into structured tasks and connecting sensor processing, meteorological data, plume inversion, visualization, and reporting tools, the proposed framework reduced the manual effort required to move from field measurements to interpretable results. The field evaluations further showed that this tool-grounded design can maintain traceable quantitative outputs while making complex methane analysis more accessible to users with limited prior experience. The main contribution is therefore not the use of an LLM to calculate methane emissions, but the development of a constrained coordination layer that connects otherwise fragmented monitoring procedures. The reliability of the resulting analysis remained dependent on the quality of the measurements and the validity of the underlying dispersion model. Source localization performed better in open or less obstructed environments, whereas lower accuracy at the wastewater facility reflected the limitations of steady-state Gaussian plume inversion in built environments. Buildings, vegetation, facility structures, unstable winds, and insufficient spatial coverage can distort plume transport and reduce the ability of the model to constrain a unique source. Results produced under such conditions should therefore be interpreted as decision-support information and verified through additional field measurements rather than treated as autonomous final decisions. Future work will focus on improving the reliability of the framework under more complex monitoring conditions. The dispersion modeling and validity assessment will be further improved to help the system recognize when the available data are insufficient for reliable analysis. Additional self-checking mechanisms may also reduce errors during tool execution and result interpretation. Another important direction is to systematically evaluate how the choice of LLM affects overall agent performance. Future studies will compare different LLM families and model sizes to separate the contribution of the underlying model from that of the agent scaffolding and integrated tools. Domain-specific post-training may also improve tool usage and reduce the failure cases observed in this study. The framework will also be optimized for smaller and more portable computing platforms. This could support its integration with remotely operated aerial or ground-based sensing systems in hazardous or difficult-to-access environments. Beyond methane monitoring, the same architecture may be adapted to other gas-sensing applications, including our previously developed CO sensing platform 38, as well as H2S detection and other atmospheric monitoring applications. Data and Code Availability The code and data supporting this study are available from the corresponding author upon reasonable request for research and validation purposes. The source code and selected anonymized test datasets will be made publicly available upon publication. Some facility coordinates and site-identifying information are not publicly available due to site and partner confidentiality requirements. During the preparation of this manuscript, the author used ChatGPT (GPT-5.5, OpenAI, 2026 version) and Gemini (Google, 2026 version) for language editing, figure generation and refinement, and figure caption editing. The author carefully reviewed, verified, and edited all AI-assisted outputs and takes full responsibility for the content of this publication. References Achiam et al. (2023) J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774. Cited by: §1. Albertson et al. (2016) J. D. Albertson, T. Harvey, G. Foderaro, P. Zhu, X. Zhou, S. Ferrari, M. S. Amin, M. Modrak, H. Brantley, and E. D. Thoma A mobile sensing approach for regional surveillance of fugitive methane emissions in oil and gas production. Environmental science & technology 50 (5), p. 2487–2497. Cited by: §2.2. Alvarez et al. (2018) R. A. Alvarez, D. Zavala-Araiza, D. R. Lyon, D. T. Allen, Z. R. Barkley, A. R. Brandt, K. J. Davis, S. C. Herndon, D. J. Jacob, A. Karion, et al. Assessment of methane emissions from the us oil and gas supply chain. Science 361 (6398), p. 186–188. Cited by: §1. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §2. Ball et al. (2025) D. Ball, N. Eichenlaub, and A. Lashgari Performance evaluation of fixed-point continuous monitoring systems: influence of averaging time in complex emission environments. Sensors 25 (9), p. 2801. Cited by: §1, §1. Bell et al. (2023) C. Bell, C. Ilonze, A. Duggan, and D. Zimmerle Performance of continuous emission monitoring solutions under a single-blind controlled testing protocol. Environmental science & technology 57 (14), p. 5794. Cited by: §1. Brandt et al. (2014) A. R. Brandt, G. Heath, E. Kort, F. O’Sullivan, G. Pétron, S. M. Jordaan, P. Tans, J. Wilcox, A. Gopstein, D. Arent, et al. Methane leaks from north american natural gas systems. Science 343 (6172), p. 733–735. Cited by: §1. Brown et al. (2020) T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language models are few-shot learners. Advances in neural information processing systems 33, p. 1877–1901. Cited by: §1. Chen et al. (2026) C. Chen, N. Li, J. Qi, H. Chang, W. Shi, J. Xie, J. Yuan, H. Yang, J. Guo, C. Xu, et al. Leveraging llms for environmental complexity: structured fine-tuning data sets and deployment strategies. Environmental Science & Technology 60 (1), p. 497–509. Cited by: §1. Chen et al. (2023) Q. Chen, C. Schissel, Y. Kimura, G. McGaughey, E. McDonald-Buller, and D. T. Allen Assessing detection efficiencies for continuous methane emission monitoring systems at oil and gas production sites. Environmental science & technology 57 (4), p. 1788–1796. Cited by: §1, §1. Cheng et al. (2026) F. Cheng, Q. Li, L. He, H. Li, B. W. Brooks, Z. Yu, and J. You Leveraging large language models for contextual prioritization of contaminants of emerging concern in chemical mixtures. Environmental Science & Technology 60 (15), p. 11380–11391. Cited by: §1. Coburn et al. (2018) S. Coburn, C. B. Alden, R. Wright, K. Cossel, E. Baumann, G. Truong, F. Giorgetta, C. Sweeney, N. R. Newbury, K. Prasad, et al. Regional trace-gas source attribution using a field-deployed dual frequency comb spectrometer. Optica 5 (4), p. 320–327. Cited by: §1, §1. Duan and Wang (2024) Z. Duan and J. Wang Exploration of llm multi-agent application implementation based on langgraph+ crewai. arXiv preprint arXiv:2411.18241. Cited by: §1. Duncan (2015) I. J. Duncan Does methane pose significant health and public safety hazards?—a review. Environmental Geosciences 22 (3), p. 85–96. External Links: Document Cited by: §1. Erland et al. (2022) B. M. Erland, A. K. Thorpe, and J. A. Gamon Recent advances toward transparent methane emissions monitoring: a review. Environmental Science & Technology 56 (23), p. 16567–16581. Cited by: §1, §1. Grattafiori et al. (2024) A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783. Cited by: §2. Guo et al. (2024) T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang Large language model based multi-agents: a survey of progress and challenges. arXiv preprint arXiv:2402.01680. Cited by: §1, §2. Han et al. (2024) S. Han, Q. Zhang, W. Jin, and Z. Xu LLM multi-agent systems: challenges and open problems. arXiv preprint arXiv:2402.03578. Cited by: §1. Hong et al. (2023) S. Hong, M. Zhuge, J. Chen, X. Zheng, Y. Cheng, J. Wang, C. Zhang, Z. Wang, S. K. S. Yau, Z. Lin, et al. MetaGPT: meta programming for a multi-agent collaborative framework. In The twelfth international conference on learning representations, Cited by: §2. Hu et al. (2026) X. Hu, Y. Ding, Y. Yan, C. Wang, B. Weng, S. Hardeman, and M. Xue Multi-sensor and multi-model investigation of methane plumes from a wastewater treatment plant to improve emission inversion during the morning boundary layer transition. In 106th AMS Annual Meeting, Cited by: §1, §3. Huang et al. (2025) H. Huang, Y. Li, B. Jiang, B. Jiang, L. Liu, Z. Liu, R. Sun, and S. Liang A middle path for on-premises llm deployment: preserving privacy without sacrificing model confidentiality. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, p. 8332–8370. Cited by: §2. IJzermans et al. (2024) R. IJzermans, M. Jones, D. Weidmann, B. van de Kerkhof, and D. Randell Long-term continuous monitoring of methane emissions at an oil and gas facility using a multi-open-path laser dispersion spectrometer. Scientific Reports 14 (1), p. 623. Cited by: §1. Jackson et al. (2024) R. Jackson, M. Saunois, A. Martinez, J. Canadell, X. Yu, M. Li, B. Poulter, P. Raymond, P. Regnier, P. Ciais, et al. Human activities now fuel two-thirds of global methane emissions. Environmental Research Letters 19 (10), p. 101002. Cited by: §1. Jacob et al. (2016) D. J. Jacob, A. J. Turner, J. D. Maasakkers, J. Sheng, K. Sun, X. Liu, K. Chance, I. Aben, J. McKeever, and C. Frankenberg Satellite observations of atmospheric methane and their value for quantifying methane emissions. Atmospheric Chemistry and Physics 16 (22), p. 14371–14396. Cited by: §1. Li et al. (2023) G. Li, H. Hammoud, H. Itani, D. Khizbullin, and B. Ghanem Camel: communicative agents for” mind” exploration of large language model society. Advances in neural information processing systems 36, p. 51991–52008. Cited by: §2. Li and Liang (2026) M. E. Li and S. Liang Natural language to DGGS-aware methane insights with a multi-LLM-agent framework. In Proceedings of the 9th Conference on Spatial Knowledge and Information (SKI Canada 2026), Banff, Alberta, Canada. Cited by: §1. Metzger et al. (2025) N. Metzger, A. Lashgari, U. Esmail, D. Ball, and N. Eichenlaub A framework for optimizing continuous methane monitoring system configuration for minimal blind time: application and insights from over 100 operational oil and gas facilities. ACS ES&T Air 2 (8), p. 1439–1453. Cited by: §1, §1. Minaee et al. (2024) S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao Large language models: a survey. arXiv preprint arXiv:2402.06196. Cited by: §1. Nisbet et al. (2020) E. Nisbet, R. Fisher, D. Lowry, J. France, G. Allen, S. Bakkaloglu, T. Broderick, M. Cain, M. Coleman, J. Fernandez, et al. Methane mitigation: methods to reduce emissions, on the path to the paris agreement. Reviews of Geophysics 58 (1), p. e2019RG000675. External Links: Document, Link Cited by: §1. Ollama (2026) Ollama Ollama: get up and running with large language models. Note: https://github.com/ollama/ollama Cited by: §2. Saunois et al. (2025) M. Saunois, A. Martinez, B. Poulter, Z. Zhang, P. A. Raymond, P. Regnier, J. G. Canadell, R. B. Jackson, P. K. Patra, P. Bousquet, et al. Global methane budget 2000–2020. Earth System Science Data 17 (5), p. 1873–1958. Cited by: §1. Schwietzke et al. (2017) S. Schwietzke, G. Pétron, S. Conley, C. Pickering, I. Mielke-Maday, E. J. Dlugokencky, P. P. Tans, T. Vaughn, C. Bell, D. Zimmerle, et al. Improved mechanistic understanding of natural gas methane emissions from spatially resolved aircraft measurements. Environmental Science & Technology 51 (12), p. 7286–7294. Cited by: §1. Sherwin et al. (2022) E. D. Sherwin, J. S. Rutherford, Y. Chen, S. Aminfard, E. A. Kort, R. B. Jackson, and A. R. Brandt Single-blind validation of space-based point-source methane emissions detection and quantification. Cited by: §1. Thorpe et al. (2023) A. K. Thorpe, R. O. Green, D. R. Thompson, P. G. Brodrick, J. W. Chapman, C. D. Elder, I. Irakulis-Loitxate, D. H. Cusworth, A. K. Ayasse, R. M. Duren, et al. Attribution of individual methane and carbon dioxide emission sources using emit observations from space. Science advances 9 (46), p. eadh2391. Cited by: §1. Vaswani et al. (2017) A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin Attention is all you need. Advances in neural information processing systems 30. Cited by: §1. Wang et al. (2022) J. L. Wang, W. S. Daniels, D. M. Hammerling, M. Harrison, K. Burmaster, F. C. George, and A. P. Ravikumar Multiscale methane measurements at oil and gas facilities reveal necessary frameworks for improved emissions accounting. Environmental science & technology 56 (20), p. 14743–14752. Cited by: §1. Weller et al. (2018) Z. D. Weller, J. R. Roscioli, W. C. Daube, B. K. Lamb, T. W. Ferrara, P. E. Brewer, and J. C. von Fischer Vehicle-based methane surveys for finding natural gas leaks and estimating their size: validation and uncertainty. Environmental science & technology 52 (20), p. 11922–11930. Cited by: §1. [38] B. Weng, B. Mijiddorj, T. Beringer, L. Mijiddorj, Y. Yang, A. Ho, and E. Hassan An environment-adaptive low-power iot architecture for portable smoking detection with real-time data quality assurance. Cited by: §4. Weng (2024) B. Weng The road to climate change mitigation via methane emissions monitoring. Nature Reviews Electrical Engineering 1 (2), p. 69–70. Cited by: §1. Wu et al. (2024) Q. Wu, G. Bansal, J. Zhang, Y. Wu, B. Li, E. Zhu, L. Jiang, X. Zhang, S. Zhang, J. Liu, et al. Autogen: enabling next-gen llm applications via multi-agent conversations. In First conference on language modeling, Cited by: §1, §2. Yan et al. (2025) Y. Yan, L. Mijiddorj, T. Beringer, B. Mijiddorj, A. Ho, and B. Weng Machine learning-enhanced ndir methane sensing solution for robust outdoor continuous monitoring applications. Sensors 25 (24), p. 7691. Cited by: §1. Yan (2025) Y. Yan Development of a machine-learning enhanced high performance methane sensing instrument for field applications. Master’s Thesis, University of Oklahoma–Graduate College. Cited by: §1. Yang et al. (2025) A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Cited by: §2. Yang and Ravikumar (2025) S. L. Yang and A. P. Ravikumar Assessing the performance of point sensor continuous monitoring systems at midstream natural gas compressor stations. ACS ES&T Air 2 (4), p. 466–475. Cited by: §1. Yuvaraj et al. (2026) R. Yuvaraj, T. Lauvaux, C. Abdallah, P. Ciais, J. Akani Guery, J. Bonne, A. Groshenry, N. M. Hoang, and L. Joly High-resolution modeling of methane plumes: validation and sensitivity experiments to explore emission quantification approaches. Environmental Science & Technology 60 (5), p. 4029–4041. Cited by: §3. Zhai et al. (2026) M. Zhai, Q. Zeng, R. Qiu, J. Li, Q. Zhu, T. D. Waite, B. Ni, and H. Duan WaterRAG: a multiagent retrieval-augmented generation framework to support water industry transitions to net-zero. Environmental Science & Technology 60 (15), p. 11529–11541. External Links: Document Cited by: §1. Zhou et al. (2021) X. Zhou, X. Peng, A. Montazeri, L. E. McHale, S. Gaßner, D. R. Lyon, A. P. Yalin, and J. D. Albertson Mobile measurement system for the rapid and cost-effective surveillance of methane and volatile organic compound emissions from oil and gas production sites. Environmental science & technology 55 (1), p. 581–592. Cited by: §1. Zhou et al. (2026) Z. Zhou, X. Wang, Y. Yan, L. Mijiddorj, Y. Ding, T. Beringer, P. M. Khiabani, W. G. Jentner, X. Hu, C. Wang, et al. AIMNET: an iot-empowered digital twin for continuous gas emission monitoring and early hazard detection. IEEE Internet of Things Magazine. Cited by: §1, §2.1. Supporting Information A Locally Deployable Tool-Grounded LLM Multi-agent Framework for Automating Methane Emission Analysis and Reporting Yang Yan,† Zifan Zhou,‡ Xuan Wang,‡ Erum Hassan,† Bilguunzaya Mijiddorj,† Jie Cao,¶ Bin Li,‡ and Binbin Weng∗,† †School of Electrical and Computer Engineering, University of Oklahoma, Norman, OK, 73019 USA ‡Department of Electrical Engineering, Pennsylvania State University, University Park, PA, 16801 USA ¶School of Computer Science, University of Oklahoma, Norman, OK, 73019 USA E-mail: binbinweng@ou.edu S1.Scoring Rubric The LLM-agent outputs were evaluated using a 0–10 rubric. The score was assigned based on task completion, scientific validity, data handling, visualization and reporting quality, and decision-support usefulness. Table S1: General scoring rubric for evaluating LLM-agent outputs. Score range Description 9–10 Excellent: accurately presents methane data, meteorological inputs, source location, model-derived emission rate, plume results, uncertainty, and validity information. Figures and conclusions are consistent and require only minor edits. 7–8 Good: main analytical results are correct and usable, with only minor omissions or unclear details in weather sources, uncertainty, validity, figures, or recommendations. 5–6 Fair: major outputs are identifiable, but important information is incomplete, inconsistent, or requires expert revision before use. 3–4 Poor: major omissions or errors in methane units, meteorological inputs, source localization, plume results, figures, or conclusions substantially limit usability. 1–2 Very poor: mostly incorrect or incomplete, with unsupported conclusions, physically inconsistent results, or unreadable report content. 0 Failed: no valid report, failed execution, unreadable output, or content unrelated to the methane-monitoring task. S2. Example Prompts Fifty representative prompts were used to evaluate task interpretation, parameter extraction, workflow selection, and tool coordination. The prompts covered field planning, data processing, source and plume analysis, emission estimation, and report generation. Actual coordinates were replaced with “[latitude, longitude]” to protect facility locations. Table S2: Example prompts for evaluating methane field-test route planning No. Example prompt 1 I want to conduct a methane field test around [latitude, longitude] tomorrow morning. What is the best time to start? 2 I want to conduct a methane field test around [latitude, longitude] this week. Which day and time would be most suitable? 3 I plan to test methane emissions around [latitude, longitude] on July 5. What would be the best time of day? 4 I want to conduct a field test around [latitude, longitude] this afternoon. Are the expected weather and wind conditions suitable? 5 I want to conduct a methane survey around [latitude, longitude] next week. Recommend the best date and time based on the weather forecast. 6 I want to conduct a methane field test around [latitude, longitude]. What route should I follow? 7 I want to test methane emissions around the landfill sometime this week. Which monitoring route would be most suitable based on the expected wind direction? 8 I want to conduct a methane field test around the oil and gas facility on July 5. Which route should I take? 9 I plan to conduct a mobile methane survey around [latitude, longitude] tomorrow morning. Generate a route that covers the most likely downwind detection areas. 10 I want to conduct a methane field test around [latitude, longitude] sometime this week. Recommend the best day and time and generate an appropriate monitoring route. Table S3: Example user prompts for methane-data retrieval, preprocessing, and input validation No. Example prompt 11 Plot CH4, humidity, and temperature from the uploaded CSV file on the same time axis. 12 Plot CH4, CO2, and H2O measurements from the uploaded CSV file for comparison. 13 Filter out abnormal methane measurements from the uploaded dataset and plot the cleaned data. 14 Process the uploaded raw AIMNet sensor data and plot the methane measurements before and after preprocessing. 15 Collect ten minutes of data from AIMNet devices 001, 002, and 003 and plot their methane measurements on the same figure. 16 Compare the methane measurements from AIMNet devices 001, 002, and 003. Identify which device reports the highest values and whether any device shows abnormal measurements. 17 The LI-COR LI-7810 measurements are 56 seconds behind the AIMNet measurements. Correct the time offset and plot the aligned data. 18 The LI-COR LI-7810 measurements are 56 seconds behind, and the LI-COR LI-7700 measurements are 21 seconds behind. Correct both time offsets and plot the aligned LI-7810, LI-7700, and AIMNet device measurements on the same figure. 19 Check the uploaded AIMNet dataset for missing values, duplicate timestamps, invalid GPS records, and abnormal methane measurements. 20 Compare the processed AIMNet methane measurements with the corresponding LI-COR measurements and summarize the differences among the sensors. Table S4: Example prompts for known-source plume reconstruction and emission-rate estimation. No. Example prompt 21 The methane source is located at [latitude, longitude]. The wind speed was 6 mph and the wind direction was 150 degrees. Methane enhancement was 1 ppm at 5 m and 200 ppb at 10 m. Reconstruct the plume. 22 I tested a methane source at [latitude, longitude]. The wind speed was 3 m/s and the wind direction was 220 degrees. Methane enhancement was 1.2 ppm at 5 m, 500 ppb at 10 m, and 150 ppb at 20 m. Estimate the emission rate and plot the plume. 23 I tested a methane source at [latitude, longitude] on July 5, 2026, at 10:00 a.m. Methane enhancement was 800 ppb at 8 m and 300 ppb at 15 m. Retrieve the weather conditions and reconstruct the plume. 24 I tested a methane source at [latitude, longitude] on January 20, 2026, at 2:00 p.m. Methane concentrations were 3.0 ppm at 5 m and 2.3 ppm at 10 m, with a background of 2.0 ppm. Retrieve the weather conditions and plot the plume. 25 The methane source is located at [latitude, longitude]. The wind speed was 3.1 m/s and the wind direction was 295 degrees. Methane enhancement was 2400 ppb at 6 m and 700 ppb at 12 m. Reconstruct the plume. 26 I tested a methane source at [latitude, longitude]. The wind speed was 2.5 m/s and the wind direction was 45 degrees. Methane concentrations were 3.4 ppm at 4 m and 2.5 ppm at 9 m, with a background of 2.1 ppm. Estimate the emission rate. 27 I tested a methane source at [latitude, longitude] on August 3, 2026, at 9:00 a.m. Methane enhancement was 1.5 ppm at 5 m, 600 ppb at 10 m, and 200 ppb at 20 m. Retrieve the weather conditions and reconstruct the plume. 28 I uploaded methane and GPS measurements collected around a known source at [latitude, longitude]. Use the measurement timestamps to retrieve the weather conditions and generate a plume map. 29 The source is located at [latitude, longitude]. The wind speed was 4 m/s and the wind direction was 180 degrees. Methane enhancement was 900 ppb at 5 m, 450 ppb at 10 m, and 120 ppb at 20 m. Plot the plume and estimate the emission rate. 30 I tested a methane source at [latitude, longitude] on June 15, 2026, at 11:00 a.m. Methane concentrations were 2.9 ppm at 5 m, 2.4 ppm at 10 m, and 2.1 ppm at 20 m, with a background of 2.0 ppm. Retrieve the weather conditions and reconstruct the plume. Table S5: Example user prompts for inverse methane-source localization and emission analysis No. Example prompt 31 I conducted a vehicle-based methane survey. Plot the methane measurements along the survey route, identify the suspected source location, and generate a plume map. 32 I uploaded a manual-survey CSV file containing methane concentrations and GPS coordinates. Estimate the most likely source location and reconstruct the plume. 33 I conducted a vehicle survey around [latitude, longitude] on July 5, 2026, at 10:00 a.m. Retrieve the corresponding weather conditions and identify the most likely methane source. 34 I uploaded vehicle-survey data with a wind speed of 3 m/s and a wind direction of 220 degrees. Estimate the source location, emission rate, and plume distribution. 35 Compare the possible source locations around the survey route and identify the location that best matches the observed methane measurements. 36 I uploaded a vehicle-survey CSV file containing methane concentrations, GPS coordinates, and timestamps. Identify the suspected source location, estimate the emission rate, and generate a plume map. 37 Estimate the most likely methane source coordinate and show the model-based source-likelihood region around the predicted location. 38 Determine whether the uploaded methane survey represents a single source or multiple sources. If multiple sources are detected, estimate each source location separately. 39 The uploaded survey contains two spatially separated methane-enhancement regions. Localize the suspected source associated with each region and generate separate plume maps. 40 Analyze the uploaded vehicle-survey data, estimate the source location and emission rate, assess the validity of the localization result, and recommend additional measurements if needed. Table S6: Example user prompts for validity assessment, report generation, and report revision No. Example prompt 41 Generate a methane-monitoring report for the monitoring point at [latitude, longitude]. 42 Generate a report summarizing the results from the last three monitoring points. 43 Generate a report for the field-test route planned on July 10. 44 Prepare a report for the most recent vehicle-based methane survey, including the concentration map, predicted leak location, plume visualization, and emission-rate estimate. 45 Generate a short field report for the manual methane test conducted at [latitude, longitude]. 46 Combine the last three source-localization results into one report and compare their predicted source locations and validity assessments. 47 Add the route plan generated on July 10 to the corresponding methane-monitoring report. 48 Increase the font size of the text, figure labels, and table content in the previous report to improve readability. 49 Change the title of the previous report to “Methane Monitoring and Emission Analysis Report” while keeping the remaining content unchanged. 50 Regenerate the previous report with a shorter conclusion and clear separation between measured results and model-derived predictions. 5 S2. Example of a Successful Methane Monitoring Report This figure shows an additional successful example of automated report generation. The framework extracted the location information and processed the uploaded CSV measurements. It generated the methane concentration plot and identified the major concentration peaks above the defined threshold. The framework also used multiple measurement points and wind information to generate a Gaussian plume map. The estimated methane release rate was 8.96 g/h. These results were then summarized in the final report with the corresponding figures and interpretations. This example demonstrates the successful execution of a multi-step methane monitoring task. Figure 10: Successful Example of Automated Methane Report Generation. S4. Failure Cases Representative failure cases were included to illustrate the main limitations observed during evaluation. These failures primarily involved misinterpretation of the user’s request, routing to an inappropriate downstream workflow, preprocessing errors such as incorrect methane-column identification, and plume-analysis errors when the retrieved wind speed and direction did not represent the transient local conditions during field measurement.Figures 11–13 present representative examples of these failure modes. Figure 11: Representative failure in automatic methane-column identification and visualization. Because the CSV file contained ambiguous column headers and multiple candidate columns with similar numerical distributions, the preprocessing routine selected the wrong methane field and generated an incorrect methane concentration plot. Figure 12: Representative failure caused by incomplete time-series data. Missing timestamps and data gaps in the original files caused the agent to incorrectly connect discontinuous measurements during visualization, resulting in a misleading methane concentration plot. Figure 13: Representative plume-reconstruction failure caused by nonrepresentative meteorological inputs. The wind speed and direction retrieved by the Weather Agent did not capture the transient local wind conditions during field measurement, resulting in an inaccurate Gaussian plume visualization and unreliable source-localization and emission-rate estimates.