Paper deep dive
Designing probabilistic AI monsoon forecasts to inform agricultural decision-making
Colin Aitken, Rajat Masiwal, Adam Marchakitus, Katherine Kowal, Mayank Gupta, Tyler Yang, Amir Jina, Pedram Hassanzadeh, William R. Boos, Michael Kremer
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 99%
Last extracted: 3/13/2026, 12:37:57 AM
Summary
The paper introduces a decision-theory framework for designing probabilistic monsoon onset forecasts tailored to heterogeneous smallholder farmers. By blending AI weather prediction models (NGCM, AIFS) with a novel 'evolving-expectations' Bayesian statistical model, the authors created a system that outperforms static climatology and individual models. This system was operationally deployed in 2025 to provide subseasonal forecasts to 38 million Indian farmers, demonstrating a scalable approach for climate adaptation.
Entities (6)
Relation Signals (3)
Evolving-expectations model â blendedwith â NGCM
confidence 100% ¡ We create a blended model that combines probabilistic information from the evolving-expectations model with rainfall forecasts from two AIWP models
Blended model â deployedby â India Ministry of Agriculture and Farmersâ Welfare
confidence 100% ¡ used in a 2025 program of Indiaâs Ministry of Agriculture and Farmersâ Welfare that distributed probabilistic local onset forecasts
Blended model â predicts â Indian Monsoon
confidence 100% ¡ The blended system yields more skillful Indian monsoon forecasts at longer lead times
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Hundreds of millions of farmers make high-stakes decisions under uncertainty about future weather. Forecasts can inform these decisions, but available choices and their risks and benefits vary between farmers. We introduce a decision-theory framework for designing useful forecasts in settings where the forecaster cannot prescribe optimal actions because farmers' circumstances are heterogeneous. We apply this framework to the case of seasonal onset of monsoon rains, a key date for planting decisions and agricultural investments in many tropical countries. We develop a system for tailoring forecasts to the requirements of this framework by blending systematically benchmarked artificial intelligence (AI) weather prediction models with a new "evolving farmer expectations" statistical model. This statistical model applies Bayesian inference to historical observations to predict time-varying probabilities of first-occurrence events throughout a season. The blended system yields more skillful Indian monsoon forecasts at longer lead times than its components or any multi-model average. In 2025, this system was deployed operationally in a government-led program that delivered subseasonal monsoon onset forecasts to 38 million Indian farmers, skillfully predicting that year's early-summer anomalous dry period. This decision-theory framework and blending system offer a pathway for developing climate adaptation tools for large vulnerable populations around the world.
Tags
Links
- Source: https://arxiv.org/abs/2603.07893v2
- Canonical: https://arxiv.org/abs/2603.07893v2
Trouble viewing inline? Open PDF directly â
Full Text
104,702 characters extracted from source content.
Expand or collapse full text
Designing probabilistic AI monsoon forecasts to inform agricultural decision-making Colin Aitken 1,* , Rajat Masiwal 2 , Adam Marchakitus 2 , Katherine Kowal 3 , Mayank Gupta 4 , Tyler Yang 5 , Amir Jina 7 , Pedram Hassanzadeh 2,3 , William R. Boos 5,8 , and Michael Kremer 1,6,* 1 Development Innovation Lab, University of Chicago, IL, 60637 2 Department of the Geophysical Sciences, University of Chicago, IL, 60637 3 Data Science Institute, University of Chicago, IL, 60637 4 Development Innovation Lab India, University of Chicago Trust, India, 560025 5 Department of Earth and Planetary Science, University of California, Berkeley, California, 94720 6 Kenneth C. Griffin Department of Economics, University of Chicago, IL, 60637 7 Harris School of Public Policy, University of Chicago, IL, 60637 8 Climate and Ecosystem Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, California, 94720 * Corresponding authors: caitken@uchicago.edu, kremermr@uchicago.edu ABSTRACT Hundreds of millions of farmers make high-stakes decisions under uncertainty about future weather 1, 2 . Forecasts can inform these decisions 3â7 , but available choices and their risks and benefits vary between farmers 8â11 . We introduce a decision-theory framework 12â14 for designing useful forecasts in settings where the forecaster cannot prescribe optimal actions because farmersâ circumstances are heterogeneous. We apply this framework to the case of seasonal onset of monsoon rains, a key date for planting decisions and agricultural investments in many tropical countries. We develop a system for tailoring forecasts to the requirements of this framework by blending systematically benchmarked 15 artificial intelligence (AI) weather prediction models with a new âevolving farmer expectationsâ statistical model. This statistical model applies Bayesian inference to historical observations to predict time-varying probabilities of first-occurrence events throughout a season. The blended system yields more skillful Indian monsoon forecasts at longer lead times than its components or any multi-model average. In 2025, this system was deployed operationally in a government-led program that delivered subseasonal monsoon onset forecasts to 38 million Indian farmers, skillfully predicting that yearâs early-summer anomalous dry period 16 . This decision-theory framework and blending system offer a pathway for developing climate adaptation tools for large vulnerable populations around the world. Introduction Farmers worldwide make high-stakes decisions about planting, crop choice, and input investments under considerable uncertainty about future weather 1, 17 . Randomized experiments find that exposure to weather forecasts meaningfully shapes these decisions 2, 5â7 . In tropical regions, for example, key decisions hinge on the timing of rainy-season onset: crops may fail if planted before dry spells transition to continuous rain, while expectations of rainy season length shape crop choice and investment intensity 18â20 . However, many smallholder farmers lack access to forecasts of key weather events, particularly at longer lead times 21â23 . Recent advances in artificial intelligence weather prediction (AIWP) offer an opportunity to close this gap, with AIWP models matching or exceeding the skill of state-of-the-art numerical weather prediction (NWP) models 24â30 at a fraction of the computational cost. Many AIWP models are also open-source and easy to use and tailor. If AIWP forecasts can be translated into comprehensible and actionable information, these advances could deliver substantial benefit to hundreds of millions of smallholder farmers. Realizing these benefits, howeverâwhether through AIWP or NWP or a blend of bothârequires confronting a fundamental design problem. Although weather forecasts can help farmers improve decisions 3, 4 , the appropriate agricultural response to a forecast may depend on information the forecaster does not observe, including individual farmersâ land characteristics, access to inputs, and experience with different crops and technology. These attributes can be heterogeneous across farmers 8â11, 17 . For example, a risk-averse farmer may require greater certainty before planting, even if waiting reduces expected yields; a farmer with access to irrigation or drought-tolerant seed varieties may be more willing to plant before a potential dry spell after initial rains. Similarly, a farm household with a member with outside employment may be more willing to take risks than one fully dependent on farm income. This may make it infeasible for central authorities to prescribe optimal agricultural decisions for arXiv:2603.07893v2 [cs.LG] 10 Mar 2026 each farmer. We address the question of how to design forecasts to empower farmers to reach their own decisions. The value of a forecast also depends on how it updates farmersâ expectations relative to what they would have believed without the forecast. Many forecasts are conventionally evaluated against a static âclimatologyâ baseline consisting of the historical median onset date at a location 31â33 . In a number of important practical cases, like rainy season onset, such a baseline can be highly unrealistic from the perspective of a farmer making agricultural decisions. Consider a year in which onset is delayed; when a static climatology forecast is examined after the historical median onset date but before the actual onset has occurred, that static forecast would predict that onset has likely already happened. In contrast, a farmer would be aware that continuous rains have not yet started and would expect onset to occur in the future. The static nature of the baseline means that using it in a benchmark will exaggerate a forecasting modelâs value to farmers, whose beliefs about likely onset dates evolve throughout the season. We address these problems here, leveraging decision theory to guide forecast design in settings where end users have varying objectives, constraints, and risk tolerances. We develop an âevolving-expectationsâ statistical model in which the distribution of future onset probabilities is dynamically updated throughout the season to reflect the fact that, as the season progresses, a farmer will observe for some period of time that onset has not yet occurred. We then combine the evolving-expectations model with AIWP models, selected using a decision-oriented benchmark 15 , to construct blended probabilistic statisticalâAI forecasts of local Indian monsoon onset. This blending not only combines information from multiple forecast models found to be individually skillful and operationally practical 15 , but dynamically weights the contributions of its components by lead time. This makes the blending approach more appropriate than a traditional multimodel ensemble for a setting in which AIWP or NWP model skill declines more quickly with lead time than that of the evolving-expectations model. This work builds on evidence that farmers can use information about local rainy season onset to adjust input and planting decisions 2 . Many of these decisions benefit from predictions with lead times longer than one week 3 , but publicly available onset forecasts either provide a single date for the entire country or are available only at shorter lead times. The India Meteorological Department (IMD) releases a skillful forecast with multiple weeks of lead time for monsoon onset over Kerala, in southern India 34, 35 , but this has little correlation with rainy season onset dates in the rest of the country 36 . The IMD also releases maps denoting the current extent of monsoon progression, as well as short-term forecasts of progression over the next 1-3 days. Farmers thus lack local predictions of monsoon onset at lead times that can help with longer-term planning or planting decisions. The framework and models presented here and in Masiwal et al. 15 were developed to fill this gap, and were used in a 2025 program of Indiaâs Ministry of Agriculture and Farmersâ Welfare that distributed probabilistic local onset forecasts weekly to 38 million farmers across 13 states 16 . Decision-Theory Framework Informing Forecast Design We use the tools of decision theory 13, 14 , starting with Blackwellâs Informativeness Theorem 12 , to develop a model to guide a forecaster choosing a weather prediction to send to a heterogeneous group of farmers. The farmers make agricultural decisions whose payoffs depend both on future weather and additional information known to each farmer but not to the forecaster. The forecaster obtains signals about future weather from prediction models. These assumptions allow derivation of three key considerations for forecast design and evaluation, summarized in Table 1. First, probabilistic forecasts allow heterogeneous farmers to tailor decisions to their constraints in ways that deterministic forecasts or advisories do not. We show that a well-calibrated probabilistic forecast (i.e., with probabilities aligning with empir- ical frequencies) leads to better outcomes than a deterministic message built from that forecast (Proposition 1, Supplementary Fig. S1) 14, 37 . For example, a farmer whose livelihood is at stake if crops fail may consider a 30% chance of the rainy seasonâs false onset too risky to plant, while planting early may pose less of a risk to a farmer with secure income outside of agriculture. A deterministic forecast that flattens the probability into â true onset is not expected in the next 2 weeks" will not meet both farmersâ informational needs. More generally, the decision-theory model implies that forecasters should avoid âcoarseningâ messages by removing potentially relevant information. In practice, while end users are capable of interpreting probabilistic information, comprehension varies with phrasing and context. 38â40 . Careful testing is therefore essential to arrive at appropriate message framing and text. Second, forecasts will more robustly benefit farmers if they incorporate information to which farmers already have access. Some farmers may perfectly incorporate new forecasts with their existing knowledge, and they will benefit as long as the forecasts contain new information. However, other farmers may take forecasts at face value, and they will only benefit if the forecast is better than their priors, i.e., information they would have used in the absence of forecasts. A corollary is that forecasts should be evaluated relative to a baseline containing relevant information already available to both forecasters and farmers. For forecasts of an eventâs timing, e.g., monsoon onset, this motivates the âevolving-expectationsâ statistical model presented in the next section. Baselines that do not represent farmer expectations can substantially overstate the value of new predictive models and lead to dissemination of information worse than what farmers already have. To guarantee farmers are better off in 2/22 expectation, forecasts need to fully incorporate farmersâ existing information (Proposition 2). In practice, incorporating every farmerâs private information about the weather would be infeasible, but forecasters can approximate it by blending relevant, publicly available information into forecast models. Third, the decision-theory model implies that as long as farmers correctly understand the forecast, it is possible to identify whether the forecast benefits farmers in expectation by testing whether farmers change their decisions. This motivates measuring forecast value by whether individuals change their decisions in response to information, rather than on whether researchers deem those decisions to be âcorrect.â When farmers make decisions to manage potential risks, an evaluation of forecast impacts on yields or profits in a single season cannot, in general, determine whether the forecast was ex-ante beneficial. This is because the evaluator will not observe the full distribution of potential weather events and other variables that affect these outcomes each season (Proposition 3). Indeed, risk-averse farmers may respond to accurate forecasts in ways that make them better off but reduce expected profits (Proposition 4). For example, a farmer may reasonably hedge against a potential dry event by purchasing drought-tolerant seeds with lower expected yields. Observing that yields were lower in response to forecasts will not allow the evaluator to determine whether forecasts helped the farmer manage their risk. A valid test of forecast relevance does not then depend on researchersâ knowledge of best practices for local farming or measures of impact on variables such as yield or farmer profits (Proposition 5). Changes in decisions can generally be measured more precisely than changes in outcomesâsuch as crop yield or income, which are affected by factors beyond farmersâ decisionsâand therefore are practical for testing forecast relevance when sample sizes are limited (Proposition 6). Model ResultImplication for Forecast Design and Evaluation Probabilistic forecasts better allow heterogeneous farmers to make their own decisions based on in- formation relevant to farming that the forecaster does not have (Proposition 1). ⢠Provided probabilistic forecasts. â˘Designed messages in part based on focus groups with farmers aimed at increasing comprehension of probabilistic information. If some farmers take forecasts at face value, guar- antees on benefits of forecasts require forecasts to be well-calibrated and incorporate farmersâ other sources of information about weather outcomes (Proposition 2). â˘Designed a new âevolving expectations" statistical model as both a baseline and a component of the blended model. ⢠Used the blended model, which met the criteria for dissemination, rather than AIWP model outputs on their own, which did not. ⢠Tested calibration of blended model across multiple time periods. Experiments measuring yield and profit outcomes do not capture the value of forecasts if farmers are risk-averse. Whether forecasts have value to farmers can be measured by observing whether farmers change decisions in response to forecasts (Propositions 3 - 6). â˘Identified monsoon onset as valuable target based on previous research finding that farmers change deci- sions after receiving monsoon onset forecasts 2 . â˘Chose an agriculturally-relevant monsoon onset defini- tion, capturing the start of continuous local rainfall not followed by a potentially crop-damaging dry spell, to specifically target farmer decisions. Table 1. Implications of decision-theory framework and operational decisions informed by the results, and actions we took as a result. An Evolving-Expectations Model The decision-theory framework establishes that forecasts should be evaluated relative to a baseline (i.e., reference forecast) that includes a model of farmersâ prior beliefs in the absence of forecasts. For monsoon onset, this means the baseline should include a simple but powerful piece of information: whether or not rains have already started locally. Following the frameworkâs implications for forecast design, we use an existing, well-studied agronomic definition of onset as the first wet spell of the season that is not soon followed by a prolonged dry period that could prove harmful to crops 36 (see Methods). We modify the conventional climatology predictor (i.e., median or mean historical onset date) by developing an âevolving-expectationsâ statistical model that applies Bayesian updating to the historical distribution of onset dates, calculated using IMD rain-gauge data, conditioning on whether onset has occurred in a given season. As the season progresses with onset unobserved, the model 3/22 shifts probability mass forward in time, concentrating it in the remaining plausible onset weeks (Fig. 1). This modified baseline has significant implications. For example, for a region around18 ⌠N (Fig. 1C), if onset has not occurred by June 26, the static (unconditional) climatology predicts the probability of onset to be9%in the following week, while the evolving-expectations model predicts a42%probability. The decision-theory framework implies that forecasts should not be disseminated if they cannot outperform this new baseline, which captures information about onset that is available to farmers (Table 1). We evaluate probabilistic forecasts of onset produced by this model using three complementary metrics: the Brier Score, which measures the accuracy of predicted probabilities for onset in specific weeks; the Ranked Probability Score (RPS), which additionally accounts for how close a predicted week is to the true onset date; and the area under the receiver operating characteristic (ROC) curve (AUC), which measures the ability to discriminate between likely and unlikely onset weeks irrespective of the modelâs calibration (see Methods). We primarily assess the 2000â2024 cross-validation period, but also include results from an additional 1965â1978 hold-out period, noting that initial conditions used in the AIWP models may be less accurate in that pre-satellite era 15 . Figure 2 shows the AUC and Brier Skill Score (BSS) for static climatology, the evolving expectations model, a calibrated hybrid AI model (NGCM; calibration is described below), and a model that blends two AI models and the evolving-expectations model (described in the next section). The evolving-expectations model outperforms static climatology at all individual weekly lead times (Fig. 2A, B) and in both validation periods (Extended Data Figs. 1, 2) for all metrics. It outperforms raw (i.e., uncalibrated) probabilistic forecasts from a hybrid AI model (NGCM) shown to be skillful on its own for monsoon onset predictions 15 (Fig. 4). Raw NGCM predictions are overconfident when comparing observed frequencies with predicted probabilities (Fig. 3A), but the evolving-expectations model even outperforms the calibrated NGCM for some metrics at longer lead times (Fig. 2B) and when all lead times out to four weeks are considered jointly (Extended Data Fig. 1). At these longer lead times, the AIWP models on their own therefore do not meet the second criterion outlined in the decision-theory framework for dissemination, which requires that forecasts improve on a farmerâs prior beliefs. This motivates a blended approach that combines both information sources. Blending AI and Statistical Forecasts We create a blended model (see Methods) that combines probabilistic information from the evolving-expectations model with rainfall forecasts from two AIWP models selected as part of a decision-relevant benchmarking 15 : Googleâs NGCM 41 and the European Centre for Medium-Range Weather Forecastsâ AIFS 42 . AIFS is a purely data-driven model, while NGCM is a hybrid model combining differentiable dynamical core with neural network parameterizations of physical processes. Intuitively, the blended model upweights AIWP-predicted wet spells when they are likely to correspond to true onset and downweights them on calendar dates when they are more likely to be a false onset (i.e., followed by a prolonged dry spell). The relative influence of each information source can vary with lead time: a modelâs prediction of rainfall in the coming week may receive greater weight than climatology, even as its prediction of dry conditions four weeks ahead receives less weight than climatology. The flexible weighting allows each information source to contribute most strongly where it is most informative, which is not achieved by traditional multimodel ensembles or simple averaging methods. The blended modelâs multinomial logistic regression specification also âcalibratesâ the probabilistic forecasts (see reliability diagrams in Fig. 3 and Extended Data Fig. 3). The blended model outperforms both the static (unconditional) climatology baseline and the evolving-expectations (condi- tional climatology) baseline across all three metrics and all evaluation periods, including the cross-validation model selection period (2000â2024; see next paragraph), the pre-satellite-era hold-out period (1965â1978), and the 2025 dissemination period (Fig. 2 and Extended Data Figs. 1-2). Note that the 2025 comparison uses operational forecasts produced for dissemination and publicly archived on creation (see Data Availability section). Cross-validation produces forecasts for each year by training the model on all other years during the period, ensuring that predictions are scored on held-out data. When pooled across lead times up to four weeks, the blended model yields a 5â10% improvement in Brier Score, a 20â25% improvement in RPS, and a 3â5 percentage point increase in AUC relative to the static climatology baseline. In 2025, when climatology performed poorly due to an unusual pause in the northward progression of monsoon onset 15 , the blended modelâs forecasts yield aâź 20%improvement in Brier Score and RPS as well as a20point increase in AUC relative to the static climatological baseline. The blended modelâs skill is highest at a one-week lead time, showing roughly a 15% improvement in Brier Score relative to the evolving-expectations baseline (25% relative to static climatology) during 2000â2024, and declines gradually but remains positive out to four weeks (Fig. 2). Skill patterns are qualitatively similar but smaller in magnitude in the 1965â1978 hold-out period (Extended Data Fig. 2), with improvements over the evolving-expectations model persisting to three-week lead times. Forecasts remain well calibrated in both the cross-validation and hold-out periods (Fig. 3 and Extended Data Fig. 3) and skill gains occur over a wide geographic region (Supplementary Figs. S2-S4). 4/22 The blended model does more than just combine and calibrate forecasts from different models. Forecasting the dry spell criterion in the agronomic onset definition requires a forecast that is 30 days longer than the lead time; e.g., 45-day and 60-day rainfall forecasts are needed for onset predictions with 15-day and 30-day lead times, respectively (Methods). Current AIWP models do not have such long-range skill for rainfall 15 and cannot be used to directly forecast the full agronomic onset, with its dry spell criterion. Blending with the evolving-expectations model addresses this problem by weighting AIWP rainfall forecasts according to the climatological probability of onset. Intuitively, an AIWP prediction of a five-day wet spell gets more weight when that prediction is more likely to reflect a true onset, and less weight when onset is unlikely and hence the predicted five-day wet spell might by followed by a dry period. In addition, since rainfall variables are included as continuous variables, the blended model implicitly performs simple bias correction, adjusting the five- and ten-day cumulative rainfall forecasts by a different constant for each week of lead time. The evolving expectations model is trained on 124 years of Indian rain-gauge data 43 , raising the question of whether it could be used in other monsoon regions with more limited data availability (e.g., East Africa). Although the statistical model alone loses skill when trained on fewer years of data, these losses do not meaningfully propagate to the blended model, indicating that the approach remains viable under data sparsity (Extended Data Fig. 5). This may be because the AIWP forecasts and climatology have implicit information in common, so that the blended modelâs structure allows other inputs to compensate when the quality of one declines. The value of the blending approach is further explored by comparing it with a standard multimodel ensemble approach, often used as a default for combining NWP model predictions 33, 44 . The predicted probabilities of the multimodel ensemble are taken to be a weighted average of those implied by individual models. Special cases of this approach, such as Bayesian model averaging, correspond to different methods of choosing the weights. During the 2000â2024 cross-validation period, the blended model outperforms every fixed-weight ensemble, even when ensemble weights are selected ex post (Fig. 4). In the 1965â1978 hold-out period, this remains true for two of the three metrics, although the best ensemble achieves a slightly higher RPSS (Extended Data Fig. 4). Discussion The AI revolution has the potential to change hundreds of millions of lives by providing skillful forecasts of consequential weather phenomena. However, onset forecasts from these AIWP models do not, on their own, meet the criteria our decision- theory model gives for dissemination at lead times past one week (Table 1 and Fig. 2). We instead satisfy these criteria by combining multiple AIWP model forecasts with a model of farmersâ evolving-expectations based on historical rain-gauge data (Fig. 1). While this framework can also use NWP forecasts, AIWP models offer key practical advantages. In addition to having accuracy comparable to or better than that of NWP models, AIWP models can more easily produce large hindcast ensembles with the same model specification used for real-time operation. This facilitates decision-oriented benchmarking and addresses the small test-sample size problem for infrequent phenomena such as monsoon onset 15 , and also enables a blending approach that leverages the individual strengths of multiple models. Forecasters can improve skill for particular phenomena by adding existing models to a blend even when they cannot improve or tailor the underlying weather models themselves. To illustrate the magnitude of the skill improvement provided by the blended model, we place our results within the broader forecasting literature while noting that these are not direct comparisons. At one-week lead times, the blended modelâs BSS of 30% is broadly comparable to skill scores reported for the Global Ensemble Forecast System (GEFS) 45 and the European Centre for Medium-Range Weather Forecasts (ECMWF ensemble) 46 precipitation forecasts at 4-5 day lead times. This is representative of the upper range of skill for major operational forecasts of precipitation-related metrics over the continental United States and Europe/North Africa respectively. Also in the first week, the evolving-expectations model improves on static climatology by about 10 percentage points and the blended model improves on calibrated NGCM forecasts by about 7 percentage points. These are roughly comparable to the increase in BSS between GEFS version 11 and GEFS version 12 precipitation forecasts at up to one-week lead times, a major development 45 . Skill necessarily declines with lead time; however, the positive skill out to four weeks is noteworthy in the context of subseasonal-to-seasonal forecasting where state-of-the-art AIWP and NWP models show near-zero or negative BSS for weekly precipitation at weeks 3â4 47 . The blended model had large skill in 2025 even though that yearâs monsoon onset was atypical (Fig. 2). The monsoon first reached mainland India eight days earlier than normal on May 24, but by May 29 its northward progression stalled for more than two weeks. This unusual progression meant the climatological baseline performed poorly, but the blended modelâs accuracy in 2025 was similar to its accuracy in a typical year, so its skill relative to the climatological baseline was high. Most existing monsoon onset forecasts are produced once at the start of the season 33, 34 . In contrast, weekly updating allows forecast skill to increase as the true onset date approaches, enabling users to incorporate progressively more informative guidance throughout the planning cycle. Forecasts of monsoon onset and other periodic phenomena (such as first frost or the first hurricane of a season) should consider adopting evolving climatological baselines, constructed by conditioning historical distributions on the event not having occurred by the forecast date. Similar principles might apply to other quasi-periodic 5/22 climate phenomena, including the El NiĂąoâSouthern Oscillation and the MaddenâJulian Oscillation. More generally, forecast evaluation should employ baselines that include information already available to users, ensuring that reported skill corresponds to true added value (Table 1). One practical limitation of the 2025 implementation was that forecasts were produced at weekly resolution. This arose from both the multinomial logit functional form, in which the number of coefficients scales quadratically with the number of forecast bins, and the variability of onset timing within a grid cell. Future work might retain skill gains from the blending procedure while enabling daily forecast resolution, and also potentially enhance spatial resolution finer than2 ⌠. Additional models (including NWP models) and dynamical information could be added to the blended framework. The best models and predictors to include in a blended framework may differ from those that perform best in benchmarking exercises on their own, as the relevant question is which models and variables provide the most information not already contained in other modelsâ forecasts 48 .Traditional predictor selection procedures, including sparsity-promoting techniques such as LASSO 49 , may be further effective in conjunction with cross-validation to avoid overfitting. In this case, the main role of decision-oriented benchmarking 15 in a blending procedure may be to identify models with flaws that would make them unsuitable for inclusion, rather than to select the best models to be blended. Realizing the potential benefits of AIWP for applications such as farming requires more than building accurate weather pre- diction models: it requires identifying knowledge that decision-makers may already have, and designing statistical frameworks to produce decision-relevant information by flexibly combining multiple sources of data. It also requires careful message design to ensure users can understand and make use of this information. The principles developed here offer a replicable pathway for building climate adaptation tools tailored to large, vulnerable populations across the tropics and beyond. Methods The blended model is designed to predict the probability that onset in each2 ⌠à 2 ⌠grid cell occurs in each of four weeks following forecast initialization. A multinomial logistic regression specification combining probabilistic information from the evolving-expectations model with AIWP rainfall forecasts was evaluated using leave-one-year-out cross-validation over 2000â2024 using three probabilistic metrics: Brier Skill Score (BSS), area under the Receiver Operating Characteristic curve (AUC), and Ranked Probability Skill Score (RPSS). This procedure mitigates overfitting by evaluating models only on data they were not specifically trained on. Data Daily rainfall data from the India Meteorological Department (IMD) observational rain-gauge network 43 regridded to a 2 ⌠by 2 ⌠horizontal resolution was used as both a ground truth data source and an input into the climatology models. Precipitation forecasts from NGCM (30 ensemble members) and AIFS were regridded to the same resolution and initialized twice weekly, beginning in May of each year through the date of monsoon onset for each grid cell. Overall, the model was trained on 32 grid cells selected for potential dissemination (shown in Fig. 1). Grid cells were selected for dissemination based on several factors, including the skill of the underlying AIWP weather models, the climatological likelihood of dry spells following an initial five day wet spell, and the variability of the Moron- Robertson onset date within the two degree grid cells. Onset Definition Forecasts were for a modified Moron-Robertson 36 definition of monsoon onset for a2 ⌠à 2 ⌠grid cell, which defines local monsoon onset as the first wet day (⼠1m) of the first 5-day wet sequence after April 1 whose cumulative rainfall is at least the amount of the local climatological 5-day wet spell, which is not followed by any 10-day dry spell (< 5 mmin total) within the subsequent 30 days. This precipitation-based onset definition is more appropriate than circulation-based indices for agricultural applications. 2 Focus groups indicated that farmers generally did not consider April or May rain to be monsoonal, even if the following dry spells did not meet the Moron-Robertson threshold. As a result, forecasts of very early triggers of the Moron-Robertson onset definition would be unlikely to affect farmer decision-making and hence have limited value. We therefore further restricted our onset definition to consider dates after the monsoon reaches Kerala, the first part of India to receive the monsoon, as declared by the India Meteorological Department (IMD) 34, 50 . Climatology and Evolving-Expectations Model Specifications The evolving-expectations model is used both as a baseline for benchmarking and an input into the blending algorithm. To avoid small-sample problems creating instability during periods where onset is less likely, a Kernel Density Estimator was fit to convert past onset dates for each grid cell into a smooth probability distribution. The Kernel Density Estimator (with Gaussian kernel) estimates the probability distribution as a mixture of normal distributions, with one normal distribution in the mixture centered at each previous onset date. In other words, if a grid cell has 6/22 observed onset datesd 1 , d 2 ,¡ , d n over the pastnyears, the Kernel Density Estimate of the probability density function for onset dates is: f(x) = 1 n n â i=1 N Ď (xâ d i ), whereN Ď (x)denotes the probability distribution of a normal distribution with mean 0 and standard deviationĎ.HereĎis selected via the Sheather-Jones method 51 , a data-driven approach designed to minimize the out-of-sample mean integrated squared error of the estimator. We call the resulting probability distribution (static or unconditional) climatology, for the given grid cell, and use it as a baseline comparator. As input into the blended model and an additional baseline, we also produced evolving-expectations probability distributions for each grid cell and each forecast date, which is constructed by conditioning the climatological distribution on the fact that onset has not occurred by the forecast initialization date. Equivalently, the evolving expectations distribution is the Bayesian posterior for a farmer whose prior belief was the climatological distribution and has observed onset not to have occurred yet. This provides the model with additional information: for example, if onset usually occurs June 15 and the forecast is being made June 30, the model âknowsâ that onset may be likely to come very soon even if the unconditional probability of an onset in early July is quite low (Fig. 1). Blended Model Specification The blended model uses a multinomial logit specification to predict whether onset will occur in the first, second, third, fourth, or after the fourth week following the forecast initialization date. We refer to each of these time units as a âbinâ. For a given forecasti(whereiincorporates both the grid cell and the initialization date) and binj(e.g. week 1, week 2, etc.), we define the following: 1. p i j is the probability the evolving expectations model assigns to binj, conditional on not having occurred at the time of the forecast, winsorized 1 to lie between .0001 and .99. 2. Ď i j = log p i j 1âp i j is the logit-transformed probability of onset in binj. We perform this logit transformation so the probabilities are on a logit scale when being fed into the multinomial logit procedure. 3. Îą i j andν i j are the maximum amount of rainfall predicted by AIFS and NGCM respectively in a five-day period beginning in binj, minus the five-day onset threshold for the grid cell corresponding to forecasti. The NGCM rainfall prediction for each day is defined as the mean rainfall predicted across all ensemble members. 4. β i j andÎź i j are the minimum amount of rainfall predicted by AIFS and NGCM respectively in a ten-day period beginning in bin j, computed analogously. Interaction terms are included betweenĎ,Îą, andνvariables within each week to capture correlations between the various onset predictors. For example, a negative coefficient on someĎ i j Îą i j term indicates a positive correlation between AIFS onset predictions and climatological onset probabilities, which we correct for to avoid model overconfidence. LetY i j indicate whether observed onset occurs in binjfor forecasti. We estimate a multinomial logistic regression ofY i j on the following variables Ď i j , Îą i j , ν i j , Ď i j Îą i j , Ď i j ν i j , Îą i j ν i j , Ď i j Îą i j ν i j , β i j , Îź i j for j = 1, 2, 3, 4. Note that we allow information from any lead-week bin j to affect probabilities for each outcome bin j Ⲡ. More explicitly, we estimate coefficients t â, j, j Ⲡââ0,..., 9, jâ1,..., 4, and j Ⲡâ1,..., 4 such that log P(Y i j Ⲡ= 1) P(Y i5 = 1) = 4 â j=1 t 0 j j Ⲡ+ t 1 j j â˛ Ď i j + t 2 j j Ⲡι i j + t 3 j j Ⲡν i j + t 4 j j â˛ Ď i j Îą i j + t 5 j j â˛ Ď i j ν i j + t 6 j j Ⲡι i j ν i j + t 7 j j â˛ Ď i j Îą i j ν i j + t 8 j j Ⲡβ i j + t 9 j j ⲠΟ i j These ratios, along with the condition that â 5 j Ⲡ=1 P(Y i j Ⲡ= 1) = 1 determine the probabilities the forecast assigns to onset occurring within each bin. The interactions are intended to utilize the modelsâ strengths in predicting five-day wet spells, weighting these according to the underlying climatological onset probabilities to down-weight wet spells during times when dry spells are likely. The ten-day rainfall sums are intended to model forecasts relevant to predicting dry spells in later weeks, so they are not interacted with the probability of onset as the week of the dry spell is not likely to be the same as the week of onset. 1 For a < b, the winsorization of a variable x to lie between a and b is defined as max(a, min(x, b)). 7/22 Potential Sources of Bias This modeling procedure is subject to several potential sources of bias. First, the years 2000-2018 overlap with the training period for both AIWP models (with AIFSâ training and fine-tuning extending through 2022), which could overstate their apparent accuracy and lead the blended model to put too much weight on them. However, this overlap is not necessarily a significant source of bias, as the Indian monsoon onset was not an explicit target of the global models, and when benchmarking raw individual model outputs we did not find evidence that AIWP models performed significantly better during the training period than testing period 15 . Second, because the AIWP models were selected based in part on strong benchmarking performance, there is the risk of a âwinnerâs curseâ 52 in which their accuracy may be artificially high in years that were included in the benchmarking work and the blended model may similarly overweight them. Although the 1965-1978 hold-out period lies outside the AIWP modelsâ training data, benchmarking during these years contributed to the selection of AIFS and NGCM. This bias is likely small, as the wider set of models examined during benchmarking generally exhibited similar levels of skill. Finally, while the cross-validation procedure accounts for overfitting in a particular model specification, the selection of the best-performing model specification is still subject to a winnerâs curse and measures of its skill may be biased upwards. 53 This bias is generally small when comparing a limited number of candidate models and does not affect performance estimates on data not used for model selection. Evaluation Metrics Models were evaluated across three standard probabilistic metrics: the Ranked Probability Score (RPS), the Brier Score, and the Area Under the Receiver Operating Characteristic Curve (AUC). Overall results are provided separately for the 2000-2024 model selection period, the 1965-1978 hold-out period, and the 2025 dissemination period. As above, for each forecastiand binjletp i j denote the forecast probability of onset in this bin, and setY i j = 1if onset did occur in this bin, and 0 if not. The following metrics were used: ⢠Brier score (BS): BS = 1 n n â i=1 m â j=1 (Y i j â p i j ) 2 (1) where n is the number of forecasts and m is the number of bins per forecast. ⢠Ranked Probability Score (RPS): The RPS is a generalization of the Brier Score which takes into account the distance between bins: for example, a prediction of onset in week 2 is âcloserâ to correctly predicting a week 3 onset than a prediction of onset in week 1. It is defined as RPS = 1 n n â i=1 m â k=1 k â j=1 (Y i j â p i j ) ! 2 (2) ⢠Area Under the Receiver Operating Characteristic Curve (AUC): The Receiver Operating Characteristic (ROC) curve is formed by using a probability thresholdtâ [0, 1]for a binary classifier based on the model output and plotting the true positive rate against the false positive rate for all choices oft. These points define a curve, called the ROC curve, and the area under the curve is bounded by 0 and 1. For constructing this curve, we treat each binâs forecast as a separate binary forecast. The area under this curve can be interpreted as the probability that, selecting a random forecast and bin in which onset occurred and selecting a random forecast and bin in which onset did not occur, the model assigned higher probability to the bin in which onset occur than the bin in which onset did not occur. More formally, this can be computed as AUC = â i, j,i Ⲡ, j ⲠY i j (1â Y i Ⲡj Ⲡ)¡ 1[p i j > p i Ⲡj Ⲡ] â i, j Y i j ! â i, j (1â Y i j ) ! .(3) In general, improving the resolution of a forecast will increase its AUC, but improving its calibration without improving its resolution will not. Note that a higher AUC corresponds to a better forecast. 8/22 Brier scores and ranked probability scores are then converted to skill scores (BSS and RPSS), defined as the percent improvement a given metric shows relative to unconditional climatology. Formally, for a metricxand set of forecastsI, we define Skill(x, I) = 1â x I x climatology, I (4) We did not adjust for ensemble size, as standard adjustments do not apply to calibrated outputs and so would not affect any of our main results. Model Evaluation Model skill was calculated according to the three metrics above during the 2000-2024 model selection period, a 1965-1978 hold-out period, and finally during the 2025 dissemination period. All skill scores are computed relative to static climatology. Note that the decision-theory criteria for dissemination are then that a modelâs skill scores should exceed those of the evolving expectations model, not that they should exceed zero. After model selection was complete, the blended model was compared to raw NGCM model outputs, raw NGCM model outputs subject to a simple calibration scheme, or a traditional multimodel ensemble on the fuller 2000-2024 period. Calibration was performed via Platt Scaling 54 on individual bins and renormalizing so that probabilities add up to one, and cross-validated by year. Since the AIWP models do not explicitly forecast the MOK component of our onset definition, we evaluated these outputs using a climatological median MOK date of June 2. Sensitivity analyses were conducted to check that results were not driven by this component of the definition. When a model output identified a potential onset without a full 30-day followup period, we counted this as a forecasted onset unless the model forecasted a 10-day dry spell before the end of the forecasting window. To make the comparison to a multi-model ensemble as favorable as possible for the ensemble, we identified the post- hoc ensemble weights that maximized the RPSS in the 2000-2024 period, using grid search on weights in 10 percentage point increments to find coarsely estimated optima, and then using a standard BroydenâFletcherâGoldfarbâShanno (BFGS) algorithm 55â58 initialized at these weights to refine these estimates into the best post-hoc weights overall. To understand whether model skill was likely due to overfitting, we also evaluated the skill of various submodels (excluding the 10-day predictors or only including a subset of five-day predictors) on both the model selection period and the additional hold-out period. (Fig. 4, Extended Data Fig. 4) Evaluation of 2025 Dissemination Only 28 grid cells actually received forecasts, because the monsoon had progressed past the southernmost grid cells before dissemination began. Following discussions with stakeholders, dissemination in each grid cell continued until the IMD had declared onset over at least half of the cell. Since the IMDâs meteorology-based declaration and the purely rainfall-based agricultural definition of onset can differ, this led to situations in which forecasts went out after the Moron-Robertson onset had already triggered but before the meteorological onset had been declared. In seven grid cells the Moron-Robertson onset occurred before any forecasts went out, while five additional grid cells had a single forecast go out 1-3 days after the Moron-Robertson onset. These forecasts are excluded from the evaluation as they are not directly related to the performance of the forecasting model, but highlight operational challenges inherent to large-scale dissemination. Decision-Theory Model In this section we adapt standard tools in decision theory 12â14 to apply to the problem of a weather forecaster choosing what information to provide in order to benefit a heterogeneous group of farmers. The model laid out below is designed to capture a setting in which: 1.The forecaster has information about future weather that farmers do not have. Farmers have to make decisions for which the optimal choice depends both on the weather and on information about their own circumstances the forecaster does not have. The model also applies in cases where the forecaster has access to private information, but cannot perfectly tailor messages to individual farmersâ circumstances due to e.g. limitations of the dissemination platform. 2.Communication is restricted to (potentially probabilistic) forecasts: forecasters cannot fully explain the forecasting process or provide individualized advice, but may convey uncertainty through probabilities. 3. Farmers have some information about the weather (for example, what happened in past years). Some farmers believe the forecaster has already incorporated this information into the forecasts and therefore they take the forecasted probabilities at face value. In other words, these farmers adopt the forecasted probabilities as their own beliefs without further modification. 9/22 The model allows for risk aversion and does not place restrictions on the value of information, actions, or payoffs. All of our results hold for farmers who have information about the weather which they correctly incorporate into their posterior beliefs via Bayesian updates; we focus on farmers who take the forecasts at face value to illustrate the stronger assumptions needed for forecasts to provide value to these users. Formally, assume farmerifrom a setIchooses an actionθ i â Î i , whose expected payoffg i (θ i , Ď i )depends on future weather stateĎ i â ⌠i . This payoff differs between farmers, since their circumstances differ, and each farmer tries to maximize their own payoff. As an example,θ i may represent a crop or fertilizer choice or planting date, whileĎ i may denote the onset date of sustained rainfall or the expected amount of rainfall next week in farmeriâs location. We assume farmersâ decisions affect their payoffs but not those of other farmers. Some information, given by the value of a random variableS i , is known to both farmeriand the forecaster. The forecaster has access to a additional weather information in the form of a stochastic forecast schemeÎ i . A realization ofÎ i is a probabilistic forecastΞ i : ⌠i â [0, 1]. For tractability, we assume that⌠i is finite and that the forecaster must commit ex ante to whether a forecast is sent, prior to observing the realizationΞ i . Farmers have private information about agriculture and their own circumstances, so the payoffs g i are known to individual farmers but not the forecaster. The forecaster is benevolent in the sense that they want to maximize farmersâ payoffs. A forecast is defined to be useful if no farmer is worse off (in expectation) for having received forecasts, and at least one farmer is better off in the sense of being able to make a decision with higher expected payoff. Farmer iâs prior belief p i over ⌠i is given by p i (t) =P(Ď i = t| S i = Ď i ). We say that the forecast schemeÎ i is well-calibrated if the probabilities produced by its forecasts match the actual weather variableâs empirical frequencies; that is, for any potential forecast Ξ i , P(Ď i = t| Î i = Ξ i ) = Ξ i (t) âtâ ⌠i . We say that the forecast scheme incorporates all information farmeribelieves to be shared if conditioning on the information farmeribelieves to be shared does not give any information beyond the forecasted probabilities; in other words, if for all possible forecasts Ξ i and shared information Ď i P(Ď i = t| Î i = Ξ i , S i = Ď i ) =P(Ď i = t| Î i = Ξ i ). On the other hand, the forecast scheme contains information not previously known to the farmer if there is a piece of information the the farmer believes to be shared, and a forecast in the forecasting scheme, such that the probability of some weather event is changed by knowing the piece of shared information. In other words, if there exist Ξ i , Ď i , and Ď i such that P(Ď i = t| Î i = Ξ i , S i = Ď i )̸=P(Ď i = t| S i = Ď i ). The farmer adopts the forecast as their posterior Ěp(t) = Ξ(t). Theoretical Results In this section we state the main results of the decision-theory model. Proofs can be found in Supplementary Note 1. We first show that each farmer is weakly better off if the forecaster provides probabilistic forecasts rather than deterministic forecasts. This is a straightforward application of Blackwellâs Theorem 12 . Intuitively, different farmers may have different thresholds for action (depending on payoffs, inputs, access to outside options, etc.) and so it may not be possible to find a single deterministic threshold which is optimal for all farmers. Providing probabilities allows individual farmers to make use of the information in different ways depending on their circumstances. Proposition 1. Suppose forecasts are well-calibrated. Then each farmer is (weakly) better in expectation to if the forecaster provides them with the probabilities of each potential outcome rather than coarsening them to deterministic forecasts. We then examine when farmers benefit from receiving forecasts. As is common in decision-theory literature, we can give very general conditions under which farmers weakly benefit (i.e. are not harmed), but will need to assume the set of decisions is large and diverse to make general statements about strict benefits. This is plausible in a complex and heterogeneous setting like agriculture and is empirically testable (Proposition 5.) Proposition 2. If the forecasting scheme is well-calibrated and incorporates all information each of the farmers believes to be shared, then: 10/22 a. All farmers weakly benefit in expectation from receiving forecasts b. Individual farmers can only strictly benefit if the forecast scheme contains information not previously known to them. c. For each farmerifor whom the forecast scheme contains previously unknown information, there exists a payoff function g i (θ i , Ď i ) such that they would strictly benefit in expectation if their payoff function were g i . Informally, we interpret the third statement as saying at least one farmer strictly benefits from a forecast containing new information as long as the set of decision problems (objectives, inputs, outside options, risk thresholds, etc.) is sufficiently diverse. Finally, we give some results related to the question of measurement. We first show that, even if policymakers could measure farmer payoffs perfectly, the realized value of the forecasts depends on the weather realization, so an individual randomized controlled trial will not allow an analyst to estimate the overall expected effects of providing forecasts. We then show that under the assumptions we have made so far, and an additional assumption that decisions are impactful, the number of farmers who have changed a decision in response to forecasts provides a lower bound for the number of farmers who benefit from the program. Finally, we show that if the forecaster can identify all decisions impacted by the forecast, measuring changes in decision-making has higher statistical power for detecting whether forecasts are useful than measuring changes in outcomes. For simplicity, we assume a finite number of farmersi = 1, 2,¡ , n,although similar statements hold for more general sets. Throughout what follows we assume forecasts are well-calibrated and fully incorporate shared information. Proposition 3. Suppose the forecaster has access to a measure of farmer payoffsg, and creates an experiment by which farmers are randomly assigned to receive or not receive forecasts. The confidence interval produced by this experiment will in general not be a confidence interval for the average expected utility gain from forecasts at any confidence level. To illustrate, suppose the forecast concerns a large-scale natural disaster (e.g., a regional flood) that affects all farmers in the experiment simultaneously. In some years, the forecasted probability of a flood may be close to the baseline probability, in which case the forecasts might have limited value to farmers who already had this information. In other years, the forecast may assign a high probability (say, 80 percent) to a flood. The expected value of such information depends on how farmers trade off the costs of precaution against the expected losses from flooding, and hence on the full distribution of outcomes induced by the forecast. However, in any realized experiment only a single aggregate outcome is observed: either the flood occurs or it does not. Consider the case where it does not. Even with an arbitrarily large sample of farmers in a single season the experiment compares the costs of precautions farmers took given an80%risk of flooding with the costs incurred due to the actual non-occurrence of a flood. These will not match the expected payoff to farmers of receiving forecasts unless the experiment is repeated over a number of years and geographic regions and results are pooled across different realized outcomes 59 . Risk-averse farmers may be modeled as having a payoff functiong i (θ i , Ď i ) = u(h i (θ i , Ď i )),wherehis an observable function (e.g. income) anduis a concave utility function. In this case, measurements of the impact of forecasts onhmay be negative even when the effect on farmers is positive. Intuitively, this is because farmers may use forecasts to manage risk, paying a small cost to insure themselves against large negative shocks. Proposition 4. If farmers are risk-averse in the sense of the preceding paragraph, then a forecast can make them better off in the sense of increasing their expected utility, reducing their expected observed income. Rather than measuring payoffs directly, we find that evaluators can learn about the value of forecasts by measuring whether or not forecasts change farmersâ decisions Proposition 5. If forecasts are well-calibrated and fully incorporate all information farmers believe to be shared, and that the utility-maximizing decision conditional on either the prior or the received forecast is unique with probability one. Then, the number of farmers who strictly benefit in expectation from a forecasting scheme is greater than or equal to the expected number (across all draws from the scheme) of farmers who changed a decision as a result of a single instantiation of the forecast. The new assumption added in this proposition, that utility-maximizing decisions are unique with probability one, is primarily a technical assumption to remove cases where forecasts spuriously change a decision with no impact on payoff. Finally, we show that measuring decisions can improve power over measuring yield or income outcomes. Proposition 6. Suppose a forecaster can perfectly measure farmer choicesθ i and payoffsg i and runs a randomized controlled trial to determine whether any farmers strictly benefit from forecasts. Under the assumptions of Proposition 5, it follows that: 1. If the distribution of either θ i or g i is changed by the forecasts, then at least one farmer strictly benefits. 2.For any test of whether the distribution ofg i is changed by the forecasts based on the experiment, there is a test of whether the distribution of θ i is changed by the forecasts with weakly lower Type I and Type I error rates in expectation. 11/22 Acknowledgments This work grew out of a collaboration with Indiaâs Ministry of Agriculture and Farmersâ Welfare. We thank Joy Dada and Stephanie Dragoi for excellent research assistance. We are grateful for fruitful discussions with Shreya Agrawal, Medha Deshpande, Josh Deutschmann, Genevieve Flaspohler, Subimal Ghosh, Ankur Gupta, Neil Hausmann, Stephan Hoyer, Peter Huybers, Mrutyunjay Mohapatra, Vincent Moron, Raghu Murtugudde, Yaw Nyarko, Sivananda Pai, D.R. Pattanaik, Thara Prabhakaran, Satya Prakash, V.S. Prasad, Suryachandra Rao, Muthalagu Ravichandran, Sakha Sanap, Sahadat Sarkar, K. K. Singh, Gokul Tamilselvam, Witold Wi ̨ecek, and Janni Yuval. We are grateful to Google and the European Centre for Medium- Range Weather Forecasts for making the AI models used in this work publicly available, to Shreya Agrawal and Stephan Hoyer for assistance in running NGCM with real-time initial conditions, and to the India Meteorological Department for publishing 125 years of rain gauge data. We thank the many people at the Development Innovation Lab, the Development Innovation Lab - India, and Precision Development for their work preparing forecasts for dissemination. This work is partially supported by AIM for Scale, the Gates Foundation, Wellspring Philanthropic Fund, and by the University of Chicagoâs Human-centered Weather Forecasts Initiative (a program of the Institute for Climate and Sustainable Growth, ICSG), Development Innovation Lab, and AI for Climate Initiative (a joint program of the Data Science Institute, DSI, and ICSG). We thank the University of Chicagoâs DSI, Research Computing Center, and Social Science Computing Services for computational resources. Author contributions PH, WB, AJ, and MK conceived and planned the study and supervised the research. AM generated the AIWP forecasts. CA and MK conceptualized the decision-theory framework, and CA proved the associated theorems. CA, RM, AM, MG, and TY analyzed data. CA designed the âevolving expectationsâ and blended models. CA, PH, WB, AJ, and MK drafted the manuscript. CA and AJ created the figures. All authors discussed and interpreted the results. All authors reviewed and edited the manuscript. Data availability ERA5 data used in this study for initializing the AI models is obtained from the Copernicus Climate Change Service (C3S) Climate Data Store. Real-time initial conditions used in the 2025 monsoon season forecasts can be sourced from NCEP GDAS/FNL product inventory (https://w.nco.ncep.noaa.gov/pmb/products/gfs/#GDAS) and the ECMWF IFS Open Data Portal (https://data.ecmwf.int/forecasts/). The India Meteorological Depart- ment (IMD) 1-degree gridded rainfall data are available from the IMD data portal (https://w.imdpune.gov. in/cmpg/Griddata/Rainfall_1_NetCDF.html ). ECMWFâs AIFS model weights are obtained fromhttps: //huggingface.co/ecmwf/aifs-single-1.0. Pre-trained weights for Google Researchâs NGCM (stochastic, IMERG precipitation) can be accessed athttps://neuralgcm.readthedocs.io/en/latest/checkpoints. html . The re-gridded forecasts from NGCM and AIFS and the code used in the study can be accessed athttps: //doi.org/10.5281/zenodo.18894299. During the 2025 season, blended model outputs were uploaded to Zen- odo as they were produced; seehttps://zenodo.org/records/15477259,https://zenodo.org/records/ 15490559,https://zenodo.org/records/15527948,https://zenodo.org/records/15588265 https: //zenodo.org/records/15632841.https://zenodo.org/records/15651248,https://zenodo.org/ records/15683783,https://zenodo.org/records/15731043, andhttps://zenodo.org/records/ 15756597 . Not all forecasts uploaded to Zenodo were disseminated. Blended model outputs will also be placed on Zenodo during the 2026 season. Figures 12/22 A 98°E94°E90°E86°E82°E 78°E74°E70°E 6°N 10°N 14°N 18°N 22°N 26°N 30°N 34°N 38°N Jul 10Jun 26 Jun 12May 29May 15 0.08 0.04 0.06 0.02 0.00 0.04 0.06 0.02 0.00 B C May 24 May 28 Jun 02 Jun 07 Jun 12 Jun 16 Jun 21 Jun 26 Jul 01 Jul 06 Median Onset Date Date B C Probability Probability within shaded region: 0.42 0.23 0.14 Probability within shaded region: 0.31 0.24 0.21 Unconditional probability Conditional probability of evolving expectations model at specific date 26°N, 78°E 18°N, 80°E Advance of monsoon onset Fig. 1. Static climatology and the evolving-expectations model. A) Median onset date for the grid cells used in dissemination. B) The distribution of possible onset dates predicted by the evolving-expectations model for a grid cell centered at latitude 26 ⌠N and longitude 78 ⌠E. Shaded distribution shows the unconditional probability of onset; lines show the probability from the evolving expectations model if the onset has not occurred by the date signified by the dot. Illustrative probabilities are calculated over the shaded grey region. As the season progresses without an onset, the probability of onset in any particular future week increases. C) The distribution of possible onset dates predicted by the evolving-expectations model for a grid cell centered at 18 ⌠N and 80 ⌠E. The decision-theory framework implies that models that cannot outperform this baseline should not be used on their own for dissemination. 13/22 Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model A B C D 0.0 0.1 0.2 0.3 0.6 0.7 0.8 0.9 -0.1 0 0.1 0.2 0.9 0.8 0.7 0.6 0.5 0.3 -0.2 Area under ROC curve Brier skill score 200020052010201520202025 Year Area under ROC curve Brier skill score 1.0 4 weeks3 weeks2 weeks1 week Lead time 4 weeks3 weeks2 weeks1 week Lead time Fig. 2. Evaluation of temporal components of modelsâ skill for the 2000-2024 period. A) Area under ROC curve (AUC) by lead time. The baseline for the bars is chosen to be 0.5, indicating the AUC of a forecast with no ability to distinguish onsets from non-onsets. B) Brier skill score by lead time, computed relative to a traditional (static) climatology model. C) AUC by year across all lead times. The 2000-2024 scores are computed via cross-validation. The 2025 scores are only for forecasts that were actually disseminated before Moron-Robertson onset in each grid cell. Dissemination began in late May, so the set of initialization dates is smaller than in other years. D) Brier skill score by year across all lead times, calculated relative to a traditional (static) climatology model. A deterministic version of AIFS was used in both the blended model and benchmarking, so NGCM is shown here as a representative AIWP model as it performs better on these probabilistic metrics (Fig. 4). See Extended Data Fig. 2 for results for the 1965-1978 period. 14/22 0 0.25 0.50 0.75 1.00 Forecast probability Obser ved frequency CA B 0.250.500.751.00 0 0.250.500.751.00 0 0.250.500.75 1.00 Blended Model NGCM (calibrated)NGCM (raw) Fig. 3. Reliability diagram and histogram of probabilities assigned by different models during the 2000-2024 cross-validation period. If the model is well-calibrated the points should line up with the dashed line. Each point represents a decile of probabilities assigned by the model, and compares the average predicted probability in that decile and the fraction of events which occurred in observation. Histograms are normalized so that the bin capturing probabilities between 0 and 10% has height equal to one. A deterministic version of AIFS was used for both benchmarking and the blended model, so NGCM is used as a representative AIWP model for this figure. See Extended Data Fig. 3 for the results from the 1965-1978 period. BSS RPSS AUC 00.020.05 NGCM (raw) AIFS (calibrated) NGCM (calibrated) NGCM : E NGCM : AIFS : E Post-Hoc Best MME Blended Model â0.08â0.0400.040.080.120.160.20.24 AUC Difference BSS/ RPSS 0.010.030.040.06-0.01-0.02 0.26 0.08 Evolving Exp. Fig. 4. Performance of models during the primary 2000-2024 period, including submodels of the blended model. Vertical lines represent the evolving expectations model, which the decision-theory model implies should be a baseline for dissemination. The interaction models, indicated with colons, are submodels of the final blended model combining the evolving expectations model (E) with 5-day rainfall variables from one or more AIWP models but not 10-day rainfall variables. The best multimodel ensemble (MME) is selected as a linear combination of probabilities from AIFS, NGCM, and the evolving expectations model maximizing the RPSS post-hoc on the 2000-2024 period. All other models (including calibration) are trained via leave-one-year-out cross-validation. Skill scores and differences in AUC are computed relative to unconditional climatology, which has an AUC of .810, a Brier Score of 0.583, and a RPS of 0.611. See Extended Data Fig. 4 for the results from the 1965-1978 period. 15/22 0.00 0.05 0.10 0.15 0.20 2000-20241965-19782025 BSS Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model 0.0 0.1 0.2 2000-20241965-19782025 RPSS 0.5 0.6 0.7 0.8 2000-20241965-19782025 AUC Extended Data Fig. 1. Skill scores aggregated by test period. The blended model is cross-validated during the 2000-2024 period, and trained on 2000-2024 data during the other periods. The climatology model and the evolving expectations model are cross-validated by year using 1900-2024 IMD data. Skill scores are all computed relative to static climatology. 16/22 0.5 0.6 0.7 0.8 0.9 4 weeks3 weeks2 weeks1 week Lead time A Area under ROC curve 0.00 0.05 0.10 0.15 0.20 4 weeks3 weeks2 weeks1 week Lead time Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model B Brier skill score C 0.6 0.7 0.8 0.9 1.0 196519701975 Area under ROC curve D -0.1 0.0 0.1 196519701975 Year Brier skill score Extended Data Fig. 2. Temporal components of model skill, 1965-1978. A) Area under ROC curve (AUC) by lead time during the 1965-1978 pre-satellite-era period. The baseline for the bars is chosen to be 0.5, indicating the AUC of a forecast with no ability to distinguish onsets from non-onsets. B) Brier skill score by lead time, computed relative to a traditional (static) climatology model. C) Area under ROC curve by year across all lead times. The 2000-2024 scores are computed via cross-validation. D) Brier skill score by year across all lead times, calculated relative to a traditional (static) climatology model. 17/22 A NGCM (raw) B NGCM (calibrated) C Blended Model 00.250.500.751.0000.250.500.751.0000.250.500.751.00 0 0.25 0.50 0.75 1.00 Forecast probability Observed frequency Extended Data Fig. 3. Reliability diagram and histogram of probabilities assigned by different models during the 1965-1978 pre-satellite-era period. If the model is well-calibrated the points should line up with the dashed line. Each point represents a decile of probabilities assigned by the model, and compares the average predicted probability in that decile and the fraction of events which occurred in observation. Histograms are normalized so that the bin capturing probabilities between 0 and 10% has height equal to one. -0.03-0.02-0.0100.010.020.030.040.050.06 NGCM (raw) AIFS (calibrated) NGCM (calibrated) NGCM : E NGCM : AIFS : E Best MME Blended Model -0.12-0.08-0.0400.040.080.120.160.20.24 AUC Difference BSS/RPSS Evolving Exp. AUC RPSS BSS Extended Data Fig. 4. Performance of models during the 1965-1978 pre-satellite-era period, including submodels of the blended model. Vertical lines represent the evolving-expectations (E) model, which our decision-theory model implies should be a baseline for dissemination. The interaction models, indicated with colons, are submodels of the final blended model using 5-day rainfall variables from one or more AIWP models but not 10-day rainfall variables. The best multimodel ensemble is selected as a linear combination of probabilities from AIFS, NGCM, and the evolving expectations model maximizing the RPSS post-hoc on the 2000-2024 period. All other models (including calibration) are trained via leave-one-year-out cross-validation. Skill scores and differences in AUC are computed relative to unconditional climatology, which has an AUC of 0.822, a Brier Score of 0.567, and a Ranked Probability Score of 0.594. 18/22 BSS Evolving Expectations Model Blended Model 0.0000.0250.0500.0750.100 RPSS Evolving Expectations Model Blended Model 0.00.10.2 AUC bar = 2000-2024 â = 1965-1978 Evolving Expectations Model Blended Model 0.000.050.100.150.20 Climatology training period 1900-2024 1910-2024 1920-2024 1930-2024 1940-2024 1950-2024 1960-2024 1970-2024 1980-2024 1990-2024 2000-2024 Extended Data Fig. 5. Skills of the evolving expectations model and the blended model using different training periods for the evolving expectations model. Skill scores are computed relative to static climatology trained on all years 1900-2024. Bars indicate scores on the 2000-2024 period, and circles represent scores in the 1965-1978 period. 19/22 References 1. Davis, B. et al. Estimating global and country-level employment in agrifood systems (Food & Agriculture Org., 2023). 2.Burlig, F., Jina, A., Kelley, E. M., Lane, G. & Sahai, H. The value of forecasts: Experimental evidence from India. Working Paper 32173, National Bureau of Economic Research (2024). DOI: 10.3386/w32173. 3.GinĂŠ, X., Townsend, R. M. & Vickery, J. Forecasting when it matters: Evidence from semi-arid India. Tech. Rep., Working paper (2017). 4.Rosenzweig, M. R. & Udry, C. R. Assessing the benefits of long-run weather forecasting for the rural poor: Farmer investments and worker migration in a dynamic equilibrium model. Tech. Rep., National Bureau of Economic Research (2019). 5.Rudder, J. & Viviano, D. Learning from weather forecasts and short-run adaptation: Evidence from an at-scale experiment. Tech. Rep., Working Paper (2024). 6.Yegbemey, R. N., Aloukoutou, A. M. & AĂŻhounton, G. B. D. The impact of short message services (SMS) weather forecasts on cost, yield and income in maize production. Afr. Dev. 46, 163â188 (2021). 7.Fosu, M., Karlan, D., Kolavalli, S. & Udry, C. Disseminating innovative resources and technologies to smallholders in Ghana (dirts). Work. Pap. (2018). 8.Mulungu, K. et al. One size does not fit all: Heterogeneous economic impact of integrated pest management practices for mango fruit flies in Kenyaâa machine learning approach. J. Agric. Econ. 75, 261â279 (2024). 9.Kuivanen, K. et al. Characterising the diversity of smallholder farming systems and their constraints and opportunities for innovation: A case study from the northern region, Ghana. NJAS: Wageningen J. Life Sci. 78, 153â166 (2016). 10. Akzar, R., Umberger, W. & Peralta, A. Understanding heterogeneity in technology adoption among Indonesian smallholder dairy farmers. Agribusiness 39, 347â370 (2023). 11.Krishna, V. V. & Veettil, P. C. Gender, caste, and heterogeneous farmer preferences for wheat varietal traits in rural India. PloS one 17, e0272126 (2022). 12. Blackwell, D. Equivalent comparisons of experiments. The annals mathematical statistics 265â272 (1953). 13. Epstein, E. S. A Bayesian approach to decision making in applied meteorology. J. Appl. Meteorol. 1, 169â177 (1962). 14. Murphy, A. H. The value of climatological, categorical and probabilistic forecasts in the cost-loss ratio situation. Mon. Weather. Rev. 105, 803â816 (1977). 15.Masiwal, R. et al. Decision-oriented benchmarking to transform AI weather forecast access: Application to the Indian monsoon. arXiv preprint arXiv:2602.03767 (2026). 16. Vallangi, N. AI-driven weather forecasts for climate adaptation in India. Nat. Clim. Chang. 16, 102â102 (2026). 17.Hansen, J. W., Mason, S. J., Sun, L. & Tall, A. Review of seasonal climate forecasting for agriculture in sub-saharan africa. Exp. agriculture 47, 205â240 (2011). 18. Ati, O. F., Stigter, C. J. & Oladipo, E. O. A comparison of methods to determine the onset of the growing season in northern Nigeria. Int. J. Climatol. A J. Royal Meteorol. Soc. 22, 731â742 (2002). 19. Wakjira, M. T. et al. Rainfall seasonality and timing: implications for cereal crop production in Ethiopia. Agric. For. Meteorol. 310, 108633 (2021). 20. Gadgil, S. & Gadgil, S. The Indian monsoon, GDP and agriculture. Econ. political weekly 4887â4895 (2006). 21.Yegbemey, R. N., Bensch, G. & Vance, C. Weather information and agricultural outcomes: Evidence from a pilot field experiment in benin. World Dev. 167, 106178 (2023). 22.Tamru, S. et al. Climate and weather services can enhance ethiopian farmersâ resilience to climate change: Economy-wide impact analysis. Clim. Risk Manag. 100725 (2025). 23.Linsenmeier, M. & Shrader, J. G. Global inequalities in weather forecasts. SocArXiv 7e2jf, Center for Open Science (2023). DOI: 10.31219/osf.io/7e2jf. 24. Allen, A. et al. End-to-end data-driven weather prediction. Nature 641, 1172â1179 (2025). 25.Ben Bouallègue, Z. et al. The rise of data-driven weather forecasting: A first statistical assessment of machine learningâ based weather forecasts in an operational-like context. Bull. Am. Meteorol. Soc. 105, E864âE883 (2024). 26. Bi, K. et al. Accurate medium-range global weather forecasting with 3D neural networks. Nature 619, 533â538 (2023). 20/22 27. Lam, R. et al. Learning skillful medium-range global weather forecasting. Science 382, 1416â1421 (2023). 28.Pathak, J. et al. Fourcastnet: A global data-driven high-resolution weather model using adaptive fourier neural operators. arXiv preprint arXiv:2202.11214 (2022). 29. Price, I. et al. Probabilistic weather forecasting with machine learning. Nature 637, 84â90 (2025). 30.Nath, S. et al. Predicting east africaâs rainfall extremes with calibrated, hybrid physical and ai systems (2026). Preprint, accessed March 4, 2026. 31.Bombardi, R. J., Pegion, K. V., Kinter, J. L., Cash, B. A. & Adams, J. M. Sub-seasonal predictability of the onset and demise of the rainy season over monsoonal regions. Front. Earth Sci. 5, 14 (2017). 32.Goswami, P. & Gouda, K. Evaluation of a dynamical basis for advance forecasting of the date of onset of monsoon rainfall over india. Mon. Weather. Rev. 138, 3120â3141 (2010). 33.Scheuerer, M., Bahaga, T. K., Segele, Z. T. & Thorarinsdottir, T. L. Probabilistic rainy season onset prediction over the greater horn of Africa based on long-range multi-model ensemble forecasts. Clim. Dyn. 62, 3587â3604 (2024). 34.Pai, D. & Nair, R. M. Summer monsoon onset over Kerala: New definition and prediction. J. Earth Syst. Sci. 118, 123â135 (2009). 35.Pattanaik, D. & Bushair, M. Objective method of predicting monsoon onset over Kerala in medium and extended range time scale using numerical weather prediction models. Discov. Appl. Sci. 6, 368 (2024). 36.Moron, V. & Robertson, A. W. Interannual variability of Indian summer monsoon rainfall onset date at local scale. Int. journal climatology 34 (2014). 37. Krzysztofowicz, R. The case for probabilistic forecasting in hydrology. J. Hydrol. 249, 2â9 (2001). 38.Ripberger, J. et al. Communicating probability information in weather forecasts: Findings and recommendations from a living systematic review of the research literature. Weather. Clim. Soc. 14, 481â498 (2022). 39.Patt, A., Suarez, P. & Gwata, C. Effects of seasonal climate forecasts and participatory workshops among subsistence farmers in zimbabwe. Proc. Natl. Acad. Sci. 102, 12623â12628 (2005). 40.Luseno, W. K., McPeak, J. G., Barrett, C. B., Little, P. D. & Gebru, G. Assessing the value of climate forecast information for pastoralists: Evidence from southern ethiopia and northern kenya. World development 31, 1477â1494 (2003). 41. Kochkov, D. et al. Neural general circulation models for weather and climate. Nature 1â7 (2024). 42. Lang, S. et al. AIFS-ECMWFâs data-driven forecasting system. arXiv preprint arXiv:2406.01465 (2024). 43.Pai, D. S. et al. Development of a new high spatial resolution (0.25 ⌠à 0.25 ⌠) long period (1901â2010) daily gridded rainfall data set over India and its comparison with existing data sets over the region. Mausam 65, 1â18 (2014). 44. Vannitsem, S. et al. Statistical postprocessing for weather forecasts: Review, challenges, and avenues in a big data world. Bull. Am. Meteorol. Soc. 102, E681âE699, DOI: 10.1175/BAMS-D-19-0308.1 (2021). 45.Zhou, X. et al. The development of the ncep global ensemble forecast system version 12. Weather. Forecast. 37, 1069â1084 (2022). 46.European Centre for Medium-Range Weather Forecasts. Brier skill score of weather parameters â 24h precipitation, forecast day 4. ECMWF Charts. https://charts.ecmwf.int/products/plwww_m_eps_wp_brier_ts?forecast_day=4¶meter=24h% 20precipitation (2026). Accessed: 2026-03-03. 47.Chen, L. et al. A machine learning model that outperforms conventional global subseasonal forecast models. Nat. Commun. 15, 6425 (2024). 48. Kang, Y., Cao, W., Petropoulos, F. & Li, F. Forecast with forecasts: Diversity matters. Eur. J. Oper. Res. 301, 180â190 (2022). 49. Tibshirani, R. Regression shrinkage and selection via the lasso. J. Royal Stat. Soc. Ser. B: Stat. Methodol. 58, 267â288 (1996). 50. Joseph, P., Sooraj, K. & Rajan, C. The summer monsoon onset process over South Asia and an objective method for the date of monsoon onset over Kerala. Int. J. Climatol. A J. Royal Meteorol. Soc. 26, 1871â1893 (2006). 51.Sheather, S. J. & Jones, M. C. A reliable data-based bandwidth selection method for kernel density estimation. J. Royal Stat. Soc. Ser. B (Methodological) 53, 683â690 (1991). 52.Zhong, H. & Prentice, R. L. Correcting âwinnerâs curseâ in odds ratios from genomewide association findings for major complex human diseases. Genet. Epidemiol. The Off. Publ. Int. Genet. Epidemiol. Soc. 34, 78â91 (2010). 21/22 53.Varma, S. & Simon, R. Bias in error estimation when using cross-validation for model selection. BMC Bioinforma. 7, 91, DOI: 10.1186/1471-2105-7-91 (2006). 54.Platt, J. et al. Probabilistic outputs for support vector machines and comparisons to regularized likelihood methods. Adv. large margin classifiers 10, 61â74 (1999). 55.Broyden, C. G. The convergence of a class of double-rank minimization algorithms: 2. the new algorithm. IMA journal applied mathematics 6, 222â231 (1970). 56. Fletcher, R. A new approach to variable metric algorithms. The computer journal 13, 317â322 (1970). 57. Goldfarb, D. A family of variable-metric methods derived by variational means. Math. computation 24, 23â26 (1970). 58. Shanno, D. F. Conditioning of quasi-Newton methods for function minimization. Math. computation 24, 647â656 (1970). 59.Rosenzweig, M. R. & Udry, C. External validity in a stochastic world: Evidence from low-income countries. The Rev. Econ. Stud. 87, 343â381 (2020). 22/22 Supplementary Information for: Designing probabilistic AI monsoon forecasts to inform agricultural decision-making Colin Aitken 1 , Rajat Masiwal 2 , Adam Marchakitus 2 , Katherine Kowal 3 Mayank Gupta 4 , Tyler Yang 5 , Amir Jina 7 , Pedram Hassanzadeh 2,3 , William R. Boos 5,8 , Michael Kremer 1,6 1 Development Innovation Lab, University of Chicago, IL, 60637 2 Department of the Geophysical Sciences, University of Chicago, IL, 60637 3 Data Science Institute, University of Chicago, IL, 60637 4 Development Innovation Lab India, University of Chicago Trust, India, 560025 5 Department of Earth and Planetary Science, University of California, Berkeley, California, 94720 6 Kenneth C. Griffin Department of Economics, University of Chicago, IL, 60637 7 Harris School of Public Policy, University of Chicago, IL, 60637 8 Climate and Ecosystem Sciences Division, Lawrence Berkeley National Laboratory, Berkeley, California, 94720 ⢠Supplementary Notes S1 to S2 ⢠Supplementary Figure S1 to S17 2 Supplementary Note 1: Proofs of Theoretical Results This section records proofs of the propositions in the "Theoretical Results" part of the method section. We first establish a Lemma that will underlie the strategy of most of the proofs. Lemma 1. If Ξ(t) is well-calibrated and fully incorporates all information farmer i believes to be shared, then adopting the forecast at face value represents an accurate Bayesian update for the farmer. Proof of Lemma 1. The given assumptions imply P(Ď i = t| Î i = Ξ i ,S i = Ď i ) = P(Ď i = t| Î i = Ξ i ) = Ξ i (t). Proof of Proposition 1. This is essentially one direction of Blackwellâs informativeness theorem [1]. Let Ë Îž denote a deterministic forecast computed from a forecast instance Ξ. The definition of θ i,f implies that E ĎâźÎž [g i (θ i,Ξ ,Ď)]⼠E ĎâźÎž [g i (θ i, Ë Îž ,Ď)] In particular, E Ξ [E ĎâźÎž [g i (θ i,Ξ ,Ď)]| Ξ]⼠E Ξ [E ĎâźÎž [g i (θ i, Ë Îž ,Ď)]| Ξ] so the result follows from the Law of Iterated Expectations. The proof of Proposition 2 consists of Lemma 1 applied to Lemma 2. Lemma 2. The statement of Proposition 2 holds for farmers who perform perfect Bayesian updates without the assumption that forecasts are well-calibrated or incorporate shared information. Proof. For any distribution f on ⌠i , let θ i,f be farmer iâs decision maximizing expected utility E Ď i âźf [g i (θ i,f ,Ď i )] under the assumption that Ď i is distributed according to f. The expected benefit of a forecast scheme Î i to farmer i is given by E[g i (θ i, Ěp i ,Ď i )â g i (θ i,p i ,Ď i )], which can be rewritten as E Î i ,S i E Ď i âź Ěp i |Î i ,S i [g i (θ i, Ěp i ,Ď i )â g i (θ i,p i ,Ď i )] , This is guaranteed to be non-negative by definition θ i, Ěp i is the choice of θ maximizing expected utility when Ď i is distributed according to Ěp i . If the forecast does not contain information outside the farmerâs decision set, then Ěp i (t) = p i (t) for all i and t, and no farmer will strictly benefit. By Blackwellâs informativeness theorem [1], if the forecasts contain information outside the prior information set at least one decision problem will be improved by the forecast, so the converse follows as long as the set of decision problems is sufficiently broad to contain the one guaranteed by Blackwellâs Theorem. Proof of Proposition 3. To provide a counterexample, we assume a large number of identical farmers who receive the same Ξ i and Ď i , and whose g i âs are identical but whose measured Ëg i may differ due to noise. (This latter assumption is not necessary but prevents the confidence interval from having size zero.) The average expected utility gain from forecasts is E Î i ,S i E Ď i âź Ěp i |Î i ,S i [g i (θ i, Ěp i ,Ď i )â g i (θ i,p i ,Ď i )] , which by assumption does not depend on i. Given a particular draw of Ξ i and Ď i , the measured treatment effect will be normally distributed with mean g i (θ i, Ěp i ,Ď i )â g i (θ i,p i ,Ď i ). and variance proportional to 1/ â n. As nââ, the confidence interval around this quantity will converge to length zero, so if the expected utility gain has any dependence on Ξ i or Ď i it will exclude the true average expected utility gain. 3 Proof of Proposition 4. It suffices to give an example. Suppose u(x) = â x, and suppose the farmer is deciding whether to purchase crop insurance against a potential drought. Let Ď i be an indicator for whether or not a drought occurs, and let θ i be an indicator for whether or not the farmer purchases insurance. We assume that a priori the probability of a drought is 10%, and the insurance is priced in a way that the farmer would not want to buy it without further information. Specifically, we set: h(0, 0) = 100 h(1, 0) = 0 h(0, 1) = 81 h(1, 1) = 16, so that the farmerâs expected income without insurance is 90 and with insurance is 74.5, while their expected utility without insurance is 9 and with insurance is 8.5. Now, suppose the forecasting system issues an accurate system about the likelihood of a drought. In particular, assume that 80% of the time the system indicates a drought will not occur, and 20% of the time the system indicates a drought will occur with probability 50%. In this case the farmersâ expected income without insurance is the same, while their expected income if they buy insurance when the drought risk is elevated is 89.7. However, their expected utility if they buy insurance when the drought risk is elevated rises to 9.3. In particular, a farmer with access to forecasts will manage their risk by buying the insurance plan, in- creasing their utility but decreasing their expected income. An experiment which only measures expected income would show that farmers were harmed, failing to account for their risk aversion. Proof of Proposition 5. Again we can assume farmers are Bayesian, appealing to Lemma 1. We saw above that the expected benefit of a forecast scheme Î i to such a farmer i is E Î i ,S i E Ď i âź Ěp i |Î i ,S i [g i (θ i, Ěp i ,Ď i )â g i (θ i,p i ,Ď i )] . Let A i be a random variable equal to 1 if farmer i changes a decision after seeing forecast Ξ i and 0 otherwise. The assumption on θ i, Ěp i implies that for each i: P(A i = 1) = P(g i (θ i, Ěp i ,Ď i ) > g i (θ i,p i ,Ď i )) Markovâs inequality then implies that if P(A i = 1) > 0, then E Î i ,S i E Ď i âź Ěp i |Î i ,S i [g i (θ i, Ěp i ,Ď i )â g i (θ i,p i ,Ď i )] > 0. so a farmer strictly benefits from the forecast scheme in expectation if their probability of changing a decision is positive. The result then follows because E " X i A i # ⤠X i 1[E[A i ] > 0] where the left hand denotes the expected number of farmers changing a decision from a single forecast instantiation and the right hand side is the number of farmers who benefit from the forecast scheme. Proof of Proposition 6. The first part follows from Proposition 5, noting that g i can only be changed by the forecasts if θ i is. The second follows from Blackwellâs informativeness theorem, since g i is a garbling of θ i . The specific translation from informativeness to Type-I and Type-I errors can be found in [2] with further explanation in [3]. Supplementary Note 2: Sensitivity Analyses One potential concern is the gap in performance between NeuralGCM (and models built on it, including the multimodel ensemble) where ensemble member onsets are defined using a climatological MOK filter, and a variant where the true MOK is used (even for forecasts well before it occurs.) To address this, we provide alternate versions of Figures 2 through 4 as well as Extended Data Figure 1 for alternate onset definitions, first by replacing the condition that onset must happen after MOK with a condition 4 that it must happen after the climatological June 1 MOK date, and second by removing the MOK filter altogether and counting any date after May 1. The resulting figures (SI Figures S5 - S17) look generally similar to the figures in the main text. In particular, the observation that during the 2000-2024 period NeuralGCM does not outperform the evolv- ing expectations model beyond one- or two- week lead times even after calibration, and the observation that the hybrid model outperforms any multimodel ensemble in the 2000-2024 period do not appear to be the result of an unfair comparison. A few comparisons, which are closer in the main text, appear to have some dependence on chosen definitions. 5 Supplementary Figures Is p(onset) ⼠90%? Pl antDel ay I will have trouble bearing th e costs of a dry s pell, so need lots of certainty to p lant Farmer A Is p(onset) ⼠80%? Pl antDel ay I have some savings, and am willing to ta ke ri sks to get a l onger gro wing season Farmer B Is p(onset) ⼠90%? Plant High-yield Seeds I have access to high-y ield seeds with more stri ngent wate r needs Is p(onset) ⼠80%? Plant Regular Seeds Del ay Farmer C Is p( onset) ⼠70%? Plan tDel ay My fa mily member can supplement our income with non- agricultu ra l labor if needed Is p( onset) ⤠90% Outs ide Job No Outside Job Farmer DFarmer Dâ Farmer s In Dier ent Circumstances Face Dier ent Deci si on Problems Mod el ind icates 80 % chance of on set thi s wee k Prov ide ad vice tai lored to a pa rticular farmerâs si tua tion (Farmer B) Farmer A Farmer B Farmer C Farmer DFarmer Dâ DelayPlant Plant Regular Seeds Plant Outside Jo b Commun icate probabilit y di rectly to farmers Farmer A Farmer B Farmer C Farmer DFarmer Dâ Exposed to unwanted ris k No inf ormation to inf orm seed choice No inf ormation to inf orm labor decis ion Each fa rmer uses pri vate info rmati on to make decision best fo r t heir own situ ati on âO ne size fits allâ advice poorly suite d to individual fa rmers â needs Probabil is ti c forecasts l et farmer s make decis ions that t their own unique object iv es and constr aints Supplementary Figure S1: An illustration of the decision-theory modelâs implication that probabilis- tic forecasts provide more value than deterministic forecasts. 6 10 20 30 708090 lon lat Skill -0.1 0.0 0.1 Brier Skill (Blended Model vs Evolving Expectations Model) Supplementary Figure S1: Map of Brier skill scores of blended model relative to the evolving ex- pectations model by grid cell during the 2000-2024 cross-validation period. Grid cells used for model training and intended for dissemination are depicted with thick outlines. 7 10 20 30 708090 lon lat Skill -0.1 0.0 0.1 RPS Skill (Blended Model vs Evolving Expectations Model) Supplementary Figure S2: Map of RPS skill scores of blended model relative to the evolving ex- pectations model by grid cell during the 2000-2024 cross-validation period. Grid cells used for model training and intended for dissemination are depicted with thick outlines. 8 10 20 30 708090 lon lat Difference -0.025 0.000 0.025 0.050 0.075 AUC Difference (Blended Model - Evolving Expectations Model) Supplementary Figure S3: Map of differences between the AUC of the blended model and that of the evolving expectations model by grid cell during the 2000-2024 cross-validation period. Grid cells used for model training and intended for dissemination are depicted with thick outlines. 9 A NGCM (raw) B NGCM (calibrated) C Blended Model 00.250.500.751.0000.250.500.751.0000.250.500.751.00 0 0.25 0.50 0.75 1.00 Forecast probability Observed frequency Supplementary Figure S4: Reliability diagram and histogram of probabilities assigned by different models during the 2000-2024 cross-validation period, removing the MOK filter from the onset definition. If the model is well-calibrated the points should line up with the dashed line. Each point represents a decile of probabilities assigned by the model, and compares the average predicted probability in that decile and the fraction of events which occurred in observation. Histograms are normalized so that the bin capturing probabilities between 0 and 10% has height equal to one. A NGCM (raw) B NGCM (calibrated) C Blended Model 00.250.500.751.0000.250.500.751.0000.250.500.751.00 0 0.25 0.50 0.75 1.00 Forecast probability Observed frequency Supplementary Figure S5: Reliability diagram and histogram of probabilities assigned by different models during the 1965-1978 pre-satellite-era period, removing the MOK filter from the onset definition. If the model is well-calibrated the points should line up with the dashed line. Each point represents a decile of probabilities assigned by the model, and compares the average predicted probability in that decile and the fraction of events which occurred in observation. Histograms are normalized so that the bin capturing probabilities between 0 and 10% has height equal to one. 10 A NGCM (raw) B NGCM (calibrated) C Blended Model 00.250.500.751.0000.250.500.751.0000.250.500.751.00 0 0.25 0.50 0.75 1.00 Forecast probability Observed frequency Supplementary Figure S6: Reliability diagram and histogram of probabilities assigned by different models during the 2000-2024 cross-validation period, using a climatological MOK date as part of the onset definition. If the model is well-calibrated the points should line up with the dashed line. Each point represents a decile of probabilities assigned by the model, and compares the average predicted probability in that decile and the fraction of events which occurred in observation. Histograms are normalized so that the bin capturing probabilities between 0 and 10% has height equal to one. A NGCM (raw) B NGCM (calibrated) C Blended Model 00.250.500.751.0000.250.500.751.0000.250.500.751.00 0 0.25 0.50 0.75 1.00 Forecast probability Observed frequency Supplementary Figure S7: Reliability diagram and histogram of probabilities assigned by different models during the 1965-1978 pre-satellite-era period, using a climatological MOK date as part of the onset definition. If the model is well-calibrated the points should line up with the dashed line. Each point represents a decile of probabilities assigned by the model, and compares the average predicted probability in that decile and the fraction of events which occurred in observation. Histograms are normalized so that the bin capturing probabilities between 0 and 10% has height equal to one. 11 -0.02-0.0100.010.020.030.040.050.060.07 NGCM (raw) AIFS (calibrated) NGCM (calibrated) NGCM : E NGCM : AIFS : E Best MME Blended Model -0.08-0.0400.040.080.120.160.20.240.28 AUC Difference BSS/RPSS Evolving Exp. AUC RPSS BSS Supplementary Figure S8: Performance of models during the 2000-2024 cross-validation period, removing the MOK filter from the onset definition. Vertical lines represent the evolving expectations model, which our decision-theory model implies should be a baseline for dissemination. The interaction models, indicated with colons, are submodels of the final blended model using 5-day rainfall variables from one or more AIWP models but not 10-day rainfall variables. The best multimodel ensemble is selected as a linear combination of probabilities from AIFS, NGCM, and the evolving expectations model maximizing the Ranked Probability Skill Score post-hoc on the 2000-2024 period. All other models (including calibration) are trained via leave-one-year-out cross-validation. Skill scores and differences in AUC are computed relative to unconditional climatology, which has an AUC of .790, a Brier Score of .612, and a Ranked Probability Score of .682. 12 -0.03-0.02-0.0100.010.020.030.040.050.06 NGCM (raw) AIFS (calibrated) NGCM (calibrated) NGCM : E NGCM : AIFS : E Best MME Blended Model -0.12-0.08-0.0400.040.080.120.160.20.24 AUC Difference BSS/RPSS Evolving Exp. AUC RPSS BSS Supplementary Figure S9: Performance of models during the 1965-1978 pre-satellite-era period, removing the MOK filter from the onset definition. Vertical lines represent the evolving expectations model, which our decision-theory model implies should be a baseline for dissemination. The interaction models, indicated with colons, are submodels of the final blended model using 5-day rainfall variables from one or more AIWP models but not 10-day rainfall variables. The best multimodel ensemble is selected as a linear combination of probabilities from AIFS, NGCM, and the evolving expectations model maximizing the Ranked Probability Skill Score post-hoc on the 2000-2024 period. All other models (including calibration) are trained via leave-one-year-out cross-validation. Skill scores and differences in AUC are computed relative to unconditional climatology, which has an AUC of .807, a Brier Score of .582, and a Ranked Probability Score of .646. 13 -0.02-0.0100.010.020.030.040.050.060.07 NGCM (raw) AIFS (calibrated) NGCM (calibrated) NGCM : E NGCM : AIFS : E Best MME Blended Model -0.08-0.0400.040.080.120.160.20.240.28 AUC Difference BSS/RPSS Evolving Exp. AUC RPSS BSS Supplementary Figure S10: Performance of models during the 2000-2024 cross-validation period, using a climatological MOK date as part of the onset definition. Vertical lines repre- sent the evolving expectations model, which our decision-theory model implies should be a baseline for dissemination. The interaction models, indicated with colons, are submodels of the final blended model using 5-day rainfall variables from one or more AIWP models but not 10-day rainfall variables. The best multimodel ensemble is selected as a linear combination of probabilities from AIFS, NGCM, and the evolving expectations model maximizing the Ranked Probability Skill Score post-hoc on the 2000-2024 period. All other models (including calibration) are trained via leave-one-year-out cross-validation. Skill scores and differences in AUC are computed relative to unconditional climatology, which has an AUC of .821, a Brier Score of .574, and a Ranked Probability Score of .595. 14 -0.03-0.02-0.0100.010.020.030.040.050.06 NGCM (raw) AIFS (calibrated) NGCM (calibrated) NGCM : E NGCM : AIFS : E Best MME Blended Model -0.12-0.08-0.0400.040.080.120.160.20.24 AUC Difference BSS/RPSS Evolving Exp. AUC RPSS BSS Supplementary Figure S11: Performance of models during the 1965-1978 pre-satellite- era period, using a climatological MOK date as part of the onset definition. Vertical lines represent the evolving expectations model, which our decision-theory model implies should be a baseline for dissemination. The interaction models, indicated with colons, are submodels of the final blended model using 5-day rainfall variables from one or more AIWP models but not 10-day rainfall variables. The best multimodel ensemble is selected as a linear combination of probabilities from AIFS, NGCM, and the evolving expectations model maximizing the Ranked Probability Skill Score post-hoc on the 2000- 2024 period. All other models (including calibration) are trained via leave-one-year-out cross-validation. Skill scores and differences in AUC are computed relative to unconditional climatology, which has an AUC of .827, a Brier Score of .561, and a Ranked Probability Score of .590. 15 0.5 0.6 0.7 0.8 0.9 4 weeks3 weeks2 weeks1 week Lead time A Area under ROC curve 0.0 0.1 0.2 0.3 4 weeks3 weeks2 weeks1 week Lead time Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model B Brier skill score C 0.6 0.7 0.8 0.9 1.0 200020052010201520202025 Area under ROC curve D -0.2 -0.1 0.0 0.1 0.2 200020052010201520202025 Year Brier skill score Supplementary Figure S12: Temporal components of model skill, using a climatological MOK date as part of the onset definition. (A) Area under ROC curve (AUC) by lead time during the 2000-2024 cross-validation period. The baseline for the bars is chosen to be 0.5, indicating the AUC of a forecast with no ability to distinguish onsets from non-onsets. (B) Brier skill score by lead time, computed relative to a traditional (static) climatology model. (C) Area under ROC curve by year across all lead times. The 2000-2024 scores are computed via cross-validation. The 2025 scores are only for forecasts that were actually disseminated before MR onset in each grid cell. Dissemination began in late May, so the set of initialization dates is smaller than in other years. (D) Brier skill score by year across all lead times, calculated relative to a traditional (static) climatology model. 16 0.5 0.6 0.7 0.8 0.9 4 weeks3 weeks2 weeks1 week Lead time A Area under ROC curve 0.0 0.1 0.2 4 weeks3 weeks2 weeks1 week Lead time Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model B Brier skill score C 0.6 0.7 0.8 0.9 1.0 200020052010201520202025 Area under ROC curve D -0.1 0.0 0.1 0.2 200020052010201520202025 Year Brier skill score Supplementary Figure S13: Temporal components of model skill, removing the MOK filter from the onset definition. (A) Area under ROC curve (AUC) by lead time during the 2000-2024 cross-validation period. The baseline for the bars is chosen to be 0.5, indicating the AUC of a forecast with no ability to distinguish onsets from non-onsets. (B) Brier skill score by lead time, computed relative to a traditional (static) climatology model. (C) Area under ROC curve by year across all lead times. The 2000-2024 scores are computed via cross-validation. The 2025 scores are only for forecasts that were actually disseminated before MR onset in each grid cell. Dissemination began in late May, so the set of initialization dates is smaller than in other years. (D) Brier skill score by year across all lead times, calculated relative to a traditional (static) climatology model. 17 0.5 0.6 0.7 0.8 0.9 4 weeks3 weeks2 weeks1 week Lead time A Area under ROC curve 0.00 0.05 0.10 0.15 0.20 4 weeks3 weeks2 weeks1 week Lead time Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model B Brier skill score C 0.6 0.7 0.8 0.9 1.0 196519701975 Area under ROC curve D 0.0 0.1 0.2 196519701975 Year Brier skill score Supplementary Figure S14: Temporal components of model skill, using a climatological MOK date as part of the onset definition, 1965-1978. (A) Area under ROC curve (AUC) by lead time during the 1965-1978 pre-satellite-era period. The baseline for the bars is chosen to be 0.5, indicating the AUC of a forecast with no ability to distinguish onsets from non-onsets. (B) Brier skill score by lead time, computed relative to a traditional (static) climatology model. (C) Area under ROC curve by year across all lead times. (D) Brier skill score by year across all lead times, calculated relative to a traditional (static) climatology model. 18 0.5 0.6 0.7 0.8 0.9 4 weeks3 weeks2 weeks1 week Lead time A Area under ROC curve 0.00 0.05 0.10 0.15 0.20 4 weeks3 weeks2 weeks1 week Lead time Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model B Brier skill score C 0.6 0.7 0.8 0.9 1.0 196519701975 Area under ROC curve D -0.10 -0.05 0.00 0.05 0.10 0.15 196519701975 Year Brier skill score Supplementary Figure S15: Temporal components of model skill, removing the MOK filter from the onset definition, 1965-1978. (A) Area under ROC curve (AUC) by lead time during the 1965-1978 pre-satellite-era period. The baseline for the bars is chosen to be 0.5, indicating the AUC of a forecast with no ability to distinguish onsets from non-onsets. (B) Brier skill score by lead time, computed relative to a traditional (static) climatology model. (C) Area under ROC curve by year across all lead times. (D) Brier skill score by year across all lead times, calculated relative to a traditional (static) climatology model. 19 0.00 0.05 0.10 0.15 0.20 2000-20241965-19782025 BSS Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model 0.0 0.1 0.2 2000-20241965-19782025 RPSS 0.5 0.6 0.7 0.8 2000-20241965-19782025 AUC Supplementary Figure S16: Skill scores aggregated by test period, using a climatological MOK date as part of the onset definition. The blended model is cross-validated during the 2000- 2024 period, and trained on 2000-2024 data during the other periods. The climatology model and the evolving expectations model are cross-validated by year using 1900-2024 IMD data. Skill scores are all computed relative to static climatology. 20 0.00 0.05 0.10 0.15 0.20 2000-20241965-19782025 BSS Static Climatology Evolving Expectations Model NGCM (calibrated) Blended Model 0.0 0.1 0.2 2000-20241965-19782025 RPSS 0.5 0.6 0.7 0.8 2000-20241965-19782025 AUC Supplementary Figure S17: Skill scores aggregated by test period, removing the MOK filter from the onset definition. The blended model is cross-validated during the 2000-2024 period, and trained on 2000-2024 data during the other periods. The climatology model and the evolving expectations model are cross-validated by year using 1900-2024 IMD data. Skill scores are all computed relative to static climatology. 21 References [1] David Blackwell. Equivalent comparisons of experiments. The annals of mathematical statistics, pages 265â272, 1953. [2] David A Blackwell and Meyer A Girshick. Theory of games and statistical decisions. Courier Corpo- ration, 1979. [3] Xiaosheng Mu, Luciano Pomatto, Philipp Strack, and Omer Tamuz. From blackwell dominance in large samples to rĂŠnyi divergences and back again. Econometrica, 89(1):475â506, 2021. 22