<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <title>weatherml</title>
  <subtitle>New papers on machine learning for weather and climate</subtitle>
  <link href="https://weatherml.github.io/"/>
  <link rel="self" href="https://weatherml.github.io/feed.xml"/>
  <id>https://weatherml.github.io/</id>
  <updated>2026-10-04T05:12:33Z</updated>
  <entry>
    <title>AI Emulation of Stochastic Sudden Stratospheric Warming with Interpretable Latent Structure</title>
    <link href="https://arxiv.org/abs/2610.02069"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2610.02069"/>
    <id>https://arxiv.org/abs/2610.02069</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>C. Daniel Boscu, Daniel Hernandez, Fabio Alvarez Ventura, Justin Finkel, Ashesh Chattopadhyay, Pedram Hassanzadeh, Dorian S. Abbot</name></author>
    <category term="Other"/>
    <summary>Rare weather regime transitions pose a challenge for data-driven modeling due to class imbalance. In this study, we develop a probabilistic deep learning emulator for a prototypical system with regime transitions, the stochastic Holton--Mass model of stratospheric variability, and analyze the structure of its learned latent space. The Holton--Mass model exhibits two metastable regimes, a strong and a weak polar vortex, maintained by nonlinear wave--mean flow interactions, with weak stochastic forcing intermittently triggering rare transitions between these regimes that qualitatively represent SSW events. We employ a ResNet-inspired Conditional Variational Autoencoder with six-layer encoder and decoder layers and explicit current-state conditioning to model the distribution of the system's state at the next time step (one day). The emulator accurately reproduces short-term dynamics, steady-state probability distributions, regime persistence statistics, rare transition rates, the transition committor function, and the transition expected lead time of the physical model. Beyond emulation fidelity, we interrogate the learned latent representation to understand how the model internalizes the underlying metastable structure of the dynamics. Principal Component Analysis of the 32-dimensional latent space reveals a clear and unsupervised separation into four physically interpretable clusters corresponding to strong versus weak vortex regimes and stable versus transition-prone configurations. Such emergent regime separation in latent space is hard to identify for deep generative models applied to high-dimensional stochastic systems. Our results show that carefully designed probabilistic emulators can uncover physically meaningful manifolds governing extreme-event dynamics, potentially aiding the development of improved operational advanced warning systems.</summary>
  </entry>
  <entry>
    <title>Unsupervised Domain Adaptation for Enhanced Radiometer Image Precipitation Estimation using Conditional Flow Matching</title>
    <link href="https://arxiv.org/abs/2610.01890"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2610.01890"/>
    <id>https://arxiv.org/abs/2610.01890</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>Victor Enescu, Assaad Zeghina, Matthieu Meignin, Nicolas Viltard, Cécile Mallet</name></author>
    <category term="Other"/>
    <summary>Deep generative networks have recently achieved unprecedented performance in precise image and video editing using sophisticated textual prompts. However, the effectiveness of such models heavily depends on access to very large supervised and annotated image datasets, which can be very difficult to obtain. This is particularly true for satellite instruments, which very rarely overlap with labelled data, and suffer from domain shifts in the rare occasions they do. In this paper, we investigate the potential of flow matching models for unsupervised domain adaptation of satellite radiometer images. Our main contribution is a novel unsupervised method that achieves precise domain alignment by leveraging parts of the deterministic ordinary differential equations in flow matching models, conditioned on different satellite instruments. A key strength of our approach is its ability to preserve essential information while adapting across any domains since the perturbations are in theory bijective. Extensive experiments conducted on the GPM-Core constellation show the benefit of our conditional domain adaptation, particularly in improving rain precipitation estimation from radiometer imagery.</summary>
  </entry>
  <entry>
    <title>Varda-single-1.0: deterministic data-driven weather forecasting at 1 km resolution over Switzerland's complex topography</title>
    <link href="https://arxiv.org/abs/2610.01835"/>
    <link rel="related" href="https://weatherml.github.io/papers/regional-models/#2610.01835"/>
    <id>https://arxiv.org/abs/2610.01835</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>Alberto Pennino, Francesco Zanetta, Michele Cattaneo, Claire Merker, Radi Radev, Jonas Bhend, Louis Frey, Hugues de Laroussilhe, Ophélia Miralles, Carlos Osuna, Daniele Nerini, Andreas Pauling, Daniel Hupp, Ulrich Hamann, Mary McGlohon, Marti Bosch, Luca Lanzilao, Marco Arpagaus, Lukas Jansing, Daniel Leuenberger, Mark A. Liniger, Katrin Ehlert, Matthew Chantry, Håvard Homleid Haugen, Gert Mertes, Ana Prieto Nemesio, Mario Santa Cruz, Jasper Wijnands, Gabriel Moldovan, Harrison Cook, Oliver Fuhrer</name></author>
    <category term="Regional Models"/>
    <summary>We present Varda-single-1.0, a medium-range data-driven weather prediction system built for the Alpine domain. It provides hourly deterministic regional forecasts on a mesh of 1 km resolution and global forecasts on a 31 km mesh. The system comprises two independently trained stretched-grid Graph Transformer models with encoder-processor-decoder architecture, developed in the Anemoi framework: a 6-hourly autoregressive forecaster and a temporal downscaler reconstructing hourly forecasts between the forecaster's steps. Its training curriculum includes pre-training on ERA5 reanalysis data, followed by training on a 20-year kilometre-scale regional reanalysis, and finally fine-tuning on operational kilometre-scale analyses. Verified over one year against operational analyses and surface station observations, Varda-single is competitive with or improves on MeteoSwiss' operational numerical weather prediction baselines for most headline scores and variables. It broadly matches the skill of the high-resolution 1 km ICON-CH1-EPS control at lead times up to +33 h and generally outperforms the 2 km ICON-CH2-EPS control at lead times up to +120 h. Despite competitive aggregate scores, Varda-single underestimates some local wind maxima and produces overly smooth convective precipitation fields, consistent with the smoothing associated with squared-error training. To gain insight into the model's behaviour, we investigate three case studies beyond the aggregated headline scores, and find particular weaknesses in Varda-single's representation of local winds over complex terrain. Varda-single represents an important step in the development of high-resolution ML forecasting over complex terrain, in complementing the operational regional numerical weather prediction models of MeteoSwiss with data-driven models and in providing a pretrained model for researchers and user-specific applications.</summary>
  </entry>
  <entry>
    <title>Explaining El Niño Forecasts with the Average Gradient Outer Product</title>
    <link href="https://arxiv.org/abs/2610.01095"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2610.01095"/>
    <id>https://arxiv.org/abs/2610.01095</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>Yuan Hui, Dorian S. Abbot, Robert J. Webber</name></author>
    <category term="Other"/>
    <summary>An important and unresolved problem in the physical sciences is explaining the predictions made by neural networks. Several explainable artificial intelligence (XAI) methods have been proposed to address this problem, including gradient XAI, Integrated Gradients, and GradientSHAP. We evaluate the baseline XAI methods according to four scores: sensitivity (XAI patterns strongly affect predictions), attribution (XAI patterns reproduce the change in prediction relative to a baseline), robustness (XAI patterns remain stable for nearby inputs), and coherence (XAI patterns are spatially smooth). We also introduce a new method, average gradient outer product (AGOP) XAI, that uses global gradient information to identify an important direction for a specific input. We apply XAI to neural network predictions of the El Niño-Southern Oscillation (ENSO) based on data from the Zebiak-Cane model.   AGOP XAI achieves the highest attribution, robustness, and coherence scores in the architecture and lead-time comparisons reported here. Its sensitivity is surpassed by gradient XAI, which is maximally sensitive by definition. Beyond diagnosing neural-network behavior, AGOP XAI can generate candidate hypotheses about physical mechanisms. The method highlights an equatorial thermocline-depth signal consistent with recharge oscillator physics, together with a southeastern-Pacific lobe that may be specific to the Zebiak-Cane model. Finally, we test the physical relevance of AGOP using optimized perturbations that move the Zebiak-Cane model along AGOP explanation coordinates. Such perturbations can suppress the selected extreme events or, from a near-neutral ensemble, generate strong El Niño or La Niña events 10 months later.</summary>
  </entry>
  <entry>
    <title>Weather Jiu-Jitsu: Exploring the Feasibility of Control Paradigms in Weather Foundation Models</title>
    <link href="https://arxiv.org/abs/2610.00792"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2610.00792"/>
    <id>https://arxiv.org/abs/2610.00792</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>Prakriti Biswas, Kobi Abayomi, Upmanu Lall</name></author>
    <category term="Global Models"/>
    <summary>Weather Jiu-Jitsu is a control paradigm for extreme climatological events, inspired by chaos theory. As a proposition, small, precise, targeted, and cost-inexpensive perturbations can redirect trajectories of a large dynamical system. This strategy has been demonstrated analytically in the Lorenz-63 system, where a naturally chaotic trajectory switching between two attractors can be confined to a single attractor, indefinitely, via arbitrarily small perturbations. This paper examines the feasibility of Microsoft's Aurora -- a 1.3 billion parameter global atmospheric model -- as a test bed for this strategy. This paper explores three questions: (1) Is Aurora a reliable enough simulation environment to serve as a meaningful testbed? (2) Are the perturbations required to redirect its trajectories small enough to be physically plausible? (3) Does Aurora's learned latent space (the parametric estimators on climatological attributes) yield any apparent, structured, and/or perhaps interpretable features that can convey a geo/atmospheric response to initial conditions? We find evidence consistent with all three: Aurora's modeled trajectories respond to perturbations beyond measurement drift, the perturbation magnitudes required are small relative to the model's own forecast uncertainty, and its latent representations exhibit directional structure that responds to Jiu-Jitsu-type interventions, even though that structure does not separate extreme from normal states outright. These results should be read as feasibility diagnostics rather than a demonstration of control: we do not implement or test an actual steering intervention on Aurora, and several of our findings, particularly around the model's latent-space geometry, are exploratory.</summary>
  </entry>
  <entry>
    <title>Benchmarking Generative Models for Weather Data Assimilation on Real Station Observations</title>
    <link href="https://arxiv.org/abs/2610.00728"/>
    <link rel="related" href="https://weatherml.github.io/papers/data-assimilation/#2610.00728"/>
    <id>https://arxiv.org/abs/2610.00728</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>Ruizhe Huang, Qidong Yang, Jonathan Giezendanner, Sherrie Wang</name></author>
    <category term="Data Assimilation"/>
    <summary>Weather reanalysis products rely on computationally intensive numerical weather predictions followed by data assimilation that corrects the forecast toward observations. Deep generative models offer a cheaper alternative that shifts much of this cost from inference to offline training. However, existing generative approaches have been evaluated on synthetic observations or under different datasets and evaluation schemes, making it unclear which design choices actually improve real-world data assimilation. We present the first controlled benchmark of generative weather data assimilation on real weather station observations. Using 11,849 NOAA MADIS stations across the contiguous United States and four weather variables, we evaluate methods while holding the dataset, observation operator, and deep learning architecture fixed. The benchmark compares the major design choices, including diffusion versus flow matching, pixel versus latent-space formulations, and multiple inference-time conditioning strategies, against a classical 3D-Var baseline. The benchmark reveals three clear conclusions. First, learned generative priors outperform the Gaussian prior of 3D-Var (35.7% vs. 33.3% RMSE reduction over ERA5) despite using no ERA5 background field at inference. Second, full-gradient guidance consistently outperforms stop-gradient and initial-noise optimization. Third, other choices provide little measurable benefit: diffusion and flow matching perform nearly identically under matched conditions, and latent-space variable mixing does not help. We further evaluate both dense and sparse station settings and find advantages from generative AI and full-gradient guidance more pronounced under sparsity. Together, these results identify which components of generative weather data assimilation improve performance on real station observations and establish a standardized benchmark for future work.</summary>
  </entry>
  <entry>
    <title>STCFormer: Adaptive Spatio-Temporal Modeling with Dynamic Cluster Transformer for Station-based Weather Forecasting</title>
    <link href="https://arxiv.org/abs/2610.00377"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2610.00377"/>
    <id>https://arxiv.org/abs/2610.00377</id>
    <updated>2026-10-01T00:00:00Z</updated>
    <author><name>Rongwen Li, Haixin Xie, Mingyang Wang, Hongwu Liu, Kun Fang, Changjian Chen, Zhuo Tang, Kenli Li</name></author>
    <category term="Global Models"/>
    <summary>Station-based weather forecasting supports daily life and economic activity, yet accurate forecasts require modeling complex spatial dependencies among stations. Recent clustering-based selective modeling offers a promising alternative to dense inter-station interactions. However, a grouping shared across an observation window may obscure local changes in station relationships, while intra-cluster interactions alone may miss important global context. The theoretical advantages of selective interactions over dense connectivity also remain insufficiently understood. We therefore propose STCFormer, an adaptive spatio-temporal Transformer that dynamically groups stations according to their local evolution within each temporal patch. Its Cluster-Guided Attention Block combines fine-grained local attention within clusters and global attention over regional state summaries, allowing each station to access information beyond its own cluster. We further show that a derived Lipschitz upper bound for cluster-conditioned local attention is no larger than its fully connected counterpart, explaining a potential robustness benefit and motivating the design of InfoLoss. Experiments on three real-world weather datasets spanning eight temperature and wind forecasting tasks show that STCFormer achieves the lowest 24-hour mean squared error on all eight tasks and ranks first or second in 47 of 48 comparisons across metrics and forecasting horizons. Ablations and case studies further confirm the benefits of locally adaptive grouping and complementary local-global interactions. Our code can be obtained at https://github.com/hnu-vis/STCFormer.</summary>
  </entry>
  <entry>
    <title>Less is more: error-distance scaling relation for data-efficient kilometer-scale downscaling of extreme heat</title>
    <link href="https://arxiv.org/abs/2609.40140"/>
    <link rel="related" href="https://weatherml.github.io/papers/regional-models/#2609.40140"/>
    <id>https://arxiv.org/abs/2609.40140</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Ahmed Marey, Henry Lu, Abhishek Gaur, Sherif Goubran, Malek Aloui, Theodore Potsis, David Rolnick, Alex Hernandez-Garcia, Liangzhu Leon Wang</name></author>
    <category term="Regional Models"/>
    <summary>Extreme heat is where urban adaptation needs kilometer-scale data the most, but the simulations training a downscaler can cost more than they save, and how much is needed has not been identified. We measured it with CASPER, a U-Net with a structure-preserving loss downscaling 32 km reanalysis to 1 km temperature, humidity and wind, across 24 configurations of one to eight months. Held-out error grows linearly with climatological distance to the training data, RMSE = 0.83 + 2.95 d, explaining 90% of its variance against 7% for volume and predicting unseen months in advance. On held-out extreme summer weeks CASPER preserves the fine-scale structure and cross-variable physics that matched-budget baselines degrade, and matches station observations during documented heat waves to within 1.8 K. Transfer to a new region degrades geographically; 11 days of local simulation cuts Vancouver's held-out error from 3.8 to 1.3 K. Training periods should span the target climate: the same accuracy for four times less simulation, putting kilometer-scale downscaling of extreme heat within reach of groups without large computing facilities.</summary>
  </entry>
  <entry>
    <title>RainAtlas: A Multi-Continental Dataset for Precipitation Downscaling</title>
    <link href="https://arxiv.org/abs/2609.39833"/>
    <link rel="related" href="https://weatherml.github.io/papers/regional-models/#2609.39833"/>
    <id>https://arxiv.org/abs/2609.39833</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Pierre-Louis Lemaire, Luca Schmidt, Wietze Suijker, Alex Hernandez-Garcia, David Rolnick</name></author>
    <category term="Regional Models"/>
    <summary>Extreme rainfall events are increasing in intensity and frequency as climate change accelerates. While kilometer-scale precipitation forecasts are critical for supporting local decision-making, the limited availability of high-resolution precipitation observations hinders their accuracy, especially in under-resourced regions. Machine learning models are widely used to downscale precipitation data to km-scale, but their application to unseen geographies presents challenges. First, processing raw high-resolution precipitation datasets across regions requires significant engineering and domain expertise. Second, generalization across regions remains difficult. To help overcome these barriers, we release RainAtlas, a large-scale, ML-ready and multi-continental dataset for precipitation downscaling. Covering three continents, RainAtlas harmonizes heterogeneous hourly km-scale observations to a common 2-km grid. Each regional partition contains around 210,000 aligned low- and high-resolution precipitation pairs, respectively from ERA5 reanalysis and direct observations. We benchmark state-of-the-art ML-based downscaling models across RainAtlas using a wide range of metrics. Our evaluation reveals substantial variance in out-of-domain generalization depending on the training regions. This underscores the need for cross-regional, multi-source km-scale evaluation, establishing RainAtlas as a well-positioned benchmark for precipitation downscaling research.</summary>
  </entry>
  <entry>
    <title>A library for differentiable signal processing and machine learning on the sphere</title>
    <link href="https://arxiv.org/abs/2609.39737"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.39737"/>
    <id>https://arxiv.org/abs/2609.39737</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Thorsten Kurth, Max Rietmann, Mauro Bisson, Andrea Paris, Alberto Carpentieri, Jean Kossaifi, Anima Anandkumar, Christian Hundt, Boris Bonev</name></author>
    <category term="Other"/>
    <summary>The two-dimensional sphere embedded in three-dimensional Euclidean space S2, plays a central role in a variety of scientific and engineering domains, including geophysics, planetary science, geodesy, atmospheric physics, quantum chemistry, cosmology, and virtual reality, among many others. As machine learning increasingly permeates these fields, the demand grows for robust tools that process and model functions on the sphere, while respecting the inherent topological and symmetry properties of the domain. We present torch-harmonics, a comprehensive library that offers efficient, differentiable implementations of advanced signal processing and machine learning (ML) methods for spherical data. These include the spherical harmonic transform (SHT), the spherical analogue of the Fourier transform, vector spherical harmonics, discrete-continuous and spectral convolutions, as well as both global and neighborhood spherical attention mechanisms. Beyond traditional representations, torch-harmonics provides the building blocks for state-of-the-art spherical ML architectures such as spherical transformers in order to enable scalable, rotationally-aware learning and inference in modern scientific and engineering applications.</summary>
  </entry>
  <entry>
    <title>Butterfly Effect Confirmed in Global AI Weather Models: Evidence from Tropical Cyclone Forecasting</title>
    <link href="https://arxiv.org/abs/2609.39379"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.39379"/>
    <id>https://arxiv.org/abs/2609.39379</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jeremy Cheuk-Hin Leung, Daosheng Xu, Weiye Yu, Shaojing Zhang, Xiaodong Zeng, Gaozhen Nie, Jie Feng, Jingchen Pu, Yi Li, Kaijun Ren, Qingcun Zeng, Banglin Zhang</name></author>
    <category term="Global Models"/>
    <summary>A paradox recently emerged in artificial intelligence (AI) weather prediction research. While some claim AI weather models cannot simulate atmospheric butterfly effect, this conflicts with AI models' limited predictability and advances in AI ensemble forecasting. This study demonstrates via counterexamples that the butterfly effect does exist in AI weather predictions. For Super Typhoon Khanun, AI predictions are constrained by a double-attractor system. Minor initial perturbations confined to two regions trigger state transitions between two local attractors, causing a 1006-km difference in the predicted storm position on Day 7. This behavior is consistent with numerical weather prediction models and observed in ~12% of tropical cyclones in the past 5 years. These findings verify AI's ability to capture atmospheric chaos and provide the physical basis for AI ensemble forecasting.</summary>
  </entry>
  <entry>
    <title>Proper Scoring Rule-based Diffusion for Probabilistic Weather Forecasting</title>
    <link href="https://arxiv.org/abs/2609.38632"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.38632"/>
    <id>https://arxiv.org/abs/2609.38632</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Joonhyeong Park, Giung Nam, Hyungi Lee, Kyunghyun Cho, Byoungwoo Park, Juho Lee</name></author>
    <category term="Global Models"/>
    <summary>Recent probabilistic weather forecasters train stochastic predictors with the continuous ranked probability score (CRPS) to generate each ensemble member in a single forward pass. These models learn the predictive distribution from the forecast context alone, which becomes difficult at longer forecast horizons where uncertainty is high. To learn the predictive distribution more effectively, we introduce auxiliary conditional denoising tasks that predict the same future state from the context and its corrupted version, which provides partial future information that can reduce prediction ambiguity. Building on distributional diffusion models, we learn the conditional distributions of these tasks with a single stochastic predictor by minimizing a proper scoring rule across noise levels. At inference, the predictor can still generate each ensemble member in a single forward pass at the fully corrupted endpoint. Standard CRPS training is recovered as the endpoint-only special case of our formulation, so our framework extends existing CRPS-based forecasters with only additional conditioning inputs. Controlled experiments show that the auxiliary tasks improve one-step forecasting across architectures, with larger gains at longer forecast horizons. The gains extend to high-dimensional global weather forecasting under both training from scratch and fine-tuning, along with improved calibration and potential benefits for generalization under distribution shift.</summary>
  </entry>
  <entry>
    <title>Methodological Changes to the Attention ResUNet Hourly Precipitation Postprocessor</title>
    <link href="https://arxiv.org/abs/2609.38609"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.38609"/>
    <id>https://arxiv.org/abs/2609.38609</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Thomas M. Hamill</name></author>
    <category term="Other"/>
    <summary>This note is a technical companion to a previously published preprint describing an Attention Residual U-Net that postprocesses deterministic forecasts from The Weather Company's Global and Regional Atmospheric Forecast (GRAF) model into probabilistic hourly precipitation forecasts. It documents what has changed in that method since publication. Feature-wise Linear Modulation conditioning on calendar season and forecast lead time is used to produce a single trained model for each season, replacing 192 separately trained per-month, per-lead checkpoints. Lead time is extended from 48 to 72 h. Two new input channels are used, per-pixel local solar hour and a static, monthly-varying precipitation climatology. During verification, the climatological reference against which the Brier Skill Score is computed now has an added diurnal dimension, on top of the monthly resolution it already had. Brier Skill Score and reliability are compared between the new vs. the previous training. Forecasts generated with the new training show a modest, consistent improvement of the current training over the original.</summary>
  </entry>
  <entry>
    <title>An Input-Frugal Deep Learning Framework for Weather-Driven National Crop-Yield Forecasting: A Case Study of Brazilian Soybean</title>
    <link href="https://arxiv.org/abs/2609.38447"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.38447"/>
    <id>https://arxiv.org/abs/2609.38447</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Fernando Dupin da Cunha Mello, Prashant Kumar, Erick G. Sperandio Nascimento</name></author>
    <category term="Global Models"/>
    <summary>Reliable, timely crop-yield forecasts are essential for market stability and risk management, yet many approaches rely on costly or hard-to-scale inputs. We present a frugal, transferable, and architecture-agnostic deep learning framework that uses routine weather as the only time-varying input plus two lightweight static context inputs (crop year and an agro-environmental label) to capture long-run change and regional heterogeneity, while supporting multiple sequence encoders under identical data requirements. Using a 20-season Brazilian soybean case study (2001/02-2020/21) with leave-one-year-out cross-validation, we benchmark MLP, CNN, LSTM, CNN-LSTM, a Transformer encoder and the Mamba state-space model against linear ridge regression and a five-year moving-average "farmer" baseline. All deep learning variants outperform ridge, and all sequential encoders surpass the non-sequential MLP. The Transformer achieves the best national accuracy (RMSE 149 kg ha^-1; rRMSE 5.3%; R^2 = 0.784), reducing error by 47.6% relative to the farmer baseline. In-season forecasts improve monotonically from early- to late-season issuance, reaching approximately 50% lower error than the baseline at the latest forecast point. Ablations indicate that the agro-environmental label and spatial instance expansion (multiple grid-node weather sequences per municipality-year) contribute positively without increasing input complexity. SHAP diagnostics suggest crop year explains most of the long-run trajectory, whereas within-season weather and agro-environmental context primarily drive interannual deviations, with moisture/cloud and thermal-demand variables dominating. Overall, the framework is straightforward to deploy across other crops and geographic regions and is naturally compatible with operational weather forecasts for routine monitoring.</summary>
  </entry>
  <entry>
    <title>Physics-Guided Flow-Map Matching for Precipitation Nowcasting</title>
    <link href="https://arxiv.org/abs/2609.37487"/>
    <link rel="related" href="https://weatherml.github.io/papers/nowcasting/#2609.37487"/>
    <id>https://arxiv.org/abs/2609.37487</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Shunya Nagashima, Takumi Bannai, Makoto Misaizu, Keisuke Maeda, Takahiro Ogawa, Miki Haseyama</name></author>
    <category term="Nowcasting"/>
    <summary>Precipitation nowcasting, generating future radar fields from past observations, is critical for flood warning and disaster response. It is also a demanding benchmark for spatiotemporal generative modeling, with chaotic dynamics, heavy-tailed intensities, and rare high-intensity structures that matter most. Deterministic models minimize a pixel loss and are driven toward the conditional mean, which blurs exactly those structures, while generative models that add a stochastic residual on top of a deterministic backbone inherit the same blur. We propose Physics-Guided Flow-Map Matching (PG-FMM), a conditional flow-map model that decouples predictable advection from uncertain small-scale detail. A frozen Lagrangian advection prior transports the radar field and supplies an explicit motion forecast, and a flow-map generative head, conditioned on the past frames and the prior rollout rather than summed onto it, produces sharp stochastic detail in four sampling steps. The prior serves only as guidance, so the head replaces blurred structure instead of inheriting it. Extensive experiments on four radar benchmarks show that PG-FMM outperforms state-of-the-art methods on 18 of 24 metrics, with the largest gains at heavy-rain thresholds, where the critical success index improves by up to 58.9%. The project page can be found at https://neurogica.github.io/PG-FMM.</summary>
  </entry>
  <entry>
    <title>NowcastDiT: Diffusion Transformers are Effective Precipitation Nowcasters</title>
    <link href="https://arxiv.org/abs/2609.37038"/>
    <link rel="related" href="https://weatherml.github.io/papers/nowcasting/#2609.37038"/>
    <id>https://arxiv.org/abs/2609.37038</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Haoran Xu, Xingzhuo Guo, Yuchen Zhang, Jincheng Zhong, Jianmin Wang, Mingsheng Long</name></author>
    <category term="Nowcasting"/>
    <summary>Precipitation nowcasting demands accurate short-term forecasts under strong spatiotemporal variability. Diffusion models are well suited to modeling complex precipitation distributions, yet existing approaches often introduce increasingly specialized designs, leaving the capability of a standard diffusion architecture underexplored. We show that a standard Diffusion Transformer already provides a simple and scalable foundation for precipitation nowcasting, with domain-specific requirements accommodated naturally within its design space. Based on this principle, we develop NowcastDiT and instantiate this flexibility through two complementary adaptations: a dynamics-aware noise prior for temporally coherent forecasts, and end-to-end reinforcement learning with timestep-aware rewards for meteorological skill. Experiments on SEVIR and MRMS benchmarks show that NowcastDiT achieves state-of-the-art performance in both perceptual quality and meteorological skill. These results suggest that standard DiT can serve as an effective foundation for precipitation nowcasting.</summary>
  </entry>
  <entry>
    <title>A neural network-based Universal Thermal Climate Index for reliable global thermal-stress classification across extreme weather</title>
    <link href="https://arxiv.org/abs/2609.35949"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.35949"/>
    <id>https://arxiv.org/abs/2609.35949</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Bikem Pastine, Milan Klöwer, Tianning Tang, Sarah Wilson Kemsley, Louise Slater</name></author>
    <category term="Global Models"/>
    <summary>Extreme temperatures are the leading cause of climate-related mortality world-wide. Climate-health research and operational weather forecasting require accurate estimates of human thermal stress. The Universal Thermal Climate Index (UTCI) is among the most sophisticated and widely used feels-like temperature metrics. However, its ubiquitous polynomial approximation does not generalize well to extreme weather conditions. Here, we introduce Neural-UTCI, a neural network that calculates UTCI with substantially higher accuracy across global conditions at a lower computational cost for operational use. Neural-UTCI reduces the polynomial approximation RMSE from 2.78°C to 0.36 °C, an 87% improvement, and lowers thermal stress misclassification rates from 5.3% to 1.7%, with consistent performance across resampling experiments. These differences affect thermal exposure metrics. For example, during the 2003 European heatwave summer in Rome, Italy, the number of very strong heat stress days increases from 15 to 35 days when using Neural-UTCI compared to operational products like ERA5-HEAT. Simultaneously, Neural-UTCI reliably classifies extreme cold stress conditions, allowing continuous global application. By improving UTCI accuracy, Neural-UTCI can strengthen climate-health risk assessments and public weather warning systems, especially as global warming increases the incidence of extreme events.</summary>
  </entry>
  <entry>
    <title>Explainable Deep Learning for Probabilistic Nowcasting of Radar Reflectivity in Tornadic Storms</title>
    <link href="https://arxiv.org/abs/2609.35675"/>
    <link rel="related" href="https://weatherml.github.io/papers/nowcasting/#2609.35675"/>
    <id>https://arxiv.org/abs/2609.35675</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Nathan Erickson, Amy McGovern, Aaron Hill</name></author>
    <category term="Nowcasting"/>
    <summary>Tornadoes pose substantial risk to human life and property in the United States, causing more than 50 fatalities and \$100 million of property damage on average annually. When tornadoes are likely, weather radar provides critical information for forecasters by providing information on storm morphology, storm motion, and intensity trends. Additional tools such as satellite and numerical weather prediction model runs can provide useful short-term information for understanding changes in storm characteristics. This work demonstrates a U-Net deep-learning system for nowcasting the evolution of radar reflectivity following tornadogenesis, which can provide value to forecasters by synthesizing large amounts of input data (e.g., radar imagery, near-storm environment data) and generating predictions of radar reflectivity from its inputs. Inputs to the model are radar imagery from the Multi-Radar Multi-Sensor (MRMS) dataset and near-storm environment data from the High-Resolution Rapid Refresh (HRRR) numerical weather prediction model. The U-Net is trained on a dataset of tornadic storms to produce 30 minutes of probabilistic predictions of radar reflectivity following tornadogenesis, with probabilistic predictions obtained by predicting parameters of the SinhArcSinh, or SHASH, distribution. The model produces physically realistic predictions of radar evolution, achieves comparable skill to next-hour forecasts from the HRRR, demonstrates reasonable probabilistic calibration and is accompanied by a variety of explainability methods to improve understanding by end users. Additionally, predictions from the model can be obtained much more quickly than those from a numerical weather prediction model. With further development, this model could be extended to nowcast radar reflectivity evolution in an operational setting.</summary>
  </entry>
  <entry>
    <title>Safe Greenhouse Climate Control Using Lagrangian-Constrained PPO with Kolmogorov-Arnold Networks</title>
    <link href="https://arxiv.org/abs/2609.34966"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.34966"/>
    <id>https://arxiv.org/abs/2609.34966</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Hangzun Liu, Yuling Fan, Fang Tian, Zhilong Bie, Zaiwen Feng, Yongliang Qiao</name></author>
    <category term="Climate Modeling"/>
    <summary>Greenhouse climate control balances economic return with maintaining temperature, humidity and CO2 within crop-adapted growth ranges. Conventional reinforcement learning (RL) greenhouse controllers use fixed reward penalties to limit climate constraint violations, yet such heuristic penalties cannot explicitly constrain long-term cumulative violations. Poorly tuned weights either lead to overly conservative policies and lower yields, or fail to suppress persistent climate deviations that harm photosynthesis and induce crop diseases. To address this issue, we formulate greenhouse climate regulation as a Constrained Markov Decision Process (CMDP) and use a Lagrangian safe RL framework RCPO-PPO to separate economic optimization and cumulative safety constraints, enabling adaptive penalty adjustment without manual tuning. To handle strong nonlinear, time-varying coupling between greenhouse microclimate and crop growth, Kolmogorov-Arnold Networks (KANs) replace Multi-Layer Perceptrons (MLPs) as policy and value approximators for improved nonlinear representation. Sinusoidal cyclic time features are embedded in observations to capture diurnal environmental periodicity. Simulations use a classic winter lettuce greenhouse model driven by 40-day real weather disturbances. Compared with vanilla penalty-based PPO, our method cuts cumulative climate violations by 18.65% and raises lettuce economic profit by 2.91%, keeping violations stable near the safety threshold. This decoupled CMDP optimization with KAN-based policy representation mitigates long-term climate risks and boosts planting profits, offering a constraint-aware control strategy for precision greenhouse cultivation.</summary>
  </entry>
  <entry>
    <title>MW-Nowcast: Six-hour ensemble nowcasting of extreme precipitation</title>
    <link href="https://arxiv.org/abs/2609.34836"/>
    <link rel="related" href="https://weatherml.github.io/papers/nowcasting/#2609.34836"/>
    <id>https://arxiv.org/abs/2609.34836</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Ning Wang, Zuliang Fang, Weixin Jin, Zhongjian Lv, Shuang Qin, Pengcheng Zhao, Siqi Xiang, Jiang Bian, Haoyi Xiong, Nan Guan, Bin Zhang, Liangjie Zhang, Denvy Deng, Qi Zhang, Matt Corey, Jitu Keshri, Sridhar Iyer, Hongyu Sun, Kit Thambiratnam, Jonathan Weyn, Richard E. Turner, Haiyu Dong</name></author>
    <category term="Nowcasting"/>
    <summary>Extending reliable nowcasting of extreme precipitation could provide critical additional time for warnings and emergency response during high-impact events such as flash floods. Radar-based generative machine-learning models have enabled skilful hyperlocal precipitation nowcasting, but accurate prediction of intense precipitation remains confined to the first few hours. Because storm-scale structure is predictable for longer than individual cells, a natural strategy is to predict that structure while generatively modelling only the uncertain local growth, decay, reorganisation and initiation of storms. Here we present Microsoft Weather Nowcast (MW-Nowcast), a six-hour ensemble radar nowcasting model that jointly learns a deterministic predictor to capture organised precipitation structure shared across ensemble members, and a generator to produce diverse local residuals around this shared prediction. Across independent test data from the United States, Europe and China, MW-Nowcast achieves higher detection skill than leading methods for heavy and extreme precipitation throughout the 6 h horizon. For the most intense rainfall, MW-Nowcast doubles the available warning time across all three regions, delivering 6 h forecasts with skill previously limited to 3 h for the leading generative baseline. A cost-loss decision analysis shows that MW-Nowcast retains substantial value for a broad range of applications even at 4-6 h, where alternative methods offer little benefit. These additional hours can give forecasters and emergency managers the time to warn and act before extreme rainfall strikes, helping to protect lives and property.</summary>
  </entry>
  <entry>
    <title>Predicting Delayed Train Trajectories on the Dutch Railway Network: Explainable AI Evaluation of Topological, Operational and Weather Features with Tree Based Ensemble Methods</title>
    <link href="https://arxiv.org/abs/2609.34692"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.34692"/>
    <id>https://arxiv.org/abs/2609.34692</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jia Long Bao, Ali Mohammed Mansoor Alsahag, Seyed Sahand Mohammadi Ziabari</name></author>
    <category term="Other"/>
    <summary>The reliable prediction of passenger train delays is a critical component of railway management. While contemporary research frequently attempts to maximize absolute accuracy by deploying opaque deep learning architectures, the underlying data mechanics driving longitudinal predictive decay remain underexplored. Consequently, this study provides an explainable temporal robustness analysis of network-wide railway delay prediction. Focusing on the Dutch railway network, this research utilizes interpretable tree-based ensembles to integrate granular topological, environmental, and operational features. The overarching finding establishes that while feature-rich tree-based models improve simultaneous (within-month) prediction, predictive performance systematically degrades when evaluated across non-simultaneous (future) months. Furthermore, multi-horizon SHAP and dispersion analyses explicitly link this degradation to environmental feature volatility and instability within the statistical target definition. Ultimately, this thesis demonstrates that richer feature sets alone are insufficient to resolve long-term forecasting constraints, underscoring the necessity to transition toward dynamic, season-aware architectures anchored by absolute operational boundaries.</summary>
  </entry>
  <entry>
    <title>Low latency global carbon budget reveals strong land sink recovery in 2025</title>
    <link href="https://arxiv.org/abs/2609.34226"/>
    <link rel="related" href="https://weatherml.github.io/papers/remote-sensing/#2609.34226"/>
    <id>https://arxiv.org/abs/2609.34226</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Philippe Ciais, Piyu Ke, Xiangjun Tian, Stephen Sitch, Wei Li, Xiaomeng Du, Xiaofan Gui, Ben Poulter, Thomas Colligan, Auke M. van der Woude, Anne-Wil van den Berg, Wouter Peters, Zhu Liu, Zhu Deng, Zhe Jin, Yilong Wang, Junjie Liu, Sudhanshu Pandey, Chris O'Dell, Jiang Bian, John Miller, Xin Lan, Jefferson Goncalves De Souza, Michael O'Sullivan, Pierre Friedlingstein, Youngryel Ryu, Helin Zhang, Julien Alléon, Yi Xi, Daniel S. Goll, Lei Zhu, Guido R. van der Werf, Yitong Yao, Shilong Piao, Frédéric Chevallier</name></author>
    <category term="Remote Sensing"/>
    <summary>The atmospheric CO2 growth rate fell sharply in 2025, from a record 3.76 $\pm$ 0.09 ppm yr-1 in 2024 to 2.06 $\pm$ 0.09 ppm yr-1 (NOAA marine boundary layer observations), below the 2015-2022 mean of 2.47 ppm yr-1, even as fossil CO2 emissions rose by 0.7% to 10.38 GtC yr-1. Here we present a low-latency global and regional carbon budget for 2025, combining three dynamic global vegetation models (DGVMs) and ocean model emulators with four atmospheric inversions constrained by OCO-2 satellite retrievals. The global net land sink reached 2.36 $\pm$ 0.16 GtC yr-1 in 2025 (DGVMs: 2.04 $\pm$ 0.24; inversions: 2.68 $\pm$ 0.20 GtC yr-1), strengthening by 2.81 $\pm$ 0.31 GtC yr-1 from 2024 and exceeding the 2015-2022 mean by 0.71 $\pm$ 0.13 GtC yr-1. Ocean uptake (3.11 $\pm$ 0.36 GtC yr-1) remained similar to 2024, making the land sink rebound the dominant driver of the slowdown in CO2 growth. Tropical lands shifted from net sources in 2024 to net sinks in 2025, with enhanced uptake across much of Africa and northern Eurasia, and land flux anomalies covaried with GRACE terrestrial water storage. Where the sink had weakened substantially in 2023-2024, about 80% of the area showed some recovery, with overall recovery of 87.3% (DGVMs) to 99.5% (inversions). Recovery exceeded 100% in the tropics but remained incomplete in the northern extratropics, indicating a strong but spatially uneven rebound of the land carbon sink.</summary>
  </entry>
  <entry>
    <title>StatD2GAN: When Calibration Masks Generator Quality in Held-Out Evaluation of Synthetic Weather Sequences</title>
    <link href="https://arxiv.org/abs/2609.33761"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.33761"/>
    <id>https://arxiv.org/abs/2609.33761</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Mustafa Ozaytac, Ozge Karadag Atas</name></author>
    <category term="Other"/>
    <summary>Generative models for multivariate weather series are routinely evaluated with pooled distributional metrics computed after marginal calibration. We show this practice can invalidate architectural conclusions, and rebuild the evaluation of StatD2GAN, a three-discriminator GAN with evolutionary weight adaptation, around a held-out protocol: the final two calendar years of each dataset are held out behind a 168 hour embargo, calibration is fitted on the training block only, and all metrics are computed on the held-out block. Evidence comes from 25 matched (location, seed) pairs across five Koppen-Geiger climates, tested with Wilcoxon signed-rank tests under Holm correction. Four results follow. First, isotonic calibration drives the Kolmogorov-Smirnov distance to within 2% of a per-location noise-and-shift floor for every architecture tested, including a deliberately weak RCGAN baseline, so calibrated marginal metrics cannot discriminate between architectures. Second, the sorted-representation discriminator is the only component whose removal significantly degrades cross-variable dependence (Kendall tau MAE +0.080, Holm p = 0.009), with a regime-dependent effect: near zero in Ankara, above 115% in Dubai and Yakutsk. A rank-transformed variant isolates the mechanism as quantile supervision of the marginals rather than copula matching. Third, physical constraint violations are injected by calibration, not the generator; projection removes them at negligible cost (deltaKS &lt;= 0.003). Fourth, pooled metrics conceal a collapse of between-sequence weekly-mean variability, a proxy for seasonal and regime diversity, in TimeGAN that only sequence-level statistics expose. We recommend floor-referenced marginal evaluation, matched-pair testing, and sequence-level variance decomposition as minimum requirements for calibrated generative pipelines.</summary>
  </entry>
  <entry>
    <title>Suitable Measures for the Potential Operational Utility of AI NWP Rainfall Forecasts Over Africa</title>
    <link href="https://arxiv.org/abs/2609.31775"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.31775"/>
    <id>https://arxiv.org/abs/2609.31775</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Shruti Nath, Docko Sow, Koomi Toussaint Amoussouvi, Fenwick Cooper, Josiah Kiarie Kimani, John Bagiliko, Florian Pappenberger, Rendani Mbuvha</name></author>
    <category term="Global Models"/>
    <summary>Artificial intelligence (AI)-based weather prediction is approaching the skill of physical numerical weather prediction (NWP) systems at a fraction of the computational cost. This is particularly promising for Africa, where rainfall extremes are intensifying and many forecasting centres lack the infrastructure to run physical models at extended lead times. We present a calibrated comparison of GraphCast, GenCast and the Functional Generative Network (FGN) against the physical NWP model IFS for rainfall prediction across Africa. Deterministic and probabilistic forecasts are postprocessed using Isotonic Distributional Regression and evaluated with the Continuous Ranked Probability Score against IMERG, RFEv2 and CHIRPS across seasons, wet and dry regimes, elevation zones and lead times. All models retain skill beyond climatology across most seasons and at extended lead times. AI models generally outperform IFS in wet regions, whereas IFS performs better in dry, high-elevation areas, where its finer resolution better represents orographic controls on rainfall. Across observational datasets and seasons, AI models achieve a median improvement of approximately 5% over IFS. GraphCast achieves calibrated skill comparable to the ensemble-based FGN, although FGN provides greater significant skill at longer lead times. These results highlight the potential of calibrated AI weather prediction to provide accessible and computationally efficient rainfall forecasts across Africa, while demonstrating the continuing importance of spatial resolution, ensemble design and regional characteristics.</summary>
  </entry>
  <entry>
    <title>Learning Hierarchical Causal Representations of the Effects of Forcings on Temperature in Climate Models</title>
    <link href="https://arxiv.org/abs/2609.30995"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.30995"/>
    <id>https://arxiv.org/abs/2609.30995</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Shan Zhao, Ilija Trajkovic, Julia Kaltenborn, Yaniv Gurwicz, Peer Nowack, David Rolnick, Julien Boussard</name></author>
    <category term="Climate Modeling"/>
    <summary>Machine learning (ML) emulators provide a fast and cost-effective method to simulate climate change scenarios after being trained on Earth System Models projections. However, the black-box nature of those data-driven approaches limit the usability and trustworthiness of their outputs and in particular their use as causal attribution tools. Here, we develop a hierarchical causal representation learning framework applied to sea surface temperature fields from a state-of-the-art global climate model. As a key advance over previous work, our framework explicitly models both atmospheric dynamical interactions arising from internal climate variability and forced responses due to changes in atmospheric greenhouse gas and aerosol concentrations. When trained on future climate change scenarios, our method accurately predicts the long-term global mean and regional temperature evolution and shows physically realistic responses to perturbations in greenhouse gas and aerosol concentrations when evaluated on unseen scenarios. Our results underline the potential of causal representation learning frameworks for advancing climate model emulation.</summary>
  </entry>
  <entry>
    <title>Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events</title>
    <link href="https://arxiv.org/abs/2609.30746"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.30746"/>
    <id>https://arxiv.org/abs/2609.30746</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Isabella S. Thiel, Juan Bello-Rivas, Yannis G. Kevrekidis, Themistoklis P. Sapsis</name></author>
    <category term="Other"/>
    <summary>Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only \(50\) time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with $20$ times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.</summary>
  </entry>
  <entry>
    <title>On the Limits of Univariate Deep Learning for Significant Wave Height Forecasting</title>
    <link href="https://arxiv.org/abs/2609.30688"/>
    <link rel="related" href="https://weatherml.github.io/papers/ocean-sea-ice/#2609.30688"/>
    <id>https://arxiv.org/abs/2609.30688</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Yilin Zhai, Hongyuan Shi, Zaijin You</name></author>
    <category term="Ocean &amp; Sea Ice"/>
    <summary>This study conducts a systematic hyperparameter search across five deep learning architectures, DLinear, LSTM, PatchTST, ResAttLstm, and Mamba2, and nine context lengths (1-168 h) for single-station significant wave height (Hs) forecasting on NDBC buoy 41009, followed by re-evaluation of the best configurations on a 47-buoy, 37-year corpus. The five families converge to a common performance level on the multi-buoy evaluation (between-family SD = 0.0014 m^2, 0.8% of the grand mean), a spread dwarfed by the 4.83x cross-dataset MSE shift between buoy corpora. All multi-buoy trials beat persistence (mean skill +0.062), but no architecture consistently outperforms the others. On the single-buoy experiment, skill peaks at 12-24 h where five trials fall below persistence, per-family Q4/Q3 test MSE ratios range from 2.4 to 2.6, and deep models underperform persistence for the most extreme 1% of waves. These findings are consistent with the interpretation that persistence already captures the dominant linear-inertial signal in univariate Hs, and that architecture engineering under this univariate input setting has reached diminishing returns: cross-buoy variance, not model class, dominates forecast error. Future work should prioritise atmospheric covariates, zero-shot cross-buoy transfer, and decomposition of Hs into swell and wind-sea components. By establishing a rigorous reference baseline for what univariate Hs models can and cannot achieve, this study provides a benchmark against which future multivariate and physics-informed approaches can be calibrated, and offers practical guidance for lightweight buoy-level forecasting in mid-latitude storm-dominated and swell-mixed environments.</summary>
  </entry>
  <entry>
    <title>Understanding Perturbed Parameter Ensemble Sensitivities Using A Contrastive Learning Approach</title>
    <link href="https://arxiv.org/abs/2609.30420"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.30420"/>
    <id>https://arxiv.org/abs/2609.30420</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Da Fan, David John Gagne, Gregory S Elsaesser, Brian Medeiros, Addisu G Semie, Qingyuan Yang, Akila Sampath, Subashree Venkatasubramanian</name></author>
    <category term="Climate Modeling"/>
    <summary>Perturbed parameter ensembles (PPEs) reveal how physics parameters affect climate simulations, but interpreting parameter sensitivities across multivariate, spatially structured outputs remains challenging, particularly when calibrating models against observations. We develop an explainable contrastive learning model that maps 5 monthly cloud and radiation fields into a shared representation space. We train the model on the fields of two 100-member Community Atmosphere Model version 6 (CAM6) PPEs, spanning 34 parameters, that only differ in the warm rain microphysics scheme: KK2000, the default bulk microphysics scheme, and TAU-ML, a neural network emulator of a bin microphysics scheme. The learned representations separates two PPEs with over 94\% linear classification accuracy while preserving the seasonal variability and ensemble spread due to parameter perturbations. In the shared representation space, the representations of satellite observations occupy the same low-dimensional manifold as the PPEs but are displaced from them most strongly during boreal spring and autumn. TAU-ML PPE has a lower distance to observations compared to KK2000 in the representation space. Integrated Gradients attributions highlights the contributions in subtropical low-cloud regions, Northern and Southern Hemisphere storm track regions, and tropical convection regions to differences between PPEs and observations. Regional attributions correlate most strongly with parameters associated with cloud microphysics, boundary layer turbulence, and deep convection. These results demonstrate that explainable representations of climate fields can attribute model differences to specific variables, regions, seasons, and physical parameters.</summary>
  </entry>
  <entry>
    <title>Lightweight Probabilistic Downscaling from a Deterministic Base Model</title>
    <link href="https://arxiv.org/abs/2609.29383"/>
    <link rel="related" href="https://weatherml.github.io/papers/downscaling/#2609.29383"/>
    <id>https://arxiv.org/abs/2609.29383</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Joseph McLean, Tiffany Vlaar, Sigrid Passano Hellan, Linus Ericsson</name></author>
    <category term="Downscaling"/>
    <summary>Climate data downscaling is the task of increasing the spatial resolution of climate data, typically by generating fine-resolution regional climate data from coarse global model output. Recent machine learning (ML) work in the related task of weather forecasting has seen significant improvements due to newly devised training methods and architectural components, but these have not yet benefited downscaling. We adapt two of these methods to create a family of lightweight probabilistic ML downscaling models built on a modified U-Net backbone and evaluate them on the CORDEX-ML-Bench suite for daily maximum temperature and precipitation across three geographic regions: the Alps, New Zealand and South Africa. We find that a two-stage training curriculum, combining deterministic pretraining with probabilistic tuning, transfers well to downscaling, beating the state-of-the-art for RMSE. Our work provides an advancement towards lightweight, probabilistic downscaling models, reducing the current trade-off between computational intensity and distributional fit.</summary>
  </entry>
  <entry>
    <title>Generative Atmospheric Super-Resolution from Heterogeneous In Situ Observations through Composable Interfaces</title>
    <link href="https://arxiv.org/abs/2609.29027"/>
    <link rel="related" href="https://weatherml.github.io/papers/downscaling/#2609.29027"/>
    <id>https://arxiv.org/abs/2609.29027</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Yang Xu, Dibyajyoti Chakraborty, Haiwen Guan, Sen Wang, Romit Maulik</name></author>
    <category term="Downscaling"/>
    <summary>Atmospheric observations are sparse, heterogeneous, and unevenly distributed, whereas many generative atmospheric models learn distributions over regularly gridded multivariate states. Once pretrained, diffusion models can supply atmospheric priors that can be combined with observation-derived likelihood factors in a Bayesian formulation. However, these observation sources differ substantially in geometry and sampling density, complicating the consistent use of their observations within a common inference framework. Here, we formulate this reconstruction problem as generative atmospheric super-resolution and introduce composable observation interfaces for conditioning a single pretrained 13-variable atmospheric diffusion model. The interfaces convert sparse radiosonde (R), clustered aircraft (A), and dense irregular surface-station (S) observations into source-specific likelihood factors that specify where observations constrain the gridded state, how residuals are counted under uneven sampling, and how strongly each source guides posterior sampling. We developed the aircraft and surface observation interfaces using 2019 observations and evaluated the selected interfaces throughout 2020 without further tuning. Compared with reconstructions conditioned only on radiosonde observations, the composed R+A+S interface reduces RMSE evaluated against ERA5 by $9.24\%$ across all 13 state variables over the CONUS domain. The aircraft and surface factors provide complementary improvements in upper-air and surface variables. The R+A+S combination also lowers the Continuous Ranked Probability Score (CRPS), while evaluations at held-out aircraft and surface-station observations show reduced prediction errors. Together, these results demonstrate a modular route for conditioning a pretrained atmospheric generative prior on heterogeneous in situ observations without retraining the underlying model.</summary>
  </entry>
  <entry>
    <title>Evaluating Cross-region Generalization for Wavelet-Diffusion Precipitation Downscaling</title>
    <link href="https://arxiv.org/abs/2609.28749"/>
    <link rel="related" href="https://weatherml.github.io/papers/downscaling/#2609.28749"/>
    <id>https://arxiv.org/abs/2609.28749</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Weikang Qian, Yixin Wen, Chugang Yi, Zhi Li, Lingcheng Li, Haizhao Yang</name></author>
    <category term="Downscaling"/>
    <summary>Diffusion models have shown strong potential for kilometer-scale precipitation downscaling, but their performance in geographically unseen regions and event regimes remains insufficiently understood. Building on the wavelet diffusion model (WDM) framework, this study evaluates cross-region and cross-event generalization. Six 3 x 3 deg U.S. regions represent convective, winter, tropical, and atmospheric-river precipitation regimes. Low-resolution inputs are generated by block averaging NOAA Multi-Radar/Multi-Sensor (MRMS) composite reflectivity fields. A WDM trained only on Oklahoma (OK) samples and a WDM trained on all six regions are compared with nearest-neighbor and Bicubic interpolation. Model performance is evaluated using three metric families that measure image-domain reconstruction, spectral and distributional fidelity, and bin-wise precipitation detection. The OK-trained WDM remains competitive outside OK. Although the all-region WDM delivers the best and most consistent overall image-domain and detection performance, its gains are uneven across precipitation intensities. Bin-wise critical success index (CSI) over 5-dBZ reflectivity bins shows that WDM improvements concentrate in localized higher-reflectivity structures, which image-domain metrics partly obscure. In addition, the performance differences among samples are strongly associated with the spatial organization of the precipitation field, quantified by Moran's I as the spatial autocorrelation of each reflectivity bin. The sample-level Moran's I-CSI correlation stratified by sample intensity reaches 0.901 in all six regions, including regions unseen during training. Overall, these findings support future efforts to transfer downscaling models to regions with limited local training data and to generate globally consistent, high-resolution precipitation products.</summary>
  </entry>
  <entry>
    <title>HClimRep-Ocean: A Global Ocean Emulator on an Unstructured Mesh</title>
    <link href="https://arxiv.org/abs/2609.28601"/>
    <link rel="related" href="https://weatherml.github.io/papers/ocean-sea-ice/#2609.28601"/>
    <id>https://arxiv.org/abs/2609.28601</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Kacper Nowak, Aleksei Koldunov, Nikolay Koldunov, Savvas Melidonis, Ankit Patnala, Simon Grasse, Julius Polz, Christian Lessig, Martin Schultz, Thomas Jung</name></author>
    <category term="Ocean &amp; Sea Ice"/>
    <summary>Machine-learning (ML) emulators for atmospheric processes have advanced rapidly in recent years, transforming weather forecasting. Although early ML ocean forecasting models now exist, they remain less developed than their atmospheric counterparts. Unlike the atmosphere, much of the ocean's kinetic energy resides in mesoscale eddies whose characteristic spatial scales are approximately an order of magnitude smaller than those of comparable atmospheric features. Moreover, complex coastlines, narrow straits, and ice-covered seas make boundary representation a central challenge that atmospheric models do not face. Consequently, numerical ocean simulations commonly use locally refined or even completely unstructured meshes. However, their data-driven counterparts have so far been built around latitude-longitude grids. We present HClimRep-Ocean, an ocean emulator that operates directly on the native unstructured mesh of FESOM2. The emulator is trained on a 209-year AWI-CM3 control integration and is run without atmospheric forcing, receiving the atmospheric state only at initialisation time, which isolates the predictability carried by the ocean state itself. Skill is strongly field-dependent: for currents, HClimRep-Ocean outperforms every reference at 30 day forecast, whereas for temperature and salinity a damped-anomaly persistence forecast remains the more accurate estimator. This behaviour is physically interpretable: current variability is largely geostrophic and internally generated, whereas sea-surface temperature and salinity fluctuations are driven by atmospheric forcing through weather state. Evaluated independently on the OceanBench benchmark, a reanalysis-trained variant of HClimRep-Ocean achieves the lowest RMSE against GLORYS reanalysis among all assessed systems, confirming the competitiveness of the native-mesh approach.</summary>
  </entry>
  <entry>
    <title>PISCES: Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather Anomaly Detection and Early Warning</title>
    <link href="https://arxiv.org/abs/2609.28022"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.28022"/>
    <id>https://arxiv.org/abs/2609.28022</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Kevin Lee, Alison J. March</name></author>
    <category term="Other"/>
    <summary>Space weather early warning depends on detecting solar wind transients in in-situ measurements at the first Sun-Earth Lagrange point (L1), before they reach Earth. Fixed thresholds can miss combined magnetic and plasma structure, and many learning methods provide a single anomaly score. We present the Physics-Informed Solar-wind Convolutional autoEncoder for Space-weather (PISCES), a convolutional autoencoder trained without catalog labels on OMNI solar wind measurements under physics constraints. Its loss includes magnetic field consistency, an empirical relation between temperature and velocity, the Parker spiral angle, and penalties on changes between consecutive one-minute samples in derived quantities calculated from the reconstruction. At inference, PISCES separates the anomaly score into magnetic and plasma reconstruction errors, physics relations, and residual corrections, and reports the magnitude of each contribution. Attenuation of the skip connections, selected on validation data, improves average precision for the trained models, while the untrained scores remain nearly the same. The trained models also give a more consistent ordering of these physical contributions. After smoothing with a trailing median, the alarms can precede independently observed sudden commencements, including positive sudden impulses.</summary>
  </entry>
  <entry>
    <title>Sparse-Observation Atmospheric Thermal Forecasting with Physics-Informed Neural Networks for Climate-Aware Digital Twins</title>
    <link href="https://arxiv.org/abs/2609.27290"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.27290"/>
    <id>https://arxiv.org/abs/2609.27290</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Tannaz Goodarzvand Chegini, Elyas Shivanian, Behzad Karimi, Faraz Dadgostari</name></author>
    <category term="Other"/>
    <summary>Short-horizon forecasts of atmospheric temperature are needed to support climate-aware digital-twin systems, but such forecasts must be produced where thermal observations are incomplete. This study evaluates a physics-informed neural network for potential-temperature forecasting, constrained by a pressure-coordinate thermodynamic advection-source equation and a diabatic-source closure fit from the preceding 12-hour period and frozen before future-time training. Using hourly ERA5 reanalysis at three pressure levels, the model is evaluated as a conditional hindcast at lead times of one, two and three hours against persistence, local-trend, and two matched neural-network baselines, one of which receives the same future meteorological forcing as the PINN, helping distinguish the physical constraint from access to future forcing. In an Oklahoma development case, mean RMSE improvement over the strongest baseline grew from 8.1\% at one hour to 23.8\% at three hours; under an observation-density sweep down to 5\% of candidate locations, this 3-hour advantage remained 14.6--16.9\%, with no evidence that lower density improves performance. Under a fixed protocol transferred to an Alabama heat event with three virtual-observation layouts, three-hour improvement ranged 19.7-24.4\% with consistent origin-level wins. A parallel Montana stress test, in which fixed pressure levels intersected complex terrain, produced a three-hour degradation of roughly 17.5\%, identifying a terrain-related applicability limit of the formulation. Together, these results indicate that the physics constraint's benefit grows with forecast horizon, persists under severe observation sparsity, and transfers across regions, but is bounded by the validity of a fixed vertical-coordinate representation over complex terrain, evidence relevant to physics-constrained components of climate-aware forecasting and digital-twin systems.</summary>
  </entry>
  <entry>
    <title>Analysis of trade-offs in urban heat mitigation using a Bayesian Optimization framework for an urban canopy layer model</title>
    <link href="https://arxiv.org/abs/2609.25953"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.25953"/>
    <id>https://arxiv.org/abs/2609.25953</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Rebekka Walter, Johanna Gelhaus, David Anton, Henning Wessles, Stephan Weber</name></author>
    <category term="Climate Modeling"/>
    <summary>To mitigate the challenges of climate change and intensifying heat stress in urban areas, local adaptation strategies are discussed and introduced in cities worldwide. To understand processes and potential trade-offs of these strategies a Bayesian optimization and surrogate modeling framework was employed to investigate urban parameter ranges of heat mitigation strategies with focus on three thermal metrics: daytime air temperature, Universal Thermal Climate Index (UTCI), and nighttime air temperature. Based on an urban street canyon configuration, it was shown that heat mitigation measures that reduce daytime air temperature and UTCI are often associated with higher nighttime temperatures. This results in a curved Pareto front that reflects the trade-off between daytime and nighttime thermal comfort. Under identical forcing conditions, different urban configurations, varying in geometry, vegetation, and surface characteristics, are shown to alter peak canyon air temperature by up to 5.2~$^\circ$C during the day and 2.6~$^\circ$C at night, while UTCI varies by up to 7.9~$^\circ$C, demonstrating that favorable urban configurations can substantially mitigate microclimatic heat stress. These findings suggest that combining multi-objective Bayesian optimization and surrogate modeling can help bridge the gap between computationally intensive climate simulations and practical decision-making in urban planning. An interactive visualization tool was developed to explore these trade-offs, making the often opposing relationships between urban parameters and the three thermal metrics directly accessible to planners.</summary>
  </entry>
  <entry>
    <title>A dataset of one-dimensional idealized probabilistic fields</title>
    <link href="https://arxiv.org/abs/2609.25720"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.25720"/>
    <id>https://arxiv.org/abs/2609.25720</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Gregor Skok, Romain Pic</name></author>
    <category term="Global Models"/>
    <summary>Verification of probabilistic weather forecasts remains a crucial aspect of numerical weather prediction, as new AI-based models become more widely used alongside the more traditional physics-based ensemble forecasting systems that continue to be developed and improved. We present a first-of-its-kind idealized probabilistic dataset composed of one-dimensional cases aimed at analyzing the behavior and properties of verification methods for probabilistic forecasts and comparing their behavior. It covers a wide range of probabilistic cases, such as constant, localized events, gradients, fronts, noisy, bimodal, and limiting cases. Moreover, the code associated with the dataset provides great flexibility for customizing the experiments it covers. The dataset represents the first building block of the more extensive comparison dataset of the Bridging The Gap project, which aims to facilitate the development and comparison of spatial verification methods for probabilistic forecasts.</summary>
  </entry>
  <entry>
    <title>West-WRF AI 2-km: High-Resolution Prediction of Integrated Vapor Transport and Precipitation</title>
    <link href="https://arxiv.org/abs/2609.25512"/>
    <link rel="related" href="https://weatherml.github.io/papers/regional-models/#2609.25512"/>
    <id>https://arxiv.org/abs/2609.25512</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Nazak Rouzegari, Vesta Afzali Gorooh, Agniv Sengupta, Phu Nguyen, Kuo-Lin Hsu, Amir AghaKouchak, Soroosh Sorooshian, F. Martin Ralph, Luca Delle Monache</name></author>
    <category term="Regional Models"/>
    <summary>We introduce a stretched-grid artificial intelligence (AI) weather forecasting model with 2-km resolution over the western United States and part of the Northeast Pacific and approximately 31-km resolution elsewhere globally. Forecasting over the western U.S. is challenging because complex topography and atmospheric rivers (ARs) strongly influence orographic precipitation. West-WRF AI 2-km builds on a global model pretrained with a 40-year European Centre for Medium-Range Weather Forecasts Reanalysis v5 (ERA5) dataset and is fine-tuned with the Center for Western Weather and Water Extremes (CW3E) 2-km regional reanalysis to produce autoregressive 6-hourly forecasts of precipitation and integrated vapor transport (IVT). Forecasts are evaluated over winters 2020-2023 using gridded precipitation observations, rain gauges, and AR Reconnaissance dropsondes and are benchmarked against coarser-resolution AI forecasts and regional and global numerical weather prediction (NWP) systems. West-WRF AI 2-km reproduces observed precipitation-intensity distributions, retains fine-scale spectral variability, and produces sharper narrow coastal precipitation bands and localized, terrain-sensitive extremes. Its broader-scale performance remains comparable to coarser-resolution configurations while preserving large-scale skill despite higher resolution. Dropsonde verification shows lower errors and improved categorical skill at the most extreme IVT threshold. Overall, West-WRF AI 2-km provides its greatest value for localized precipitation extremes and intense AR-related moisture transport.</summary>
  </entry>
  <entry>
    <title>FAST-ML: A Hybrid Physics-Machine Learning Framework for Tropical Cyclone Intensity Forecasting</title>
    <link href="https://arxiv.org/abs/2609.25505"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.25505"/>
    <id>https://arxiv.org/abs/2609.25505</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Shijie Xiao, Jonathan Lin, Thomas Ehrmann, Ali Sarhadi</name></author>
    <category term="Global Models"/>
    <summary>Rapid intensification (RI) remains one of the most consequential and difficult aspects of tropical cyclone (TC) forecasting. Although full-physics numerical weather prediction models can represent the processes governing RI, resolving storm-environment interactions remains computationally expensive, while purely data-driven approaches often lack physical interpretability. We present FAST-ML, a hybrid framework that bridges data-driven efficiency with physical constraints. A physically informed dual-stream neural parameterization ingests 3D ERA5 fields to diagnose ventilation controls---environmental wind shear and mid-level entropy deficit. By optimizing these parameters end-to-end through a differentiable FAST intensity model, this architecture establishes a robust new paradigm for observation-driven parameter optimization, ensuring storm evolution remains strictly governed by thermodynamic principles. By better capturing the storm's continuous intensity evolution, FAST-ML improves upon its physical baseline, reducing ensemble CRPS across forecast lead times, with a reduction of approximately 31% at 60 h and nearly halving the RI false alarm ratio without sacrificing detection skill. In a 100-member ensemble configuration, FAST-ML produces intensity forecasts comparable to FNV3 for selected storms under the evaluated input configurations. Furthermore, zero-shot tests on selected Eastern Pacific storms provide encouraging evidence of cross-basin transferability. FAST-ML provides a modular intensity forecasting framework that can be coupled with externally supplied storm tracks and environmental fields. It demonstrates that observation-driven parameter learning within physically constrained dynamics simultaneously enhances accuracy, interpretability, and computational efficiency.</summary>
  </entry>
  <entry>
    <title>Learning Prognostic Variables for AI Convective Parameterizations via Symbolic Distillation</title>
    <link href="https://arxiv.org/abs/2609.24882"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.24882"/>
    <id>https://arxiv.org/abs/2609.24882</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jurij Schönfeld, Tom Beucler, Julien Savre, Steven Sherwood, Veronika Eyring</name></author>
    <category term="Climate Modeling"/>
    <summary>Hybrid AI-physics climate modeling aims to improve coarse (~100km-resolution) Earth system models by learning to parameterize subgrid processes from high-fidelity data. However, this so far mostly involves local-in-time, diagnostic parameterizations, in which the subgrid state depends only on the current coarse state with no memory of previous states, which is unrealistic for processes such as convection that have intrinsic persistence. To address this, we enhance local-in-time parameterizations by learning prognostic variables that compactly carry important, additional past information where no explicit sub-grid information is available. First we compress past information into a low-dimensional latent space using an autoencoder, which then informs a neural network trained to parameterize targeted subgrid-scale processes. We then replace the autoencoder with symbolic equations that govern the time evolution of the latent variables, yielding additional prognostic memory variables that can be integrated alongside the resolved atmospheric state. We evaluate this approach on two systems: the Lorenz-96 model (online) and surface precipitation from high-resolution atmospheric simulations (offline). A forced multivariate linear ordinary differential equation recovers most of the added value achieved by the autoencoder-based approach in both experiments. Benchmarked against diagnostic parameterizations without memory, our memory-informed approach improves climate statistics and temporal structure, including a realistic diurnal cycle of tropical land precipitation.</summary>
  </entry>
  <entry>
    <title>Inference of Unknown Dynamical Components Using Next Generation Reservoir Computing: From Chaotic Systems to Climate Data</title>
    <link href="https://arxiv.org/abs/2609.24754"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.24754"/>
    <id>https://arxiv.org/abs/2609.24754</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jule Budnick, Andrew Keane, Serhiy Yanchuk</name></author>
    <category term="Other"/>
    <summary>We investigate next generation reservoir computing (NGRC) as a data-driven approach for inferring unseen components of dynamical systems. We compare NGRC with traditional reservoir computing (RC) using the Lorenz and Rössler system, where two unknown components are inferred from one given component. For both systems, NGRC achieves accurate results while requiring fewer training data and less computational time than RC. We identified an inverse proportional behavior between the number of time-delayed steps needed for NGRC and the temporal resolution, indicating that the physical time span covered by the delay interval is an important factor in determining the required number of delayed steps. Finally, we apply NGRC to the observational climate data of ENSO (El Niño--Southern Oscillation) and infer one observable from the remaining variables. Despite the noise and complexity of the real-world data, the NGRC shows promising results. Our findings demonstrate the potential of NGRC for efficient inference of unseen components in both controlled dynamical systems and real-world data.</summary>
  </entry>
  <entry>
    <title>Climate Variability Modulates the Impact of Price Spikes on Food Insecurity</title>
    <link href="https://arxiv.org/abs/2609.24394"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.24394"/>
    <id>https://arxiv.org/abs/2609.24394</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jordi Cerdà-Bautista, Vasileios Sitokonstantinou, Homer Durand, Gherardo Varando, Michele Ronco, Gustau Camps-Valls</name></author>
    <category term="Climate Modeling"/>
    <summary>Climate variability influences whether a market disruption escalates into a food crisis, yet broad climate patterns like El Niño, tracked months before they alter hydro-climatic conditions, are still not incorporated as an early-warning component in food-security responses. We address this gap by introducing sensitivity regimes, a stratification of regions by the direction and strength of their vegetation response to the El Niño Southern Oscillation, and using them to estimate how food price spikes affect acute food insecurity across sub-Saharan Africa. Integrating remote sensing, socioeconomic data, and causal machine learning, we find that in regions where ENSO systematically suppresses vegetation, a price spike raises the share of the population at acute risk by 5.4 percentage points in the following month. In regions where vegetation is unaffected by or positively linked to ENSO, the estimated effect is smaller (around 2 percentage points) and statistically insignificant. These results demonstrate that climate context is critical for understanding food security vulnerabilities. Sensitivity regimes can be combined with operational price-spike triggers to stage anticipatory action: the ENSO state flags vulnerable regions months ahead, and a pre-positioned response in those regions to a price spike would avert the largest jump in acute food insecurity.</summary>
  </entry>
  <entry>
    <title>From Regional to Global: Transfer Learning for Atmospheric Transport Emulators</title>
    <link href="https://arxiv.org/abs/2609.23838"/>
    <link rel="related" href="https://weatherml.github.io/papers/air-quality-composition/#2609.23838"/>
    <id>https://arxiv.org/abs/2609.23838</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jeff Clark, Elena Fillola, Nawid Keshtmand, Raul Santos-Rodriguez, Matthew Rigby</name></author>
    <category term="Air Quality &amp; Composition"/>
    <summary>Greenhouse gas emissions estimates can be derived using inverse methods by combining atmospheric concentration observations with chemical transport models. The latter traditionally use physics-driven simulators such as Lagrangian Particle Dispersion Models (LPDMs), which are expensive to run and do not scale well to modern satellites' high resolution data. Previously we developed a performant atmospheric transport emulator that approximates LPDM outputs ("footprints") over South America ~1,000X faster than the UK Met Office's LPDM. Expanding towards global emulation is not straightforward, as atmospheric transport is regionally heterogeneous. This paper evaluates spatial transferability capabilities of models across four world regions: South America, East Asia, South Asia, North Africa using both region-specific and multi-region models, and leave-one-region-out experiments. Regional differences are characterised in the context of input variable and output footprint distributions. This work builds intuition in cross-region generalisation and transfer learning, aiding regional performance towards efficient global emissions estimates.</summary>
  </entry>
  <entry>
    <title>ClimTip-GML: A global bias-corrected and downscaled dataset for assessing impacts of climate tipping events</title>
    <link href="https://arxiv.org/abs/2609.23149"/>
    <link rel="related" href="https://weatherml.github.io/papers/post-processing/#2609.23149"/>
    <id>https://arxiv.org/abs/2609.23149</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Philipp Hess, Sebastian Bathiany, Lucas Ferreira Correa, Laura C. Jackson, Casey R. Patrizio, Niklas Boers</name></author>
    <category term="Post-processing"/>
    <summary>Assessing the impacts of future climate scenarios including tipping events of major Earth system components such as the Amazon rainforest (ARF) or the Atlantic meridional overturning circulation (AMOC), requires accurate and high-resolution simulations. Here, we present ClimTip-GML, the first globally bias-corrected and downscaled climate dataset for impact assessment of large-scale tipping scenarios, comprising eight key variables at 0.25° spatial resolution from three general circulation models (GCMs): CESM1-CAM5, HadGEM3-GC31-MM, and MPI-ESM1-2-HR. The dataset includes 100-year-long climate simulations with preindustrial and historical conditions, as well as scenarios at a +2°C warming level with and without tipping transitions of the AMOC or ARF. We apply generative machine learning (GML) techniques trained on reanalysis data to bias-correct and downscale the GCMs in a manner that is physically consistent across space, time, and all eight variables. Comprehensive validation shows substantially reduced biases, improved small-scale spatial variability, multivariate correlations, and consistent long-term climate responses to the external forcing and tipping events. The results hence permit substantially improved impact assessments of tipping transitions of the ARF and AMOC, directly informing mitigation and adaptation policies.</summary>
  </entry>
  <entry>
    <title>Diffusion-Based Super-Resolution of Adriatic Sea Oceanographic Fields</title>
    <link href="https://arxiv.org/abs/2609.22574"/>
    <link rel="related" href="https://weatherml.github.io/papers/downscaling/#2609.22574"/>
    <id>https://arxiv.org/abs/2609.22574</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Rajat Srivastava, Muhammad Sarmad, Emanuele Mele, Massimo Cafaro, Marco Pulimeno, Italo Epicoco</name></author>
    <category term="Downscaling"/>
    <summary>High-resolution oceanographic fields are critical for resolving mesoscale and sub-mesoscale coastal dynamics, yet their generation remains constrained by both computational cost and observational sparsity. We present OcDiffSR, a conditional denoising diffusion probabilistic model (DDPM) for oceanographic super-resolution that reconstructs high-resolution sea-surface fields from coarse-resolution reanalysis inputs. The model is trained on ten years (2011-2020) of paired low-resolution (GLORYS12V1, 1/12) and high-resolution (Mediterranean Sea Physics Reanalysis, Med MFC, 1/24) data, and evaluated on an independent test year (2009) over the Adriatic Sea. OcDiffSR employs a conditional U-Net augmented with multi-scale low-resolution encoders, cross-attention bottleneck layers, and sinusoidal seasonal embeddings via Feature-wise Linear Modulation (FiLM), enabling joint super-resolution of sea-surface temperature (SST), salinity (SSS), and horizontal velocity components with visually coherent circulation patterns. Benchmarked against bilinear interpolation and the state-of-the-art residual diffusion model CorrDiff, OcDiffSR achieves substantially lower reconstruction errors for scalar fields (RMSESST=0.477 C, RMSESSS=0.346 psu), near-unity Pearson correlation (PCC &gt;= 0.999), and high structural similarity (SSIM &gt;= 0.964). For dynamical vector fields, OcDiffSR outperforms both baselines in absolute error and spatial coherence, though moderate correlation (PCC = 0.64) reflects the intrinsic stochasticity of oceanic velocity fields. Daily and monthly evaluations confirm temporal robustness across all seasons. These results establish OcDiffSR as a reliable framework for high-fidelity oceanographic downscaling and reanalysis enhancement, producing fields that are visually consistent with known ocean dynamics.</summary>
  </entry>
  <entry>
    <title>Spatial Aggregation of ROC and Precision-Recall Curves</title>
    <link href="https://arxiv.org/abs/2609.19517"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.19517"/>
    <id>https://arxiv.org/abs/2609.19517</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Romain Pic, Zhongwei Zhang, Sebastian Engelke, Johanna Ziegel</name></author>
    <category term="Global Models"/>
    <summary>Receiver Operating Characteristic (ROC) and Precision-Recall (PR) curves are widely used to assess the discrimination ability of forecasts for binary events, such as threshold exceedances or warnings of extreme events. In weather forecasting, forecasts are provided as spatial fields, yielding location-wise ROC and PR curves that are often aggregated to facilitate comparison. However, the effect of the aggregation strategy on performance assessment remains poorly understood.   We investigate how different aggregation strategies for ROC and PR curves affect the assessment of discrimination ability. In particular, we identify conditions under which aggregation strategies satisfy two desirable properties for fair comparison: preservation of dominance between forecasts and preservation of concavity or achievability of the curves. We obtain sufficient conditions and propose two strategies satisfying them. They are compared with existing strategies from the literature, and we analyze their properties and highlight potential pitfalls that may lead to misleading interpretations. Based on these findings, we provide practical guidelines for the interpretation of aggregated ROC and PR curves. The proposed framework is illustrated with AI-based global weather forecasts, showing how different aggregation strategies can yield different rankings of competing forecasts.</summary>
  </entry>
  <entry>
    <title>A more predictable Madden-Julian Oscillation index derived from Koopman spectral analysis</title>
    <link href="https://arxiv.org/abs/2609.19435"/>
    <link rel="related" href="https://weatherml.github.io/papers/climate-modeling/#2609.19435"/>
    <id>https://arxiv.org/abs/2609.19435</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Claire Valva, Edwin P. Gerber</name></author>
    <category term="Climate Modeling"/>
    <summary>The Madden-Julian oscillation (MJO) is a major source of subseasonal-to-seasonal (S2S) predictability. The MJO is commonly defined and tracked with indices such as the Real-time Multivariate MJO (RMM) index. Although the RMM provides a useful description of the MJO, its evolution can be noisy and difficult to predict. We define an MJO index using a data-driven approximation of the Koopman operator. The Koopman index captures similar tropical circulation and convection patterns to the RMM but evolves more smoothly and predictably. Skillful prediction extends to 46 days for the Koopman index compared to 11 days for the RMM under the same prediction framework. While this new approach does not recover the RMM as well as operational S2S models, which provide skillful forecasts up to 35 days, the Koopman index could complement existing MJO diagnostics in evaluating and developing extended-range forecast systems.</summary>
  </entry>
  <entry>
    <title>Butterfly Effect and the Kinetic Energy Cascade in Probabilistic Machine Learning Weather Prediction Models</title>
    <link href="https://arxiv.org/abs/2609.18489"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.18489"/>
    <id>https://arxiv.org/abs/2609.18489</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Jiakai Chen, Joel Oskarsson, Simon Driscoll, Sebastian Schemm</name></author>
    <category term="Global Models"/>
    <summary>This study analyses kinetic energy (KE) spectra, difference kinetic energy (DKE) spectra, and signatures of KE transfer across spatial scales in four state-of-the-art probabilistic machine learning weather prediction (MLWP) models: NeuralGCM-ENS, FourCastNet 3, AIFS-ENS, and GenCast. Results are compared with those from the physics-based numerical weather prediction model IFS-ENS. While NeuralGCM-ENS successfully reproduces the expected upscale transfer of KE, noise injection at its encoder stage underestimates mesoscale KE. Conversely, AIFS-ENS, GenCast, and FourCastNet 3 produce realistic KE spectral magnitudes but do not capture the expected upscale transfer of KE. In particular, AIFS-ENS and GenCast, which employ spatially uncorrelated stochastic perturbations, exhibit enhanced accumulation of KE at high wavenumbers. All examined models exhibit upscale error growth, reflected by the progressive shift of the DKE spectral peak toward larger wavelengths over time. However, the MLWP models struggle to reproduce the rapid initial growth of ensemble spread at small spatial scales associated with the butterfly effect. The results show that MLWP models can misrepresent the known scale transfer of kinetic energy despite producing skilful weather forecasts.</summary>
  </entry>
  <entry>
    <title>Every Fixed Metric Has a Blind Spot: A Learned Atmospheric Critic for Scoring Forecast Realism</title>
    <link href="https://arxiv.org/abs/2609.18381"/>
    <link rel="related" href="https://weatherml.github.io/papers/global-models/#2609.18381"/>
    <id>https://arxiv.org/abs/2609.18381</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Younes Elberkennou, Dmitri Demler, Thierry Meier, Luca Rispoli, Fanny Lehmann, Joel Oskarsson</name></author>
    <category term="Global Models"/>
    <summary>Despite their high accuracy on point-wise metrics, machine learning weather forecasting models can exhibit different failure modes such as blurring, periodic irregularities, and other unphysical spatial artifacts. This has motivated a variety of metrics to detect known failure cases. Existing metrics fix a representation or transformation in advance, and that choice limits the artifacts they can detect. We propose to train a discriminator for separating reference data from the model's output, and using its output logit to obtain a divergence-like realism score. The discriminator learns whatever separates the model's fields from real weather, adapting to whichever failure mode that model exhibits. We compare our learned atmospheric critic to existing metrics using various synthetic corruptions applied to ERA5 reanalysis data. Our method successfully identifies the corruptions and ranks their severity, while existing metrics fail on at least one corruption. Additionally, we evaluate forecasts from real weather models, and find that the realism score degrades with longer lead times and the metric generally assigns higher realism to numerical models than to machine learning models.</summary>
  </entry>
  <entry>
    <title>IRENE: A Convolutional GRU Ensemble Model for Radar Precipitation Nowcasting over Italy</title>
    <link href="https://arxiv.org/abs/2609.17175"/>
    <link rel="related" href="https://weatherml.github.io/papers/nowcasting/#2609.17175"/>
    <id>https://arxiv.org/abs/2609.17175</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Alessandro Camilletti, Gabriele Franch, Elena Tomasi, Marco Cristoforetti</name></author>
    <category term="Nowcasting"/>
    <summary>We present IRENE (Italian Radar Ensemble Nowcasting Experiment), a deep learning model for probabilistic short-range precipitation nowcasting over the Italian domain at \SI{1}{km} spatial and 5 min temporal resolution. IRENE adopts an encoder--forecaster architecture built on multi-scale Convolutional Gated Recurrent Units (ConvGRUs), trained on the national radar composite produced by the Italian Civil Protection Department (DPC). An importance-sampling scheme focuses training on precipitation-relevant events, while the almost-fair Continuous Ranked Probability Score (afCRPS) is adopted as the primary probabilistic loss function. Two additional training configurations are proposed: an adversarial (GAN) variant, IRENE-GAN, designed to improve the spatial sharpness of the generated forecasts, and a spectrally constrained variant, IRENE-GAN-RAPSD, in which the adversarial objective is complemented by an explicit penalty on the radially averaged power spectral density. The three configurations are evaluated against the stochastic extrapolation method STEPS and the pre-trained deep learning model DGMR. All IRENE configurations attain a lower Continuous Ranked Probability Score than both benchmarks at every lead time and rank histograms closer to uniformity, indicating better probabilistic skill and ensemble calibration. In terms of ensemble-mean mean absolute error the advantage is confined to the first 90 min, beyond which the strongly damped DGMR fields and, to a lesser extent, STEPS become competitive. Spectral analysis shows that the adversarial training removes the progressive loss of small-scale variance exhibited by IRENE, at the cost of an excess of fine-scale power at long lead times that the spectral penalty only partially controls.</summary>
  </entry>
  <entry>
    <title>Predictability-Guided Multiscale Probabilistic Forecasting of Wind Direction under Extreme Shear</title>
    <link href="https://arxiv.org/abs/2609.16707"/>
    <link rel="related" href="https://weatherml.github.io/papers/other/#2609.16707"/>
    <id>https://arxiv.org/abs/2609.16707</id>
    <updated>2026-09-01T00:00:00Z</updated>
    <author><name>Hailong Shu</name></author>
    <category term="Other"/>
    <summary>Accurate multi-horizon wind direction forecasting is critical for turbine yaw control and grid security. Rapid directional shear (turning $\ge 90^\circ$) challenges models via non-Euclidean geometry on $S^1$, multiscale dynamics, and regime-dependent uncertainty. Conventional discrete models and foundation models suffer from mid-frequency phase lag and turning misalignments. We show that directional predictability decays at disparate rates across frequency subbands, rendering monolithic mechanisms suboptimal. We propose a predictability-guided paradigm: slow synoptic drift $\to$ deterministic regression; intermediate turning $\to$ continuous latent differential flows; unresolved turbulence $\to$ conditional residual diffusion; followed by causal recalibration. On a 10,000-sequence multi-year benchmark, our framework maintains calm-weather accuracy (Test MCE $38.48^\circ$) while reducing extreme-turning error (Case 1 MCE $60.69^\circ$ vs $70.42^\circ$ for zero-shot foundation models). The circular CRPS reaches $22.36^\circ$, with 93.88\% coverage at nominal 95\% (91.01\% out-of-distribution). Density estimation further reveals near-antipodal bimodal structure under severe shear (13.39\%--15.43\% tail mass $\ge 135^\circ$), exposing a geometric bound where single-center calibration under-covers (81.56\%), motivating multimodal circular manifold learning.</summary>
  </entry>
</feed>
