City With High Rise Buildings Under Blue Sky With Lightning

AI Forecasts Geomagnetic Storm Risk Across 66,935 Substations With 30–60 Minutes of Warning

The May 2024 geomagnetic storm was a wake-up call that arrived in shades of pink and green. Auroras spilled far beyond their usual polar range, visible to millions who had never seen them. Less photogenic was the damage to satellite operations and GPS accuracy, which rippled into precision agriculture as positioning signals drifted. Utilities spent the week on alert, watching for currents that can cook transformers from the inside.

The hard part is not knowing a storm is coming. It is estimating when and where its effects will be most severe, with enough lead time for grid operators to actually do something about it.

A machine learning system built during a Microsoft Research summer internship takes direct aim at that second problem. It forecasts space-weather risk across 66,935 substations in the continental United States, producing location-specific estimates 30 to 60 minutes ahead of potential impact. That substation count comes from the GridSFM-derived grid data used as input, and it covers the continental United States specifically. That is the geographic scope for the whole project, and it is worth stating once: everything below applies to that grid, not to transmission networks in general.

Why space weather resists forecasting

Space-weather prediction spans several coupled systems, and each link in the chain adds uncertainty. The solar wind changes rapidly. Its interaction with Earth’s magnetosphere is irregular. The ground effects that follow depend on local conditions that have nothing to do with the Sun.

Geomagnetically induced currents, or GICs, are the mechanism of concern. They flow through transmission networks when a changing magnetic field drives current into the ground and along conductors. Regions with resistive bedrock can experience stronger GICs than regions with more conductive geology, because the current has nowhere easy to go and prefers the wires. Transmission-line orientation matters too, since a line’s angle relative to the induced electric field influences how much current it picks up. Latitude stacks on top of both: northern assets see stronger forcing, consistent with the higher geomagnetic activity observed at high latitudes.

The result is a forecasting problem with a geophysical half and a geographic half. Getting the storm right is necessary but not sufficient. You also have to know which of tens of thousands of assets sits on rock that amplifies the effect.

Thunderstorm over Hills
Thunderstorm over Hills

Inside the system

The pipeline was assembled during a summer internship at Microsoft Research and covers the continental United States at substation resolution. It combines solar-wind observations, forecasts of the Auroral Electrojet (AE) and Disturbance Storm Time (Dst) indices, physics-informed constraints, local geological conductivity, and grid-infrastructure data.

The architecture runs in three stages.

First, solar-wind measurements from the L1 Lagrange point feed forecasts of the AE and Dst indices. In parallel, geological conductivity and location features are assembled for each substation. L1 sits sunward of Earth, which is what buys the 30 to 60 minute horizon: the solar wind takes about that long to travel from the measurement point to the magnetosphere.

Second, a gradient-boosting model combines the forecasted indices with the location-specific inputs to estimate dB/dt, the rate of magnetic-field change that drives GIC risk.

Third, those predictions are converted into location-specific risk estimates and aggregated into a continental risk assessment.

A system of 50 AI agents helped explore features, validation strategies, and model configurations across the pipeline. The agents did not replace the modeling decisions, but they widened the search over configurations considerably faster than a single researcher could.

Training on public data only

Every input in the pipeline comes from a public source: NASA OMNI and NASA-aggregated Kyoto World Data Center data, INTERMAGNET and U.S. Geological Survey magnetometer observations, and GridSFM-derived grid data. No proprietary utility telemetry, no closed datasets.

That choice matters for reproducibility. Anyone with an internet connection and enough compute can attempt to rebuild the pipeline, audit the preprocessing, or test the model on a different grid. It also means the work inherits the limitations of its sources. Magnetometer coverage is uneven, and the Kyoto and INTERMAGNET records have their own quality flags and gaps. Public data is a floor for reproducibility, not a guarantee of it. As a parallel case in the arXiv ecosystem shows, even well-documented public infrastructure can behave unexpectedly: see When an arXiv API Query Returns Nothing: A Reproducibility Note on the cs.LG Listing for how thin a retrieved record can turn out to be. The lesson generalizes. Public availability and practical reproducibility are different claims.

Benchmarks: AE, Dst and the Burton equation

The evaluation period runs from 2020 to 2026. The AE predictor was designed specifically to forecast rare, high-intensity geomagnetic activity, the kind that actually drives infrastructure risk and that standard metrics often smooth away. It produced forecasts spanning nearly the full observed range of AE activity and outperformed several empirical solar-wind-based approaches.

The Dst predictor supplied a complementary signal describing large-scale storm strength. During the most geomagnetically active periods, the machine learning model outperformed the Burton equation on 62.2% of individual hours. That figure needs context to be useful. The Burton equation is a classical physics-based baseline whose errors grow precisely when storms intensify, so a 62.2% hourly edge concentrates in the regime operators care about rather than spreading evenly across quiet conditions. Burton-style approaches also tend to under-disperse: the model produced a substantially wider prediction range, which matters when the events you care about live in the tail.

The two signals also combine. Adding the Dst forecast improved severe-event detection in the end-to-end system by 1.2 percentage points when paired with the AE forecasts. That is a modest number, but in rare-event forecasting, modest gains at the margin are often where the operational value sits.

GIC risk stage: detection rates and their cost

The final stage converts predicted geomagnetic activity into per-substation dB/dt estimates, combining storm conditions with each location’s latitude and geological factor. Because no equivalent widely deployed operational system provides a direct industry benchmark, the team compared against simple linear regression. That comparison is a floor, not a verdict.

Detection rates came in at 76.5% for major events (≥10 nT/min), 81.2% for severe events (≥20 nT/min), and 64.1% for extreme events (≥50 nT/min). False-alarm rates increased with storm severity, which reflects the unavoidable trade-off between missed events and cautious alerts. The extreme-event bucket is the hardest: those events are rarest, so the model has the least data to learn from, and the cost of a miss is highest. Performance also varied by latitude, with the highest detection rates at northern stations where geomagnetic activity is strongest.

Those three figures are not good or bad in the abstract. They are points on a sensitivity curve, and where an operator wants to sit depends on how expensive a precautionary alert is relative to a damaged transformer.

From one continental alert to a risk map

The system does not issue a single alert for the continental United States. It produces continuous risk estimates that distinguish lower-risk locations from areas where resistive geology can amplify ground-level effects. Figure 3 in the work illustrates output for a representative major-storm scenario. Worth stating plainly: the map is a demonstration of the model’s continental-scale output, not a record of a specific event.

That distinction is easy to blur when a risk map looks like a weather radar image. A forecast map shows what the model would say given a scenario. It is not evidence that the scenario occurred or that the model got it right.

The bottleneck has shifted

Two caveats deserve weight. First, there is no deployed operational benchmark for substation-level GIC forecasting, so the comparison against linear regression is a floor, not a verdict. Second, generalizing to grids with different geology, different line orientations and different magnetometer density is an open question, not a solved one.

What the project demonstrates is that the ingredients for substation-level space-weather forecasting are publicly available, and that the bottleneck has shifted. Storm detection is largely a solved problem with decades of satellite infrastructure behind it. The unsolved part is translating a global geomagnetic disturbance into a ranked list of specific assets at risk, with enough lead time to act. That translation is where machine learning earns its place, and where the next round of validation will decide whether this stays a promising internship project or becomes something operators actually lean on. For readers tracking how research moves from preprint to practice, the same dynamics shape the field more broadly, as explored in What arXiv’s Category Pages Actually Tell Us About AI Research – And What They Don’t.

Similar Posts