Bellwether

When grid demand departs from what was expected, what caused it? Bellwether forecasts demand with an uncertainty band, flags the hours that fall outside it, and attributes each of those episodes to a cause computed from stored history. Three US balancing authorities, two years of hourly data. The forecast is the instrument rather than the argument: a band you cannot trust cannot tell you an hour is unusual, so sections 1 to 5 establish that it can, and section 6 is what the project is for.

Snapshot generated 2026-08-31. Three balancing authorities out of 83, chosen to contrast rather than to sample.

1. The model beats the statistical baselines everywhere

Table view: skill against baselines
marketreductionbaselinemase
CISO43.350Seasonal naive, daily0.390
ERCO25.516Seasonal naive, daily0.380
PACE29.243Seasonal naive, daily0.506

Zero-shot, with no weather input and no training on these series. MASE below 1.0 means it beat the naive baseline on its own scale.

The daily seasonal lag beats the weekly one in every market, which contradicts the textbook expectation for electricity load. Weather persistence is the likely reason: yesterday resembles today because yesterday's weather does.

2. A reference point: the operator's own day-ahead forecast

marketarmsMAPE (%)MASE
ERCOChronos-Bolt3.1220.369
ERCOOperator day-ahead2.1510.252
ERCOSeasonal naive, daily4.4370.514
PACEChronos-Bolt2.8960.516
PACEOperator day-ahead4.4420.767
PACESeasonal naive, daily4.0670.714

ERCOT beats us by 31%. PacifiCorp East does not beat us, or even seasonal-naive. That reads as forecasting investment rather than method: ERCOT runs a large, weather-driven market and forecasts it accordingly.

CISO is absent on purpose. Its DF series diverges from its own D series at midday in a way nobody here has explained, and an operator baseline nobody understands is worse than no operator baseline.

3. A forecast, drawn. CISO

Market for this section

Blue is what happened. Orange is the forecast median, and the shaded band is the 80% interval, which should contain the actual on 4 days in 5.

Table view: this forecast window

4. Never report coverage without width

Table view: coverage and width
marketarmcoveragewidth_pct
CISOChronos-Bolt78.968.15
CISO+ calendar79.268.02
CISO+ scale77.977.90
CISO+ holiday (pooled)78.127.90
CISO+ holiday (by class)78.197.90
CISO+ volatility77.417.90
CISO+ weather77.847.69
CISO+ weather + volatility77.147.66
ERCOChronos-Bolt76.409.67
ERCO+ calendar79.2210.25
ERCO+ scale80.9911.26
ERCO+ holiday (pooled)80.9111.26
ERCO+ holiday (by class)80.8911.26
ERCO+ volatility78.4810.25
ERCO+ weather77.438.87
ERCO+ weather + volatility77.438.71
PACEChronos-Bolt77.628.47
PACE+ calendar75.748.08
PACE+ scale77.928.59
PACE+ holiday (pooled)77.918.59
PACE+ holiday (by class)77.928.59
PACE+ volatility77.148.39
PACE+ weather77.518.18
PACE+ weather + volatility77.978.24

A model can buy coverage by widening its interval and sharpen its way back out again, and the two are indistinguishable in a coverage column. This is not hypothetical: it is what separated the calendar arm from the weather arm below, and reading coverage alone reversed the conclusion.

The base model's defect is a level error, not a conditioning one. Chronos-Bolt conditions its intervals well, widening them 43 to 83% between its narrowest and widest month. It simply runs a few points under nominal. Scaling its own spread about its own median fixes the level and keeps the conditioning; rebuilding the distribution from residual quantiles fixes the average by flattening the seasons, which is worse for anyone acting on it.

A second model, and the defect stays where it was. Every distributional claim on this page rested on one checkpoint, and a level error a few points under nominal is exactly the kind of thing that could belong to the model or to the grid. TimesFM 2.5 (200M parameters, trained by different people on a different corpus) was run against Chronos-Bolt (205M) on the same 702 windows, the same 24 hour horizon, and the same 2,048 hours of context. Matching the context matters: TimesFM accepts far more, and letting it read years where the other reads days would compare two amounts of evidence and report the difference as a difference of method.

Chronos-Bolt wins on accuracy in all three markets, by 4.0% of sMAPE on PACE, 8.3% on ERCO and 10.7% on CISO, with WQL and MASE agreeing everywhere. That is the smaller half of the result. The larger half is that both models land under nominal on ERCO and PACE and both land closest on CISO. The ordering of the three markets by calibration difficulty survives a change of checkpoint, which is what a property of the grid looks like and not what a property of a model looks like.

Two further arms are on the chart below: Chronos-Bolt at 48M parameters against its own 205M sibling, and TimesFM re-run on its own context ceiling rather than the matched one. Both are drawn here because they were scored on the same windows and are read with the same rule, and each is the subject of a passage underneath. The matched context turns out not to be what lost, which is the last passage in this section.

Table view: the three checkpoints
marketarmcoveragewidth_pct
CISOChronos-Bolt base (205M)79.648.70
CISOChronos-Bolt small (48M)81.459.57
CISOTimesFM 2.5 (200M)80.189.82
CISOTimesFM 2.5 (200M, 16k context)81.9010.00
ERCOChronos-Bolt base (205M)77.379.38
ERCOChronos-Bolt small (48M)77.319.84
ERCOTimesFM 2.5 (200M)74.919.98
ERCOTimesFM 2.5 (200M, 16k context)77.399.93
PACEChronos-Bolt base (205M)79.109.10
PACEChronos-Bolt small (48M)79.099.35
PACETimesFM 2.5 (200M)77.109.21
PACETimesFM 2.5 (200M, 16k context)78.329.38

And this section's own rule earns its keep on California. Read from the coverage column alone, TimesFM's 80.2% beats Chronos-Bolt's 79.6% and CISO is the one market where the challenger calibrates better. Read with the width beside it, CISO is the market where it pays most for the appearance: that half point costs a band 12.9% wider. On ERCO no such care is needed, where it covers 2.5 points less on a band 6.4% wider, which is worse in both directions at once.

Four fifths of the parameters buy a tenth of the margin. The small checkpoint carries 48M against base's 205M and read the same 702 windows with the same 2,048 hours of context. Measured as the share of base's gain over the daily seasonal naive that it keeps, rather than as a ratio of the metrics themselves, it holds 90.8% on ERCO, 93.0% on CISO and 95.3% on PACE, with WQL agreeing to within a point everywhere. It also beats TimesFM 2.5 on MASE, WQL and sMAPE in all three markets on a quarter of the parameters, which removes the equal-budget condition the comparison above was careful to impose and leaves that ordering standing anyway.

MarketMASE gain kept (%)WQL gain kept (%)80% width vs base (%)
CISO93.093.910.1
ERCO90.889.24.8
PACE95.395.22.7

On the interval it buys nothing, and this section's rule catches it a third time. CISO's 81.4% is the first above-nominal coverage figure anywhere on this page, and it costs a band 10.1% wider. ERCO and PACE are flat on coverage, 77.3% against 77.4% and 79.1% against 79.1%, on bands 4.8% and 2.7% wider: less sharp everywhere and no better conditioned. The market ordering survives the third checkpoint as well, ERCO worst and CISO best, so the level error above has now held across two training corpora, two architectures and a 4x parameter range.

None of this is a reason to change the shipped model. The small checkpoint ran 3.1x faster in the session that measured it, and nothing here is compute bound: a day-ahead forecast has hours to make a prediction that takes under a second. It would matter somewhere that is bound, which the weekly refresh may turn out to be. That speed figure is a within-session ratio on purpose, because accuracy in this project has reproduced exactly three times and runtime has never reproduced once.

The context was not what lost. The comparison above matched TimesFM to Chronos's 2,048 hours and said plainly that this left a question open: a model held to a rival's limit might be losing on evidence rather than on method. So it was re-run on its own ceiling, 16,256 hours, 7.9 times the matched context and more history than Chronos-Bolt is able to read at all, over the same 702 windows.

MarketMASE better by (%)Coverage change (pts)80% width change (%)
CISO2.21.71.8
ERCO4.32.5-0.6
PACE1.11.21.9

The extra history is worth something real, small, and different in every market. MASE improves 4.3% on ERCO, 2.2% on CISO and 1.1% on PACE, with WQL and sMAPE moving the same way and about as far inside each market. Three metrics agreeing within a market is what makes it an effect; the spread between markets is what stops it being one number, and nothing measured here predicts which end a market lands on.

And this section's rule finally catches a model in its favour. On ERCO coverage rises 2.5 points onto a band that is 0.6% narrower, the only place on this page where coverage has not been paid for with width. That is what running short of evidence looks like: given more, the model became both more accurate and more certain. On CISO and PACE the bands widen instead, 1.8% and 1.9%, for 1.7 and 1.2 points of coverage, which is the ordinary purchase every other arm here has made; on CISO it buys a figure that now sits 1.9 points above nominal rather than closer to it. So the narrowing is a rescue of one badly under-covering market, not a property of longer context.

It still does not win. Against Chronos-Bolt base it closes between a quarter and a half of the accuracy gap and stops, in every market. On ERCO it arrives at exactly base's coverage, 77.4% against 77.4%, and needs a band 5.8% wider to get there. It does not beat the 48M checkpoint either. Eight times the evidence, read by a model four times that size, lands behind both: the ordering in this section is a property of the models and not of what they were allowed to read.

5. Where the forecast actually fails

Table view: error by local hour
marketbucketvalue
CISO0.0001.834
CISO1.0001.689
CISO2.0001.490
CISO3.0001.475
CISO4.0001.434
CISO5.0001.325
CISO6.0001.305
CISO7.0001.550
CISO8.0002.000
CISO9.0002.543
CISO10.0002.885
CISO11.0003.060
CISO12.0003.423
CISO13.0003.866
CISO14.0004.197
CISO15.0004.250
CISO16.0004.211
CISO17.0003.665
CISO18.0003.557
CISO19.0003.357
CISO20.0003.273
CISO21.0002.989
CISO22.0002.561
CISO23.0002.019
ERCO0.0002.800
ERCO1.0002.404
ERCO2.0002.495
ERCO3.0002.621
ERCO4.0002.787
ERCO5.0002.977
ERCO6.0002.719
ERCO7.0002.854
ERCO8.0003.031
ERCO9.0002.852
ERCO10.0002.730
ERCO11.0002.939
ERCO12.0003.005
ERCO13.0003.120
ERCO14.0003.692
ERCO15.0004.101
ERCO16.0004.440
ERCO17.0004.622
ERCO18.0004.280
ERCO19.0003.600
ERCO20.0003.575
ERCO21.0003.477
ERCO22.0003.363
ERCO23.0003.202
PACE0.0002.464
PACE1.0002.193
PACE2.0002.241
PACE3.0002.269
PACE4.0002.266
PACE5.0002.317
PACE6.0002.073
PACE7.0002.045
PACE8.0002.222
PACE9.0002.291
PACE10.0002.609
PACE11.0003.023
PACE12.0003.261
PACE13.0003.131
PACE14.0003.406
PACE15.0003.598
PACE16.0003.659
PACE17.0003.677
PACE18.0003.550
PACE19.0003.084
PACE20.0003.168
PACE21.0003.037
PACE22.0002.934
PACE23.0002.788

Peak error lands at 17:00 local in ERCO and 15:00 in CISO, the evening ramp in both cases. Aggregate metrics hide this completely.

A trap worth naming. Hour of day and horizon step are the same variable in a single backtest run, because origins advance by exactly the horizon, so every hour is always forecast at the same lead time. The first version of this analysis reported a diurnal profile that was line for line a horizon-step profile, looked entirely reasonable, and named the wrong hours. These numbers pool four staggered origin sets so the two cross.

6. What caused it: anomalies attributed to stored evidence

This is the part the rest of the project exists to support. Section 5 says where the forecast fails; this says why, episode by episode, with numbers a reader can check.

Table view: attribution rate by market
marketepisodesattributedshareleading
CISO10.010.0100.0temperature
ERCO10.010.0100.0temperature
PACE10.08.080.0temperature
Market for this section
Table view: CISO largest anomalies and their causes
start (UTC)hoursdirectionpeak (MW outside band)causewhat the data says
2025-05-09 16:0020above3,475temperaturePopulation-weighted temperature averaged 24.7 C during the episode against 15.3 C over the preceding 14 days, an anomaly of +9.4 C. Unusual heat raises cooling load.
2024-12-25 02:0022below2,431holidayThe episode covers a US federal holiday (2024-12-25 local). Holidays depress commercial and industrial load, and the forecaster sees only demand history with no calendar, so it cannot anticipate one.
2024-11-28 12:0025below1,678holidayThe episode covers a US federal holiday (2024-11-28 local). Holidays depress commercial and industrial load, and the forecaster sees only demand history with no calendar, so it cannot anticipate one.
2025-03-14 13:0011above3,012temperaturePopulation-weighted temperature averaged 9.0 C during the episode against 12.1 C over the preceding 14 days, an anomaly of -3.1 C. Unusual cold raises heating load.
2025-07-31 21:001below16,916data qualityThe reported demand at 2025-07-31 21:00 UTC is 11,819 MW, against about 29,884 MW in the hours either side. A single-hour move of this size with immediate recovery is a reporting artifact, not grid behaviour, so this episode should not be explained as an event.
2025-05-26 12:0012below2,534holidayThe episode covers a US federal holiday (2025-05-26 local). Holidays depress commercial and industrial load, and the forecaster sees only demand history with no calendar, so it cannot anticipate one.
2024-11-02 10:0014above2,776temperaturePopulation-weighted temperature averaged 14.6 C during the episode against 17.6 C over the preceding 14 days, an anomaly of -3.0 C. Cooler than usual, so less cooling load. Note this points to demand below the interval and the episode ran above it, so the anomaly does not explain this episode.
2025-03-18 19:005below2,866temperaturePopulation-weighted temperature averaged 16.2 C during the episode against 11.7 C over the preceding 14 days, an anomaly of +4.5 C. Milder than usual, so less heating load.
2024-11-17 07:0018above1,240temperaturePopulation-weighted temperature averaged 11.6 C during the episode against 14.8 C over the preceding 14 days, an anomaly of -3.2 C. Unusual cold raises heating load.
2025-08-06 17:007above2,267temperaturePopulation-weighted temperature averaged 28.9 C during the episode against 21.0 C over the preceding 14 days, an anomaly of +7.9 C. Unusual heat raises cooling load.

Each of the ten largest episodes per market is put to three screens computed in Python from stored data: an unusual temperature against the preceding fortnight, a US federal holiday inside the window, and a data-quality check for values no grid could physically produce. A cause was found for 28 of the 30 episodes. Temperature leads with 23, holidays account for 4, and one episode is not a grid event at all.

Nothing here is generated. Every quantity is measured, and the written brief for each episode is checked token by token against the evidence it was given: a number that does not trace to a measurement means the brief is rejected rather than shown. All 30 briefs pass that check. An explanation that invents a figure is worse than no explanation, because it reads exactly like one that did not.

The most severe episode in the analysis is not a grid event. It is an EIA value of 11,819 MW sitting between two hours near 29,900. A grid does not shed and recover 60% of its load in two hours, so the screen marks it an artifact and disqualifies it from being explained. Without that screen this page would carry a confident account of a blackout that never happened.

What this does not claim. Thirty episodes is a small sample, the three screens are the ones the detector's own output justified rather than a complete taxonomy, and strength orders candidates rather than estimating a probability. Two PACE episodes have no candidate at all and are reported as such.

7. Temperature buys accuracy and does not fix calibration

Against a calendar-only control, not against the raw model. That control is what makes the number mean anything: rebuilding a predictive distribution from residual quantiles is itself a recalibration, so a weather corrector scored only against the base model collects credit for work that has nothing to do with weather.

MarketsMAPE change (%)WQL change (%)Coverage, calendar (%)Coverage, weather (%)
CISO-1.7-1.879.377.8
ERCO-12.7-13.879.277.4
PACE-3.3-4.275.777.5

ERCOT is the most temperature-coupled market of the three by a wide margin, and that was measured before any model was fitted: summer correlation 0.925, winter -0.694, against CISO's 0.559 and positive 0.254. It predicted the ordering above correctly.

California's demand barely tracks its own weather and its winter correlation has the wrong sign. Gas heating plus behind-the-meter solar, which is the duck curve restated.

Everything above was measured with observed temperature, which hands the corrector perfect knowledge of tomorrow and makes every weather number a ceiling. The table below replaces it with the forecast a forecaster would actually have had: NOAA's archived NDFD grids, restricted per window to the freshest run published at or before that window opened.

Archived forecast temperature is three-hourly and everything else here is hourly, so a straight swap would differ from the arm above in two ways at once, forecast error and resolution, and attribute the difference to whichever the reader already believed. The middle arm is the observed series put through the coarseness and none of the error.

All three are scored on one set of 362 origins, bounded by the archive rather than by the forecast, and each carries its own calendar-only control which comes out identical in all three passes. That is the check that they saw the same windows. It is also why the observed arm below reads a few tenths off the table above, which scores more windows: the comparison that matters is within a table, never across the two.

MarketTemperaturesMAPE change (%)Coverage (%)
CISOObserved-1.677.8
CISOObserved, 3-hourly-1.877.8
CISONDFD forecast-2.078.0
ERCOObserved-12.377.2
ERCOObserved, 3-hourly-12.277.3
ERCONDFD forecast-9.876.0
PACEObserved-3.577.6
PACEObserved, 3-hourly-3.677.5
PACENDFD forecast-2.977.2

Four fifths of the ceiling survives contact with a real forecast. On ERCO, the only market with a large weather effect, perfect foresight cuts sMAPE 12.3% against the calendar control and a forecast available in advance cuts it 9.8%. PACE behaves the same way on a smaller effect, 3.5% falling to 2.9%. CISO is not evidence either way: its whole weather effect is 1.6% and the three arms sit inside 0.01 points of each other, so the forecast arm coming out nominally ahead of the observed one is noise.

Resolution costs nothing. The three-hourly arm matches the hourly one in all three markets to within 0.003 points of sMAPE, and sometimes on the better side. Demand responds to temperature slowly enough that sampling it eight times a day loses nothing a corrector can use, so the entire shortfall is forecast error. The concern that motivated the arm was unfounded, which is only knowable because it was measured.

Where the forecast does cost is calibration, which is the same shape as the rest of this section: ERCO's coverage falls from 77.2% on observed temperature to 76.0% on forecast, against a nominal 80%. The honest summary is not that weather helps only with perfect foresight, but that weather helps and most of the help is available in advance.

8. Holidays: the effect is real and it is confined to six days

Table view: learned offsets
marketobservanceoffset
CISOfederal only-527
CISOwidely observed-1,017
ERCOfederal only126
ERCOwidely observed-689
PACEfederal only23
PACEwidely observed-130

One offset applied to all eleven federal holidays improved 28 of 33 widely-observed holidays and only 10 of 27 federal-only ones. Below a coin flip on the second group is the signature of a correction being applied where nothing needs correcting.

Splitting the offset by whether private employers actually close confirms it. In ERCO and PACE the federal-only offset does not shrink, it changes sign. Demand on Veterans Day in Texas is not depressed at all, and the pooled version had been pushing those days down because Christmas dragged the shared estimate with it.

Market for this section
Table view: CISO per-holiday change
datenameobservancechangedirection
2024-11-112024-11-11federal only34worse
2024-11-282024-11-28widely observed-413better
2024-12-252024-12-25widely observed-984better
2025-01-012025-01-01widely observed309worse
2025-01-202025-01-20federal only-98better
2025-02-172025-02-17federal only-71better
2025-05-262025-05-26widely observed-71better
2025-06-192025-06-19federal only-265better
2025-07-042025-07-04widely observed-191better
2025-09-012025-09-01widely observed-281better
2025-10-132025-10-13federal only-334better
2025-11-112025-11-11federal only299worse
2025-11-272025-11-27widely observed-449better
2025-12-252025-12-25widely observed-576better
2026-01-012026-01-01widely observed-803better
2026-01-192026-01-19federal only-155better
2026-02-162026-02-16federal only406worse
2026-05-252026-05-25widely observed-324better
2026-06-192026-06-19federal only-316better
2026-07-032026-07-03widely observed110worse

It was measured twice and shipped neither time, and the stated reason was that a single scalar shift over 24 hours is the wrong shape: load barely moves overnight and falls hard through the working day, so one number over-corrects the small hours to reach the large ones. That was a diagnosis rather than a measurement, so it was tested.

Table view: the learned hour profile
markethouroffset
CISO0-712
CISO1-698
CISO2-654
CISO3-670
CISO4-835
CISO5-896
CISO6-1,060
CISO7-1,474
CISO8-1,959
CISO9-2,080
CISO10-1,849
CISO11-1,886
CISO12-2,070
CISO13-1,969
CISO14-1,823
CISO15-1,759
CISO16-1,009
CISO17-535
CISO18-317
CISO19-393
CISO20-543
CISO21-505
CISO22-594
CISO23-741
ERCO0-78
ERCO1-77
ERCO2-38
ERCO3-133
ERCO4-333
ERCO5-596
ERCO6-1,198
ERCO7-2,153
ERCO8-2,315
ERCO9-2,113
ERCO10-1,765
ERCO11-1,122
ERCO12-728
ERCO13-1,381
ERCO14-2,076
ERCO15-2,101
ERCO16-2,145
ERCO17-2,207
ERCO18-189
ERCO19-113
ERCO20-225
ERCO21-180
ERCO22-154
ERCO235
PACE0-32
PACE1-66
PACE2-68
PACE3-72
PACE4-86
PACE5-136
PACE6-200
PACE7-258
PACE8-310
PACE9-270
PACE10-265
PACE11-250
PACE12-222
PACE13-198
PACE14-220
PACE15-243
PACE16-149
PACE17-121
PACE18-54
PACE19-53
PACE20-40
PACE21-38
PACE22-65
PACE23-70

The diagnosis was right, and it was not what was blocking anything. A third arm learns an offset per hour rather than per day, and the shape it finds is larger than the scalar implied. On a widely-observed holiday ERCO is 132 MW below normal overnight and 2,315 below at 08:00, a 17.6x swing that no single number can express; CISO runs 714 against 2,080. The arm beats the one it replaces in all three markets, which is more than that arm could say for itself, and improves 27 of 33 widely-observed holidays.

It still does not ship, and the third measurement is the one that explains why. Against no calendar at all, CISO wins clearly, at 26.4% of its holiday error. ERCO and PACE improve on 12 of 20 holidays each, which is a coin flip. The giveaway is that on ERCO the entire calendar family fails the same way: 11 of 20 for the flat arm, 11 for the split one, 12 for this one. What is missing there is not a better corrector but a holiday effect worth correcting, and no amount of shaping manufactures one.

Three arms, three declines, and the reason changed. The first two were declined because they might have been the wrong correction. This one is the right correction and is declined because two of three markets do not need it, which leaves exactly one live option: a holiday arm for California alone.

9. Two published claims that turned out to be wrong

Both were caught by re-measurement rather than by review, and both are kept here because a results page that reports only what survived is not reporting.

"Chronos-Bolt's interval is unconditional." It is not. That measurement was taken on a residual corrector, which rebuilds the predictive distribution from scratch, so every distributional property it showed belonged to the corrector. The base model conditions its intervals well. The rule that came out of it: profile the base model, not only the arm you happen to be using.

"The duck curve shows up as a breach pattern in CISO." Below-bound breaches at the 10:00 to 11:00 solar ramp were an artifact of the same corrector. Peak error by hour survived re-measurement; worst coverage by hour did not.

A third prediction was pre-registered and falsified cleanly: a volatility-conditioned interval was supposed to fix the seasonal miscoverage. It did not, and its discriminating check passed, so the target turned out to be an artifact of the method rather than a property of the grid.

Method and sources

Scope, and what these numbers are not
  • Three balancing authorities out of 83, chosen to contrast rather than to sample. No state is a unit in this data.
  • The weather ceiling is measured, and so is the distance to it. The headline weather numbers use observed temperature, which is perfect foresight. Section 7 also scores the same correction against NOAA's archived forecasts, restricted to runs published before each window opened, and four fifths of the gain survives.
  • NCEI's archive ends about eleven months before EIA's data does, so weather work covers roughly half the demand grid. That limit was accepted rather than patched with a third-party source.
  • Rolling-origin backtest, 24 hour horizon, scored on MASE, WQL, sMAPE, and 80% interval coverage with width.