Power plants must ask permission to plug into the grid. They then wait on a public list, often for years.
Forecasts of the AI buildout quote that list as a construction schedule. Most of it never happens. I give each waiting application a percentage chance of being built, published before anyone knows the result.
The four grid operators are MISO, CAISO, NYISO and ISO-NE. Texas (ERCOT), PJM and the central plains (SPP) are not in these figures. Every figure is rebuildable from public sources.
This measures one thing: whether a power project gets built. Not prices, not who wins. When it says 20%, that should happen about 20% of the time. A single prediction cannot be judged that way. You can only judge the whole list, once the results come in. The list is public, so anyone can check whether I was right.
Each prediction carries a deadline for its result. The grid operators replace the list with a new one and delete the old, so I save a copy every morning. Nobody can go back and buy the queue as it once stood.
Last time I measured my own work it failed, and I published that rather than bury it. TirraMind collected 34 public data sources and 375,657 observations to find market patterns. Then I tested whether it predicted anything. Zero of fifty-one ideas survived. Try fifty-one at once and a few look like winners by pure luck. So I set a harder pass mark (the Benjamini–Hochberg correction, named so you can repeat it). None of mine cleared it.
I said up front how weak the test was. Even if a real effect existed, the study had about a 10% chance of spotting it. Finding nothing is honest. There may still be something there. My test was too weak to see it.
Two things went wrong, both written up. My data-processing code gave wrong or empty results thirteen documented times. Every time, its own checks said everything was fine. And 90% of the score I was about to sell came from the date alone. The model had largely learned what day it was.
That habit killed the product. I withdrew it and published the failure log. It is the same habit behind the work above: check what a number really means before anyone bets on it. It already applies to this model: I have measured where its percentages are wrong — in five of seven groups — and that failure is published beside the predictions, not buried.