- Home
- How we check
How we check the forecast
We compare the forecast against 5,447 independent field records and publish the result in full — including the part that shows most of our score comes from the calendar rather than from the weather.
What we measure
We take 5,447 occurrence records from GBIF, the international biodiversity database fed by herbaria, institutes and amateur naturalists. Nobody collected them for us, which is the whole point: a model tested on data its own authors gathered is not tested at all.
For each record we compute our index on that day in that place, and then on other days in the same place. If the model is right about timing, the day of the real find should rank above a random day. Of the 5,447 records, 709 had weather data available for the comparison.
Two questions, not one
Asking «is the forecast good» produces a number that hides which part of the model earned it. So we ask twice: did we get the week right, and did we get the month right. The second is much easier, and separating them is what makes the result readable.
The result
For the week — did the model rank the day of the find above another day in the same place — the score is 0.610, where 0.5 is a coin toss. The confidence interval is 0.556–0.655, over 7,187 comparisons.
For the month the score is 0.847, interval 0.812–0.882. High, and we now know why: that is the seasonal calendar doing the work, not the weather model.
In plain terms: we are good at the month and not yet good at the week. We set the bar at 0.750 before running the test, and 0.610 does not reach it. We are not going to call it «61% accurate», which would be a different and more flattering claim than the one the data supports.
The uncomfortable part
To find out how much the weather actually contributes, we ran the same test twice more: once with the dates shuffled between records, and once with dates drawn at random.
| Run | Week | Month |
|---|---|---|
| As it stands | 0.610 | 0.847 |
| Dates shuffled | 0.576 | 0.798 |
| Dates random | 0.496 | 0.488 |
Higher is better; 0.5 is chance. The bottom row is the control.
What that table says
The bottom row is good news: on random dates the measurement gives 0.496 against an expected 0.5. The measuring instrument is honest.
The top two rows are the bad news. Between «as it stands» and «dates shuffled» the difference is 0.034, and the confidence intervals overlap. The share has to be counted against what we have ABOVE the coin toss: 0.5 is chance, so the model's own is 0.610 − 0.5 = 0.110. Of that, the calendar accounts for 0.076 — roughly 69%. The rest we cannot credit to the weather: 0.034 sits inside the confidence intervals, and on this data its contribution is not distinguishable from zero.
Why we publish this
Because a forecast that never says where it is weak is not a forecast. Anyone can compute a number; showing the number that falls below your own stated threshold is the part nobody does, and it is the only reason to believe the numbers that do look good.
It is also a work list rather than a confession. Knowing that the weather contribution is small tells us exactly where the next improvement has to come from, and the honest 0.610 is the baseline any change will have to beat.
