Evan Atlas Metamodern philosophy

Hudson Valley
New York

Evan Atlas

Research

Hudson Valley · NY

Research · Fenocosm

A 24-hour M-class flare forecast built from the GOES/XRS event record alone is worth about what recalibration and an eleven-parameter logistic make of it

Zenodo

Read the PDFDOI: 10.5281/zenodo.22760517

Abstract

Solar flare forecasting is dominated by magnetogram-derived predictors, and the operational systems that use them have repeatedly been found hard to separate from each other or from cheap baselines. That makes a narrower question worth answering precisely: how much 24-hour M-class forecast skill is in the flare event list by itself, and what part of it is bought by model structure rather than by scaling? A companion paper fitted an exact hierarchical regime filter — a 2⁶-state Markov-switching multifractal on a Poisson counting channel — to fifty years of GOES/XRS daily flare counts and turned it into a causal, filter-only now-cast. This paper verifies that forecast against the field's conventions, on the same sealed trace, with every difference carried by a paired 90-day-block bootstrap on one common panel of resamples rather than by comparing two marginal intervals. The ordering of the answer is the result.

(i) Recalibration is worth more than the hierarchy's entire margin over persistence. A two-parameter Platt map fitted on the training block alone — a = −0.369440, b = +0.924339 — moves the sealed forecast's Brier Skill Score on the 7,305-day verification tail from +0.3385 to +0.3645 against that tail's own climatology, a paired +0.026 [+0.018, +0.035], where the filter's whole paired margin over a five-day recency rate is +0.021 [+0.012, +0.030]. It buys skill, not a flat reliability diagram: the worst bin still misses by 0.075.

(ii) An eleven-parameter ridge logistic on the same daily counts beats the calibrated 64-state filter. Seven log1p count lags, the trailing 30- and 90-day event-day frequencies and log1p of the trailing 7-day sum score +0.3734; the paired difference (calibrated cascade − logistic) is −0.0089 [−0.0165, −0.0019]. Stacking the 64-state posterior onto the logistic moves the score by +0.0015 [−0.0022, +0.0051].

(iii) Depth buys forecast skill, mostly through calibration, and the hierarchy buys none of the rest. Raw, k = 6 − k = 2 is +0.0729 [+0.0527, +0.1069]; give every entrant its own map and the cheapest depth indistinguishable from k = 6 is k = 4 (+0.0032 [−0.0009, +0.0070]), while a free five-state Poisson hidden Markov model with no hierarchy at all matches the six-layer cascade at +0.0009 [−0.0090, +0.0097] — and loses the likelihood race on the same block by ΔBIC 97.

(iv) The slow rungs were standing in for the solar cycle. Sunspot number and F10.7 in the exposure, on a supply repaired so that no verification bin is beyond issue-time coverage, are worth ΔBSS +0.0507 [+0.0382, +0.0665] and repair the quiet-year floor; under the refit the fitted ladder span collapses from S = 146.8 to S = 4.16.

(v) The margin over recency grows with lead time because recency decays faster, not because the filter holds up. Against h-matched persistence it widens from +0.0210 [+0.0124, +0.0294] at h = 1 to +0.0861 [+0.0515, +0.1230] at h = 14, by which horizon persistence has fallen to −0.0043, below the verification climatology. The h-matched logistic still wins at every horizon (−0.0349 [−0.0499, −0.0230] at h = 1), and a 365-day moving rate overtakes the raw filter past about three days.

(vi) At the X threshold there is a real signal, and it is a rate effect. Read against real X1.0+ labels the frozen M-ladder's own 64-state belief scores +0.1022 and beats X persistence by +0.0718 [+0.0256, +0.1274] — and a pure size lottery reproduces it, one constant flare-level X fraction q = 0.0713 applied to the filter's rate, paired difference (state route − rate-only) +0.0170 [−0.0097, +0.0514]. The marked channel's primary per-event gain is +0.0029 [−0.0006, +0.0074] nats, covering zero, and a power battery recovers a planted ×5.2 swing in the X fraction on only 3/6 seeds: an estimability statement, not a demonstration that flare sizes are regime-independent. What the marks do show is a completeness signature — the M1.0–M1.9 share of M1.0+ events falls with the regime's activity, slope −0.0789 [−0.1059, −0.0477], while flat against the lagged sunspot number (−0.0559 [−0.1748, +0.0548]).

(vii) The negative result. Applied to a second domain — the daily count of 3-hour intervals at Kp ≥ 5− over 34,577 calendar days, whose lag-27 storm-day autocorrelation of 0.193 encodes Bartels' solar-rotation recurrence — the same instrument returns UNDECIDABLE AT THIS N ON P1, S1, S2 on a synthetic two-world atlas, and on the real record THE CASCADE DOES NOT CARRY SKILL WORTH SERVING: the calibrated filter beats persistence, recurrence and a seasonal climatology at h = 1, then falls below a flat 27-day recurrence baseline at h = 2 by −0.0412 [−0.0560, −0.0238], with its BIC-selected depth flagged degenerate at S = 1.00. The lag-features logistic wins at every horizon and an additive decomposition absorbs its entire long-lead edge: beyond h = 2, count-only Kp forecasting reduces to 27-day recurrence plus a trailing activity level — an independent replication, from a different model family, of the lead-time structure Shprits et al. (2019) reported for solar-wind-driven Kp prediction.

We claim a verification, in the conventions of the field, of what the event record supports, and the finding that scaling and the feature set outrank the hierarchy on every axis we could measure. We do not claim an operational forecast, a physical mechanism, or any comparison against magnetogram-based systems, which see information no entrant here sees.

Keywords

  • solar flare forecasting
  • calibration
  • Markov-switching multifractal
  • Brier skill score
  • geomagnetic activity

Cite this

Canonical deposit: doi.org/10.5281/zenodo.22760517. Select the BibTeX below to copy it.

@misc{atlas_event_record_flare_forecast,
  author       = {Atlas, Evan Tabak},
  title        = {A 24-hour M-class flare forecast built from the GOES/XRS event record alone is worth about what recalibration and an eleven-parameter logistic make of it},
  year         = {2026},
  month        = {sep},
  howpublished = {Zenodo preprint},
  doi          = {10.5281/zenodo.22760517},
  url          = {https://doi.org/10.5281/zenodo.22760517}
}

← All research