Aggregated regime classification accuracy under methodology v2.1. Every figure includes sample, period, horizon, and version. This is not your strategy's return.
Version = p{prompt}.s{scoring}.f{fusion}. Shadow = p2.s2.f1 (does not drive bots).
Rows are models + Outcome; columns are prompt versions + Total. Click a cell for drill-down.
Outcome is the bot decision (model agreement + evidence threshold), not an arithmetic average of model accuracy.
So Outcome % is often higher than Claude / Gemini / OpenAI alone: weak or disputed signals become OFF.
Outcome n is counted per run; model n is per vote — the two n’s need not match.
Accuracy = (correct + 0.5×partial) ÷ graded. N/A and waiting are excluded from the denominator.
| When | Model | Field | Old | New |
|---|