Accuracy & baselines

In plain words

Aggregated regime classification accuracy under methodology v2.1. Every figure includes sample, period, horizon, and version. This is not your strategy's return.

Version = p{prompt}.s{scoring}.f{fusion}. Shadow = p2.s2.f1 (does not drive bots).

Rows are models + Outcome; columns are prompt versions + Total. Click a cell for drill-down.

Outcome is the bot decision (model agreement + evidence threshold), not an arithmetic average of model accuracy.

So Outcome % is often higher than Claude / Gemini / OpenAI alone: weak or disputed signals become OFF.

Outcome n is counted per run; model n is per vote — the two n’s need not match.

Accuracy = (correct + 0.5×partial) ÷ graded. N/A and waiting are excluded from the denominator.

Confusion matrix (all models)

Model rotation log

When Model Field Old New