Rebased to 100 at the start. The band behind them is the regime the agent is calling that day.
The dollar factor extracted from those four pairs, and its 20-day realised volatility — the signal the regime call is built on.
Sixty runs. Click anywhere to jump there.
Four prices go in. A regime label comes out every day. Between those two things sit a factor model, a hidden Markov model, and — on 14% of days — five language-model agents arguing about whether the statistics can be trusted.
In the regime-switching literature these are usually called risk-on, risk-off and range. The names here describe what actually separates them in the data — one state is about volatility, the other two about whether there is a direction — and avoid bull/bear language, which does not transfer to FX where every rise is someone else's fall.
EURUSD, GBPUSD, USDJPY and USDCAD from the Fed's H.10 release, all oriented USD-per-foreign so they move together against the dollar. Days where any pair is missing are dropped rather than filled — a fabricated flat day reads to the model as a calm market.
A 250-day rolling PCA collapses the four into one common direction, refit every day. That is PC1: when it rises, the dollar is rising against everything. Its 20-day realised volatility and the share of variance it explains become the detector's inputs.
Fitted on everything up to 31 December of the previous year, then frozen and run forward. States are unnamed until in-sample statistics name them: highest mean volatility becomes Turbulent, the remaining state with better carry-to-vol becomes Trending, and the last is Choppy.
733 of 5,261 days qualify: the label flips, confidence drops below its in-sample 5th percentile, or the probability vector moves hard in one day. Everywhere else the statistical call passes through untouched — which is why the agent layer costs $23 rather than $2,000.
Rates, volatility, factor structure and cross-asset each see only their own slice, as z-scores with no dates and no price levels. Shared inputs would produce correlated opinions and fake consensus, so the evidence is partitioned.
It sees the detector's proposal and all four opinions, and looks for recency bias, anchoring, and what would falsify the call. It can lower confidence freely — but it can only change the label when all four specialists independently disagree with the detector. That gate blocked 27 of its 35 attempted overrides.
What the agent layer actually changes
Across 733 wake-ups it altered 8 labels. What it really does is strip out false certainty: mean confidence on those days falls from 0.897 to 0.691, and days above 0.99 confidence go from 21% to zero. Press play and watch the confidence bar sag every time the panel convenes — that, not the label, is the intervention.
The agentic indicator against the 1989 Hamilton two-state filter, on identical days and an identical overlay. Twelve rounds.
Hamilton — Panel
The whole comparison in one picture: what actually happened, against what each system called.
Seven episodes everyone remembers, and how many days before or after each onset the two systems first called stress.
Throw away every day they agree and count who the target sides with.
Twenty-one years of calls, and where each arm earned or lost its reputation.
Shaded bands are the seven hand-labelled episodes. The strip beneath counts panel wake-ups.
Against the rule-based forward-volatility target.
Identical overlay; only the label source differs.
The panel's one real contribution is stating uncertainty better than the detector it wraps. Here is how far that goes, and what it cost elsewhere.
Stated probability against observed frequency, sized by bin count. The diagonal is perfect calibration.
Days before (+) or after (−) each episode onset.
The panel's contribution is calibration — and a free alternative beats it.
Arm C cuts the detector's raw Brier by 26% on gated days, entirely by damping overconfidence. Run a twenty-line expanding-window isotonic regression on the same detector instead and it lands at a calibrated Brier of 0.699 against the panel's 0.707. Two routes to the same correction; the free one wins.
And it cost the detector its best call
Three unanimous vetoes on 10, 11 and 12 July 2007 turned risk_off back
to range, overriding the earliest warning on the 2007–09 carry unwind
and pushing it from a +22-day lead to +19. All four specialists agreed with each
other. They were wrong. Open the Transcript tab and filter to Vetoes.
All 733 days the panel acted on: what the detector proposed, what each specialist saw, what the critic objected to, and what came out. Eighty days also carry the exact anonymised prompt each specialist received.
Pick a day to read the panel's working.
The scoreline is real. These are the reasons not to read more into it than it carries.
Where the panel genuinely wins
It is dramatically steadier — 60 regime episodes against Hamilton's 156, a median regime lasting 26 days rather than 6, and 0.91 false alarms a year against 2.15. If whipsaw is what costs you, that is the arm you want.