Results, stated plainly.
P&L Performance is a real judging category, and the number is reported here in full — but a genuinely 60%-edge agent only beats a coin flip 69% of the time over 20 trades (measured, not assumed), so a one-week P&L is closer to noise than proof. The calibration and attribution below are the honest instruments for the harder question: did the agent actually know what it was doing.
Equity over time
Calibration.
Whether stated confidence matched observed frequency — a harder, more honest question than "did it make money." Brier score and the Murphy decomposition answer it directly.
49 forecast(s), 12.8 effective (26% of face value - the sample is concentrated in a few names, so it says less about NEW ones than the count suggests)
No positions have reached their thesis horizon yet, so nothing has been attributed — attribution only fires once a claim's stated resolution date has actually passed. (4 of 4 still pending.)
Attribution.
Was the view right, and was the way it was expressed right — scored separately, so a profit on a wrong view (bottom-right) is excluded from what lets the agent size up.
The competence ladder.
Position size is earned, not chosen — four rungs, gated on resolved theses, calibration reliability, and attribution rate.
Fixed 2.2% exploration allocation while the record is too thin for Kelly to mean anything.
5 resolved theses. Kelly engages, capped at 10% of the calculated fraction.
15 resolved theses, 60% attributable (view actually explicable, not just profitable).
40 resolved theses, reliability <0.04, 70% attributable — strictly enforced.
Currently Establish — 49 resolved theses, 0% attributable, Kelly ×0.09.
Book risk.
Beta-weighted, because names are not exposures.