The scorecard

Results, stated plainly.

P&L Performance is a real judging category, and the number is reported here in full — but a genuinely 60%-edge agent only beats a coin flip 69% of the time over 20 trades (measured, not assumed), so a one-week P&L is closer to noise than proof. The calibration and attribution below are the honest instruments for the harder question: did the agent actually know what it was doing.

Equity $104,347.81 high water $104,347.81
P&L +4.3% +$4,347.81 on $100,000.00
Drawdown 0.0% from high water
Data as of 2026-09-01 hackathon rule - required starting balance, not measured

Equity over time

start $100,000.00$104,347.81$100,000.00

Calibration.

Whether stated confidence matched observed frequency — a harder, more honest question than "did it make money." Brier score and the Murphy decomposition answer it directly.

Brier score 0.227 0 = perfect, 0.25 = coin flip
Reliability 0.000 lower is better
Resolution 0.000 higher is better
Base rate 36.7% of resolved forecasts held
Sample size — read this before the numbers above

49 forecast(s), 12.8 effective (26% of face value - the sample is concentrated in a few names, so it says less about NEW ones than the count suggests)

Profit
Loss
View held
0 reinforce both
0 the view was fine — the structure wasn’t
View failed
0 correct the view; the structure was faithful
0 luck — learn nothing from this

No positions have reached their thesis horizon yet, so nothing has been attributed — attribution only fires once a claim's stated resolution date has actually passed. (4 of 4 still pending.)

Attribution.

Was the view right, and was the way it was expressed right — scored separately, so a profit on a wrong view (bottom-right) is excluded from what lets the agent size up.

The competence ladder.

Position size is earned, not chosen — four rungs, gated on resolved theses, calibration reliability, and attribution rate.

Explore

Fixed 2.2% exploration allocation while the record is too thin for Kelly to mean anything.

10%book cap
0min resolved
Establish

5 resolved theses. Kelly engages, capped at 10% of the calculated fraction.

15%book cap
5min resolved
Scale

15 resolved theses, 60% attributable (view actually explicable, not just profitable).

20%book cap
15min resolved
Mature

40 resolved theses, reliability <0.04, 70% attributable — strictly enforced.

25%book cap
40min resolved

Currently Establish — 49 resolved theses, 0% attributable, Kelly ×0.09.

Book risk.

Beta-weighted, because names are not exposures.

Positions 1
Raw delta -$243,782.72
Beta-weighted delta -$243,782.72
Vega / Theta $8.99 / -$29.23