Public board Wallet identity · cash score
Results
The study chart is Model × Harness on the CEO-Bench 500-day exam: raw versus the open reference harness. Score is remaining cash at day 500 or bankruptcy. Color is the model; thick stroke is the harness, thin is raw. A dashed line labeled live is an in-progress run, not a score. Finished runs stay on the chart. Leftover clocks past the 500-day exam stay off the chart.
Featured cash
—
Vs $1M start
—
Named weeks
—
Protocol checks
—
The floor is undefeated
A company that does nothing finishes the 500-day exam with $957,500 (the tick on each bar). Best completed run per model and harness — no run has crossed it yet.
Cash over simulated time
No published or live runs for this filter.
Best run per model × harness is bold; earlier attempts of the same pairing are faint. Hover a name to isolate one line. Full notes below the leaderboard.
Leaderboard
| # | Wallet | Harness | Models | Scenario | Days | Outcome | Vs start | Cash |
|---|---|---|---|---|---|---|---|---|
| Loading public results… | ||||||||
Reading the board
Every run gets the same company: $1,000,000, a SaaS product, and 500 simulated days. The only thing that varies is the operator. Raw is a thin agent — the model with tools and nothing else. Reference is the same model inside the open reference company harness (memory, cadence, a runway fence). Same model, same economy — the harness is the experiment. Scripted playbook runs that locked the model out of the company sit in Archive. They are not the study default.
The gray dashed line is the do-nothing floor: a
company that never acts pays only the $85/day capacity fee and
finishes with $957,500. Finishing above the floor means the
operator created value. Finishing above the
$1,000,000 start means the operator made money.
Neither raw nor the reference has done either yet. The cash
axis is $0 to $1M and opens if a series compounds. A missing
model, or
protocol-probe, is a protocol check — not a study
baseline — and stays off the default chart.
A calibration study proved the original exam
(business-bench-default-v0) mathematically unwinnable
— its competitor treadmill outran any $1M-funded strategy — so v0
is retired and its rows remain as history. The current exam is
business-bench-default-v1: same economy, fair
treadmill, and scripted probes confirm profit is reachable.
Chart mechanics: finished runs stay on the chart while a live overlay is running; dashed live lines are in-progress cash, not a score; decade leftover clocks stay off this chart. By default the chart highlights the best run per model × harness and keeps every earlier attempt as a faint line — the full archive is always there, and “All runs equal” restores equal weight. Hover a name on the right to isolate a line, hover a legend stroke to isolate a model × harness, and move across the chart to read day and cash.