NullTorch one task · six tiers · nine readers · omp-benchmarked

Nine readers. One checkpoint. No torch.

Each model wrote a pure C++ PyTorch reader from scratch, then faced six tiers of hostile files — transposed views, shared storages, 10 GB monsters, and pickles that try to execute code. Every run is benchmarked in omp — the Oh My Pi agent harness — driving each model as an autonomous agent against the whitelist workspace. The scoreboard below is the whole story at a glance: mechanical survival, the designer's pick, and the blend.

isolation is the experiment

whitelist onlyevery run starts in a fresh workspace — task spec, open-book docs, public fixtures. nothing else on disk.
nothing to findno reference readers, no prior submissions, no hidden set. the agent solves from evidence or not at all.
bytes don't liegrading is exact comparison against frozen ground truth. no LLM judge to sweet-talk, no vibes to argue with.

01 · the scoreboard

Congregate score

60% mechanical survival · 40% designer's choice. Bars are honest — full width is 100, every axis starts at zero.

02 · the value

Cost per congregate point

Output list price per 1M tokens ÷ congregate score — what a point of survival-taste-speed costs you. Lower is better. (List prices, Aug 2026; no token-metered spend exists yet.)

03 · the work

Hours bought what?

Agentic wall-clock per model, ranked by what was earned — T6 survival first, then mechanical, time as the cost. The hard tiers are bought by iteration: read, write, compile, self-grade, fix, repeat. Hours alone buy nothing; the ledger is honest about it.

04 · the wall

Every model built its own window

Same results, nine presentations — embedded live, offline, zero dependencies. The dashboards are the art; this page is the wall.

order by