one prompt · three models · three harnesses · 10,800 games
Same models.
Different harnesses.
Compare Codex CLI, Pi, and Claude Code across five independent Terra, Sol, and Luna generations each. The primary leaderboard treats every harness and model combination as one entrant while preserving each generated engine, checksum, and replay.
All 45 engines used the same prompt and requested high reasoning in isolated harness environments. The replay bundles have zero voids and verified checksums; independent Harbor reproduction is not yet claimed.
verification levels →Harness × model leaderboard
Each row is one harness and model combination, pooling its 5 independently generated engines. Rankings use the full cross-harness schedule; the dots keep generation variance visible.
All 45 generated engines
A result should be
more than a row.
Open the complete move log, step through every board state, and inspect both competing artifacts.
watch this battle →From prompt to public result
- 01Generate
Isolated Codex CLI, Pi, and Claude Code harnesses each write executable chess agents.
- 02Probe
Known positions catch malformed output before competition.
- 03Battle
Deterministic positions, colors, seeds, and move limits are recorded.
- 04Publish
Source hashes, telemetry, standings, and replay traces travel together.