All Builds at a Glance
Compare side by sideNewest first. Click a column header to sort, click a build to jump to its details. The feature score is weighted: PASS = 1, PARTIAL = 0.5, out of 143 tests.
| Build | Date▾ | Mode | Feature Score▾ | Cost▾ | Time▾ |
|---|---|---|---|---|---|
| Jul 12 | Team | 95.8% | $120 | 2h 55m | |
| Jun 9 | Sub-agents | 99.7% | $423 | 6h 58m | |
| May 30 | Team | 95.5% | $498 | 2h 54m | |
| May 4 | Sub-agents | 92.0% | $530 | 23h 50m | |
| Apr 25 | Sub-agents | 82.9% | $19 | 48m | |
| Apr 16 | Team | 43.0% | $157 | 1h 36m | |
| Apr 16 | Single | 73.4% | $23 | 32m | |
| Mar 20 | Team | 89.9% | $132 | 3h 39m | |
| Mar 18 | Team | 87.8% | $285 | 10h 59m | |
| Feb 14 | Sub-agents | 58.0% | $28 | 3h 27m | |
| Feb 13 | Sub-agents | 66.8% | $74 | 3h 00m | |
| Feb 13 | Sub-agents | 60.5% | $62 | 2h 13m | |
| Feb 12 | Sub-agents | 65.7% | $9 | 1h 44m | |
| Feb 11 | Team | 90.6% | $73 | 1h 06m |
The Experiment
Same spec, same tech, same single prompt. Only the coding agent changes. Every run is scored against the same 143 end-to-end acceptance tests, and everything is published: prompts, session logs, analysis, and code.
The Setup
Codebase
Fresh Laravel template with Livewire
MCP Servers
Laravel Boost, Playwright
The Builds
Newest first. Each build links to its full story, session analysis, code quality report, feature results, and database schema.
#15 Codex GPT-5.6 Sol Ultra(Codex CLI v0.144.3, gpt-5.6-sol-pro, multi-agent v2 team mode)
16 named sub-agents role-playing a full engineering org. Highest feature pass rate of any Codex run, and remarkably cheap for its completeness: 98% of input was cache reads, so 16 agents cost only ~$120. Repeats the cross-tenant admin data leak. Read the full build story →
2h 55m
Duration
$119.74
API Cost
286
Classes
17
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
teammates#14 Claude Code Fable 5(Claude Code v2.1.170, Fable 5 with high reasoning, orchestration left to the agent)
Fable 5 free to choose its own orchestration self-organised into a 14-phase sub-agent pipeline. Highest feature score on the page (99.7% weighted) and a polished, feature-rich result, at the cost of the longest active runtime here. Read the full build story →
6h 58m
Duration
$422.68
API Cost
195
Classes
15
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#13 Claude Code Team 4.8 xHigh(Claude Code v2.1.158, Opus 4.8 with xHigh reasoning, thinking on, 1M context, strict team mode)
Strict team-mode run: a 6-teammate team building the whole shop from a single "do it in one go" brief. Best maintainability on the page, with only address-book and stock-cap gaps left; cost sits in the upper range. Read the full build story →
2h 54m
Duration
$497.67
API Cost
205
Classes
7
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
teammates#12 Codex GPT-5.5 Goal Mode(OpenAI Codex CLI v0.128.0, GPT-5.5 with xHigh reasoning, persistent goal mode)
23 hours of autonomous work driven by a single goal brief. Strong pass rate, but also the largest codebase and the longest tech-debt tail. Read the full build story →
23h 49m
Duration
$530.26
API Cost
268
Classes
19
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#11 Codex GPT-5.5(OpenAI Codex CLI v0.124, GPT-5.5 with high reasoning)
Smallest codebase and lowest cost on the page; clean SonarCloud gate but skipped many admin features. Read the full build story →
47m 38s
Duration
$18.85
API Cost
67
Classes
4
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#10 Claude Code Opus 4.7 xHigh(Same setup as #09, stricter prompt to enforce team-mode)
Same setup as #09 with a hardened prompt that forced team-mode; the team ran and over-built the spec on paper, but buttons shipped dead. Read the full build story →
1h 36m
Duration
$157.31
API Cost
151
Classes
37
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
teammates#09 Claude Code Opus 4.7(Same prompt as #01, latest Opus)
Opus 4.7 ignored team-mode and built everything as a single agent - the fastest and cheapest Claude run. Read the full build story →
32m 27s
Duration
$22.63
API Cost
105
Classes
1
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
single#07 Claude Code Team v4(Same Prompt, 1M Context)
Same prompt as #01 on 1M context; stable specialists and the highest feature pass rate of the Opus 4.6 builds. Read the full build story →
3h 39m
Duration
$132.06
API Cost
389
Files Created
34
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
teammates#06 Claude Code Team v3(Advanced Prompt, 1M Context)
Advanced prompt with a controller and QA teammate - broadest coverage of the early builds but the longest and most expensive run of its generation. Read the full build story →
10h 59m
Duration
$284.52
API Cost
482
Files Created
158
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
teammates#05 Codex with Sub-Agents v2(More Instructions)
Codex rerun with a quality-focused prompt and xhigh reasoning; lots of classes, lots of tech debt. Read the full build story →
3h 27m
Duration
$28.40
API Cost
53
Agents Spawned
898
Tool Calls
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#04 Claude Code Team v2(More Instructions)
Tuned prompt with explicit review agents; cleanest code-smell profile of the sub-agent runs. Read the full build story →
3h 0m
Duration
$73.92
API Cost
376
Files Created
29
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#02 Claude Code with Sub-Agents
Same prompt as #01 but with sub-agents instead of teammates; slower and less complete. Read the full build story →
2h 13m
Duration
$61.97
API Cost
358
Files Created
12
Active Agents
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#03 Codex with Sub-Agents
First Codex pass at the same challenge - by far the cheapest run, but missing many features. Read the full build story →
1h 44m
Duration
$8.79
API Cost
16
Agents Spawned
357
Tool Calls
Admin login:
Efficiency
shorter = betterFeature Completeness
out of 143 testsCode Quality
Teammates
sub-agents#01 Claude Code with Team Mode
Baseline run: Opus 4.6 in team mode, the reference point the other builds are compared against. Read the full build story →
1h 6m
Duration
$73.44
API Cost
388
Files Created
31
Active Agents
Admin login: