One-Shot Shop Challenge

Let’s build a full-featured online shop (roughly Shopify-level) from a single prompt. We’ll repeat it with different coding agents so we can compare them.

Fabian Wesner
Fabian Wesner

Last updated: 14th of July 2026


All Builds at a Glance

Compare side by side

Newest first. Click a column header to sort, click a build to jump to its details. The feature score is weighted: PASS = 1, PARTIAL = 0.5, out of 143 tests.

BuildDateModeFeature ScoreCostTime
#15Codex GPT-5.6 Sol UltraJul 12Team
95.8%
$1202h 55m
#14Claude Code Fable 5Jun 9Sub-agents
99.7%
$4236h 58m
#13Claude Code Team 4.8 xHighMay 30Team
95.5%
$4982h 54m
#12Codex GPT-5.5 Goal ModeMay 4Sub-agents
92.0%
$53023h 50m
#11Codex GPT-5.5Apr 25Sub-agents
82.9%
$1948m
#10Claude Code Opus 4.7 xHighApr 16Team
43.0%
$1571h 36m
#09Claude Code Opus 4.7Apr 16Single
73.4%
$2332m
#07Claude Code Team v4Mar 20Team
89.9%
$1323h 39m
#06Claude Code Team v3Mar 18Team
87.8%
$28510h 59m
#05Codex with Sub-Agents v2Feb 14Sub-agents
58.0%
$283h 27m
#04Claude Code Team v2Feb 13Sub-agents
66.8%
$743h 00m
#02Claude Code with Sub-AgentsFeb 13Sub-agents
60.5%
$622h 13m
#03Codex with Sub-AgentsFeb 12Sub-agents
65.7%
$91h 44m
#01Claude Code with Team ModeFeb 11Team
90.6%
$731h 06m

The Experiment

Same spec, same tech, same single prompt. Only the coding agent changes. Every run is scored against the same 143 end-to-end acceptance tests, and everything is published: prompts, session logs, analysis, and code.

The Setup

Specification

Full spec with acceptance criteria, no code snippets

tecsteps/shop/.../specs

Codebase

Fresh Laravel template with Livewire

MCP Servers

Laravel Boost, Playwright

QA Verification

143 E2E tests, two independent checks per build

View Testplan

The Builds

Newest first. Each build links to its full story, session analysis, code quality report, feature results, and database schema.

OpenAI

#15 Codex GPT-5.6 Sol Ultra(Codex CLI v0.144.3, gpt-5.6-sol-pro, multi-agent v2 team mode)

Codex CLI v0.144.3gpt-5.6-sol-proMulti-Agent v216 Named Sub-AgentsTeam Mode

16 named sub-agents role-playing a full engineering org. Highest feature pass rate of any Codex run, and remarkably cheap for its completeness: 98% of input was cache reads, so 16 agents cost only ~$120. Repeats the cross-tenant admin data leak. Read the full build story →

2h 55m

Duration

$119.74

API Cost

286

Classes

17

Active Agents

Feature Tests: 135 of 143 passed (95.8%)

Admin login:

Efficiency

shorter = better
Cost
$119.74
Time
2h 55m

Feature Completeness

out of 143 tests
Pass
135
Partial
4
Fail
4

Code Quality

LOC
11,292
Code smells
139
Tech debt
26.9 h
Duplication
0.8%

Teammates

teammates
Count
17
Claude

#14 Claude Code Fable 5(Claude Code v2.1.170, Fable 5 with high reasoning, orchestration left to the agent)

Claude Code v2.1.170Fable 5High ReasoningAgent's ChoiceSub-Agents

Fable 5 free to choose its own orchestration self-organised into a 14-phase sub-agent pipeline. Highest feature score on the page (99.7% weighted) and a polished, feature-rich result, at the cost of the longest active runtime here. Read the full build story →

6h 58m

Duration

$422.68

API Cost

195

Classes

15

Active Agents

Feature Tests: 142 of 143 passed (99.7%)

Admin login:

Efficiency

shorter = better
Cost
$422.68
Time
6h 58m

Feature Completeness

out of 143 tests
Pass
142
Partial
1
Fail
0

Code Quality

LOC
10,314
Code smells
76
Tech debt
12.6 h
Duplication
1.5%

Teammates

sub-agents
Count
15
Claude

#13 Claude Code Team 4.8 xHigh(Claude Code v2.1.158, Opus 4.8 with xHigh reasoning, thinking on, 1M context, strict team mode)

Claude Code v2.1.158Opus 4.8xHigh ReasoningThinking On1M ContextTeam Mode

Strict team-mode run: a 6-teammate team building the whole shop from a single "do it in one go" brief. Best maintainability on the page, with only address-book and stock-cap gaps left; cost sits in the upper range. Read the full build story →

2h 54m

Duration

$497.67

API Cost

205

Classes

7

Active Agents

Feature Tests: 134 of 143 passed (95.5%)

Admin login:

Efficiency

shorter = better
Cost
$497.67
Time
2h 54m

Feature Completeness

out of 143 tests
Pass
134
Partial
5
Fail
4

Code Quality

LOC
11,230
Code smells
64
Tech debt
10.5 h
Duplication
0.8%

Teammates

teammates
Count
7
OpenAI

#12 Codex GPT-5.5 Goal Mode(OpenAI Codex CLI v0.128.0, GPT-5.5 with xHigh reasoning, persistent goal mode)

Codex CLI v0.128.0gpt-5.5xHigh ReasoningPersistent GoalPlan ModeSub-AgentsCodex Pro

23 hours of autonomous work driven by a single goal brief. Strong pass rate, but also the largest codebase and the longest tech-debt tail. Read the full build story →

23h 49m

Duration

$530.26

API Cost

268

Classes

19

Active Agents

Feature Tests: 126 of 143 passed (92.0%)

Admin login:

Efficiency

shorter = better
Cost
$530.26
Time
23h 50m

Feature Completeness

out of 143 tests
Pass
126
Partial
11
Fail
6

Code Quality

LOC
13,645
Code smells
82
Tech debt
23.2 h
Duplication
3.2%

Teammates

sub-agents
Count
19
OpenAI

#11 Codex GPT-5.5(OpenAI Codex CLI v0.124, GPT-5.5 with high reasoning)

Codex CLI v0.124.0gpt-5.5High ReasoningSub-AgentsCodex Pro

Smallest codebase and lowest cost on the page; clean SonarCloud gate but skipped many admin features. Read the full build story →

47m 38s

Duration

$18.85

API Cost

67

Classes

4

Active Agents

Feature Tests: 108 of 143 passed (82.9%)

Admin login:

Efficiency

shorter = better
Cost
$18.85
Time
48m

Feature Completeness

out of 143 tests
Pass
108
Partial
21
Fail
14

Code Quality

LOC
1,743
Code smells
5
Tech debt
2.0 h
Duplication
1.3%

Teammates

sub-agents
Count
4
Claude

#10 Claude Code Opus 4.7 xHigh(Same setup as #09, stricter prompt to enforce team-mode)

Claude Code v2.1.114Opus 4.7xHigh ReasoningThinking On1M ContextTeam Mode

Same setup as #09 with a hardened prompt that forced team-mode; the team ran and over-built the spec on paper, but buttons shipped dead. Read the full build story →

1h 36m

Duration

$157.31

API Cost

151

Classes

37

Active Agents

Feature Tests: 51 of 143 passed (43.0%)

Admin login:

Efficiency

shorter = better
Cost
$157.31
Time
1h 36m

Feature Completeness

out of 143 tests
Pass
51
Partial
21
Fail
71

Code Quality

LOC
5,043
Code smells
41
Tech debt
3.6 h
Duplication
1.4%

Teammates

teammates
Count
37
Claude

#09 Claude Code Opus 4.7(Same prompt as #01, latest Opus)

Claude Code v2.1.112Opus 4.7High ReasoningThinking On1M Context

Opus 4.7 ignored team-mode and built everything as a single agent - the fastest and cheapest Claude run. Read the full build story →

32m 27s

Duration

$22.63

API Cost

105

Classes

1

Active Agents

Feature Tests: 93 of 143 passed (73.4%)

Admin login:

Efficiency

shorter = better
Cost
$22.63
Time
32m

Feature Completeness

out of 143 tests
Pass
93
Partial
24
Fail
25

Code Quality

LOC
2,667
Code smells
13
Tech debt
2.6 h
Duplication
0.4%

Teammates

single
Count
1
Claude

#07 Claude Code Team v4(Same Prompt, 1M Context)

Claude Code v2.1.81Team ModeOpus 4.6High Reasoning1M Context

Same prompt as #01 on 1M context; stable specialists and the highest feature pass rate of the Opus 4.6 builds. Read the full build story →

3h 39m

Duration

$132.06

API Cost

389

Files Created

34

Active Agents

Feature Tests: 121 of 143 passed (89.9%)

Admin login:

Efficiency

shorter = better
Cost
$132.06
Time
3h 39m

Feature Completeness

out of 143 tests
Pass
121
Partial
15
Fail
7

Code Quality

LOC
4,537
Code smells
57
Tech debt
8.4 h
Duplication
1.4%

Teammates

teammates
Count
35
Claude

#06 Claude Code Team v3(Advanced Prompt, 1M Context)

Claude Code v2.1.80Team ModeOpus 4.6High Reasoning1M Context

Advanced prompt with a controller and QA teammate - broadest coverage of the early builds but the longest and most expensive run of its generation. Read the full build story →

10h 59m

Duration

$284.52

API Cost

482

Files Created

158

Active Agents

Feature Tests: 119 of 143 passed (87.8%)

Admin login:

Efficiency

shorter = better
Cost
$284.52
Time
10h 59m

Feature Completeness

out of 143 tests
Pass
119
Partial
13
Fail
7

Code Quality

LOC
5,708
Code smells
91
Tech debt
14.3 h
Duplication
8.0%

Teammates

teammates
Count
159
OpenAI

#05 Codex with Sub-Agents v2(More Instructions)

OpenAI Codex v0.101.0Sub-Agents (experimental)GPT-5.3-codexReasoning: xhigh

Codex rerun with a quality-focused prompt and xhigh reasoning; lots of classes, lots of tech debt. Read the full build story →

3h 27m

Duration

$28.40

API Cost

53

Agents Spawned

898

Tool Calls

Feature Tests: 70 of 143 passed (58.0%)

Admin login:

Efficiency

shorter = better
Cost
$28.40
Time
3h 27m

Feature Completeness

out of 143 tests
Pass
70
Partial
26
Fail
39

Code Quality

LOC
7,178
Code smells
113
Tech debt
25.1 h
Duplication
3.0%

Teammates

sub-agents
Count
54
Claude

#04 Claude Code Team v2(More Instructions)

Claude Code v2.1.41Sub-AgentsOpus 4.6Thinking: OnReasoning: Max

Tuned prompt with explicit review agents; cleanest code-smell profile of the sub-agent runs. Read the full build story →

3h 0m

Duration

$73.92

API Cost

376

Files Created

29

Active Agents

Feature Tests: 82 of 143 passed (66.8%)

Admin login:

Efficiency

shorter = better
Cost
$73.92
Time
3h 0m

Feature Completeness

out of 143 tests
Pass
82
Partial
27
Fail
30

Code Quality

LOC
6,033
Code smells
38
Tech debt
5.2 h
Duplication
3.8%

Teammates

sub-agents
Count
30
Claude

#02 Claude Code with Sub-Agents

Claude Code v2.1.41Sub-AgentsOpus 4.6Thinking: On

Same prompt as #01 but with sub-agents instead of teammates; slower and less complete. Read the full build story →

2h 13m

Duration

$61.97

API Cost

358

Files Created

12

Active Agents

Feature Tests: 73 of 143 passed (60.5%)

Admin login:

Efficiency

shorter = better
Cost
$61.97
Time
2h 13m

Feature Completeness

out of 143 tests
Pass
73
Partial
27
Fail
43

Code Quality

LOC
6,033
Code smells
60
Tech debt
8.6 h
Duplication
3.6%

Teammates

sub-agents
Count
13
OpenAI

#03 Codex with Sub-Agents

OpenAI Codex v0.99.0Sub-Agents (experimental)GPT-5.3-codexReasoning: xhigh

First Codex pass at the same challenge - by far the cheapest run, but missing many features. Read the full build story →

1h 44m

Duration

$8.79

API Cost

16

Agents Spawned

357

Tool Calls

Feature Tests: 89 of 143 passed (65.7%)

Admin login:

Efficiency

shorter = better
Cost
$8.79
Time
1h 44m

Feature Completeness

out of 143 tests
Pass
89
Partial
10
Fail
39

Code Quality

LOC
6,037
Code smells
54
Tech debt
12.7 h
Duplication
2.8%

Teammates

sub-agents
Count
17
Claude

#01 Claude Code with Team Mode

Claude Code v2.1.39Team ModeOpus 4.6Thinking: On

Baseline run: Opus 4.6 in team mode, the reference point the other builds are compared against. Read the full build story →

1h 6m

Duration

$73.44

API Cost

388

Files Created

31

Active Agents

Feature Tests: 126 of 143 passed (90.6%)

Admin login:

Efficiency

shorter = better
Cost
$73.44
Time
1h 6m

Feature Completeness

out of 143 tests
Pass
126
Partial
7
Fail
9

Code Quality

LOC
6,108
Code smells
168
Tech debt
22.3 h
Duplication
2.9%

Teammates

teammates
Count
32
Fabian Wesner

Enthusiastic Berlin-based entrepreneur. Former CTO at Rocket Internet and Project A. Co-founded Spryker and raised millions with ROQ. Today, SMEs and enterprises book me to help them adopt agentic engineering and leverage AI across all departments. I'm also looking for an exceptional founder team to join as tech co-founder and build a unicorn.