Quick Overview
In this video, creator Nate Herk conducts a head-to-head benchmark comparing the front-end web design and coding capabilities of Claude Code and Codex. He evaluates both autonomous AI agents across eight identical website builds using varying levels of prompt specificity to compare design quality, completion time, sub-agent usage, and API token costs.
Key Points
- 1.Codex consistently produced cleaner and more visually balanced user interfaces across several landing page build tests compared to Claude Code.
- 2.Across eight test builds, Codex required nine sub-agents and five hours total, compared to twenty-five sub-agents and over fourteen hours for Claude Code.
- 3.Codex consumed approximately 550,000 output tokens totaling ninety-eight dollars in API costs, whereas Claude Code consumed 2.95 million output tokens costing 444 dollars.
- 4.When supplied with extremely specific layout and content prompts, both tools generated nearly identical functional outputs, but Codex still completed tasks faster and at lower token expense.
- 5.Claude Code repeatedly generated text-heavy pages and encountered rendering bugs during layout collapses that Codex avoided.
Summary
Nate Herk conducts a direct benchmark between two AI coding systems, Claude Code and Codex, by tasking each with building eight identical website projects using the same starting data, copy, and brand parameters. The comparison focuses not only on visual design quality and user experience, but also on hard performance metrics, including sub-agent counts, runtime, output token usage, and API billing costs.
The test begins with structured brand landing pages. For Bowl and Bloom, Codex implemented a clear navigation bar, horizontal product scroll animations, and an intuitive custom box builder, finishing in forty-nine minutes for 18.28 dollars. Claude Code produced a text-heavy layout that suffered from rendering bugs when viewed in split screen, taking two hours and costing fifty-three dollars. Similar patterns emerged with MinuteCraft and TrailLatch, where Codex organized dashboards and interactive assembly breakdowns more cleanly, while Claude Code introduced overly dense copy and UI scaling defects, although Herk preferred Claude Code's single-page layout for TrailLatch.
The comparison then moves to service pages, including Present and Clear and North Ledger Studio. In both instances, Codex used visual contrast, color blocking, and structured feature cards to make the user journey readable. Claude Code defaulted to continuous prose and raw financial tables that resembled technical documentation rather than marketing landing pages. For North Ledger Studio, Codex took fifty-five minutes and cost 18.28 dollars, whereas Claude Code took over two hours and cost fifty-two dollars.
To evaluate adaptability, Herk tests both systems under extreme prompting conditions. When given a vague, open-ended prompt to build an impressive website, Codex designed a dark-themed space observatory interface in eight minutes for 1.39 dollars, while Claude Code took two hours and forty-two minutes and spent fifty-one dollars on a cluttered sound-wave tool. When given an extraordinarily specific prompt for a database branching tool called Resalt and an incident report, both agents produced almost identical interfaces, yet Codex completed the work in six minutes compared to twenty-two minutes for Claude Code.
The aggregate results across all eight builds show Codex requiring nine sub-agents, five total hours, 550,000 output tokens, and ninety-eight dollars. Claude Code required twenty-five sub-agents, fourteen hours and nine minutes of execution time, 2.95 million output tokens, and 444 dollars.
Testing Framework and Initial Brand Builds
Nate Herk tests Claude Code and Codex by giving both tools identical prompts, copy, and brand guidelines across eight separate website projects. In initial tests for brands like Bowl and Bloom, MinuteCraft, and TrailLatch, both tools attempted to structure complete landing pages and interactive elements. Codex created cleaner visual hierarchies and navigation flows, while Claude Code tended toward dense text layouts and occasional UI layout bugs when rendering nested components.
Service and Advisory Page Layouts
Further comparisons on sites like Present and Clear and North Ledger Studio highlighted differences in visual pacing and user journeys. Codex used distinct color transitions, clear call-to-action placement, and structured summaries to guide visitors. In contrast, Claude Code produced long paragraphs and report-style tables that felt more like raw documentation than consumer-facing landing pages.
Vague Versus Highly Specific Prompting
When given open-ended prompts with minimal instructions, Codex completed a creative speculative observatory page in eight minutes for 1.39 dollars, whereas Claude Code took nearly three hours and cost fifty-one dollars to build a crowded audio wave interface. When given rigid, component-by-component prompts for a database tool named Resalt and an incident report, both tools generated virtually identical outputs, though Codex remained significantly cheaper and faster to execute.
Overall Efficiency and Cost Metrics
Summing up the data from all eight builds, Claude Code used twenty-five sub-agents, fourteen hours and nine minutes of execution time, 2.95 million output tokens, and 444 dollars. Codex required nine sub-agents, five hours total runtime, 550,000 output tokens, and ninety-eight dollars, demonstrating substantial operational efficiency advantages across different design tasks.
The Bottom Line
The video establishes that Codex currently holds a major advantage over Claude Code in speed, token efficiency, and automated visual layout design. It demonstrates that while highly prescriptive prompting forces both models into producing identical web pages, Codex consistently delivers those results at a fraction of the time and financial cost. The video leaves open how future updates to Claude Code and underlying vision loops might improve its aesthetic judgment and token management.
FAQ
What is Claude Code and Codex in the context of this web design benchmark?
Claude Code and Codex are AI coding agents tested to determine which tool produces better website designs, cleaner user interfaces, and more efficient code when given identical design prompts.
Which AI agent won the design comparison between Claude Code and Codex?
Codex won the majority of the head-to-head design matchups by producing better visual hierarchy, cleaner navigation, fewer layout bugs, and less cluttered text than Claude Code.
How did the total execution time compare between Claude Code and Codex across all builds?
Codex took a total of five hours across all eight builds, whereas Claude Code took fourteen hours and nine minutes to complete the same set of tasks.
What was the difference in total cost between Claude Code and Codex across the eight website builds?
Codex cost approximately ninety-eight dollars in API usage, while Claude Code cost 444 dollars for the identical set of projects.
How did Claude Code and Codex perform when given extremely specific, detailed prompts?
When provided with strict, highly specific instructions, both Claude Code and Codex generated nearly identical final pages, though Codex completed the build in six minutes compared to twenty-two minutes for Claude Code.
How many total output tokens did Claude Code consume compared to Codex?
Claude Code consumed 2.95 million output tokens across all eight projects, while Codex consumed approximately 550,000 output tokens.
Worth watching for
Web developers, UI designers, and software engineers evaluating autonomous AI coding agents for front-end development and web design.
- claude-code
- codex
- web-design
- ai-agents
- benchmarking