Quick Overview
This video is a comparative benchmark review presented by Igor from The AI Advantage. It evaluates and compares the front-end website creation capabilities of Anthropic's Claude Fable 5 and Fable 5.1 against OpenAI's GPT-6 Astra. The test was conducted using 50 one-shot prompts evaluated through blind human scoring, automated visual judges, and code functionality audits.
Key Points
- 1.In a benchmark of 50 website builds, Claude Fable 5.1 outperformed Claude Fable 5 in both human preference and AI visual ratings while matching its 47 out of 50 functionality score.
- 2.When comparing Claude Fable 5.1 against OpenAI's GPT-6 Astra, GPT-6 Astra won 70 percent of human blind preference tests with 35 wins to Fable 5.1's 15 wins.
- 3.Automated AI visual ratings overwhelmingly favored GPT-6 Astra over Fable 5.1, scoring Astra as the visual winner in 47 out of 50 matchups under an OpenAI model and 48 out of 50 under Fable 5.1.
- 4.GPT-6 Astra proved strongest in data dashboards, narrative and editorial sites, and browser games, while Fable 5.1 performed better on audio interfaces and experimental utilities.
- 5.GPT-6 Astra was the cheapest model to generate full sites with at 20.64 dollars for 50 sites, compared to 21.24 dollars for Fable 5 and 29.15 dollars for Fable 5.1.
Summary
The video opens with Igor exploring how modern AI models perform when tasked with building complete websites and interactive front-end applications from scratch. To determine the strongest model for web creation, Igor built a blind benchmark comparing Anthropic's Claude Fable 5, Claude Fable 5.1, and OpenAI's GPT-6 Astra. The test suite consisted of 50 distinct prompts spanning ten categories, including dashboards, data visualizations, experimental interfaces, playable games, generative art, and marketing storefronts. Each test used frozen one-shot generation directly through model APIs without manual repair, prompt refinement, or caching.
Evaluation followed three distinct rubrics. First, Igor performed a blind human preference review, judging aesthetics, structure, and user experience side by side. Second, an automated AI judge scored the visual design quality on factors like hierarchy, rhythm, color, and SVG craft. Third, a functionality audit verified whether all prompted features were present and operational in the browser.
The first contest evaluated Claude Fable 5 against the newer Claude Fable 5.1. In the human preference test, Fable 5.1 won 30 decisions, Fable 5 took 17, and three resulted in ties. The AI visual judge rated Fable 5.1 the winner in 40 cases, giving Fable 5 only seven wins and three ties. Functionality testing resulted in a tie, with both Anthropic models successfully delivering working code on 47 of the 50 builds. Fable 5.1 showed clear improvements in margin spacing, typography hierarchy, and SVG vector rendering across marketing sites and utility layouts.
The benchmark then compared Claude Fable 5.1 against GPT-6 Astra. In the human blind choice test, GPT-6 Astra won 35 matchups, representing 70 percent of all tests, while Fable 5.1 won 15. In visual evaluation, the AI judge initially awarded 47 wins to Astra, zero to Fable 5.1, and three ties. To eliminate potential OpenAI bias, Igor re-ran the visual evaluation using Fable 5.1 as the judge, which still awarded 48 wins to Astra and only two to Fable 5.1. In functional testing, Astra passed 48 builds, while Fable 5.1 passed 47.
Category breakdowns highlighted distinct model personalities. GPT-6 Astra swept dashboards with five wins to zero, producing clean, SaaS-like control rooms and global map visualizations. Astra also dominated playable games, generating detailed dungeon crawlers with distinct entity sprites, and swept narrative editorial sites five to zero. Fable 5.1 showed distinct strengths in quirky, artistic formats, winning four out of five experimental interface tests, such as terminal simulations, and three out of five audio tools.
Finally, Igor analyzed total API costs across the 50 builds. GPT-6 Astra was the cheapest at 20.64 dollars total, or roughly 41.3 cents per site. Fable 5 cost 21.24 dollars total, or 42.5 cents per site. Fable 5.1 was the most expensive at 29.15 dollars total, averaging 58.3 cents per site.
Benchmark Setup and Scoring Rubrics
The benchmark tested Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra across 50 distinct website generation prompts covering ten categories, including dashboards, storefronts, and browser games. Every test used a single one-shot prompt sent directly to each model's API without continuations, batching, or caching. Outputs were evaluated across three separate blind rubrics: Igor's personal side-by-side blind preference, an automated AI visual judge rating layout and styling, and an automated functionality audit testing whether prompted features were present and operational.
Claude Fable 5 versus Claude Fable 5.1
In the head-to-head comparison between Anthropic models, Fable 5.1 demonstrated clear design improvements over Fable 5. Fable 5.1 secured 30 human preference wins compared to 17 for Fable 5, alongside three ties. The AI visual judge rated Fable 5.1 superior in 40 out of 50 tests. Both models achieved identical functionality scores by successfully delivering working code for 47 out of 50 prompted web applications, with Fable 5.1 offering cleaner layout spacing, better visual hierarchy, and sharper SVG assets.
Claude Fable 5.1 versus GPT-6 Astra
In the comparison between Fable 5.1 and GPT-6 Astra, Astra won the human preference test decisively with 35 wins against Fable 5.1's 15. The AI visual review favored Astra even more heavily, giving it 47 wins out of 50 under an OpenAI visual evaluator and 48 wins when re-run using Fable 5.1 as the judge. In functional browser testing, Astra successfully executed 48 out of 50 builds, narrowly edging out Fable 5.1's 47 passing builds.
Category Strengths and Generation Costs
GPT-6 Astra dominated structured, professional categories, winning all five matchups in dashboards and editorial narrative sites, and leading in browser games, science simulations, and SaaS-style tools. Fable 5.1 excelled in artistic and unconventional categories, winning four out of five experimental interface tests and leading in audio synthesizers. Across all 50 generated sites, GPT-6 Astra was the most economical model at 20.64 dollars total, compared to 21.24 dollars for Fable 5 and 29.15 dollars for Fable 5.1.
The Bottom Line
The benchmark establishes that OpenAI's GPT-6 Astra outperforms Anthropic's Claude Fable 5.1 in overall visual appeal, dashboard design, game generation, and cost efficiency for one-shot web creation. Claude Fable 5.1 represents a clear aesthetic improvement over Fable 5 and retains a creative edge in experimental and audio interfaces. The comparison leaves open how multi-turn prompt refinement or framework-specific agents might alter these one-shot generation results.
FAQ
What is the website generation benchmark comparing Claude Fable and GPT-6 Astra?
It is a standardized test suite created by Igor to evaluate Claude Fable 5, Claude Fable 5.1, and GPT-6 Astra across 50 one-shot front-end website prompts spanning ten distinct categories.
How did Claude Fable 5 compare to Claude Fable 5.1 in the blind tests?
Claude Fable 5.1 won 30 out of 50 human preference matchups and 40 out of 50 AI visual reviews, while matching Fable 5 with 47 out of 50 functional builds.
Which AI model won the human preference rating between Claude Fable 5.1 and GPT-6 Astra?
GPT-6 Astra won 35 out of 50 blind human preference matchups, capturing 70 percent of Igor's choices compared to 15 wins for Fable 5.1.
How did GPT-6 Astra and Claude Fable 5.1 perform on functional browser code tests?
GPT-6 Astra successfully delivered 48 out of 50 fully functional websites, while Claude Fable 5.1 passed 47 out of 50 builds.
In which specific website categories did Claude Fable 5.1 outperform GPT-6 Astra?
Claude Fable 5.1 outperformed GPT-6 Astra in experimental interfaces with four wins to Astra's one, and in audio and music interfaces with three wins to Astra's two.
What was the total API generation cost for each AI model across the fifty website builds?
GPT-6 Astra cost 20.64 dollars at 41.3 cents per site, Claude Fable 5 cost 21.24 dollars at 42.5 cents per site, and Claude Fable 5.1 cost 29.15 dollars at 58.3 cents per site.
Worth watching for
Front-end developers, product designers, and technical founders looking to evaluate and select the best generative AI models for automated website and web application generation.
- gpt-6-astra
- claude-fable
- web-development
- ai-benchmark
- front-end-code