Skip to main content

Quick Overview

In this video, tech creator Matt Wolfe provides a first-look review and practical demonstration of OpenAI's GPT-6 Astra model. The video follows the initial announcement of the model and examines early benchmark results alongside live tests of its coding and autonomous computer-use features. It was created to demonstrate how the model performs in real-world application building compared to other leading AI systems.

Key Points

  • 1.OpenAI announced GPT-6 Astra, rolling out initially to select organizations before reaching all ChatGPT Plus, Pro, Business, and Enterprise users.
  • 2.On benchmark evaluations, GPT-6 Astra achieved 99.9 percent on ARC-AGI-3 and 84.6 percent on Terminal Bench Science 0.1, showing major reasoning improvements.
  • 3.On coding benchmarks like DeepSWE, GPT-6 Astra scored 74.1 percent, slightly below Meta Muse Spark 1.3 at 75.4 percent.
  • 4.Operating through computer-use capabilities, the model autonomously created a 3D humanoid wolf character in Blender, added a 50-bone rig, and imported it into Unreal Engine as a playable character.
  • 5.GPT-6 Astra generated full interactive Three.js web applications, including a 3D planetary simulation and a playable 3D Megabonk game clone, in under ten minutes each.

Summary

Matt Wolfe provides a hands-on review and benchmark analysis of OpenAI's newly announced GPT-6 Astra model following early access testing. OpenAI announced that GPT-6 Astra is rolling out to limited organizations before expanding to ChatGPT Plus, Pro, Business, and Enterprise users, as well as the API. While the model achieved massive score jumps on specific benchmarks such as ARC-AGI-3, where it scored 99.9 percent compared to 7.8 percent on GPT-5.6 Sol, and Terminal Bench Science 0.1 at 84.6 percent, coding benchmarks showed more competitive results. On DeepSWE, GPT-6 Astra scored 74.1 percent, matching Gemini 3.8 Flash and Claude Opus 5, but trailing Meta Muse Spark 1.3, which achieved 75.4 percent. On the Artificial Analysis Intelligence Index, GPT-6 ranked fifth overall at an estimated cost of 1.67 dollars per task.

Wolfe tests GPT-6 Astra on BuseyBench, an SVG code generator benchmark judged by an independent AI evaluator. GPT-6 Astra took first place with a 7.1 score, producing a detailed vector graphic in nine minutes using 63,858 tokens at an estimated cost of 1.94 dollars. Testing game creation capabilities, Wolfe prompted the model to build a 3D clone of the game Megabonk using Three.js. GPT-6 generated a fully playable game with multiple character classes, sound effects, level upgrades, and weapon mechanics in eight minutes, a task that previously took up to two hours with older models. Another prompt generated Orbis, an interactive 3D world simulator allowing real-time adjustments to sunlight, sea levels, rainfall, and terrain modification.

The most notable demonstrations involved GPT-6 Astra using its computer-use capabilities to manipulate professional 3D software directly. When prompted to create a 3D humanoid wolf, the model autonomously operated Blender, modeled the character, and saved it in eight minutes. In a follow-up prompt, it added a 50-bone animation rig and generated an eight-second looping running animation in six minutes. Wolfe then tasked GPT-6 with creating a game in Unreal Engine. Over 35 minutes, the model launched Unreal Engine, built an environment named Whisperwood with trees, trails, and a pond, and imported the animated wolf as a fully playable third-person character with keyboard navigation controls.

The video concludes with a showcase of demonstrations published by other early access testers on social media. These include a Little Planet Three.js app by Matt Berman, a 3D Fall Guys clone, a detailed Manhattan environment by Matt Shumer, an underwater Atlantis explorer by Pietro Schirano, and an interactive 3D historical tour of the Library of Alexandria by Ethan Mollick. Wolfe notes that despite minor animation flaws and mixed positioning across broad intelligence indexes, GPT-6 Astra represents a significant leap forward in autonomous tool execution and interactive application development.

Release Details and Benchmark Results

GPT-6 Astra marks a major release from OpenAI, launching first to a limited group before expanding to ChatGPT paid tiers and API access. Benchmark results indicate substantial jumps across multiple evaluations, including saturated scores on GPQA Diamond at 98.0 percent and ARC-AGI-3 at 99.9 percent. In coding evaluations such as DeepSWE, GPT-6 Astra achieved 74.1 percent, closely matching Claude Opus 5 and Gemini 3.8 Flash, while slightly trailing Meta Muse Spark 1.3. On the aggregated Artificial Analysis Intelligence Index, GPT-6 ranks fifth, remaining closely tied with GPT-5.6 Sol at an estimated run cost of 1.67 dollars per task.

SVG Generation and Web App Building

Testing the model on BuseyBench for SVG generation, GPT-6 took first place with an AI judge score of 7.1 out of 10, using 63,858 tokens across nine minutes. When tasked with creating games and interactive web pages via single prompts, GPT-6 built a fully functional 3D Megabonk clone using Three.js in eight minutes, compared to older workflows taking one to two hours. It also created an interactive planetary ecosystem simulator called Orbis, featuring sliders for environmental controls, terraforming tools, and population dynamics.

Computer Use in Blender and Unreal Engine

GPT-6 demonstrated autonomous computer use by controlling external software tools directly. Given text prompts, it opened Blender, modeled a 3D humanoid wolf figure, rigged it with a 50-bone structure, and generated running animations. Following up with another prompt, the model opened Unreal Engine, constructed a playable 3D forest environment called Whisperwood, and imported the rigged wolf character with full keyboard movement controls, completing the entire Unreal Engine setup in 35 minutes.

Community Creations and Early Reception

Early access users shared several complex builds completed with GPT-6 Astra, including an interactive Little Planet Three.js app, a full Manhattan street environment in Unreal Engine, a 3D horror maze game, a Fall Guys obstacle course clone, and a multi-agent simulation where characters autonomously communicated in speech. Historian Ethan Mollick also demonstrated an interactive 3D reconstruction of the ancient Library of Alexandria with narrated audio tours.

The Bottom Line

The video establishes that GPT-6 Astra provides substantial improvements in autonomous task execution, particularly when controlling desktop applications like Blender and Unreal Engine. It demonstrates that the model can generate complex Three.js web apps and 3D assets in minutes rather than hours. However, it leaves unresolved why the model ranks fifth on the aggregate Artificial Analysis index despite outperforming older models in practical software manipulation.

FAQ

What is GPT-6 Astra and what new capabilities does it introduce?

GPT-6 Astra is a flagship AI model from OpenAI featuring advanced computer-use capabilities that allow it to autonomously operate external software like Blender and Unreal Engine, generate complex Three.js applications, and perform high-level reasoning across technical benchmarks.

How did GPT-6 Astra perform on the ARC-AGI-3 reasoning benchmark?

GPT-6 Astra scored 99.9 percent on the ARC-AGI-3 benchmark, representing a significant increase over GPT-5.6 Sol's score of 7.8 percent and the average human tester score of 48 percent.

How did GPT-6 Astra compare against Meta Muse Spark 1.3 on DeepSWE?

On the DeepSWE coding benchmark, GPT-6 Astra achieved a score of 74.1 percent, slightly behind Meta Muse Spark 1.3, which scored 75.4 percent according to Meta's reported data.

How does GPT-6 Astra rank on the Artificial Analysis Intelligence Index?

GPT-6 Astra ranked in fifth place on the Artificial Analysis Intelligence Index, tying closely with GPT-5.6 Sol, with an estimated average cost of 1.67 dollars per task.

How long did GPT-6 Astra take to build the playable Whisperwood Unreal Engine environment?

GPT-6 Astra took 35 minutes to construct the Whisperwood forest world in Unreal Engine, import the custom rigged wolf character, and establish playable character movement controls.

Worth watching for

Developers, AI researchers, and technical creators interested in evaluating the real-world coding, benchmark scores, and autonomous computer-use capabilities of OpenAI's GPT-6 Astra model.

  • gpt-6-astra
  • openai
  • computer-use
  • artificial-intelligence
  • unreal-engine
  • blender