I had the chance to test GPT-6 early, and I went a little overboard. This is without a doubt the best model I’ve ever used. Games, tiny worlds, research, websites... I kept finding things I wanted to try. So I put this site together to share what I built and what I thought of it.
The 3D projects impressed me the most. There’s a whole city made of ASCII characters, seven miniature worlds you can explore, and a town simulation with jobs, families, and elections. I also gave it the kind of work I do every week: researching videos, making decks, and editing writing.
In my early testing, I often got exactly what I asked for and wished it had gone further. It can underdeliver when the prompt doesn't push the model enough, in the form of specifics or longer horizon goals. I got more ambitious results when I spelled those out. A surprising number of the harder tasks took around thirty minutes, though that wasn’t a time limit I set. Astra just had this tendency to work for 30 minutes.
You’ll also see where I had to step in. Some game controls needed fixing, the slides got repetitive, and I stopped one render job because it was slowing down my Mac. Those issues were part of the experience, even with the projects I’m excited about.
Worlds, games & systems
For Seven Little Worlds, I asked for seven miniature biomes with real geometry, water, weather, animals, and an orbiting camera. Then I asked for more detail in a second version. You can explore the original build below.
The level of detail is incredible, right down to the terrain, wildlife, and little objects scattered through each biome. It also shows a much better understanding of 3D space: the islands feel like coherent places you can move around and inspect from different angles.
Cloudtop Chaos turns a set of floating platforms into a game you can actually play. The jumps, obstacles, and camera work together, so you can judge the distance to the next landing instead of just looking at a colorful scene.
AFTERHOURS is a 3D city rendered entirely as characters. You can walk through traffic, people, and rain, with more neighborhoods appearing as you go.
I asked for an entire SimCity-style game with Newhaven. That’s a huge assignment, and it’s still unfinished. The playable build below lets you see how far it got. I set it off using /goal and it worked for 5 days straight, designing each asset individually. I had to push it a bit on the look and feel of the game before I was satisfied. The most surprising part is the depth of functionality. It has almost every feature you've come to expect from a Sim City game: utilities, roads, traffic, population increase/decrease, politics, etc.
Even with those rough edges, it feels like we’re closer than ever to going from a prompt to a playable game that’s actually enjoyable.
A lot of the old design habits are still there, too. I keep seeing familiar color schemes and flat layouts. It still needs direction to get beyond that default look.
The reported benchmarks
OpenAI’s reported results cover reasoning, professional work, coding, science, and security. The table below reflects the final published scores and evaluation notes.
In the final DeepSWE v1.1 results, Astra scores 74.1%, just ahead of Gemini 3.8 Flash at 73.8% and Claude Opus 5 at 73.7%.
Scroll across to compare all six models →
| Evaluation | GPT-6 Astra | GPT-5.6 Sol[2] | Claude Fable 5.1 | Claude Fable 5 | Claude Opus 5 | Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| ARC-AGI-3 | 99.9%[1] | 7.8% | - | - | 30.2% | - |
| FrontierMath Tier 4 (v2) | 97.6% | 80.5% | 78.0% | 87.8% | 73.2% | - |
| Agents’ Last Exam | 59.3% | 53.6% | - | 48.7% | 55.5% | - |
| AutomationBench | 41.4% | 18.1% | 31.4% | 17.4% | 26.9% | - |
| BenchCAD | 95.9% | 83.3% | 84.3%[5] | 67.5%[5] | 82.1%[5] | - |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% | 69.9% | 73.7% | 73.8% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% | 21.4% | 30.0% | - |
| GPQA Diamond | 96.0% | 94.6% | 93.7% | 92.6% | 93.7% | 95.3% |
| GeneBench Pro | 37.8% | 28.7% | - | - | - | - |
| MedChemBench(internal) | 49.3% | 47.4% | - | - | - | - |
| HealthBench Professional(length-adjusted) | 63.4% | 60.5% | 58.1%[11] | 60.9%[11] | 56.4%[11] | 52.1% |
| ExploitBench | 100.0% | 78.5% | - | - | 70% | - |
| SRE-Bench(four attempts) | 99.2% | 68.7% | - | - | - | - |
| Auto-review circumvention(internal; lower is better) | 0.00% | 0.29% | - | - | - | - |
Source & evaluation notes
Source: OpenAI, GPT-6 Astra: A new generation of intelligence, final published post, verified September 3, 2026.
A dash means no result is listed for that evaluation setting. Scores are the maximum reported at any effort. Prompts, tools, and environments can differ from production ChatGPT.
SRE-Bench shows the four-attempt results from the post’s cybersecurity section. Its summary table separately reports single-attempt scores: Astra 88.0%, GPT-5.6 Sol 55.9%, and Claude Opus 5 12.5%.
Evaluation notes
- ARC-AGI-3. Astra uses OpenAI’s Responses API harness with two settings changed to better match real-world performance; the changes do not specifically target ARC-AGI-3.
- GPT-5.6 Sol. This is the version in the API, ChatGPT Codex, and ChatGPT Work. The ChatGPT Chat version is slightly different.
- BenchCAD. Claude scores use three modified evaluation settings described in the Fable 5.1 System Card.
- HealthBench Professional. All Claude results use GPT-5.4 grading and length-adjusted, unclipped scores. Fable 5.1 uses Opus 5 fallback for provider refusals.
ExploitBench uses historical vulnerabilities, so prior exposure may affect results. Auto-review circumvention measures a specific internal evaluation. These tests cover different tasks and should not be combined into an overall score.
Research, writing & design
I noticed improvements in research and presentations early on. For a couple of my videos, I had it research the topic, build an argument with sources, and turn that into a deck with speaker notes. I could use that help every week.
I did have to nudge the model to add variety to the data-center deck. The first version repeated nearly the same layout; the revision mixes photo spreads, comparisons, diagrams, and charts. The finished deck looked great and used Forward Future's branding well.
Data centers.
The physical engine of AI.

1 / 20The physical engine of AI.
I still had to explain what would work on camera. That meant asking for actual screenshots of articles and posts, changing the reveal order in one of the decks, and leaving room for my face on screen. The speaker notes also needed another pass to be useful while recording.
I also had it create a complete illustrated manga. You can read it in the project gallery below.
Browser workflows
I asked Codex to record itself doing a few different jobs in its own browser. Each clip below shows one use case, with idle stretches cut out. The on-screen clock keeps the original elapsed time.
A diagram, a shopping comparison, and a walk through Kyoto.
1Draw a research workflow
Build a five-stage diagram in Excalidraw, starting from a blank canvas.
2Compare three Pokémon listings
Search eBay, exclude auctions, and inspect item prices, shipping, and seller feedback. No purchase was made.
3Plan a Kyoto walking route
Add Nishiki Market to a route through Kiyomizu-dera and Yasaka Shrine, then inspect a Street View preview.
The daily experience
Compared with GPT-5.6, I found myself spelling out more of what I wanted. It was less willing to just assume what I wanted and to keep going, which is surprising given GPT-5.6 was very clinical in that way. Astra is even more so.
Its 3D understanding is unparalleled. The scenes and Blender assets came together unusually well, and the games were playable. Some controls and pacing still needed follow-up, but I spent much less time trying to get the 3D basics right. I think we're closer to being able to create full, playable, fun games than ever before.
Astra is also really good at writing. It has less (although still some) AI smell in its writing than other models. Most of the common AI writing smell patterns are still there (some em dashes, rule of 3 examples, not this but that, etc). But it's also very steerable in its writing.
Astra is built for knowledge work. It's clear that's where the OpenAI team spent a lot of time. Presentations, research, and most of all browser control are all vastly improved from GPT-5.6.
































