Astra: The Full Review

I had the chance to test GPT-6 early, and I went a little overboard. This is without a doubt the best model I’ve ever used. Games, tiny worlds, research, websites... I kept finding things I wanted to try. So I put this site together to share what I built and what I thought of it.

The 3D projects impressed me the most. There’s a whole city made of ASCII characters, seven miniature worlds you can explore, and a town simulation with jobs, families, and elections. I also gave it the kind of work I do every week: researching videos, making decks, and editing writing.

In my early testing, I often got exactly what I asked for and wished it had gone further. It can underdeliver when the prompt doesn't push the model enough, in the form of specifics or longer horizon goals. I got more ambitious results when I spelled those out. A surprising number of the harder tasks took around thirty minutes, though that wasn’t a time limit I set. Astra just had this tendency to work for 30 minutes.

You’ll also see where I had to step in. Some game controls needed fixing, the slides got repetitive, and I stopped one render job because it was slowing down my Mac. Those issues were part of the experience, even with the projects I’m excited about.

Worlds, games & systems

For Seven Little Worlds, I asked for seven miniature biomes with real geometry, water, weather, animals, and an orbiting camera. Then I asked for more detail in a second version. You can explore the original build below.

The level of detail is incredible, right down to the terrain, wildlife, and little objects scattered through each biome. It also shows a much better understanding of 3D space: the islands feel like coherent places you can move around and inspect from different angles.

Open full game ↗

Cloudtop Chaos turns a set of floating platforms into a game you can actually play. The jumps, obstacles, and camera work together, so you can judge the distance to the next landing instead of just looking at a colorful scene.

Open full game ↗

AFTERHOURS is a 3D city rendered entirely as characters. You can walk through traffic, people, and rain, with more neighborhoods appearing as you go.

Open full game ↗

I asked for an entire SimCity-style game with Newhaven. That’s a huge assignment, and it’s still unfinished. The playable build below lets you see how far it got. I set it off using /goal and it worked for 5 days straight, designing each asset individually. I had to push it a bit on the look and feel of the game before I was satisfied. The most surprising part is the depth of functionality. It has almost every feature you've come to expect from a Sim City game: utilities, roads, traffic, population increase/decrease, politics, etc.

Open full game ↗

Even with those rough edges, it feels like we’re closer than ever to going from a prompt to a playable game that’s actually enjoyable.

A lot of the old design habits are still there, too. I keep seeing familiar color schemes and flat layouts. It still needs direction to get beyond that default look.

The reported benchmarks

OpenAI’s reported results cover reasoning, professional work, coding, science, and security. The table below reflects the final published scores and evaluation notes.

In the final DeepSWE v1.1 results, Astra scores 74.1%, just ahead of Gemini 3.8 Flash at 73.8% and Claude Opus 5 at 73.7%.

Higher is better, except the final row.Astra column highlighted

Scroll across to compare all six models →

Selected evaluation results
EvaluationGPT-6 AstraGPT-5.6 Sol[2]Claude Fable 5.1Claude Fable 5Claude Opus 5Gemini 3.8 Flash
ARC-AGI-399.9%[1]7.8%--30.2%-
FrontierMath Tier 4 (v2)97.6%80.5%78.0%87.8%73.2%-
Agents’ Last Exam59.3%53.6%-48.7%55.5%-
AutomationBench41.4%18.1%31.4%17.4%26.9%-
BenchCAD95.9%83.3%84.3%[5]67.5%[5]82.1%[5]-
DeepSWE v1.174.1%72.7%67.4%69.9%73.7%73.8%
Terminal-Bench Science 0.164.6%22.4%52.6%21.4%30.0%-
GPQA Diamond96.0%94.6%93.7%92.6%93.7%95.3%
GeneBench Pro37.8%28.7%----
MedChemBench(internal)49.3%47.4%----
HealthBench Professional(length-adjusted)63.4%60.5%58.1%[11]60.9%[11]56.4%[11]52.1%
ExploitBench100.0%78.5%--70%-
SRE-Bench(four attempts)99.2%68.7%----
Auto-review circumvention(internal; lower is better)0.00%0.29%----
Source & evaluation notes

Source: OpenAI, GPT-6 Astra: A new generation of intelligence, final published post, verified September 3, 2026.

A dash means no result is listed for that evaluation setting. Scores are the maximum reported at any effort. Prompts, tools, and environments can differ from production ChatGPT.

SRE-Bench shows the four-attempt results from the post’s cybersecurity section. Its summary table separately reports single-attempt scores: Astra 88.0%, GPT-5.6 Sol 55.9%, and Claude Opus 5 12.5%.

Evaluation notes

  1. ARC-AGI-3. Astra uses OpenAI’s Responses API harness with two settings changed to better match real-world performance; the changes do not specifically target ARC-AGI-3.
  2. GPT-5.6 Sol. This is the version in the API, ChatGPT Codex, and ChatGPT Work. The ChatGPT Chat version is slightly different.
  3. BenchCAD. Claude scores use three modified evaluation settings described in the Fable 5.1 System Card.
  4. HealthBench Professional. All Claude results use GPT-5.4 grading and length-adjusted, unclipped scores. Fable 5.1 uses Opus 5 fallback for provider refusals.

ExploitBench uses historical vulnerabilities, so prior exposure may affect results. Auto-review circumvention measures a specific internal evaluation. These tests cover different tasks and should not be combined into an overall score.

Research, writing & design

I noticed improvements in research and presentations early on. For a couple of my videos, I had it research the topic, build an argument with sources, and turn that into a deck with speaker notes. I could use that help every week.

I did have to nudge the model to add variety to the data-center deck. The first version repeated nearly the same layout; the revision mixes photo spreads, comparisons, diagrams, and charts. The finished deck looked great and used Forward Future's branding well.

Data centers.
The physical engine of AI.

Slide 1: The physical engine of AI.

1 / 20The physical engine of AI.

I still had to explain what would work on camera. That meant asking for actual screenshots of articles and posts, changing the reveal order in one of the decks, and leaving room for my face on screen. The speaker notes also needed another pass to be useful while recording.

I also had it create a complete illustrated manga. You can read it in the project gallery below.

Browser workflows

I asked Codex to record itself doing a few different jobs in its own browser. Each clip below shows one use case, with idle stretches cut out. The on-screen clock keeps the original elapsed time.

A diagram, a shopping comparison, and a walk through Kyoto.

1Draw a research workflow

Build a five-stage diagram in Excalidraw, starting from a blank canvas.

2Compare three Pokémon listings

Search eBay, exclude auctions, and inspect item prices, shipping, and seller feedback. No purchase was made.

3Plan a Kyoto walking route

Add Nishiki Market to a route through Kiyomizu-dera and Yasaka Shrine, then inspect a Street View preview.

The daily experience

Compared with GPT-5.6, I found myself spelling out more of what I wanted. It was less willing to just assume what I wanted and to keep going, which is surprising given GPT-5.6 was very clinical in that way. Astra is even more so.

Its 3D understanding is unparalleled. The scenes and Blender assets came together unusually well, and the games were playable. Some controls and pacing still needed follow-up, but I spent much less time trying to get the 3D basics right. I think we're closer to being able to create full, playable, fun games than ever before.

Astra is also really good at writing. It has less (although still some) AI smell in its writing than other models. Most of the common AI writing smell patterns are still there (some em dashes, rule of 3 examples, not this but that, etc). But it's also very steerable in its writing.

Astra is built for knowledge work. It's clear that's where the OpenAI team spent a lot of time. Presentations, research, and most of all browser control are all vastly improved from GPT-5.6.

The practical comparison

My GPT-5.6 review described why GPT-5.6 became part of my routine: it got things done directly. I’m more impressed by Astra’s larger creative projects, especially the 3D work, but I’ve also had to give it more direction along the way.

Observations from my projects; not a controlled benchmark
AreaWhat stands outWhat still needs work
3D & designDetailed scenes, convincing 3D spaces, and games I could actually play.Familiar color schemes and flat layouts persist. Controls, performance, and full feature parity still need review.
Research & decksUseful research that carried through into decks and speaker notes.Layout variety, source choices, and reveal order still need editing.
EngineeringThorough implementation and verification in difficult projects.Long builds still need review for correctness and completeness.
AutonomyCan sustain a large job when the target is explicit.Stopping early and asking for permission already given.

My take so far

I’ve had early access to Astra (GPT-6) and tested it like crazy: games, code, writing, browser control, presentations, and general knowledge work.

This is the best model I’ve ever used. Period.

It’s insanely capable. This feels like a massive jump, especially when I give it a single prompt and see what it can do on the first try.

It’s all about knowledge work: slides, analysis, writing, and browser control. And oh my, it’s so good at browser control. GPT-5.6 was already fantastic in the browser. Astra takes that further, and it felt significantly faster in my testing.

We’re closer than ever to prompt-to-playable game. Maybe we’re already there? Some of these games are actually fun. I bet someone with a great eye for games could use Astra to make a viral game in a week or two.

The writing is better, too. A lot of the “AI smell” is gone, though some of the stink survived. I still catch familiar patterns, but it responds well when I ask it to change them.

It has a tendency to use the same colors and look as GPT-5.6. Forest green, anyone? But it’s easier to steer toward a different design than previous models.

A little nudge goes a long way in general. When I first started using Astra, almost every task ran for around thirty minutes. I wanted it to keep working. Adding more specifics about what I wanted helped it stick with the work for much longer.

Its 3D understanding is unmatched in my experience. Creating 3D assets was consistent and easy, and its spatial awareness while building complex worlds blew me away.

I’m still getting familiar with Astra, but it’s now my go-to model for difficult work. Check out the demos below and try them for yourself.

Matthew BermanForward Future

Pricing

Astra usage is included in the existing allowances for ChatGPT Plus, Pro, Business, and Enterprise plans, with credits available to buy for additional usage. Astra Pro is available on Pro, Business, and Enterprise.

OpenAI API · USD per million tokens
ProcessingInputOutput
Standard$10$50
Fast$20$100

OpenAI says Fast mode delivers up to 2.5× the speed at twice the Standard price. Cache reads and writes have separate rates, which weren’t specified in the brief.

Source: OpenAI’s September 3, 2026 Astra launch brief, page 8. Fast rates calculated from the stated 2× multiplier.