An engineer in a striped beanie holding a 3D-printed impeller up to the light in a home workshop, CAD wireframe glowing on a monitor behind, warm work-lamp lighting
News

Engineers Say GPT-6 Astra's Viral Machine Parts Aren't Safe to Build

Astra's machine-part videos raced across social media this week. Behind them: a benchmark score earned only with self-testing, a 96.1%-accurate part that was still wrong, and engineers who say fixing the output takes longer than starting over.

OpenAIGPT-6 AstraCADBenchmarksEngineering

Since GPT-6 Astra launched on 3 September, the videos have been unstoppable: a rotating turbocharger split into working parts, a camera exploded into more than a hundred components, a car separated into several hundred individually named pieces — all generated from a single prompt. One robotics designer used it to lay out a full competition robot intake in about seven hours. Engineers who have actually tested the tool, though, tell a different story, and International Business Times reported it today: by their assessment, almost none of those viral parts are safe to manufacture as they come out of the model.

The score behind the hype had a condition

The number that set all this off was BenchCAD, a benchmark measuring how closely an AI-built part matches a target shape. Astra scored 95.9 per cent, against 83.3 for the earlier GPT-5.6 Sol and 84.3 for Anthropic’s Claude Fable 5.1 — though OpenAI noted the Claude comparison used modified evaluation settings.

The condition attached to that headline figure rarely makes it into the clips. Astra hit 95.9 per cent only in an agentic setting, where the model could render the part, measure it, and try again before submitting. In the plain setting with no tools, Astra has no published score at all, and GPT-5.6 Sol still leads. A model checking its own work is a genuine step forward. It is a different claim from “AI can now do engineering,” which is the claim the videos are doing on social media.

Where the gap shows up

The BenchCAD researchers documented a part that scored 96.1 per cent geometric similarity and was still the wrong part — a later feature quietly closed an opening the design was meant to keep. Shape-match is not function, and the metric cannot tell the difference.

A separate 2026 benchmark, neuralCAD-Edit, measured frontier models against ten professional designers on real editing tasks. The best model finished 53 percentage points behind the humans. Robotics builders reviewing Astra’s output on public forums were blunter still: the results looked impressive, several said, but would take more work to fix than drawing the part correctly from scratch.

Cost is its own constraint. OpenAI lists Astra at $10 per million input tokens and $50 per million output, and one tester estimated a single seven-hour design run burned through a few hundred dollars in usage.

What stands out to me is that this is the second time this month a headline Astra number has shrunk on inspection. Earlier this month, mathematicians flagged that two of the model’s celebrated proofs leaned on published work nobody had credited — we covered that row here. The pattern isn’t dishonesty by default; it’s that benchmark theatre travels faster than the fine print. Agentic self-testing genuinely improves results, and it also makes the base capability harder to see.

What Astra actually changes

Most experts expect the model to sit above CAD software rather than replace it: Astra reads instructions and writes code while the CAD system holds the exact geometry, dimensions and constraints. That still matters for anyone downstream. In safety-critical work — aerospace, medical devices — a part that passes every check inside the software can still fail inspection or fail in service, and no benchmark catches that. The oldest rule in engineering survives the newest model: a qualified human signs off the final drawing.

For New Zealand’s small manufacturing shops and product designers, the practical read is unglamorous. A shape-matched render is not a manufacturable part, a seven-hour generation run is not a design cycle, and the sign-off stays with a chartered engineer regardless of how good the demo looked. The model’s wider launch already came with saturated benchmarks and an access rollout most people couldn’t use; the machine-part videos are the same story wearing a workshop jacket.

— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.

FAQ

What is BenchCAD? A benchmark that measures how closely an AI-designed part matches a target shape. GPT-6 Astra scored 95.9 per cent — but only in an agentic setting where the model could render, measure and retry its own work.

Can GPT-6 Astra design parts you can actually manufacture? Not reliably, engineers say. A part scoring 96.1 per cent geometric similarity was still functionally wrong, and the best model trailed professional designers by 53 percentage points on real editing tasks in neuralCAD-Edit.

How much does GPT-6 Astra cost to run? OpenAI lists $10 per million input tokens and $50 per million output. One tester estimated a seven-hour design session consumed a few hundred dollars in usage.

Sources: https://www.ibtimes.co.uk/openai-gpt6-astra-ai-design-tool-1818412