DeepSeek's latest model, V4.1 Flash, has come remarkably close to matching OpenAI's flagship GPT-6 Astra in a battery of design tasks—while costing a fraction of the price, according to new benchmark results released by OpenDesign.
The Benchmark Findings
OpenDesign subjected 13 AI models to an identical set of design challenges, scoring each on how well it handled the assignments. GPT-6 Astra came out on top, but DeepSeek V4.1 Flash finished just a point and a half behind, a margin narrow enough to raise eyebrows across the industry.
The headline number, however, is not performance—it's price. DeepSeek's model reportedly costs roughly 70 times less to run than its OpenAI rival, meaning the Chinese lab delivered near-equivalent output at about 1.4% of the cost.
A point and a half separated the two models—but a 70-fold price gap separated their bills.
That kind of cost-to-performance ratio has become DeepSeek's calling card, and the latest results suggest the pattern is holding as newer, more capable models arrive.
Why It Matters
For businesses weighing which AI tools to adopt, the tradeoff between raw capability and operating expense is increasingly central. A model that trails the leader by a slim margin but costs dramatically less can be the more practical choice for teams running tasks at scale.
The comparison also underscores the intensifying rivalry between well-funded U.S. labs and their leaner Chinese competitors. DeepSeek has repeatedly demonstrated that aggressive efficiency can close much of the gap with the market's premium offerings.
Key takeaways from the benchmark:
- GPT-6 Astra led the field of 13 models tested by OpenDesign.
- DeepSeek V4.1 Flash finished 1.5 points behind on design tasks.
- DeepSeek's model ran at roughly 1.4% of the cost of GPT-6 Astra.
As competition heats up, benchmarks like OpenDesign's are likely to play a growing role in shaping which models companies choose to build on.
