DeepSeek V4.1 Flash’s design benchmark nearly matches GPT-6 Astra on design quality while costing only 1.4% of GPT-6 Astra’s price. In OpenDesign Arena testing this week, GPT-6 Astra scored 82.7 points and cost $1.61 per finished design, while DeepSeek V4.1 Flash scored 81.2 points and cost $0.023 per finished design. The two models were separated by 1.5 points on the OpenDesign scale.
OpenDesign Arena tested 13 AI models this week on a set of design tasks that included web apps, dashboards, mobile screens, and landing pages. The evaluation used a scoring system with a maximum of 100 points. The testing setup required a working webpage in order to assign a score, and blank or broken webpage outputs received a score of zero. These criteria determined whether a model’s output could be counted in the final score.
GPT-6 Astra, which was released on September 3, was among the models evaluated under this protocol. The same testing procedure and scoring rules were applied uniformly across all entries in the comparison. Only outputs that produced functional webpages were included in the scored results.
The three top AI models in the OpenDesign Arena comparison were GPT-6 Astra, DeepSeek V4.1 Flash, and Claude Fable 5.1. On the OpenDesign Arena scale, GPT-6 Astra scored 82.7 points, DeepSeek V4.1 Flash scored 81.2 points, and Claude Fable 5.1 scored 80.3 points. The scoring used by OpenDesign is out of 100 points.
Measured times to finish designs differed across the three models: GPT-6 Astra finished in 11.1 minutes, DeepSeek V4.1 Flash finished in 5.3 minutes, and Claude Fable 5.1 finished in 12.8 minutes. Reported cost per finished design for the same runs were $1.61 for GPT-6 Astra, $0.023 for DeepSeek V4.1 Flash, and $3.66 for Claude Fable 5.1. All three models were part of the 13-model evaluation conducted by OpenDesign Arena. The reported scores, completion times, and per-design costs were recorded under the same OpenDesign testing protocol.
DeepSeek V4.1 Flash is described as a Causal Encoder-Decoder model with a total of 552 billion parameters. During operation, the model activates 8 billion parameters for reading and 16 billion parameters for writing, reflecting its read and write activation configuration. These architecture and activation details were reported as part of the model specifications.
DeepSeek’s V4 Pro reportedly landed within 5% of Claude Fable 5 in a separate benchmark while charging a fraction of Fable’s rate. V4 Pro is an additional DeepSeek model referenced in the reporting on the company’s model lineup. DeepSeek is recruiting engineers in Beijing to build its own Code Harness software.
OpenDesign Arena testing found that DeepSeek V4.1 Flash closely rivals GPT-6 Astra in design quality while operating at a substantially lower cost. This recap concludes the benchmarking discussion and restates that the two models achieved similar design performance in the OpenDesign comparison while DeepSeek delivered its results at a markedly lower overall per-design expense level.


