DeepSeek V4 Flash 0731: New 304B Model Offers Strong Agentic Capabilities at Low Cost
A new open-weight model from DeepSeek delivers competitive intelligence at a fraction of the price, highlighting the value of comparing model cost and performance.
DeepSeek has released a new model in its V4 family, DeepSeek-V4-Flash-0731, which the company describes as having "substantially enhanced agentic capabilities." The model has 304 billion parameters and weighs 167GB on Hugging Face. Early benchmarks from Artificial Analysis rank it ahead of MiniMax M3, a 428-billion-parameter model, suggesting it punches above its weight in terms of intelligence per parameter.
The model is priced at $0.14 per million input tokens and $0.27 per million output tokens, which may make it the best value-per-intelligence model currently available. Independent testing via OpenRouter showed that output quality depends on the reasoning level used: a default setting produced a disappointing result, while a high reasoning effort setting yielded much better performance. The model is available through OpenRouter and Hugging Face.
Why it matters
This release illustrates the accelerating trend of open-weight models competing with proprietary systems on both capability and cost. For developers and organizations, the ability to run a high-performing model at a fraction of the price of alternatives changes the economics of building AI-powered applications. It also highlights that model selection is not just about raw parameter count—pricing, agentic features, and reasoning settings all play a role in real-world value.
This release illustrates the accelerating trend of open-weight models competing with proprietary systems on both capability and cost.
What you can learn from this
- Model comparison requires more than parameter count: A 304B model can outperform a 428B model on certain tasks. Learners should practise evaluating models using standardized benchmarks (like Artificial Analysis) and real-world tests relevant to their use case, not just size.
- Reasoning effort settings affect output quality: The same model can produce very different results depending on how much reasoning effort is requested. When building applications, experiment with different reasoning levels to find the best balance between quality and cost for your task.
- Agentic capabilities are a growing differentiator: Models are increasingly being evaluated on their ability to plan, use tools, and take actions—not just generate text. When choosing a model, look for documentation on agentic features like function calling, multi-step reasoning, and tool use.
- Pricing per token matters for production deployments: At $0.14/M input and $0.27/M output, this model is significantly cheaper than many alternatives. Learners should calculate total cost for their expected usage volume and compare it against performance to make informed decisions.
- Open-weight models lower the barrier to entry: Models available on Hugging Face and through providers like OpenRouter allow developers to experiment without large upfront costs. Practise pulling a model from Hugging Face and running inference locally or through an API to understand the workflow.
We teach this
Sources
- deepseek-ai/DeepSeek-V4-Flash-0731 — Simon Willison
Our reporting is an original summary; full coverage is at the links above.
Don't just read about it — build it.
Square 1 teaches the skills behind the headlines, with every line of your work graded by AI. Find your starting point in 3 minutes.
Get your free skill report