All models
DeepSeek V4 Flash 0731
deepseek/deepseek-v4-flash-0731
DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows....
Modalities
Text
In / Out price
$0.14 / $0.28 per 1M
Context
1.05M · 393K out
Released
Jul 31, 2026
Providers
Live list prices and uptime across 28 providers serving this model.
| Provider | Input /M | Output /M | Cache read /M | Context | Max output | Uptime 24h |
|---|---|---|---|---|---|---|
| Decartfp4 | $0.06 | $0.121 | $0.012 | 262K | 262K | 96.97% |
| StreamLakefp8 | $0.061 | $0.123 | $0.012 | 1.02M | 384K | 96.92% |
| DigitalOcean | $0.08 | $0.252 | $0.025 | 1.05M | — | 96.53% |
| DeepInfrafp8 | $0.08 | $0.18 | $0.016 | 1.05M | 384K | 99.87% |
| OpenInferencefp4 | $0.08 | $0.18 | $0.028 | 262K | 262K | 99.51% |
| GMICloudfp8 | $0.084 | $0.168 | $0.017 | 1.05M | — | 98.08% |
| Sail Researchfp4 | $0.09 | $0.18 | $0.02 | 262K | 262K | 93.98% |
| Relacefp4 | $0.105 | $0.21 | $0.021 | 1.05M | 1.05M | 91.32% |
| BaseTenfp8 | $0.13 | $0.26 | $0.028 | 1.05M | 1.05M | 99.88% |
| CoreWeavefp8 | $0.13 | $0.28 | $0.07 | 262K | 262K | 99.96% |
| Inceptronfp4 | $0.13 | $0.28 | $0.03 | 1.05M | 1.05M | 97.87% |
| Morphbf16 | $0.139 | $0.278 | $0.07 | 1.05M | 1.05M | 93.45% |
| Fireworks | $0.14 | $0.28 | $0.028 | 1.05M | — | 95.13% |
| AkashMLfp8 | $0.14 | $0.28 | $0.02 | 131K | 131K | 96.62% |
| Novitafp8 | $0.14 | $0.28 | $0.028 | 1.05M | 393K | 92.01% |
| Together | $0.14 | $0.28 | $0.03 | 1.05M | — | 96.73% |
| Parasailfp8 | $0.14 | $0.28 | $0.07 | 1.05M | 1.05M | 98.19% |
| AtlasCloudfp4 | $0.14 | $0.28 | $0.028 | 1.05M | 393K | 98.78% |
| SiliconFlowfp8 | $0.14 | $0.28 | $0.028 | 1.05M | 393K | 98.70% |
| Ambientfp4 | $0.14 | $0.28 | $0.028 | 1.05M | 1.05M | 97.41% |
| Baidufp8 | $0.14 | $0.28 | $0.028 | 1.05M | 131K | 99.48% |
| Io Netfp8 | $0.149 | $0.32 | $0.077 | 262K | 66K | 99.75% |
| Mancer 2fp8 | $0.15 | $0.50 | — | 1.05M | 1.05M | 93.90% |
| Venice | $0.175 | $0.35 | $0.035 | 1M | 33K | 93.91% |
| Phala | $0.20 | $0.40 | $0.07 | 1.05M | 393K | 96.20% |
| DeepSeekfp8 | $0.22 | $0.66 | $0.007 | 1.05M | 384K | 99.84% |
| Wafer | $0.28 | $0.56 | $0.07 | 1.05M | 1.05M | 98.73% |
| Cloudflare | $0.44 | $1.32 | $0.014 | 1.05M | 1.05M | 99.97% |
Implicit prompt caching available — repeated context bills at the cache read rate.
Capabilities
Tool calling
Reasoning
frequency_penalty
include_reasoning
logit_bias
logprobs
max_tokens
min_p
parallel_tool_calls
presence_penalty
reasoning
reasoning_effort
repetition_penalty
response_format
seed
stop
structured_outputs
temperature
tool_choice
top_a
top_k
top_logprobs
top_p
Pricing detail
- Input
- $0.14
- Output
- $0.28
- Cache read
- $0.028
- Cache write
- —
USD per 1M tokens.