umans/status/umans-deepseek-v4-flash-0731
Live · refreshes every 30s
← all models
Umans DeepSeek V4 Flash
umans-deepseek-v4-flash-0731 · DeepSeek-V4-Flash · DeepSeek
Operational
349.4tok/s
throughput · p50 · last 5 min
1.55s
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4 Flash: DeepSeek's fast agentic coding MoE (284B total, 13B active), served from the official 0731 release on a 1M-token context. The cheapest production model in the lineup for real agentic work. Reasoning has four modes: non-think (none), think low (low, the default), think high (high) and think max (max). Served on our own GPU infrastructure with high availability.

90 days agoin production since Aug 3, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/low/high/max
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
peak 317.7 tok/s · Aug 1now 317.4 tok/s
90 days agopre-release before Aug 3, 2026today
TTFT p50 · time to first token, lower is better
best 1.50s · Aug 1now 1.75s
90 days agopre-release before Aug 3, 2026today
Changelog

Events for Umans DeepSeek V4 Flash

incl. gateway-wide announcements
Aug 32026
Released pay-per-token: Umans DeepSeek V4 Flash Released
umans-deepseek-v4-flash-0731 joins the lineup as the cheapest way we serve real agentic work: $0.14 / $0.28 / $0.028 per 1M (input / output / cache read), a 1M context window, thinking at low effort by default (dial up high or max when a task deserves more). It is the new default for new chats and CLI setups. Founding users pay the 10x cheaper cache rate until Monday, August 10, 2026 (see /pricing). Served on our own GPU infrastructure with high availability.