umans/status/umans-deepseek-v4-pro-0813
Live · refreshes every 30s
← all models
Umans DeepSeek V4 Pro Deprecated
umans-deepseek-v4-pro-0813 · DeepSeek-V4-Pro-0813 · DeepSeek
deprecated · serving until Sep 14, 2026 · successor umans-glm-5.3
Operational
138.6tok/s
throughput · p50 · last 5 min
436ms
TTFT · p50 · last 5 min
100.00%
uptime · 24h

DeepSeek V4 Pro, served from the official 0813 release: DeepSeek's flagship coding and reasoning MoE, built for long-horizon agentic coding and demanding tool-heavy workloads on a 1M-token context window. The 0813 release supersedes the April preview with substantially stronger agentic performance. Reasoning has three modes: non-think (none), think high (high, the default) and think max (max). Billed per token ($1.32 / $3.96 / $0.044 per 1M; input / output / cache read). Served on our own GPU infrastructure with high availability. Deprecated: use `umans-glm-5.3` instead (sunset 2026-09-14).

90 days agoin production since Aug 15, 2026today
Context
1049K
Max output
393K
Recommended
393K
Vision
No
Tools
Yes
Reasoning
Toggle · none/high/max
Trends

Speed over the last 90 days

daily medians · dashed line = target
throughput p50 · output tokens per second, higher is better
90 days agopre-release before Aug 15, 2026today
TTFT p50 · time to first token, lower is better
90 days agopre-release before Aug 15, 2026today
Changelog

Events for Umans DeepSeek V4 Pro

incl. gateway-wide announcements
Aug 152026
Released pay-per-token: Umans DeepSeek V4 Pro Released
umans-deepseek-v4-pro-0813 joins the lineup as the long-context coding flagship: DeepSeek's official 0813 release of V4 Pro, the checkpoint that served here as the seat-gated pre-release lab since August 13, with a 1M context window and thinking at high effort by default (dial to max when a task deserves more). Billed per token: $1.32 / $3.96 / $0.044 per 1M (input / output / cache read). It succeeds GLM 5.2, which is deprecated and sunsets on August 23, 2026. The pre-release window's metrics stay on the model's status page as its pre-release period (before Aug 15). Served on our own GPU infrastructure with high availability.