NewDeepSeek V4.1 Flash is live

Frontier open models, engineered for AI agents.

The best open coding models, served fast from our in-house inference stack. Keep your tools. Keep your data.

Always frontier

Frontier models, the day they drop

The strongest open coding models land on our endpoint as they release, and they hold their own against the closed frontier on DeepSWE. Your setup never changes.

Serving now
Kimi K3Kimi K3GLM 5.3GLM 5.3DeepSeek V4.1 FlashDeepSeek V4.1 FlashThe next frontier model · live day one
DeepSWE · max reasoning
deepseek-v4.1-flashon Umans74.2%
gpt-5.6-sol73%
claude-fable-570%
kimi-k3on Umans69%
claude-opus-568.8%
glm-5.3on Umans66.9%

Source: DeepSWE 1.1 leaderboard, September 2026.

Best SLOs

Fast, by engineering

We do research and develop our inference stack in-house, from kernels to serving. Measured on DeepSeek V4 Flash 0731, it delivers 361.6 output tokens/s per user, 3x the next route.

Speed · output tokens/s per userZDRnon-ZDR
Umans AI361.6$0.0327
Baidu Qianfan120$0.2994
Baseten (US)100$0.0801
Baseten81$0.0750
Wafer74$0.0799
Makora71$0.0796
NovitaAI70$0.0614
Alibaba Cloud Int.65$0.0397

DeepSeek V4 Flash 0731 · 130k input + 700 output · top 8 routes by speed · effective price per 1M total tokens · ZDR per provider defaults · Umans AI measurements, 14 Sep 2026

Keep your setup

Use it in the tool you already use

Pick your tool. Setup is one command or two env vars.

OpenAI API
Any tool that takes an OpenAI-compatible endpoint. Copy, paste, add your key.
OpenAI-compatible
OPENAI_BASE_URL=https://api.code.umans.ai/v1
OPENAI_API_KEY=umans_••••••••
Or skip setup entirely

Cloud agents

Early release

A coding agent on our infra in one click, wired to our inference with the umans CLI configured. Sandboxed, reachable from any device, still running while your laptop is off.

Open cloud agents

Your data stays yours

Zero data retention, by default

We produce tokens, we don't monetize data. Prompts and code are never stored and never used for training, and independent security audits are underway.

Zero data retentionPrompts and code are never stored.Enforced
No training on your dataYour code is never used to train models.Enforced
ISO 27001Certification in progress.In progress
SOC 2 Type IICertification in progress.In progress
EU data residencyGPU capacity in Europe. Join the list below.On request

Loved by builders

Don’t take our word for it

Umans.ai is an outlier… in a good way. Other AI companies are some weird combination of very high token prices for marginally different models and/or some form of "we're harvesting your prompts… you're the product not the customer." Umans has a clear business model: selling tokens for open models. The prices are competitive, but the difference for me is trust. I'm the customer and not the product with Umans.
DPDavid PollakCEO, Spice Labs
I’ve been using your DeepSeek deployment for these past weeks. I’ve also tried Hyper, CommandCode, OpenCode and Arli, all with DeepSeek. Your work is unmatched.
Nonono· Umans Discord community
How does umans achieve high tps? I feel like it’s faster than cerebras atp
Gojim_· Umans Discord community
made use of that promo, damn, ds4f is flying, great job ❤️
blong· Umans Discord community

Pricing

Pay per token.

Run complex, long-running engineering tasks on frontier open models. Top up, create a key, and go.

ModelInput$ per 1M tokensOutput$ per 1M tokensCache read$ per 1M tokens
DeepSeek V4 FlashFeatured

Our fastest model for agentic and coding ⚡ Long agentic runs stay snappy on a 1M context window, at the lowest price in the lineup.

$0.14$0.28$0.028
Kimi K3

The first open 3T-class model, neck-and-neck with the closed frontier on agentic coding. Its 1M window reads your whole repo in one pass, and native vision takes screenshots and diagrams.

$3.00$15.00$0.30
GLM 5.3

Z.ai's flagship coding model, the GLM 5.2 successor, with large gains on complex, long-horizon tasks and a 1M context window. Always thinks; dial reasoning low to max.

$1.40$4.40$0.26

USD per 1M tokens. Every request bills exactly what it uses: input, output, and cache. Full pricing

Top up & start

On request

EU data residency

We are lighting up new GPU capacity in Europe. Same frontier open models, same zero data retention, with residency in the EU for teams that need it.

Early allocations are limited: tell us the volume you would commit to, and we will prioritize yours.

Volume commitments get priority when we allocate EU capacity.

  • Lengow
  • Billy
  • Lengow
  • Billy