Frontier open models,
engineered for AI agents.
The best open coding models, served fast from our in-house inference stack. Keep your tools. Keep your data.
Always frontier
Frontier models, the day they drop
The strongest open coding models land on our endpoint as they release, and they hold their own against the closed frontier on DeepSWE. Your setup never changes.
Source: DeepSWE 1.1 leaderboard, September 2026.
Best SLOs
Fast, by engineering
We do research and develop our inference stack in-house, from kernels to serving. Measured on DeepSeek V4 Flash 0731, it delivers 361.6 output tokens/s per user, 3x the next route.
DeepSeek V4 Flash 0731 · 130k input + 700 output · top 8 routes by speed · effective price per 1M total tokens · ZDR per provider defaults · Umans AI measurements, 14 Sep 2026
Keep your setup
Use it in the tool you already use
Pick your tool. Setup is one command or two env vars.
Cloud agents
Early releaseA coding agent on our infra in one click, wired to our inference with the umans CLI configured. Sandboxed, reachable from any device, still running while your laptop is off.
Your data stays yours
Zero data retention, by default
We produce tokens, we don't monetize data. Prompts and code are never stored and never used for training, and independent security audits are underway.
Loved by builders
Don’t take our word for it
“Umans.ai is an outlier… in a good way. Other AI companies are some weird combination of very high token prices for marginally different models and/or some form of "we're harvesting your prompts… you're the product not the customer." Umans has a clear business model: selling tokens for open models. The prices are competitive, but the difference for me is trust. I'm the customer and not the product with Umans.”
“I’ve been using your DeepSeek deployment for these past weeks. I’ve also tried Hyper, CommandCode, OpenCode and Arli, all with DeepSeek. Your work is unmatched.”
“How does umans achieve high tps? I feel like it’s faster than cerebras atp”
“made use of that promo, damn, ds4f is flying, great job ❤️”
Pricing
Pay per token.
Run complex, long-running engineering tasks on frontier open models. Top up, create a key, and go.
| Model | Input$ per 1M tokens | Output$ per 1M tokens | Cache read$ per 1M tokens |
|---|---|---|---|
DeepSeek V4 FlashFeatured Our fastest model for agentic and coding ⚡ Long agentic runs stay snappy on a 1M context window, at the lowest price in the lineup. | $0.14 | $0.28 | $0.028 |
Kimi K3 The first open 3T-class model, neck-and-neck with the closed frontier on agentic coding. Its 1M window reads your whole repo in one pass, and native vision takes screenshots and diagrams. | $3.00 | $15.00 | $0.30 |
GLM 5.3 Z.ai's flagship coding model, the GLM 5.2 successor, with large gains on complex, long-horizon tasks and a 1M context window. Always thinks; dial reasoning low to max. | $1.40 | $4.40 | $0.26 |
USD per 1M tokens. Every request bills exactly what it uses: input, output, and cache. Full pricing
On request
EU data residency
We are lighting up new GPU capacity in Europe. Same frontier open models, same zero data retention, with residency in the EU for teams that need it.
Early allocations are limited: tell us the volume you would commit to, and we will prioritize yours.

