Frontier open models, engineered for AI agents.
The best open coding models, served fast from our in-house inference stack.
Keep your tools. Keep your data.
Use it with the tools you already use
Running frontier open models
Hosted Kimi K3, GLM 5.3, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.
Why teams choose Umans AI
Frontier open coding models at the best price
Always frontier open models
The strongest open coding models, live on our endpoint the day they drop. The lineup keeps moving — you never rebuild.
Best SLOs, owned end to end
We run the inference stack in-house: our GPUs, our kernels, our capacity. Latency and uptime have a single owner — us.
Keep your setup
OpenAI- and Anthropic-compatible endpoints that drop into the tools you already use — OpenCode, Claude Code, Codex, Zed, Cursor.
Your data stays yours
Zero data retention: prompts and code are never stored, never used for training.
How it works
Use it in the tool you already use
Pick your tool. Setup is one command or two env vars.
Remote
Cloud agents
Early releaseSpin up a coding agent on our infra in one click. It lands wired up to our inference, with the umans CLI installed and configured. Reach it from your laptop or phone. Same session, any device.
Same session on your laptop, phone, or tablet. Nothing to re-set up.
Each agent runs in its own box. No access to your files, repos, or configs.
The agent runs on our servers, not yours.
Loved by builders
Don’t take our word for it
“Umans.ai is an outlier… in a good way. Other AI companies are some weird combination of very high token prices for marginally different models and/or some form of "we're harvesting your prompts… you're the product not the customer." Umans has a clear business model: selling tokens for open models. The prices are competitive, but the difference for me is trust. I'm the customer and not the product with Umans.”
“I’ve been using your DeepSeek deployment for these past weeks. I’ve also tried Hyper, CommandCode, OpenCode and Arli, all with DeepSeek. Your work is unmatched.”
“How does umans achieve high tps? I feel like it’s faster than cerebras atp”
“made use of that promo, damn, ds4f is flying, great job ❤️”
Pricing
Pay per token.
Run complex, long-running engineering tasks on frontier open models. Top up, create a key, and go.
| Model | Input$ per 1M tokens | Output$ per 1M tokens | Cache read$ per 1M tokens |
|---|---|---|---|
DeepSeek V4 FlashFeatured Our fastest model for agentic and coding ⚡ Long agentic runs stay snappy on a 1M context window, at the lowest price in the lineup. | $0.14 | $0.28 | $0.028 |
Kimi K3 The first open 3T-class model, neck-and-neck with the closed frontier on agentic coding. Its 1M window reads your whole repo in one pass, and native vision takes screenshots and diagrams. | $3.00 | $15.00 | $0.30 |
GLM 5.3 Z.ai's flagship coding model — the GLM 5.2 successor, with large gains on complex, long-horizon tasks and a 1M context window. Always thinks; dial reasoning low to max. | $1.40 | $4.40 | $0.26 |
USD per 1M tokens. Every request bills exactly what it uses: input, output, and cache. Full pricing
Coming soon
EU data residency
We are lighting up new GPU capacity in Europe. Same frontier open models, same zero data retention — with residency in the EU for teams that need it.
Tell us the volume you would commit to, and we will prioritize your allocation when capacity comes online.

