Made in EuropeNVIDIA Inception Program member

Frontier open models, engineered for AI agents.

The best open coding models, served fast from our in-house inference stack.
Keep your tools. Keep your data.

Use it with the tools you already use

OpenCodeOpenCodeClaude CodeClaude CodeπPiCodexCodexCopilotCopilot
ZedZedCursorCursorompompDeepSeek HarnessDeepSeek Harnessand more

Running frontier open models

Kimi K3Kimi K3GLM 5.3GLM 5.3DeepSeek V4 FlashDeepSeek V4 Flashand more

Hosted Kimi K3, GLM 5.3, and DeepSeek V4 Flash. Pay per token, on infrastructure we own.

  • Lengow
  • Billy

Why teams choose Umans AI

Frontier open coding models at the best price

Always frontier open models

The strongest open coding models, live on our endpoint the day they drop. The lineup keeps moving — you never rebuild.

Best SLOs, owned end to end

We run the inference stack in-house: our GPUs, our kernels, our capacity. Latency and uptime have a single owner — us.

Keep your setup

OpenAI- and Anthropic-compatible endpoints that drop into the tools you already use — OpenCode, Claude Code, Codex, Zed, Cursor.

Your data stays yours

Zero data retention: prompts and code are never stored, never used for training.

ISO 27001 · in progressSOC 2 Type II · in progress

How it works

Use it in the tool you already use

Pick your tool. Setup is one command or two env vars.

Universal
2 env vars
Works with any tool. Two env vars in any OpenAI- or Anthropic-compatible tool.
OpenAI-compatible
OPENAI_BASE_URL=https://api.code.umans.ai/v1
OPENAI_API_KEY=umans_••••••••
Anthropic-compatible
ANTHROPIC_BASE_URL=https://api.code.umans.ai
ANTHROPIC_API_KEY=umans_••••••••
Or skip setup entirely

Remote

Cloud agents

Early release

Spin up a coding agent on our infra in one click. It lands wired up to our inference, with the umans CLI installed and configured. Reach it from your laptop or phone. Same session, any device.

Move between devices

Same session on your laptop, phone, or tablet. Nothing to re-set up.

Sandboxed by default

Each agent runs in its own box. No access to your files, repos, or configs.

Keeps running while your laptop is off

The agent runs on our servers, not yours.

Loved by builders

Don’t take our word for it

Umans.ai is an outlier… in a good way. Other AI companies are some weird combination of very high token prices for marginally different models and/or some form of "we're harvesting your prompts… you're the product not the customer." Umans has a clear business model: selling tokens for open models. The prices are competitive, but the difference for me is trust. I'm the customer and not the product with Umans.
DPDavid PollakCEO, Spice Labs
I’ve been using your DeepSeek deployment for these past weeks. I’ve also tried Hyper, CommandCode, OpenCode and Arli, all with DeepSeek. Your work is unmatched.
Nonono· Umans Discord community
How does umans achieve high tps? I feel like it’s faster than cerebras atp
Gojim_· Umans Discord community
made use of that promo, damn, ds4f is flying, great job ❤️
blong· Umans Discord community

Pricing

Pay per token.

Run complex, long-running engineering tasks on frontier open models. Top up, create a key, and go.

ModelInput$ per 1M tokensOutput$ per 1M tokensCache read$ per 1M tokens
DeepSeek V4 FlashFeatured

Our fastest model for agentic and coding ⚡ Long agentic runs stay snappy on a 1M context window, at the lowest price in the lineup.

$0.14$0.28$0.028
Kimi K3

The first open 3T-class model, neck-and-neck with the closed frontier on agentic coding. Its 1M window reads your whole repo in one pass, and native vision takes screenshots and diagrams.

$3.00$15.00$0.30
GLM 5.3

Z.ai's flagship coding model — the GLM 5.2 successor, with large gains on complex, long-horizon tasks and a 1M context window. Always thinks; dial reasoning low to max.

$1.40$4.40$0.26

USD per 1M tokens. Every request bills exactly what it uses: input, output, and cache. Full pricing

Top up & start

Coming soon

EU data residency

We are lighting up new GPU capacity in Europe. Same frontier open models, same zero data retention — with residency in the EU for teams that need it.

Tell us the volume you would commit to, and we will prioritize your allocation when capacity comes online.

Volume commitments get priority when we allocate EU capacity.