Same intelligence. Up to 99% cheaper.
One API. Open-weight models. Pick your delivery window — Async or Batch — and pay only for what you use.
Same Intelligence. Fraction of the price.
Cost to process 1 billion tokens in + 1 billion tokens out at comparable intelligence.
Intelligence via Artificial Analysis Index v4.0 · Hover any bar for full pricing details · Want access to a model you don't see here — just ask us!
No credit card required · No minimum spend · Pay only for tokens used
Try Inkling-NVFP4Three speeds. One API.
Pick the delivery window that fits your workflow. All tiers use the same OpenAI-compatible API.
Real Time
Iterate on prompts with real-time responses. Full price, zero wait.
Async Inference
up to 50% off RTBackground agents that need results fast. High throughput inference.
Batch (~24 hours)
up to 80% off RTBig batch jobs where cost matters most. Deepest discounts.
Model-by-model breakdown
| Model | SLA | Input $/MTok | Output $/MTok | Cost / 1B in+out | vs Big Token | |
|---|---|---|---|---|---|---|
| Inkling-NVFP4New | Async | $0.90 | $3.00 | $3.9K | 78% cheaper | Try Inkling-NVFP4 |
| ↳ | Batch | $0.60 | $2.00 | $2.6K | 86% cheaper | |
| DeepSeek-V4-ProNew | Async | $0.98 | $1.95 | $2.9K | 90% cheaper | Try DeepSeek-V4-Pro |
| ↳ | Batch | $0.65 | $1.30 | $2.0K | 94% cheaper | |
| DeepSeek-V4-FlashNew | Async | $0.07 | $0.14 | $210 | 93% cheaper | Try DeepSeek-V4-Flash |
| ↳ | Batch | $0.05 | $0.09 | $140 | 95% cheaper | |
| Kimi-K2.6New | Async | $0.50 | $2.56 | $3.1K | 83% cheaper | Try Kimi-K2.6 |
| ↳ | Batch | $0.33 | $1.71 | $2.0K | 89% cheaper | |
| GLM-5.2-FP8New | Async | $0.70 | $2.25 | $3.0K | 84% cheaper | Try GLM-5.2-FP8 |
| ↳ | Batch | $0.47 | $1.50 | $2.0K | 89% cheaper | |
| GLM-5.1-FP8 | Async | $0.79 | $2.63 | $3.4K | 81% cheaper | Try GLM-5.1-FP8 |
| ↳ | Batch | $0.53 | $1.75 | $2.3K | 87% cheaper | |
| Qwen3.5-397B-A17B | Async | $0.29 | $1.84 | $2.1K | 93% cheaper | Try Qwen3.5-397B-A17B |
| ↳ | Batch | $0.19 | $1.23 | $1.4K | 95% cheaper | |
| Qwen3.6-35B-A3B-FP8New | Async | $0.11 | $0.75 | $860 | 95% cheaper | Try Qwen3.6-35B-A3B-FP8 |
| ↳ | Batch | $0.07 | $0.50 | $570 | 97% cheaper | |
| Qwen3.5-35B-A3B-FP8 | Async | $0.07 | $0.30 | $370 | 94% cheaper | Try Qwen3.5-35B-A3B-FP8 |
| ↳ | Batch | $0.05 | $0.20 | $250 | 96% cheaper | |
| Qwen3.5-4B | Async | $0.05 | $0.08 | $130 | 99% cheaper | Try Qwen3.5-4B |
| ↳ | Batch | $0.04 | $0.06 | $100 | 99% cheaper | |
| Qwen3.5-9B | Async | $0.08 | $0.11 | $190 | 97% cheaper | Try Qwen3.5-9B |
| ↳ | Batch | $0.05 | $0.08 | $130 | 98% cheaper | |
| Gemma-4-31BNew | Async | $0.09 | $0.26 | $350 | 94% cheaper | Try Gemma-4-31B |
| ↳ | Batch | $0.06 | $0.18 | $240 | 96% cheaper | |
| Nemotron-3-Ultra-550B-A55BNew | Async | $0.38 | $1.65 | $2.0K | 93% cheaper | Try Nemotron-3-Ultra-550B-A55B |
| ↳ | Batch | $0.25 | $1.10 | $1.4K | 96% cheaper | |
| Nemotron-3-Super-120B-A12B | Async | $0.06 | $0.34 | $400 | 93% cheaper | Try Nemotron-3-Super-120B-A12B |
| ↳ | Batch | $0.04 | $0.23 | $270 | 96% cheaper | |
| GPT-OSS-20B | Async | $0.02 | $0.10 | $120 | 99% cheaper | Try GPT-OSS-20B |
| ↳ | Batch | $0.02 | $0.07 | $90 | 99% cheaper | |
| Qwen3-VL-235B-A22B | Async | $0.16 | $1.43 | $1.6K | 29% cheaper | Try Qwen3-VL-235B-A22B |
| ↳ | Batch | $0.11 | $0.95 | $1.1K | 53% cheaper | |
| Qwen3-VL-30B-A3B | Async | $0.11 | $0.45 | $560 | 63% cheaper | Try Qwen3-VL-30B-A3B |
| ↳ | Batch | $0.08 | $0.30 | $380 | 75% cheaper | |
| Qwen3-14B-FP8 | Async | $0.03 | $0.30 | $330 | 78% cheaper | Try Qwen3-14B-FP8 |
| ↳ | Batch | $0.02 | $0.20 | $220 | 85% cheaper | |
| DeepSeek-OCR-2OCR | Async | $0.08 | $0.08 | $160 | — | Try DeepSeek-OCR-2 |
| ↳ | Batch | $0.05 | $0.05 | $100 | — | |
| olmOCR-2-7BOCR | Async | $0.15 | $0.15 | $300 | — | Try olmOCR-2-7B |
| ↳ | Batch | $0.10 | $0.10 | $200 | — | |
| LightOnOCR-2-1BOCR | Async | $0.08 | $0.08 | $160 | — | Try LightOnOCR-2-1B |
| ↳ | Batch | $0.05 | $0.05 | $100 | — | |
| Qwen3-Embedding-8B | Async | $0.03 | — | $30 | — | Try Qwen3-Embedding-8B |
| ↳ | Batch | $0.02 | — | $20 | — |
No surprises. No lock-in.
Stop overpaying for inference.
Run your background agents and workloads at a fraction of the price and double the scale.
If you can wait an hour, you can save a lot.
