Case study · UnaGo AI
How UnaGo AI cut inference costs 70% with Doubleword
70%
lower inference cost vs. Anthropic Claude
“Doubleword's batch and async pricing tiers let us match each workload to the right cost-latency tradeoff. We're paying 50 to 80% less for background work that doesn't need real-time latency.”
- Max Mednikov, CTO & Co-Founder, UnaGo AI
Background
UnaGo AI is an AI-native operations platform whose multi-agent system routes each task to a purpose-built specialist agent, coordinating handoffs and shared context across a full workflow so a small team can execute at the scale of a much larger one.
As a fast-growing startup serving UK customers, UnaGo AI runs a fast-growing mix of agentic workloads - an agentic workforce driven by customer needs, data analytics, marketing, data search, and business operations.
Its agents span very different workload profiles: some need real-time, high-quality reasoning, while many run high-volume background tasks - document processing, structured extraction, data enrichment, research - where throughput and cost matter more than latency. That split, and the rising inference spend, is what made inference strategy a problem worth solving.
Problem
Running every workload through Anthropic Claude worked well for high-quality reasoning, but applying premium real-time prices to high-volume background work made costs difficult to scale.
Much of UnaGo AI's work - async document processing, structured extraction, research, enrichment, background agent tasks - doesn't need instant responses, yet was being billed as if it did.
At the same time, European customers increasingly needed clear answers on data residency and governance. Sending sensitive operational data through external APIs slowed approvals, created procurement friction, or blocked adoption outright.
Solution
The migration was straightforward: UnaGo AI routed workloads to Doubleword through an OpenAI-compatible router with LiteLLM, with no significant re-engineering. Because Doubleword exposes dedicated batch and async inference tiers, UnaGo AI could move suitable high-volume workloads off premium real-time rates and match each agent's workload profile to the right cost-latency tradeoff. The orchestration layer handles that routing automatically, so teams don't manage inference decisions manually.
The setup is also built for resilience: to keep agentic workflows running through model blips, UnaGo AI uses automatic failover to a comparable open model on Doubleword - for example, GLM 5.2 fails over to Kimi 2.6 or DeepSeek Pro V4 - preserving high availability without interrupting the workflow.
The move also gave UnaGo AI a concrete story for European customers with a UK inference provider, with residency and auditability part of the architecture rather than an afterthought.
“The European data residency story makes compliance approvals dramatically faster for our regulated clients.”
- Max Mednikov, CTO & Co-Founder, UnaGo AI
Results
-
70% lower cost today vs. Anthropic Claude - with a roadmap to up to 95% via batch adoption later this year.
-
High availability through automatic open-model failover - e.g. GLM 5.2 to Kimi 2.6 or DeepSeek Pro V4, so agentic workflows aren't interrupted by model outages.
-
Headroom to scale the workloads that were becoming too expensive - sales-lead qualification, marketing-signal extraction, data analytics, and background agent execution.
What's next
UnaGo AI plans to extend batch inference to two more areas:
-
Large-scale data processing across agentic workflows - sales-lead qualification and extraction of marketing signals at volume.
-
"Agentic dreaming" - a platform capability where agents spend time consolidating and validating what they've learned to improve workflow quality over time. Agentic dreaming is included for customers at no extra charge, which makes its cost a factor for UnaGo - and batch inference is how they intend to make it cost-effective to run at scale.
