Meet the teams running inference at scale
Explore how Doubleword customers cut token spend, move non-urgent workloads out of real time, and run reliable batch and async inference without operational overhead.
Case study · Kelp

How Kelp cut inference cost per article 9x with Doubleword
Kelp runs its entire content pipeline on Doubleword's batch, realtime, and embedding APIs, and cut the cost of enriching an article from roughly 1.5 cents to a sixth of a cent.
9x
cheaper per article enriched
Case study · NXL

How NXL cut generation costs 79% with no drop in quality using Doubleword
NXL builds AI assistants that handle content, prospecting and coaching for enterprise sales teams.
79%
lower cost per message vs Claude Sonnet
Case study · Ken AI

How Ken AI eliminated inference bottlenecks and cut batch runtimes by 68% with Doubleword
Ken AI is a cold-email agency that runs personalized outbound for B2B SaaS companies on proprietary email infrastructure built in-house.
68%
faster batch runs
Case study · UnaGo AI

How UnaGo AI cut inference costs 70% with Doubleword
UnaGo AI is an AI-native operations platform that gives companies a team of specialist AI agents - marketing, media, research, and operations - working together in a single conversation.
70%
lower inference cost vs. Claude
Case study · OpenMed × SynthVision

119,000 medical images annotated for $452 with Doubleword
How OpenMed used Doubleword to make frontier-model knowledge distillation viable at dataset scale - at 94% lower cost than Claude Sonnet.
94%
cost savings vs. Anthropic
Case study · Dataiku

A custom PII detection model trained for $50 with Doubleword
How Dataiku's 575 Lab used Doubleword to generate the synthetic training data behind Kiji Privacy Proxy - 20× cheaper than closed-source providers.
95%
cost reduction
External write-up · NextGen Orchestration
Cutting Your LLM Inference Bill by 90%: Case Study with Doubleword.ai
An independent walkthrough by Dhruv Malik on how delayed and batched inference jobs on Doubleword can cut token costs while scheduling jobs at scale.
90%
inference bill reduction
