Case study · NXL
How NXL cut generation costs 79% with no drop in quality using Doubleword
79%
lower cost per message vs Claude Sonnet
3-5x
faster generation on the same models vs OpenRouter
98%
reliability of outputs with zero model-quality failures
Background
NXL builds an AI Revenue Team - specialist assistants that handle content, prospecting and coaching for enterprise go-to-market teams. A core workload is generating personalised outbound messaging at scale, which has to hold quality and tone consistency across very different buyer personas.
Problem
NXL was running that generation on a mix of frontier and open-weight models - Claude Sonnet directly, plus open models routed through OpenRouter. At volume, two problems surfaced:
-
Sonnet was roughly 10x more expensive - than the open-model alternatives, which made scaling personalised generation costly.
-
Data-security concern - because NXL's prompts include client knowledge-base data, routing that through frontier APIs raised a real risk - inputs that shouldn't be exposed to training or third-party handling.
Solution
NXL ran a structured, evidence-based evaluation which included a 735-generation reliability test across three personas, plus a head-to-head benchmark of five models on both OpenRouter and Doubleword using identical prompts, scored by Claude Sonnet 4 as an independent QA judge.
API integration was straightforward and the results were decisive on three fronts:
-
On speed - Doubleword ran DeepSeek V4 Flash and V4 Pro in a flat 5 seconds, versus 15s and 26s on OpenRouter - a 3-5x latency improvement on identical models.
-
On quality - every model that completed scored 9/10 on independent Sonnet QA, matching OpenRouter's output exactly.
-
On reliability - Doubleword hit a 98% success rate across the full persona set, with every failure traced to rate limiting rather than model error.
That combination - a Doubleword-hosted open model with selective QA gating - opened a path to dramatically cheaper generation without giving up quality or reliability.
“The team are super responsive, and the product works well as expected - results look good at a fraction of the cost.”
- Lewis Stock, Co-Founder, NXL
Results
-
79% lower cost per message - a Doubleword-hosted model with QA on failures only was modelled at ~$0.0012/message vs. $0.0058 on a pure-Sonnet baseline (at 1,000 messages/day), with quality held constant at 9/10.
-
3-5x faster generation - DeepSeek V4 completed in 5s on Doubleword vs. 15-26s on OpenRouter for the same models; GLM-5.2 ran ~30% faster.
-
98% reliability at scale - 718 of 735 generations succeeded across three personas, with 100% of failures caused by rate limiting, not model quality.
Modelled cost at 1,000 messages/day
| Architecture | Cost / message | Daily cost | Saving vs. baseline |
|---|---|---|---|
| Pure Sonnet (baseline) | $0.0058 | $5.80 | - |
| Open model + Sonnet QA on all outputs | $0.0045 | $4.50 | 22% |
| Open model + Sonnet QA on failures only | ~$0.0020 | $2.00 | 65% |
| Doubleword-hosted model + QA on failures only | ~$0.0012 | $1.20 | 79% |
Quality held constant - every successful generation scored 9/10 on independent Sonnet QA review.
What's next
NXL is now looking to integrate Doubleword-hosted models into its main product.
