Doubleword
    NXL logo

    Case study · NXL

    How NXL cut generation costs 79% with no drop in quality using Doubleword

    79%

    lower cost per message vs Claude Sonnet

    3-5x

    faster generation on the same models vs OpenRouter

    98%

    reliability of outputs with zero model-quality failures

    Background

    NXL builds an AI Revenue Team - specialist assistants that handle content, prospecting and coaching for enterprise go-to-market teams. A core workload is generating personalised outbound messaging at scale, which has to hold quality and tone consistency across very different buyer personas.

    Problem

    NXL was running that generation on a mix of frontier and open-weight models - Claude Sonnet directly, plus open models routed through OpenRouter. At volume, two problems surfaced:

    • Sonnet was roughly 10x more expensive - than the open-model alternatives, which made scaling personalised generation costly.

    • Data-security concern - because NXL's prompts include client knowledge-base data, routing that through frontier APIs raised a real risk - inputs that shouldn't be exposed to training or third-party handling.

    Solution

    NXL ran a structured, evidence-based evaluation which included a 735-generation reliability test across three personas, plus a head-to-head benchmark of five models on both OpenRouter and Doubleword using identical prompts, scored by Claude Sonnet 4 as an independent QA judge.

    API integration was straightforward and the results were decisive on three fronts:

    • On speed - Doubleword ran DeepSeek V4 Flash and V4 Pro in a flat 5 seconds, versus 15s and 26s on OpenRouter - a 3-5x latency improvement on identical models.

    • On quality - every model that completed scored 9/10 on independent Sonnet QA, matching OpenRouter's output exactly.

    • On reliability - Doubleword hit a 98% success rate across the full persona set, with every failure traced to rate limiting rather than model error.

    That combination - a Doubleword-hosted open model with selective QA gating - opened a path to dramatically cheaper generation without giving up quality or reliability.

    “The team are super responsive, and the product works well as expected - results look good at a fraction of the cost.”

    - Lewis Stock, Co-Founder, NXL

    Results

    • 79% lower cost per message - a Doubleword-hosted model with QA on failures only was modelled at ~$0.0012/message vs. $0.0058 on a pure-Sonnet baseline (at 1,000 messages/day), with quality held constant at 9/10.

    • 3-5x faster generation - DeepSeek V4 completed in 5s on Doubleword vs. 15-26s on OpenRouter for the same models; GLM-5.2 ran ~30% faster.

    • 98% reliability at scale - 718 of 735 generations succeeded across three personas, with 100% of failures caused by rate limiting, not model quality.

    Modelled cost at 1,000 messages/day

    Architecture Cost / message Daily cost Saving vs. baseline
    Pure Sonnet (baseline) $0.0058 $5.80 -
    Open model + Sonnet QA on all outputs $0.0045 $4.50 22%
    Open model + Sonnet QA on failures only ~$0.0020 $2.00 65%
    Doubleword-hosted model + QA on failures only ~$0.0012 $1.20 79%

    Quality held constant - every successful generation scored 9/10 on independent Sonnet QA review.

    What's next

    NXL is now looking to integrate Doubleword-hosted models into its main product.

    Cut generation costs without giving up quality.

    Get started