Doubleword

    Migration kit

    Migrate agents from GPT & Claude to open-weight models.

    Token spend is a scaling bottleneck. Transition your high-volume workloads from closed-source models to targeted open-weight models with confidence. Evaluate on the batch tier, shadow on live traffic, and deploy via our OpenAI-compatible gateway.

    $npm i -g @doubleword/cli
    terminal

    $ dw batches run batches/candidates.jsonl --watch

    [INFO] Validating batch schemas... OK

    [INFO] Enqueuing jobs against Qwen3.6-35B, Kimi-K3

    ✓Qwen3.6-35B150/150
    ✓Kimi-K3150/150

    » Batch complete. Run `dw batches analytics` to score.

    The migration architecture

    Replacing a proprietary frontier model for inference requires hard data, not vibes.

    The Doubleword migration kit is a rigorous, three-step pipeline that ensures that the selected open-weight candidate achieves performance parity before you flip the switch on live traffic.

    Step 01

    Baseline the flagship

    Extract real task traces from your current deployment to create a golden set. Capture the input, the exact JSON output, the latency, and the cost.

    If you use Arize Phoenix for tracing, export tasks directly from the logs. The migration workbook includes a 150-ticket support-triage golden set as a pre-labelled, automated scoring example.

    export_phoenix.py
    # Export from Phoenix tracing
    query = """
    SELECT
      input.value as prompt,
      output.value as baseline_response,
      attributes.llm.latency as latency_ms
    FROM spans
    WHERE attributes.llm.model_name = 'gpt-5.6'
    LIMIT 150
    """
    
    golden_set = client.query(query)
    golden_set.to_jsonl("golden_set.jsonl")
    Step 02

    Offline evaluation

    Run the golden set against candidate models on the Doubleword batch tier. Doubleword enforces schema validity via response_format.

    Candidates are scored on:

    • Category and escalation accuracy vs. golden labels.
    • Semantic similarity to the baseline, judged via a lightweight LLM-as-a-judge.
    eval-pipeline.sh
    # 1. Prepare batch files for candidates
    $ dw files prepare golden_set.jsonl \
        --model moonshotai/kimi-k3 -o kimi.jsonl
    $ dw files prepare golden_set.jsonl \
        --model qwen/qwen-3.6-35b -o qwen.jsonl
    
    # 2. Execute on the batch tier (costs cents)
    $ dw batches run kimi.jsonl --output-id .kimi-id
    $ dw batches run qwen.jsonl --output-id .qwen-id
    
    # 3. Analyze against baseline
    $ dw batches analytics --from-file .qwen-id

    Evaluation results

    150-ticket support triage

    Winner found
    Model Category Escalation Valid JSON Cost TTFB
    Claude Fable 5Current model 89.5% 99.0% 150/150 $1.200 2.1s
    Kimi K3Candidate 1 89.3% 94.0% 149/150 $0.330 4.7s
    Qwen3.6-35BCandidate 2 87.3% 98.0% 150/150 $0.005 1.7s

    Analysis: Qwen3.6-35B trails the incumbent by only 2.2% on category accuracy, but operates at a fraction of the cost with lower latency. Total evaluation cost for all candidates: ~$0.35.

    Step 03

    Parallel execution

    Test the open-weight model against real-world traffic with zero risk. Serve the user your current model's response, while fanning out the prompt for parallel execution on the Doubleword async tier for one or more models.

    The secondary call to Doubleword runs out-of-band, so the user experiences zero latency penalty while you build a production-grade comparison log. If it matches your accuracy and latency expectations, you can switch the endpoint seamlessly.

    User request
    API gateway fan-out

    Claude Fable 5

    Synchronous

    Returned to user

    Qwen3.6-35B

    Async tier

    Logged for eval

    Continuous Evaluation

    Stay on the efficient frontier.

    Open-weight models are evolving rapidly and we release them as soon as weights are available. By building this pipeline, you're creating a reusable evaluation framework to continuously adopt the most high-performing, cost-effective models the moment they drop.

    Monitor production telemetry

    Track latency, token throughput, and exact execution costs across your active deployments. Get ground-truth data on how models perform under your real-world traffic.

    Forecast zero-day savings

    When a new state-of-the-art model drops, instantly replay your historical baselines against it. Predict your exact cost savings and quality delta before updating any code.