Migration kit
Migrate agents from GPT & Claude to open-weight models.
Token spend is a scaling bottleneck. Transition your high-volume workloads from closed-source models to targeted open-weight models with confidence. Evaluate on the batch tier, shadow on live traffic, and deploy via our OpenAI-compatible gateway.
$ dw batches run batches/candidates.jsonl --watch
[INFO] Validating batch schemas... OK
[INFO] Enqueuing jobs against Qwen3.6-35B, Kimi-K3
» Batch complete. Run `dw batches analytics` to score.
The migration architecture
Replacing a proprietary frontier model for inference requires hard data, not vibes.
The Doubleword migration kit is a rigorous, three-step pipeline that ensures that the selected open-weight candidate achieves performance parity before you flip the switch on live traffic.
Baseline the flagship
Extract real task traces from your current deployment to create a golden set. Capture the input, the exact JSON output, the latency, and the cost.
If you use Arize Phoenix for tracing, export tasks directly from the logs. The migration workbook includes a 150-ticket support-triage golden set as a pre-labelled, automated scoring example.
# Export from Phoenix tracing
query = """
SELECT
input.value as prompt,
output.value as baseline_response,
attributes.llm.latency as latency_ms
FROM spans
WHERE attributes.llm.model_name = 'gpt-5.6'
LIMIT 150
"""
golden_set = client.query(query)
golden_set.to_jsonl("golden_set.jsonl")
Offline evaluation
Run the golden set against candidate models on the Doubleword batch tier. Doubleword enforces schema validity via response_format.
Candidates are scored on:
- Category and escalation accuracy vs. golden labels.
- Semantic similarity to the baseline, judged via a lightweight LLM-as-a-judge.
# 1. Prepare batch files for candidates
$ dw files prepare golden_set.jsonl \
--model moonshotai/kimi-k3 -o kimi.jsonl
$ dw files prepare golden_set.jsonl \
--model qwen/qwen-3.6-35b -o qwen.jsonl
# 2. Execute on the batch tier (costs cents)
$ dw batches run kimi.jsonl --output-id .kimi-id
$ dw batches run qwen.jsonl --output-id .qwen-id
# 3. Analyze against baseline
$ dw batches analytics --from-file .qwen-id
Evaluation results
150-ticket support triage
| Model | Category | Escalation | Valid JSON | Cost | TTFB |
|---|---|---|---|---|---|
| Claude Fable 5Current model | 89.5% | 99.0% | 150/150 | $1.200 | 2.1s |
| Kimi K3Candidate 1 | 89.3% | 94.0% | 149/150 | $0.330 | 4.7s |
| Qwen3.6-35BCandidate 2 | 87.3% | 98.0% | 150/150 | $0.005 | 1.7s |
Analysis: Qwen3.6-35B trails the incumbent by only 2.2% on category accuracy, but operates at a fraction of the cost with lower latency. Total evaluation cost for all candidates: ~$0.35.
Parallel execution
Test the open-weight model against real-world traffic with zero risk. Serve the user your current model's response, while fanning out the prompt for parallel execution on the Doubleword async tier for one or more models.
The secondary call to Doubleword runs out-of-band, so the user experiences zero latency penalty while you build a production-grade comparison log. If it matches your accuracy and latency expectations, you can switch the endpoint seamlessly.
Claude Fable 5
Synchronous
Qwen3.6-35B
Async tier
Continuous Evaluation
Stay on the efficient frontier.
Open-weight models are evolving rapidly and we release them as soon as weights are available. By building this pipeline, you're creating a reusable evaluation framework to continuously adopt the most high-performing, cost-effective models the moment they drop.
Monitor production telemetry
Track latency, token throughput, and exact execution costs across your active deployments. Get ground-truth data on how models perform under your real-world traffic.
Forecast zero-day savings
When a new state-of-the-art model drops, instantly replay your historical baselines against it. Predict your exact cost savings and quality delta before updating any code.
