A/B testing models
Not sure whethertycho-1.0, clarke-1.0, or hal-1.0 is right for your use case? Run the same prompt through multiple models and score the outputs side by side.
The script
ab_test.py
Run it
What it does
- Runs the same prompt through
tycho-1.0,clarke-1.0, andhal-1.0concurrently. - Measures latency precisely with
time.perf_counter(). - Records token usage for cost comparison.
- Scores outputs with custom heuristics — length, formatting quality, or anything else you define.
- Prints a side-by-side comparison table with latency, tokens, and truncated responses.
Adding an LLM-as-judge scorer
Use a stronger model to evaluate outputs from weaker ones:Tips
- Run at least 10 trials per model — single-shot comparisons can be noisy.
- Score on what matters to your product: accuracy, conciseness, latency, cost, or safety.
- Keep a golden dataset of prompts you run on every model update to catch regressions.
- Store results in a spreadsheet or DB so you can trend quality over time.
Next steps
Switching with fallback
Route requests to the winning model in production
Monitoring usage
Track quality and cost metrics over time