What smart routing saves
Agentlify’s cost and quality results, in production and on a public benchmark.
Most teams send every request to their strongest, most expensive model. Agentlify routes each request to the model that gives the best quality per dollar for that kind of task, and keeps learning from every eval result.
01
What does Agentlify save in production?
Half the inference bill, for about a 5% drop in eval score, on real customer traffic.
50%
Lower inference cost
average across production traffic
~5%
Eval score change
versus always using the strongest model
Source: Agentlify production dashboard.
How each request is routed
Step 1
Embed
The request is turned into a vector that captures what it asks for.
Step 2
Match
It is matched to clusters of similar past requests.
Step 3
Sample
A Thompson-sampling bandit picks the model with the best expected quality per dollar for that cluster.
Step 4
Learn
The eval result updates that model’s record for the cluster, so the next pick is better.
02
Does it hold on a public benchmark?
On RouterBench, with 11 real models and their real per-request prices, Agentlify cuts cost by 79% while staying within 5% of GPT-4’s score. A standard classifier router gets 54%, and simply mixing single models gets 28%.
Agentlify
79%
cost saved within 5% of GPT-4’s score
95% CI 77–81%
Classifier router
54%
cost saved within 5% of GPT-4’s score
95% CI 41–64%
↑ Eval score vs always using GPT-4
Cost saved vs always using GPT-4 →
RouterBench 0-shot, 11 models, 7,300 test prompts, the dataset’s real per-request costs. Each line sweeps how much the router weighs cost against quality. Mean of 20 random orderings of the test stream. Off the scale: perfect routing (the oracle) would score +17.0% while saving 93%; picking a model at random scores −33.1%.
Show as table
| Router | Cost weight | Cost saved | Eval change | GPT-4 calls |
|---|---|---|---|---|
| Agentlify | 0 | 19.4% | +0.7% | 74.5% |
| Agentlify | 0.01 | 22.5% | +0.7% | 70.9% |
| Agentlify | 0.02 | 23.8% | +0.7% | 69.2% |
| Agentlify | 0.03 | 24.6% | +0.7% | 68.6% |
| Agentlify | 0.04 | 26.7% | +0.7% | 67.8% |
| Agentlify | 0.05 | 45.8% | +0.3% | 65.4% |
| Agentlify | 0.06 | 66.1% | −0.3% | 61.7% |
| Agentlify | 0.08 | 70.8% | −1.0% | 56.3% |
| Agentlify | 0.1 | 73.9% | −2.3% | 49.1% |
| Agentlify | 0.12 | 77.9% | −4.4% | 37.8% |
| Agentlify | 0.15 | 88.3% | −9.0% | 11.1% |
| Agentlify | 0.2 | 91.1% | −10.6% | 4.2% |
| Agentlify | 0.25 | 93.1% | −11.8% | 0.0% |
| Agentlify | 0.3 | 93.1% | −11.9% | 0.0% |
| Agentlify | 0.4 | 93.2% | −11.8% | 0.0% |
| Agentlify | 0.5 | 93.2% | −11.8% | 0.0% |
| Classifier router | 0 | 5.2% | +0.4% | 93.4% |
| Classifier router | 0.01 | 6.4% | +0.6% | 91.7% |
| Classifier router | 0.02 | 7.8% | +0.4% | 90.0% |
| Classifier router | 0.03 | 9.7% | +0.3% | 87.6% |
| Classifier router | 0.04 | 11.6% | +0.1% | 84.6% |
| Classifier router | 0.05 | 14.3% | −0.2% | 81.4% |
| Classifier router | 0.06 | 18.0% | −0.7% | 77.1% |
| Classifier router | 0.08 | 23.0% | −1.8% | 67.6% |
| Classifier router | 0.1 | 28.4% | −3.2% | 55.6% |
| Classifier router | 0.12 | 59.1% | −5.4% | 34.0% |
| Classifier router | 0.15 | 85.5% | −8.3% | 14.5% |
| Classifier router | 0.2 | 91.5% | −10.8% | 4.3% |
| Classifier router | 0.25 | 93.3% | −12.3% | 0.1% |
| Classifier router | 0.3 | 93.4% | −12.4% | 0.0% |
| Classifier router | 0.4 | 93.5% | −12.5% | 0.0% |
| Classifier router | 0.5 | 93.5% | −12.5% | 0.0% |
| Always GPT-4 | – | 0.0% | 0.0% | – |
| Always Yi-34B | – | 94.4% | −17.1% | – |
| Always cheapest (Mistral-7B) | – | 98.6% | −60.3% | – |
| Random model | – | 74.7% | −33.1% | – |
| Oracle (perfect routing) | – | 92.6% | +17.0% | – |
03
What does that mean in dollars?
The same results as a monthly bill. Enter what you spend on models today.
You’d save
$5,000
a month · $60,000 a year
50% lower cost for about a 5% lower eval score, measured on production traffic. Your savings depend on your traffic mix and which models you use.
04
What if a model gets worse?
Model quality changes with every provider update. In simulation, a router trained once never recovers when the ranking flips. The bandit router recovers on its own, and with forgetting, which down-weights old results, it loses about half as much quality as a router retrained every 1,000 requests.
Simulated: synthetic models, not real LLMs or production traffic.
↑ Quality of the chosen model (rolling average of 100 requests)
Requests →
Quality lost over 5,000 requests
Cumulative regret against the best model per request; lower is better. Thin bars are 95% intervals.
When nothing changes, forgetting costs little: 71 against 66 with default settings, over the same 5,000 requests. 3 synthetic models, 5 request clusters, 20 seeds.
Show as table
| Router | Regret (drift) | 95% CI |
|---|---|---|
| Bandit router, with forgetting | 109.6 | 103.4 to 115.8 |
| Bandit router, default settings | 287.7 | 266.6 to 308.7 |
| Static router, retrained every 1,000 requests | 203.9 | 180.9 to 227.0 |
| Static router, trained once | 596.1 | 530.9 to 661.3 |
Agentlify is at agentlify.co.