Building an intelligent routing approach for analytics workloads with Claude

20% lower LLM spend in initial rollout
30% of queries routed to cost-efficient model tiers

With Claude, Tellius
  • Observed a 20% reduction in modeled LLM spend, without a measurable decline on the curated evaluation benchmark
  • Moved 30% of query volume onto cost-efficient models, reserving Claude for work that needs deep reasoning
  • Replaced a static task-to-model rules with a classifier that adapts to how users actually phrase questions
  • Build out an asynchronous Claude-as-judge workflow against a curated golden set, catching quality drift before users encounter it
  • Feeds reviewer corrections straight back into routing, so the system improves without engineering time
  • Gained per-query cost visibility across every model tier
The Challenge

Every question cost the same, no matter how simple

Tellius lets people ask questions of their data in plain language. Someone types what they want to know, and the platform interprets it, queries the underlying data, and answers. The range of what gets typed is enormous. One user wants a single number from last quarter. Another wants to know why a segment moved, which takes real analytical reasoning across several steps.

The existing platform does not distinguish between them. Both questions went to the same high-capability model. The simple one got far more reasoning capability than it needed, and the bill reflected that. As usage grows, spend scales with it in a straight line, with no obvious lever to pull.

The existing system used a pre-defined lookup table mapping task types to models, which worked until it didn't. Every new job type users invented meant another entry. The table needed constant tending otherwise it defaulted to higher reasoning models that were expensive, because nothing was watching whether its decisions were still right.

Quality control had the same shape. Checks were reactive. When an answer came back wrong, someone noticed after the fact, usually a user. Engineers spent their time triaging individual failures rather than fixing the thing that produced them.

The Solution

A decision layer in front of the models

Tellius partnered with Searce to design and build a routing layer that sits ahead of the existing dispatch. Every incoming query passes through a classifier that reads it and decides two things: what kind of task it is, across nine categories from straightforward lookup to multi-step reasoning, and how complex it is. Both decisions come with a confidence score.

Confident classifications route to a cost-efficient model. Anything the classifier is unsure about, or that scores as genuinely complex, escalates to higher reasoning model. The escalation is deliberately biased toward caution: when the system doesn't know, it spends more rather than guessing cheaply. A wrong answer costs more than the tokens saved.

That fixes the routing problem but creates a new one. Once queries fan out across several models, no single model owner is accountable for quality, and drift becomes hard to see.

Claude as the judge

The answer was to put Claude on the other side of the system as an evaluator. Asynchronously, off the live path, Claude reviews a sample of routed decisions against a golden set of curated query and response pairs that Tellius domain experts approved as correct. It returns a structured verdict: whether it agrees with the routing decision, what it would have chosen instead, its confidence, and whether a human should look.

Running this asynchronously matters. Quality assurance that sits on the request path costs the user latency on every query in order to catch problems on a few. Running it behind the scenes means the check is thorough without anyone waiting for it. The judge role is also where model choice mattered most. An evaluator has to be more reliable than the systems it grades, and it has to apply the same rubric consistently across thousands of verdicts rather than drifting as it goes.

Humans where they add the most

Flagged cases go to a review queue where a domain expert approves, corrects, or rejects the decision. Corrections write back into the playbook that guides routing, so the same misjudgment doesn't repeat.

This is a deliberate reversal of how the team spent its time before. Reviewers no longer hunt for problems. The system finds candidates and brings them the ones that need judgment, which is the part a person is actually better at.

The Outcome

Spend down, quality held

Tellius cut LLM spend by 20%, with 30% of queries handled by cost-efficient models. Answer quality has held against the golden set benchmark.

Simple questions also come back faster. A lookup that used to wait on a large reasoning model now returns from a lightweight one, and users feel that on the queries they run most often.

The less visible change is in what the team can see. The team will no longer need to maintain lookup tables and triaging individual bad answers.

The architecture was also built to outlast any particular model. Because the decision layer is abstracted from the models underneath it, Tellius can evaluate and adopt new options as they arrive without rebuilding the routing logic. As the market shifts, the system absorbs the change.

Looking Ahead

Expanding to Full Production and Scaling the Self-Learning Loop

The initial rollout proved that dynamic routing does not require sacrificing quality for cost efficiency. The immediate mission for Tellius is to expand this decision layer across 100% of analytical query traffic. As broader enterprise workloads transition to the router, the focus shifts from static policy enforcement to scaling an autonomous, self-improving feedback loop.

Sharing a Reference Blueprint with the Community

Tellius and Searce view this self-improving harness as a foundational design pattern for enterprise AI engineering. By publishing the architecture, evaluation rubrics, and feedback-loop mechanics validated in these trials, we aim to provide a practical blueprint for the broader community.

As multi-model architectures become the standard, sharing how lightweight classifiers can be safely trained and refined through frontier LLM evaluation enables other teams to move beyond static routing tables and build adaptive, cost-aware systems of their own.