Does Every AI Decision Need a Large Language Model?
Author: Monark Unadkat, Senior Solutions Engineer - Applied AI, Searce
A marketing brief experiment comparing code, a classifier, a specialized decision model, and a general-purpose LLM on the same checks.
Writing campaign copy and checking whether it includes a required disclosure are different jobs. Yet an AI workflow can send both to the same general-purpose language model. That is convenient to build. It is worth asking whether every step needs it.
We tested that question in a marketing brief-checking workflow by giving the same judgments to five different mechanisms: plain code, a conventional classifier, a specialized decision model (Jev), a general-purpose LLM (Gemini), and a combination of the two models. Could a cheaper mechanism handle selected judgments quickly and accurately enough to reduce reliance on the general-purpose model?
The decision model returned the selected checks about 4.6× faster than the LLM, detecting 36 of 47 labeled violations against the LLM's 37. But the largest shared coverage problem happened before either model answered: a keyword filter kept seven violations out of their reach.
The result is less a verdict on one product than a lesson about where each kind of mechanism belongs, and why choosing a cheaper model is only part of the business case.
What a decision model does, using Jev as the example
A decision model answers typed questions rather than generating text. We used Jev, TypeSafe's decision model: you provide context and specific questions; it returns values that software can use directly. Its three question types serve different purposes:
| Question type | In plain language | Marketing example |
|---|---|---|
| Choice | Select from a supplied list of options. | Which channel fits this request: email, social, or web? |
| Score | Rate something against a defined scale. | How closely does this concept match the brief? |
| Noul | Estimate the probability that a statement is true. | Does the brief contain an adequate disclosure? |
These are question types, not three separate models. Our experiment tested only Noul: whether a rule applied in context and, where needed, whether a disclosure satisfied it. Code translated those answers into findings and severity.
A general-purpose LLM can also return structured answers. The proposition of a decision model is specialization for that job. An answer in the correct format can still be the wrong judgment. That distinction matters when a result determines whether work moves forward.
The business task and the comparison
Our starting point was a retail marketing workflow that checked briefs against seven rules. Some required interpretation: did a restricted term describe the retailer's offer, or simply the customer's circumstances? Others checked whether a required disclosure was complete.
We evaluated 62 previously unused test briefs against all seven rules: 434 brief-and-rule decisions, including 47 labeled violations.
| Approach | How it worked | Question it helped answer |
|---|---|---|
| A: Combined Gemini check | Checked all seven rules together. | How does the starting workflow perform? |
| B: Focused Gemini checks | Used keyword matches to select narrower questions. | Does breaking down the task help? |
| C: Decision model (Jev) | Answered the same selected questions as B. | Can a specialized model replace those calls? |
| D: Local classifier | Used a conventional text classifier trained on pilot examples. | Could a simpler machine-learning approach work? |
| E: Decision model with LLM fallback | Used Jev first and consulted Gemini when answers fell in an uncertainty band. | Can we reserve the second call for harder cases? |
A and B used the same model, Gemini 3.8 Flash. The keyword filter selected questions from 45 briefs; 17 never reached a model. All 62 briefs remained in the evaluation, so skipped violations counted as misses.
How to read the evidence: The briefs were synthetic, and their labels followed a frozen, author-defined evaluation policy—not approved customer compliance guidance. The fallback results were simulated using saved responses. This is a bounded experiment, not proof of production readiness or equivalent quality.
The first limit was what the workflow chose to check
The keyword filter surfaced 40 of the 47 violations. Even a perfect model behind that filter could therefore detect only about 85% of the labeled issues. If a brief expressed a prohibited claim without a listed keyword, the model never received the question. Paying for a more capable model would not fix that omission.
The filter did not explain every difference, however. Combined Gemini detected 41 violations; focused Gemini detected 37. Of those four lost detections, one was excluded by the filter and three were incorrect judgments on candidates both approaches evaluated.
Breaking a task into smaller questions changes its behavior; it does not automatically improve it. The context supplied, the questions asked, and the rules for selecting them all affect the outcome. For a business owner, that means an overall accuracy score is not enough. Ask both: Did we check the right things? And did we judge them correctly?
The decision model was faster, with a small observed quality difference
The cleanest model comparison was B versus C, because both used the same selection process and focused questions. This is the specialized-versus-general comparison the experiment was built for.
| Result | Focused Gemini | Jev |
|---|---|---|
| Labeled violations detected, out of 47 | 37 | 36 |
| False alarms | 3 | 3 |
| Unresolved decisions, out of 434 | 0 | 3 |
| Average model response time | 1.49 seconds | 0.32 seconds |
Jev returned these judgments in roughly one-fifth of the time. It detected one fewer violation, produced the same number of false alarms, and left three decisions unresolved. That makes it worth exploring for repeated checking steps.
It does not establish that the two models are interchangeable. A tool that suggests edits to a marketer can tolerate different errors from a gate that automatically approves content.
The measured gain also applies to the model calls, not the entire marketing process. Whether it shortens a user's wait depends on how much of that process those calls occupy.
How That Compares With the Vendor's Marketing Claims
TypeSafe advertises 193.6× faster and 444.6× cheaper on its workflow evaluations. Its launch explanation describes those as upper-end gains and notes that its LLM comparison includes returning probabilities, which adds work. It also gives a 70–500 millisecond response-time range. Our average of about 320 milliseconds fits that range.
Our measured speedup was 4.6× against our Gemini configuration. The business case should use the gain demonstrated on the actual task, not a multiplier from a different workload.
Use the fast judgment first, and bring in the LLM selectively
The fallback approach offered a middle ground: use Jev's answer when its probability fell outside a predefined uncertainty band; ask Gemini when it fell inside. We selected the thresholds on development data and froze them before testing.
The diagram shows the evaluated fallback design. Model calls were measured individually; the combined fallback path was replayed from saved responses. A question excluded by the keyword filter cannot be recovered by either downstream model.
The operational detail mattered: 11 of the 45 requests that reached a model needed Gemini as well—24%. A brief could contain several questions, and one uncertain answer could trigger the second request. Counting questions instead of requests would understate how often the workflow paid for two calls.
The replay detected 37 violations, matching focused Gemini, with four false alarms instead of three. Estimated average response time was 0.72 seconds, about 2.1× faster than always using Gemini. The slower requests benefited less: on this sample, estimated p95—the time within which 95% of requests finished—was 1.99 seconds versus Gemini's measured 2.25 seconds.
Fallback reduced the average wait, but an escalated request still had to wait for both models. It was a useful tradeoff, not a free quality guarantee.
Predictable rules do not make a model deterministic
One repeated input produced Jev probabilities of 0.17, 0.21, and 0.16. With the lower threshold fixed at 0.2, the first and third runs used Jev's answer directly. The second consulted Gemini. The routing rule stayed exactly the same. The model's output changed enough to alter the action.
Gemini returned identical answers in our small repeatability check of five cases with three reruns each. That is an observation about this sample, not a universal guarantee for Gemini.
For automation, the useful test is whether repeated judgments change an operational decision. A well-defined output and a fixed threshold make a workflow easier to control and audit; they do not establish that every decision will be stable or correct.
Lower prices are promising; savings need a complete count
Jev's published rate is $0.042 per million input tokens, with no output charge. Gemini 3.8 Flash's standard paid rates, through December 31, 2026, are $0.75 for input and $3.75 for output per million tokens. These are public prices, not a reconciliation of our bill.
That is roughly an 18× difference in input-token price. It is not an 18× difference in cost per brief: Jev used about 2.2× as many input tokens in our earlier nine-call smoke test, the two Gemini approaches send different amounts of text, and fallback pays for a second call on escalated requests.
The results reported here cover quality and response time; the billing reconciliation needed for a per-brief cost comparison was not completed. What we can say is that the price gap is large and the token-usage gap partly offsets it. A cost figure per completed brief, including fallback calls and output charges, is the first number the follow-up should produce, alongside the integration, monitoring, and rework costs that any adoption decision has to include.
CChoose the least expensive approach that meets the quality requirement
The broader lesson is to evaluate each decision separately. Some need exact code, some need learned language interpretation, and some justify a more capable model.
| Option | A useful starting point when… | What this experiment established |
|---|---|---|
| Code and rules | The requirement is explicit: a calculation, required field, or known transition. | Code controlled selection and aggregation. Keyword selection missed seven violations. |
| Conventional classifier | Categories repeat and representative labeled examples are available. | Our classifier had too little training data to support a useful conclusion. |
| Specialized decision model | The task is a focused language judgment with defined answers. | Jev was faster, with close but not identical observed results. |
| Fast general-purpose LLM | The decision needs flexible interpretation or broader context. | Gemini's results changed when we changed the workflow around it. |
| More capable reasoning model | Tests show that added reasoning materially improves the outcome. | Not evaluated here. |
This is a guide to where to start testing, not a universal ordering of cost or quality. A strong conventional classifier remains an open comparison: ours had only 15 training examples for one question type and four for the other.
Two further questions fall outside this benchmark but follow from it. Cheaper checks could make broader evaluation affordable, letting a team inspect more of its output within the same budget. And the same pattern of a fast judgment, a fixed gate, and selective escalation could apply to existing automation, including RPA, wherever a system has to interpret an input before choosing an approved path. Both are worth their own trial; neither is a result of this one.
Where a decision model earns a place
A specialized decision model is worth a focused trial wherever repeated, bounded language judgments create meaningful cost or delay. Jev was our instance of that class, and our results support that level of investment in it. They do not support replacing general-purpose LLMs across a workflow, treating any one product as the answer, or assuming that typed outputs add determinism.
For this use case, the next investment should improve which issues reach the model, calculate cost per completed brief, and test on representative real examples with agreed limits for missed issues and false alarms.
The adoption decision is the same for any mechanism on the ladder: does the complete workflow meet the quality requirement at a lower cost or a shorter wait? The decision model gave us a credible candidate. The experiment showed what else must work before any candidate becomes a business improvement.