Jev and GLiNER on Classification and Policy Decisions
For multi-agent orchestration, where does Jev actually add the most value? To me, it looks like zero-shot classification on steroids. I have seen it discussed as a router, but how much does it add when the job is simply identifying intent and selecting an agent?
Routing based on conversation history, policy and agent capabilities is a different problem. I wanted to see where a smaller classifier holds up and where Jev starts to pull ahead.
I compared the zero-shot capabilities of Jev and the 340-million-parameter GLiNER2.5-Decide on public classification tasks and a separate synthetic policy test. I did not find a clear overall advantage for Jev on the public classification panel. Jev was substantially better at applying supplied policies. That test goes beyond routine intent classification, so I would treat it as a comparison of capabilities rather than a general model ranking. The synthetic labels are rule-generated and still unreviewed, so those results remain provisional.
I could see using GLiNER to classify intent, code to dispatch requests to specialist LLM agents, and Jev for decisions involving policy and context. The graph below sketches that setup. The three named specialists and Agent N represent an expandable set of agents, not a requirement to run them all. Each is a generative LLM, typically decoder-based, operating within an agent harness.
GLiNER and Jev make model judgments. Dispatch, record and policy lookups by ID, permission checks and approved tool invocation are deterministic code. After the tool returns, the selected agent composes the answer back to the client. That response box is the same agent resuming, not another specialist. Unresolved cases go to a person. I have not tested this combined pipeline.
Deterministic code does not make its inputs correct or prevent a tool call from failing. It also does not mean model inference must be stochastic. Where a rule can be written exactly, such as comparing a verified purchase date with a refund deadline, I would keep it in code. No extra LLM is needed to execute an already approved tool call.
Routine classification
GLiNER2.5-Decide is an English classification model. Its model card covers intent, routing, sentiment, urgency and handoff, but says it “does not reason, explain, or answer open questions.” That distinction matters for the policy test later.
I used the public Fast Decisions development release, not the vendor's private test benchmark. It has 1,700 documents across 17 domains and 29 question-level tasks. Think support queues, banking requests, review sentiment and product-feedback tags. Some documents have several questions. Training overlap is unknown.
GLiNER ran locally on an M3 Max using MPS. Jev used the jev-1.13.0 API. Neither was fine-tuned for these tasks. Both received the same substantive inputs and candidate meanings, but the interfaces differ. On the three multi-label tasks, GLiNER used its native interface and Jev answered a binary question for each candidate.
GLiNER led on 16 tasks, Jev on 12, with one tie. The plot uses cases where both returned valid outputs, leaving 95 to 100 cases per task. Multi-label answers count as correct only when the entire label set matches. The 95% paired bootstrap intervals use 2,000 resamples, without adjustment for multiple comparisons.
GLiNER did better on ticket queues and product-feedback types. Jev did better on restaurant aspect tags and identifying tickets containing personal information. Many intervals cross zero. Candidate sets range from 2 to 28 classes, with some classes represented by just one example. I would not turn the task-win count into an overall ranking or a claim that the models are equivalent. Class counts are included in the evidence download.
There were also 21 Jev responses whose distributions summed to 0.99. The frozen validator rejected them rather than renormalizing. Counting those cases as unsuccessful changes the tally to 18 GLiNER leads, 10 Jev leads and one tie. They are excluded from the shared-valid plot, not from the record of failures.
Applying a supplied policy
Suppose a policy allows refunds within 20 days and a purchase is 19 days old. With the other requirements met, approve it. Change the limit to 18 days, with no exception, and decline it. The request is the same. The supplied policy changes the answer.
I built 600 cases across refunds, invoice approval and agent handoff, with 100 paired families per workflow. Pairs change a policy, fact, exception, wording or the availability of required information. Wording-only changes preserve the meaning. Deterministic code generates the reference answers from factors hidden from the models. The templates were AI-assisted and have limited variety. The labels have not received human review. Both models returned valid outputs on all 600 cases.
| Workflow | GLiNER direct action | Jev direct action |
|---|---|---|
| Refunds | 10% | 99.5% |
| Invoice approval | 10% | 100% |
| Agent handoff | 10% | 100% |
The repeated 10% looked strange to me too. GLiNER chose “request information” on 598 of 600 cases. That is the reference action in exactly 20 of the 200 cases per workflow. The 10% is accuracy, not model confidence.
I checked the adapter against the installed native GLiNER API on 18 diagnostic cases. Encoded inputs matched and probabilities agreed to within 0.0000001. CPU and MPS chose the same labels. There was no truncation, and asking only the action question did not fix the sample.
I did correct a separate multi-label decoding mismatch before the full public classification run, under a recorded protocol amendment. It does not explain this result because the policy test has no multi-label questions.
That audit does not explain the collapse. An upstream issue or a different input formulation could still matter. Jev clearly did better here, but applying changing rules supplied in the input goes beyond routine intent classification. I would not use these scores to judge GLiNER's general classification ability.
Direct and derived actions
This was the more interesting result. Each policy case asks for a direct action and four atomic judgments covering completeness, the ordinary numerical limit, exceptions and urgency. I then use shared deterministic code to derive an action from the relevant model answers. Missing facts mean request information. Otherwise, allow the action if the ordinary rule or an exception permits it, and deny it if neither does. Urgency does not enter this action rule.
Jev's direct actions scored 100% on invoices and handoff. Applying the code to its atomic answers reduced those scores to 78% and 75%. Refunds went from 99.5% to 100%. These are the same 200 cases per workflow, with no abstention or filtering. Intervals use 2,000 bootstrap resamples of the 100 paired families. A zero-width interval at 100% describes this sample, not zero future risk.
The problem was the exception question. Jev got the other three atomic questions right throughout, but exception accuracy was 43.5% on invoices and 50% on handoff. In each workflow, 180 of 200 cases have no exception, so always answering no would score 90%. Not every wrong exception answer changes the action, which is why the derived-action scores are higher.
These were separately requested outputs in the same bundled call, not a trace of Jev's internal reasoning. The test measures consistency between answers. Correct final decisions did not imply reliable supporting judgments.
Where I would use each model
For ordinary routing and tagging, I do not see a clear reason from this panel to choose Jev over GLiNER. Jev looks more useful for the supplied-policy decisions tested here, subject to the unreviewed labels and limited templates. I would be much more careful about building an action from several of its answers. That procedure lost as much as 25 percentage points compared with using the direct answer.
That is why I would test GLiNER as the classifier and Jev for policy decisions, rather than assume one should do both. The combined system would need an end-to-end test that includes wrong routes, incorrect policy retrieval and missing records. Router confidence would also need validation on actual traffic. These separate results do not establish that combining the models improves accuracy, cost or latency.
I have not run a task-specific DeBERTa baseline, so this says nothing about whether either beats a trained classifier. I am expanding the tests to include CLM-v0.1-8B alongside Jev and GLiNER2.5-Decide on additional tasks. CLM is an open-weight model that scores candidate actions against the current state rather than generating a response. I want to see whether it offers a useful alternative to Jev for context-dependent decisions, and what accuracy and latency tradeoffs appear when running it locally. Those results, the broader intent evaluation and latency measurements are outside this short article.
The broader point is that effective agent orchestration requires thinking about each node in the graph. Some need a generative LLM, others a classifier, and others just code. Getting that mix right could improve accuracy, latency and cost while keeping more of the workflow deterministic.
Methods and evidence
Model outputs are learned judgments. Label decoding, probability validation, reference generation, action composition and metric calculations are deterministic code. The bootstrap uses the fixed seed 20260929. None of that makes a model answer or an unreviewed label correct.
The methods supplement covers revisions, interfaces and the post-hoc adapter audit. The evidence download includes inputs, references, saved predictions, class counts and an offline verifier. It excludes credentials, model weights and unrelated experiments.