Methods for the Jev and GLiNER comparison
This supplement describes the two completed panels in the preliminary article. It does not claim completion of the broader benchmark or of a multi-agent deployment. The orchestration diagram is a proposed architecture. No specialist generative LLM agents were evaluated in this report.
Models and interfaces
GLiNER uses fastino/GLiNER2.5-Decide at revision 5a7adf72a23b4d311abae6ce050d7f0012bb3416, through gliner2==2.0.0. The checkpoint ran in FP32 on an M3 Max using MPS. Its native mutually exclusive softmax and independent multi-label scores were retained. Jev used the pinned API model jev-1.13.0, which was also recorded in the returned responses. Neither model received task-specific fine-tuning.
The inputs preserve the same substantive state, instructions and candidate meanings, but the interfaces differ. GLiNER handles multi-label questions natively. Jev receives one binary question per candidate, with the questions bundled against the same state. The final multi-label score requires the entire predicted set to match the reference set. The decoder selects labels with probability at least 0.5. GLiNER retains its native fallback to the highest-scoring label when no label reaches the threshold. Jev can return an empty set. Single-label and ordinal questions use the maximum-probability label, with the original candidate order breaking ties.
These are deterministic decoding rules applied to learned model scores. They do not establish calibration, correctness or identical numerical outputs across devices or future API calls. The post makes no latency or cost comparison.
Public classification data
The source is the Fast Decisions public development release at revision 1a33070cabf94ce2e29105482dd2ef6c157ad7f2, licensed Apache-2.0. It contains 1,700 documents across 17 domains, with 100 documents per domain and 29 question-level tasks. Several tasks share documents. This is not the private test set used in the model card. Upstream reference labels have not been independently adjudicated, and training overlap is unknown.
The three multi-label tasks are product areas, restaurant aspects and screen genres. Class counts and positive reference counts for every task appear in the class-counts download and evidence archive. Counts for multi-label tasks can exceed 100 because a document can have several labels. Some single-label classes have only one example.
All 1,700 GLiNER records passed validation. Jev passed on 1,679 documents. On 21 documents, a distribution summed to 0.99 and the frozen validator rejected the whole response. All questions on those documents count as failed for the scheduled-case comparison, even if some individual outputs could have been recovered. The common-valid comparison excludes the same documents from both models within the affected tasks, leaving 95 to 100 cases per task. Its descriptive task-win tally is 16 GLiNER leads, 12 Jev leads and one tie. Counting rejected responses as unsuccessful gives 18, 10 and one.
Fresh policy scenarios
The fresh test panel contains 300 families with two variants each. Refunds, invoice approval and agent handoff each have 100 families. Within each domain, 20 families are allocated to each of five perturbations involving the policy, a numerical fact, an exception, a meaning-preserving paraphrase or removal of required information. A separate 90-family development panel was used in the wider experiment, but no calibration or abstention results from it are claimed in this article.
Templates and rules were AI-assisted. Their diversity is limited. Deterministic code creates the cases from seeded factors and separately derives reference answers. Model inputs contain the rendered state and questions, not evaluator labels, factor dictionaries or family identifiers. Development and test state templates differ, while question schemas are shared. All fresh reference labels remain unreviewed by humans. Automated consistency checks are not human adjudication, and no result should be described as validated real-world accuracy.
Each case asks five questions in one bundled request. Four are atomic judgments about information completeness, the ordinary numerical limit, exception presence and urgency. The fifth is a direct action choice. Both models produced valid outputs on all 600 test cases.
The reference policy and action-composition code first request information when the numerical fact is missing or the within-rule judgment is unknown. Otherwise they allow the domain action if the ordinary limit is met or an exception is present, and deny it otherwise. Urgency does not enter this action rule. In the composed procedure, the inputs to that deterministic rule are predicted atomic answers, not verified facts. The atomic answers are not reasoning traces.
Exceptions are absent in 180 of 200 reference cases per workflow. The direct-action reference counts are 108 approvals, 72 declines and 20 information requests for refunds, 94 releases, 86 holds and 20 information requests for invoices, and 100 continuations, 80 handoffs and 20 information requests for support. These counts explain why constant predictions can give misleading impressions of capability.
The generated content is offered under CC-BY-4.0 with attribution to Samuel Kahn and this benchmark project, identifying its AI-assisted, unreviewed status.
Adapter audit and amendment
The extreme GLiNER action pattern prompted a post-hoc diagnostic audit on 18 selected cases covering development and test strata. On this sample, the adapter and installed native API used identical encoded inputs. Their maximum probability difference was approximately 0.000000071. CPU and MPS selected identical labels, with maximum probability difference approximately 0.00000322. The audit found no truncation or processor fallback on the sample. Action-only calls did not rescue the sample. These diagnostic cases are not an independent accuracy estimate and do not rule out upstream issues or sensitivity to other input formulations.
Source review found a separate multi-label decoder mismatch. The original shared decoder allowed an empty set where native GLiNER forces one label. Amendment 002 corrected that behavior after fresh predictions were inspected but before the full Fast Decisions run. The change did not affect the fresh panel, which has no multi-label questions. No test prompt, state, reference, candidate or weight was tuned. The original records were preserved, and unaffected records were imported without changing their contents. The amended freeze is 57169970b82bc0e8396d877090fca55d217a4a6ac0b51449a8e689711c8d851d.
Scoring and uncertainty
Single-label accuracy compares the decoded label to the reference. Multi-label accuracy is exact set match. No cases are removed by a confidence threshold in the reported fresh scores. Invalid outputs are explicit failures rather than repaired predictions. Exclusive probabilities must be finite, nonnegative, contain the exact candidate set and sum to one within 0.0001.
The chart intervals are percentile intervals from 2,000 cluster-bootstrap resamples using seed 20260929. Related variants remain together by family. Public model differences use paired observations from the common-valid subset. These intervals do not adjust for multiple comparisons. Perfect observed scores produce zero-width empirical intervals and must not be interpreted as proof of zero future risk.
Label decoding, validation, reference generation, action composition and metric calculation are deterministic code. Resampling uses a fixed pseudorandom seed and is reproducible from the same observations. The proposed graph similarly marks code dispatch, ID-based lookups, permission checks and approved tool invocation. Its generative specialist agents and the decision models supply learned judgments. Human review is a separate activity.
Reproduction
The evidence archive contains an allowlisted export of 2,300 evaluated inputs, separate reference records within each exported case, and predictions from both models. It excludes secrets, HTTP headers, account identifiers, model weights and unrelated panels. verify_short_evidence.py recomputes the point estimates, class counts, completion counts and direct-versus-derived results using only the exported JSONL and the Python standard library. It checks them against the saved chart snapshot. Run it with python3 verify_short_evidence.py after extracting the archive.
The archive also contains the fixed chart evidence, figure builders and the relevant analysis, generation and adapter source for inspection. Regenerating charts requires NumPy and Matplotlib. The figure builders use saved intervals rather than resampling during rendering. The manifest identifies files with SHA-256 hashes. The full inference environment is not needed for this offline check.
The report is publishable only as a preliminary, explicitly unreviewed case study. Human review of the fresh references remains necessary before removing that qualification. A deployment claim would additionally require an end-to-end evaluation of the proposed agents, retrieval, dispatch and execution controls.