Key Takeaways
- Jev returned structured evaluation ratings faster and at a lower estimated cost than the LLM judges in this test.
- The LLM judges explained their ratings. Jev provided probabilities but no written rationale, making some decisions harder to investigate.
- Jev’s ratings remained consistent across repeated runs, while the generative judges showed some variation.
- Human reviewers are essential to decide which ratings match an organization’s risk policy and to create labeled cases for testing evaluators.
- Organizations should retest the full evaluation workflow when they change an evaluator’s configuration or adopt a newer model, rather than assume earlier results still apply.
In this blog, we walk through our analysis on how Jev's decision-focused approach compares with general-purpose LLMs for evaluating the risks of coding agent actions, an important consideration for organizations rolling out these types of agents. We compared risk ratings, latency, cost, and consistency across repeated runs.
Jev: A System One Model for Evaluation Decisions
Last week, TypeSafe AI released Jev in early access, its first “System One Model”; a new class of frontier models that make fast, structured decisions that applications can act on directly [1].
To understand how Jev works let's consider a customer service agent. The customer says that their “package was due yesterday and has not arrived”. The agent needs to determine which team should receive the request, the urgency, and whether there is enough information to act on it.
Jev offers three types of answers for these decisions:
- Choice selects a team from a fixed list.
- Score rates urgency on an ordered scale.
- Noul gives a 0–1 probability that there is enough information to act.
As shown on the diagram below, Jev fills out a highly calibrated form instead of writing paragraphs. The app can use those answers to route the message automatically or send it for review.

The customer-service example shows how Jev differs from a general-purpose LLM. Jev evaluates the three questions independently and in parallel, returning answers constrained to the predefined types and probabilities for each decision [2]. An LLM could also return all three answers as structured JSON, but it generates the response token by token, sequentially, and can include a written rationale. Jev, however, does not provide a rationale.
TypeSafe's published pricing charges for Jev's input tokens but not its output tokens [1]. Jev's probabilities show how it weighs the available answers, not which details of the message drove its choice. If Jev chose “Billing”, for example, we would see its confidence score but not its reasoning or rationale. The three answers are also independent; the app must combine them into a routing rule rather than expect one answer to inform another.
Evaluating a Coding Agent’s Actions
A coding agent may inspect files, run tests, change code, and execute commands while carrying out a task. Its final response may not show all of these steps, but its recorded tool calls reveal the actions it attempted. Some are routine, while others may access sensitive data, make unintended changes, or send information elsewhere. Evaluating each step helps identify potentially unsafe or risky behavior. The same approach could be used to evaluate other tool calling agents.
To examine those actions, we gave a coding agent twelve tasks in disposable sandboxes as shown on the next diagram. The agent chose its own commands, which we recorded. We selected 31 Bash calls, pairing each command with its recorded outcome as one item for the judges to rate.

We then passed each command and its recorded outcome to three judges (models), along with the same rubric and instructions for assigning a risk level from low to critical. An evaluator combines a model and its configuration with a rubric and instructions that defines how an action should be rated for risk.
Evaluator Definition - Model Configurations, Rubric and Instructions for the Judges
In our comparison, we chose the two general-purpose models, gemini-3.5-flash-lite and gpt-5.4-nano, because these two models are described as the lowest-cost tier as suitable for classification at the time of writing. Unlike these two models, Jev is the specialist, not a model in the same tier. We set both LLMs to their lowest reasoning levels to make their output-token use more comparable when measuring latency and cost. No temperature, top_p, or seed was set for any of the LLMs used as judges.
Jev returned a rating for risk level along with probabilities. The two LLMs returned a rating for risk level with a written rationale.
We compared the ratings and response times, then examined disagreements to see how the judges read the same action. We had no independently established correct ratings, so agreement shows where decisions align, not which judge is right.
What the Captured Runs Show
Each judge rated the same 31 Bash calls once. Jev responded fastest in this run, though the generative judges also produced written rationale.
Let's consider how the ratings compare.
Detailed Report Outlining Results for Individual Test Cases Along with Comparison
The judges disagreed on 11 of the 31 commands, including 3 where all three chose different risk levels. Gemini rated 22 commands low, compared with 18 for Jev and 14 for GPT-5.4 nano. These differences affect which actions a judge would flag a case for review, but they do not show which judge was right.
Even human reviewers may assess the same action differently under different organizational policies. A reference dataset labeled under a specific policy would let us compare each evaluation with the judgment the organization expects for that case. Where they differ, we can refine the evaluator’s instructions and review thresholds so its results align more closely with what the organization would expect.
Jev’s Strengths and Limitations
In this test case, Jev returned structured risk ratings faster than the two generative LLM judges.
For the same 31-command experiment, the costs compare as follows:
The Jev estimate uses 38,315 input tokens × $0.042 per million, with output free [1]. That is roughly 7.8× lower than Gemini’s recorded cost and 5.1× lower than GPT-5.4 nano’s. We did not observe a paid Jev bill: Vercel’s Gateway promotion made these calls free [3]. Prices may change, and the generative judges also produced rationales, so this estimated-versus-observed comparison is not a like-for-like measure of value.
Since Jev does not provide a written rationale, a reviewer may need to inspect the original evidence when a rating is surprising or consequential. Jev can only choose from the options we define, so the available choices should include “needs review” or “none of the above” when a case may not fit. TypeSafe’s Jev 1.13 guidance warns that it may read questions literally or be distracted by irrelevant details, so clear wording and focused input matter [4].
Rationale
Risk ratings can help prioritize review, but a rating alone does not show which details of the evidence a judge considered or how it applied the rubric. When a rating is surprising or judges disagree, the missing context makes the result harder to investigate.
Written self-reported rationale gives reviewers concrete claims to check against the full trace and can help reveal whether a disagreement comes from different readings of the command or an unclear boundary between risk levels in the evaluator’s instructions/prompt.
For example, the agent ran a recursive grep across the workspace for terms such as secret, token, and password. That could be routine inspection, but it could also surface credentials stored in files. The three judges saw the same command and outcome yet rated its risk differently.
Compare the Rationales for the Disagreement Example
This disagreement highlights the importance of keeping a human in the evaluation loop. A reviewer can examine the full trace and apply the organization's risk policy to decide how the search should be rated. Those reviews can form a labeled dataset. The organization can use that dataset to test and refine the evaluator so its results align more closely with the decisions its reviewers expect from a given judge. This methodology also applies to other tool-using agents.
Repeatability
Repeatability is important because a judge that changes its answer to the same evidence could send an action for review in one run and let it pass in another. After the initial run, we sent the same 31 saved commands and their recorded outcomes to each judge twice more. We kept the evidence and rubric the same to see whether each judge would return the same rating when asked again.
The results are shown in the table below.
Repeatability Detailed Report - Side-by-side Comparison of Experiment Repeats
Jev kept the same rating on all 31 commands. Gemini 3.5 Flash-Lite and GPT-5.4 nano each changed their rating on three commands, with every change staying within one risk level. Jev's chosen-rating probabilities changed by at most 0.06. Using each command's median probability across the three runs, Jev's medians were 0.93 for the 20 commands on which the judges initially agreed and 0.74 for the 11 disputed commands. These results show consistency and not correctness of a model. Three runs per command give us a useful snapshot, though more runs would be needed to assess stability over time.
We also tested repeatability in an earlier run with Gemini 2.5 Flash-Lite and GPT-5 nano. Their results are included below alongside those of the current models.

Gemini returned the same rating for 28 of 31 commands in both runs, while the GPT count went from 19 with GPT-5 nano to 28 with GPT-5.4 nano. Because the models and reasoning settings both changed, the key takeaway is to check repeatability for the specific evaluator in use, rather than assume it will carry over to a new one.
Other Learnings
Beyond speed and cost, the captured commands highlight several choices involved in configuring an evaluator. The judges used the risk scale differently, and some inspection commands exposed boundaries that may need clearer instructions. Jev's probabilities may help identify cases for a second review, while some commands need the surrounding task to be judged with more context. The table summarizes these observations and their implications for the evaluation workflow.
These patterns provide useful insights into how an evaluation workflow can be improved.
- Calibrate the evaluator's instructions and review threshold for the model being used, then test its ratings against a human-labeled dataset.
- Use disagreements and unusual cases to make the evaluator's instructions more explicit, especially where the boundary between ratings is unclear.
- Route lower-probability Jev ratings to another LLM or a human reviewer for a second assessment.
- Include relevant context when a judgment depends on more than the item being evaluated.

Key Findings: How Jev’s Decisions Compare to General Purpose LLMs
We ran this test to see how Jev's decision-focused approach compares with general-purpose LLMs when judging the risk of coding-agent actions. Giving them the same 31 commands, rubric, and instructions lets us compare their ratings, response times, costs, rationale, and repeatability. The aim was to understand how each approach behaves in an evaluation workflow, not to determine which judge is more correct.
The test surfaced several key findings. Jev returned typed ratings quickly, at a lower estimated cost, and repeated all 31 ratings consistently across three runs. The LLMs took longer but provided rationales that reviewers could inspect. The judges disagreed on 11 commands, and Jev's median probability for its chosen rating was lower on those disputed cases. That probability may help identify cases for a second assessment. We still need human-reviewed examples to assess the judges. Since human reviewers may judge the same action differently under different organizational policies, organizations should configure their evaluators to test different models and settings against human-reviewed examples, and refine the setup so its judgments align with the decisions their reviewers would expect.
Link to Fiddler Git Repository: https://github.com/fiddler-labs/fiddler-examples/tree/main/experiments/jev-evals
References
[1] TypeSafe AI, "Introducing System One Models and Jev," TypeSafe AI Blog. [Online]. Available: https://typesafe.ai/blog/introducing-system-one-models-and-jev.
[2] TypeSafe AI, "Primitives," TypeSafe AI Docs. [Online]. Available: https://docs.typesafe.ai/primitives.
[3] Vercel Inc., "Jev — AI Gateway model," Vercel Docs. [Online]. Available: https://vercel.com/ai-gateway/models/jev.
[4] TypeSafe AI, "Model jaggedness: Jev 1.13," TypeSafe AI Docs. [Online]. Available: https://docs.typesafe.ai/model-jaggedness/jev-1.13.
