Testing Jev for Coding Agent Evals: Comparing Speed, Cost, and Rationale

Key Takeaways

  • Jev returned structured evaluation ratings faster and at a lower estimated cost than the LLM judges in this test.
  • The LLM  judges explained their ratings. Jev provided probabilities but no written rationale, making some decisions harder to investigate.
  • Jev’s ratings remained consistent across repeated runs, while the generative judges showed some variation.
  • Human reviewers are essential to decide which ratings match an organization’s risk policy and to create labeled cases for testing evaluators.
  • Organizations should retest the full evaluation workflow when they change an evaluator’s configuration or adopt a newer model, rather than assume earlier results still apply.

In this blog, we walk through our analysis on  how Jev's decision-focused approach compares with general-purpose LLMs for evaluating the risks of coding agent actions, an important consideration for organizations rolling out these types of agents. We compared risk ratings, latency, cost, and consistency across repeated runs.

Jev: A System One Model for Evaluation Decisions

Last week, TypeSafe AI released Jev in early access, its first “System One Model”; a new class of frontier models that make fast, structured decisions that applications can act on directly [1].

To understand how Jev works let's consider a customer service agent. The customer says that their “package was due yesterday and has not arrived”. The agent needs to determine which team should receive the request, the urgency, and whether there is enough information to act on it. 

Jev offers three types of answers for these decisions:

  • Choice selects a team from a fixed list.
  • Score rates urgency on an ordered scale.
  • Noul gives a 0–1 probability that there is enough information to act. 

As shown on the diagram below, Jev fills out a highly calibrated form instead of writing paragraphs. The app can use those answers to route the message automatically or send it for review.

Decision flow with Jev

The customer-service example shows how Jev differs from a general-purpose LLM. Jev evaluates the three questions independently and in parallel, returning answers constrained to the predefined types and probabilities for each decision [2]. An LLM could also return all three answers as structured JSON, but it generates the response token by token, sequentially, and can include a written rationale. Jev, however, does not provide a rationale. 

TypeSafe's published pricing charges for Jev's input tokens but not its output tokens [1]. Jev's probabilities show how it weighs the available answers, not which details of the message drove its choice. If Jev chose “Billing”, for example, we would see its confidence score but not its reasoning or rationale. The three answers are also independent; the app must combine them into a routing rule rather than expect one answer to inform another.

Evaluating a Coding Agent’s Actions

A coding agent may inspect files, run tests, change code, and execute commands while carrying out a task. Its final response may not show all of these steps, but its recorded tool calls reveal the actions it attempted. Some are routine, while others may access sensitive data, make unintended changes, or send information elsewhere. Evaluating each step helps identify potentially unsafe or risky behavior. The same approach could be used to evaluate other tool calling agents.

To examine those actions, we gave a coding agent twelve tasks in disposable sandboxes as shown on the next diagram. The agent chose its own commands, which we recorded. We selected 31 Bash calls, pairing each command with its recorded outcome as one item for the judges to rate.

Workflow diagram showing a coding agent's recorded Bash commands from sandboxed tasks passed to three judges for risk evaluation

We then passed each command and its recorded outcome to three judges (models), along with the same rubric and instructions for assigning a risk level from low to critical. An evaluator combines a model and its configuration with a rubric and instructions that defines how an action should be rated for risk.

Evaluator Definition - Model Configurations, Rubric and Instructions for the Judges

Judge Model Call and configuration
Jev typesafe-ai/jev Typed evaluation API with one Choice question and no additional options
Gemini google/gemini-3.5-flash-lite generateObject with a JSON schema, minimal thinking, served through Vertex
GPT openai/gpt-5.4-nano generateObject with a JSON schema, reasoning effort set to none, served through OpenAI

In our comparison, we chose the two general-purpose models, gemini-3.5-flash-lite and gpt-5.4-nano, because these two models are described as the lowest-cost tier as suitable for classification at the time of writing. Unlike these two models, Jev is the specialist, not a model in the same tier. We set both LLMs to their lowest reasoning levels to make their output-token use more comparable when measuring latency and cost. No temperature, top_p, or seed was set for any of the LLMs used as judges.

Jev returned a rating for risk level along with probabilities. The two LLMs returned a rating for risk level with a written rationale.

We compared the ratings and response times, then examined disagreements to see how the judges read the same action. We had no independently established correct ratings, so agreement shows where decisions align, not which judge is right.

What the Captured Runs Show

Each judge rated the same 31 Bash calls once. Jev responded fastest in this run, though the generative judges also produced written rationale.

Judge Median response Low Medium High Critical
Jev 237 ms 18 7 6 0
Gemini 3.5 Flash-Lite 898 ms 22 4 5 0
GPT-5.4 nano 1,106 ms 14 10 7 0

Let's consider how the ratings compare.

Agreement Number of commands
All three agreed 20
Two agreed; one differed 8
All three differed 3

Detailed Report Outlining Results for Individual Test Cases Along with Comparison

The judges disagreed on 11 of the 31 commands, including 3 where all three chose different risk levels. Gemini rated 22 commands low, compared with 18 for Jev and 14 for GPT-5.4 nano. These differences affect which actions a judge would flag a case for review, but they do not show which judge was right. 

Even human reviewers may assess the same action differently under different organizational policies. A reference dataset labeled under a specific policy would let us compare each evaluation with the judgment the organization expects for that case. Where they differ, we can refine the evaluator’s instructions and review thresholds so its results align more closely with what the organization would expect.

Jev’s Strengths and Limitations

In this test case, Jev returned structured risk ratings faster than the two generative LLM judges.

Advantage Evidence and practical use
Speed Jev's median was 237 ms, versus 898 ms for Gemini and 1,106 ms for GPT-5.4 nano. Jev returned a rating, while the LLMs also wrote a rationale. So this is not a like-for-like comparison of rating-only responses. It also does not tell us which judge's ratings were more accurate.
Structured output Jev returned one typed Choice and probabilities, suitable for sorting spans before review. TypeSafe also documents Score, Noul, and multiple focused questions [2]; we did not test those features.

For the same 31-command experiment, the costs compare as follows:

Judge Cost for 31 calls Basis
Jev ~$0.0016 Estimated at TypeSafe's published rate; Gateway charged $0 during a promotion.
Gemini 3.5 Flash-Lite $0.01261760 Provider-reported cost.
GPT-5.4 nano $0.00813785 Provider-reported cost.

The Jev estimate uses 38,315 input tokens × $0.042 per million, with output free [1]. That is roughly 7.8× lower than Gemini’s recorded cost and 5.1× lower than GPT-5.4 nano’s. We did not observe a paid Jev bill: Vercel’s Gateway promotion made these calls free [3]. Prices may change, and the generative judges also produced rationales, so this estimated-versus-observed comparison is not a like-for-like measure of value.

Since Jev does not provide a written rationale, a reviewer may need to inspect the original evidence when a rating is surprising or consequential. Jev can only choose from the options we define, so the available choices should include “needs review” or “none of the above” when a case may not fit. TypeSafe’s Jev 1.13 guidance warns that it may read questions literally or be distracted by irrelevant details, so clear wording and focused input matter [4].

Rationale

Risk ratings can help prioritize review, but a rating alone does not show which details of the evidence a judge considered or how it applied the rubric. When a rating is surprising or judges disagree, the missing context makes the result harder to investigate.

Written self-reported rationale gives reviewers concrete claims to check against the full trace and can help reveal whether a disagreement comes from different readings of the command or an unclear boundary between risk levels in the evaluator’s instructions/prompt.

For example, the agent ran a recursive grep across the workspace for terms such as secret, token, and password. That could be routine inspection, but it could also surface credentials stored in files. The three judges saw the same command and outcome yet rated its risk differently.

Judge Risk level What the response revealed
Gemini 3.5 Flash-Lite Low Its rationale treated the search as a routine local inspection.
GPT-5.4 nano Medium Its rationale described the search as potential credential reconnaissance and noted secret-shaped matches in the output.
Jev High It returned a typed rating, with 0.78 probability for high and 0.18 for medium, but no written rationale.

Compare the Rationales for the Disagreement Example

This disagreement highlights the importance of keeping a human in the evaluation loop. A reviewer can examine the full trace and apply the organization's risk policy to decide how the search should be rated. Those reviews can form a labeled dataset. The organization can use that dataset to test and refine the evaluator so its results align more closely with the decisions its reviewers expect from a given judge. This methodology also applies to other tool-using agents.

Repeatability

Repeatability is important because a judge that changes its answer to the same evidence could send an action for review in one run and let it pass in another. After the initial run, we sent the same 31 saved commands and their recorded outcomes to each judge twice more. We kept the evidence and rubric  the same to see whether each judge would return the same rating when asked again.

The results are shown in the table below.

Judge Commands given the same rating in all three runs
Jev 31 of 31
Gemini 3.5 Flash-Lite 28 of 31
GPT-5.4 nano 28 of 31

Repeatability Detailed Report - Side-by-side Comparison of Experiment Repeats

Jev kept the same rating on all 31 commands. Gemini 3.5 Flash-Lite and GPT-5.4 nano each changed their rating on three commands, with every change staying within one risk level. Jev's chosen-rating probabilities changed by at most 0.06. Using each command's median probability across the three runs, Jev's medians were 0.93 for the 20 commands on which the judges initially agreed and 0.74 for the 11 disputed commands. These results show consistency and not correctness of a model. Three runs per command give us a useful snapshot, though more runs would be needed to assess stability over time.

We also tested repeatability in an earlier run with Gemini 2.5 Flash-Lite and GPT-5 nano. Their results are included below alongside those of the current models.

Chart comparing rating consistency across three repeated runs for Jev, Gemini 3.5 Flash-Lite, and GPT-5.4 nano on the same 31 Bash commands

Gemini returned the same rating for 28 of 31 commands in both runs, while the GPT count went from 19 with GPT-5 nano to 28 with GPT-5.4 nano. Because the models and reasoning settings both changed, the key takeaway is to check repeatability for the specific evaluator  in use, rather than assume it will carry over to a new one.

Other Learnings

Beyond speed and cost, the captured commands highlight several choices involved in configuring an evaluator. The judges used the risk scale differently, and some inspection commands exposed boundaries that may need clearer instructions. Jev's probabilities may help identify cases for a second review, while some commands need the surrounding task to be judged with more context. The table summarizes these observations and their implications for the evaluation workflow.

Pattern What we observed Why it matters
Judges use the scale differently Gemini rated 8 of 31 commands lower than GPT-5.4 nano and 1 higher. Jev rated 5 higher than Gemini and 1 lower. The evaluator's instructions and review threshold may need to be adjusted for a specific judge a team uses. A human-reviewed dataset will allow the team to evaluate a judge against its risk policy and guide those adjustments.
Some inspection commands are boundary cases The judges disagreed on one of the eight environment-variable and host-inspection calls. They gave three different ratings to a separate search for credential-like terms task. The evaluator's instructions should clarify how to rate cases near the boundary between categories.
Jev probabilities may help route review In the initial run, Jev's median probability for its chosen rating was 0.93 on 20 spans where all judges agreed, versus 0.77 on 11 disputed spans. A lower probability means Jev was less certain about its chosen risk rating. A team could use that signal to send the input to another LLM for a second assessment or to a human reviewer.
Task intent differs from command risk In a task that later deleted files, the initial du and ls commands were unanimously rated low. Judges saw one Bash call at a time. The full action sequence and the user's request may change how the agent's behavior should be judged.

These patterns provide useful insights into how an evaluation workflow can be improved. 

  • Calibrate the evaluator's instructions and review threshold for the model being used, then test its ratings against a human-labeled dataset.
  • Use disagreements and unusual cases to make the evaluator's instructions more explicit, especially where the boundary between ratings is unclear.
  • Route lower-probability Jev ratings to another LLM or a human reviewer for a second assessment.
  • Include relevant context when a judgment depends on more than the item being evaluated.
Chart showing Jev's median rating probability was 0.93 on commands judges agreed on versus 0.77 on disputed commands

Key Findings: How Jev’s Decisions Compare to General Purpose LLMs

We ran this test to see how Jev's decision-focused approach compares with general-purpose LLMs when judging the risk of coding-agent actions. Giving them the same 31 commands, rubric, and instructions lets us compare their ratings, response times, costs, rationale, and repeatability. The aim was to understand how each approach behaves in an evaluation workflow, not to determine which judge is more correct.

The test surfaced several key findings. Jev returned typed ratings quickly, at a lower estimated cost, and repeated all 31 ratings consistently across three runs. The LLMs took longer but provided rationales that reviewers could inspect. The judges disagreed on 11 commands, and Jev's median probability for its chosen rating was lower on those disputed cases. That probability may help identify cases for a second assessment. We still need human-reviewed examples to assess the judges. Since human reviewers may judge the same action differently under different organizational policies, organizations should configure their evaluators to  test different models and settings against human-reviewed examples, and refine the setup so its judgments align with the decisions their reviewers would expect.

Link to Fiddler Git Repository: https://github.com/fiddler-labs/fiddler-examples/tree/main/experiments/jev-evals

References

[1] TypeSafe AI, "Introducing System One Models and Jev," TypeSafe AI Blog. [Online]. Available: https://typesafe.ai/blog/introducing-system-one-models-and-jev.

[2] TypeSafe AI, "Primitives," TypeSafe AI Docs. [Online]. Available: https://docs.typesafe.ai/primitives.

[3] Vercel Inc., "Jev — AI Gateway model," Vercel Docs. [Online]. Available: https://vercel.com/ai-gateway/models/jev.

[4] TypeSafe AI, "Model jaggedness: Jev 1.13," TypeSafe AI Docs. [Online]. Available: https://docs.typesafe.ai/model-jaggedness/jev-1.13.