# What a Decision Model Taught Our Agent Harness

> We gave Jev one small, frequent decision in our agent harness: which governed data an agent may use for a question. It was quick and usually right, and its one kind of mistake showed us what to fix in the harness itself.

Published: 2026-09-29
Updated: 2026-09-29
Author: Archie Norman (Founder, MLX)
Category: Research
Tags: jev, typesafe, agent-harness, decision-models, evaluations, calibration, governed-data
Canonical URL: https://mlx.systems/blog/what-a-decision-model-taught-our-agent-harness

## TL;DR

- A decision model suits the harness's data selection: in our tests Jev chose in a few hundred milliseconds, against one to two seconds for a general-purpose model.
- Every Jev miss was a near-tie that our code read as a firm no, so a cut-off, not the model, left out evidence the answer needed.
- Our rule-based harness had the same habit, and harness v3 offers its best-ranked Data Sets instead, lifting offline coverage on a synthetic ledger from 58% to 82%.
- Jev is not in the live harness; it runs in workflow judgment steps, and any decision model that returns to the harness will start in shadow.

---

Before an agent in the Twin writes a word, its harness has already made several decisions. The first is which governed Data Sets it may use for the question in front of it. Get that wrong and everything after it is harder: the agent explores tables it does not need, guesses at what the data means, or answers confidently from the wrong source.

It is also exactly the kind of decision a decision model is built for. The options are known in advance, the answer is a choice among them, and it has to be quick. So we gave it to Jev.

## A decision, not a paragraph

[Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev) is TypeSafe's decision model. Instead of generating text, it answers typed questions: pick one of these options, give this a score, yes or no, each with probabilities attached. George [took it apart for three cents](/blog/fingerprinting-jev-for-three-cents) earlier this month. What matters here is the interface. A selection step that returns probabilities can tell you how sure it is, not only what it chose.

We put Jev on the selection step and compared it with a general-purpose model, Gemini 3.8 Flash, doing the same job. Both saw the same frozen questions, the same list of executable operations and the same permissions. Neither was shown the expected answers.

## What it got right

Jev was fast. It made each selection in a few hundred milliseconds, where the general-purpose model took one to two seconds.

It was usually right, too. It declined the questions the data could not answer, as it should, and it never selected an operation a question did not need. On four new questions frozen in advance, both models made every choice correctly.

## A near-tie is not a no

Every miss was the same question: one about records with no match. The operation that answered it was on the list, and it worked. Jev did not pick it.

Its probabilities showed why. The needed operation came back at 0.45 to 0.47 for "select": close to an even call, not a clear no. Our code took Jev's top answer as final, so a near-tie between select and skip became a firm decision to leave out evidence the answer needed.

The model had told us it was unsure. Our code threw that away.

One fix suggested itself: when Jev is unsure, hand the choice to the general-purpose model. Replayed over the study's trials, that policy caught all three misses and would have cut average selection time roughly in half. We chose its threshold after seeing the errors, though, so we treat it as a hint rather than a result.

## The same habit in our own harness

That changed the question we asked of the harness itself: where else was a cut-off throwing away something a ranking already knew?

To answer it, we built an offline evaluation that replays the harness's selection with its own code, in about ten seconds, over 87 labelled questions and 13 test builds of a synthetic ledger. Each label says what a correct answer needs.

The rule-based harness had the same habit. It offered a Data Set only when its score cleared a fixed floor, so short questions scored too low and, a fifth of the time, nothing was offered at all. When a phrase in a question matched a Data Set, it locked the question to that one Data Set, even when it was the wrong one. Meanwhile, the harness's own ranking had the needed Data Sets in its top four almost nine times in ten.

Harness v3 trusts the ranking. It offers the four best-ranked Data Sets that share a term with the question, and it locks onto one only when someone asks for it. On the synthetic ledger:

- Offline, on held-out questions, it offered all the data a question needed 82% of the time, up from 58%. On new questions frozen before any results were seen, it did so 78% of the time, up from 46%.
- Wrong shortcuts, where a matching phrase locked a question to a single Data Set, fell from 30 to 1 across the 13 builds.
- In live test runs, questions where v3 newly offered the right data took about 40% fewer steps and 30% fewer tokens. Accuracy across the run was unchanged at 91%.

We are not claiming that this makes answers faster: the time comparison was not reliable. Harness v3 is in testing on staging and off for customers.

## What else the study changed

- **Keep the probabilities.** Our selector returned only the options it picked. Any selector we trust in future has to keep every probability, including those for the options it skipped.
- **Let code do mechanical work.** A model was retyping tables the agent had already selected, and formatting failures followed. In a replay, rendering those tables in code handled every one of them without a model call.
- **Better descriptions are not a fix.** Precise, typed descriptions of each operation made selection faster, but they did not stop the miss.
- **Measure the whole answer.** Selection took a second or two. Most of an answer's time went into the agent's loop of model and tool turns, and that loop is what harness v3 is designed to shorten.

## Where Jev sits now

Jev is not in the live harness. The study concluded that a decision model should make bounded decisions whose uncertainty is kept, that mechanical work belongs in code, and that unresolved selection stays with the general-purpose model.

Jev does run in one place: judgment steps in Workflows, for organisations we have enabled. Each step calls a pinned Jev version, and the request and response are kept with the run.

If a decision model returns to the harness, it will start in shadow: it will record its choice beside the harness's own decision without changing it, and it will be judged on the offline evaluation before it touches a live answer.

The most useful thing Jev did for our harness was not a decision. It was a probability we could read, pointing at the rule that was throwing evidence away.
