← All posts

Meet Jev: A Second Opinion That Answers in Types, Not Prose

Agent! now asks Jev (TypeSafe System One) how destructive a shell command is before running it. Here's why Jev is an advisor and not another LLM provider.

Today Agent! gets a new safety layer: Jev, short for TypeSafe System One. Before a shell command runs, Agent! can now ask Jev one question: how likely is this command to irreversibly destroy data? Above your threshold, the command is refused.

It landed in commit 337c94a2, "Add Jev (TypeSafe System One) decision layer." This post covers the design decision that matters most: Jev isn't another LLM provider.

Why not just ask another model?

Agent! already talks to 23 LLM providers. The easy route would have been to add a "safety model" as one more provider and ask it, in plain English, whether a command looks dangerous. The trouble is what comes back: prose. "This command could be risky depending on…" is not a verdict. You end up parsing sentences to decide whether to run rm.

Jev works differently. It answers typed questions about a state, with a Choice, a Score or a Noul, a typed "none" answer. It doesn't generate text, and it doesn't generate tool arguments. The commit message draws the conclusion: Jev "advises the existing tool loop rather than acting as an APIProvider."

That split is the whole design:

  • Your chosen model acts. It plans, calls tools and writes code.
  • Jev judges. It returns a number that code can compare against a threshold.

Where it sits

Jev doesn't replace anything. The order of checks for every shell command is:

  1. ShellSafetyService.check: hard-coded rules that refuse catastrophic commands. No model involved.
  2. Jev: only for commands that already passed step 1, as a second opinion on what the patterns can't see.
  3. Run the command.

On day one, the gate covers both shell execution paths in ShellTools (executeTCC and executeTCCStreaming).

Built to never stall your task

An advisor that can block your agent is dangerous in its own way. If the service is down, does your overnight task hang? The answer here is no. JevAdvisor is fail-open. No key, the toggle off, or an outage means "no opinion," and the command proceeds.

Fail-open is only safe if it's visible, though. Within the first day, two follow-ups made sure of that:

  • "Make the advisor observable" (e25384f3) added a usage callback and error reporting, so a failed check is logged instead of silently looking like "safe."
  • "Log the verdict, not just that Jev answered" (d3c47e67) made every check print the actual result: the destructive-risk percentage and whether the command was allowed or refused.

The client is a real package

The first version used a hand-rolled HTTP client: POST /v1/systemone for questions, GET /v1/models for the model list, Bearer auth, and backoff on 429 and 529 responses. Within hours, 9e330357 swapped that for the TypeSafeKit package, and f3a817b8 vendored the package inside the Agent repo so a build never depends on fetching it.

Setting it up

Jev lives in Settings with its own API key, stored in the Keychain in a dedicated non-provider slot. It has a model picker, loaded from /v1/models, and an advisory toggle. Turn the toggle off and Agent! behaves exactly as before.

Why this matters

Agents are getting shell access everywhere, and the standard safety answer is "the model will be careful." Pattern rules are better, because they're checkable. But they only know the patterns someone wrote down. A typed second opinion adds judgment without adding ambiguity: it returns a number, and code makes the call.

We'll be tuning the default threshold and watching the logs. If Jev refuses something it shouldn't, or misses something it should have caught, open an issue on GitHub and include the log line.