README.mdpreviewREADME.mdsource148 lines · 8.5 KB · raw

Chapter 12: jevhooks-daemon/src, the questions and the verdicts

Here, at last, is what Jev is actually asked. It is the heart of the whole plugin, and it lives in one file, judge.rs.

Jev cannot be asked "is this command dangerous?" Or rather it can, and it will answer, and the answer will be about the word "dangerous". A command that merely prints rm -rf ~/ as text was once stopped as if it ran it. So judge.rs never asks with an adjective. It writes down a short list of concrete things a command could do, defines each in plain acts, and asks Jev which one this is. Then it decides what to do with the answer itself, in code, where a test can hold it still.

A command: one request, three questions

Before a Bash command runs, the daemon sends Jev a state (the command, first 6,000 characters; the working directory; the project root) and three questions in one request:

  1. act, a Choice: which of these describes what the command does? If it does several, the one hardest to undo.

    LabelDefined as (abridged)Consequential?
    readonly reads or prints; nothing is different afterwardsno
    buildbuilds, tests, formats, lints, installs dependencies: regenerable files onlyno
    editcreates or changes files a person wrote, or records them in version control; recoverable by an ordinary commandno
    deleteremoves or overwrites data no build regenerates and version control does not holdyes
    historydiscards or rewrites version-control state: reset --hard, clean, rebase, force pushyes
    systemchanges the machine or the account: packages, services, dotfiles, sudoyes
    remotepublishes or sends something elsewhere: push, deploy, an API call that writesyes
    unreadruns code nobody has read: a script piped from the network into a shellyes
  2. undo, a Score of four levels: how hard would it be to put everything back? 0 nothing to put back; 1 one ordinary command; 2 only with care or luck; 3 not from this machine.

  3. load, a Score of four levels: how much of the machine does it take while it runs? 0 negligible; 1 light; 2 heavy (compiles a project, a whole test suite); 3 very heavy (several heavy jobs, a large release build).

All three end with WHAT_RUNS, a paragraph that says what "the command does" means: only what the shell would execute. Text carried as data (a here-document written to a file, a quoted string printed or searched for or sent) is not executed unless it is handed to something that runs it.

Jev answers each with a probability for every option (and for a Score, an expected level from 0 to 3). Then bash_verdict decides:

flowchart TD
  A["Jev's answers"] --> R{"one consequential act at 60% or more,<br/>or expected undo at 1.6 or more?"}
  R -- yes --> ASK["ask"]
  R -- no --> S{"memory available known,<br/>and below what this load needs?"}
  S -- yes --> ASK
  S -- no --> O{"read + build + edit at 90% or more,<br/>and expected undo at most 1.2?"}
  O -- yes --> ALLOW["allow"]
  O -- no --> PASS["pass:<br/>the usual permission check"]

What a load "needs" rises from 0 MB (negligible) through 500 MB and 3,000 MB to 6,000 MB (very heavy), interpolated between levels: an expected load of 2.5 wants 4,500 MB free.

Aside: one confident act, never a sum. An early version added up the consequential acts' probabilities, and a command Jev was unsure about (its likeliest act at 36%) was stopped. Take a spread like 36% unread, 20% system, 14% delete, 30% edit: summed, that is 70% "consequential". But it is Jev saying "I don't know", and an unsure answer should change nothing. Now only the single likeliest consequential act counts, and that answer is a pass. The test doubt_spread_over_several_consequential_acts_is_not_a_flag pins exactly that case.

A real answer, from the log: a long command that read some figures and called out to another program came back read 56%, system 39%, undo 0.4, load 1.3. No single consequential act reached 60%, nothing was hard to undo, memory was ample; but the ordinary acts came to 57%, short of 90%. So: Jev: only reads (56%); nothing to undo (0.4 of 3); left to the usual permission check. Not waved through, not flagged; Claude Code's own permission rules decided. 237 ms, 1,739 input tokens, $0.00007.

A turn's end: one question

When a turn is about to end, the state is the session's last five prompts (each first 3,000 characters) and the end of the assistant's final message (last 6,000 characters), and the one question is a Choice of endings: finished, waiting (needs something only the developer can give), blocked (names an obstacle), running (work still going, will report), and stopped-early (work asked for is left undone and no obstacle is named). Only at 85% or more for stopped-early is the end refused: a wrong refusal costs the user a wasted turn.

Three checks in serve.rs come first, and none of them asks Jev: a Stop that follows this plugin's own refusal (stop_hook_active) is let through, so one refusal can never become a loop; a turn with background tasks or scheduled wakeups still pending has paused, not ended; and with no prompt or final message there is nothing to judge.

Aside: the 1.5 seconds. Every request waits at most ANSWER_WITHIN, 1.5 s, and past that the verdict is pass. The request itself runs in its own task, so an answer that arrives late still settles its hold in the spend ledger; only the waiting is abandoned.

Try it. cargo test -p jevhooks-daemon. The verdict tests in judge.rs feed bash_verdict and stop_verdict made-up probabilities: ordinary_work_is_allowed, consequential_acts_are_put_to_the_user, an_unsure_answer_is_neither_allowed_nor_asked, a_heavy_command_is_put_to_the_user_when_memory_is_short, stop_blocks_only_when_confident. Change a threshold and watch which ones notice.

For the people who maintain it

The files, in the order a request meets them:

PathWhat
main.rsThe jevhooks command line (clap): serve, mcp, status, recent, stop.
paths.rsWhere things live: state_dir, socket, lock, daemon_log, decisions, and identity, the build's fingerprint.
mcp.rsThe per-session MCP server (rmcp, stdio): the status and recent tools, and the call to client::ensure.
client.rscall, one HTTP request on the socket; ensure, which makes sure a daemon of this build serves; spawn, the one place one is started.
serve.rsThe daemon: the lock, the socket, the routes, per-session state (the last five prompts, the cost, the judged tool calls), decide, observe, status. The model is pinned here (MODEL, jev-1.13.0).
judge.rsThe questions, every term defined, the thresholds, bash_verdict, stop_verdict, and their tests.
headroom.rsThe config file (Config), and where the available-memory figure comes from (Source: a configured command, MemAvailable, or none), measured at most every 5 s.
log.rsdecisions.jsonl: DecisionRecord, OutcomeRecord, append, and recent, which reads the last 512 KiB.

The thresholds, all in judge.rs and all reported by status:

ConstantValueMeaning
BASH_ASK_AT_CONSEQUENTIAL0.60One consequential act this likely: ask.
BASH_ASK_AT_UNDO1.60Expected undo level this high: ask.
BASH_ALLOW_AT_ORDINARY0.90Ordinary acts together this likely ...
BASH_ALLOW_UNDO_AT_MOST1.20... and undo at most this: allow.
LOAD_NEEDS_MB0, 500, 3000, 6000Memory each load level should find free.
STOP_BLOCK_AT_LEAST0.85stopped-early this likely: refuse the end.
ANSWER_WITHIN1.5 sLongest wait for Jev.

A decision record in decisions.jsonl has record: "decision", the session, kind (bash or stop), the subject (a command's first 500 characters, or a final message's last 500), the tool_use_id, the verdict and line, answers with every option's probability, and ms, jev_ms, input_tokens, usd, attempts and request_id. An outcome record has record: "outcome", the session, the tool_use_id and ran, failed or denied, and is written only for a tool call that was judged.

← Previous: Chapter 11, jevhooks-daemon/ · Up: jevhooks-daemon · Next: Chapter 13, third-party/ →