Which promises are enforced: Zetetic Agents, traced through its source
The claim on the tin is unusually testable: not a prompt library, a methodology with commit-time enforcement.1 A claim like that can be checked, and this dossier checks it, one guarantee at a time. The criterion is simple: a guarantee is enforced when a hook, a script or a test can return a non-zero exit that actually stops the action, and it is requested when it exists only as prose addressed to a model. Four are enforced and fail closed. Four more are enforced until someone changes a setting. One — the method the product is named after — is report-only until you opt in. And ninety-seven documented refusals are instruction, which the project states itself, twice.23
A hundred and twenty agents, and three different ways in
The roster is not meant to be browsed. You name the shape of your problem and the shape selects the pattern,3 and what ships for each one is a method rather than a persona: the citations, the canonical moves, the documented blind spots, and the conditions under which it must refuse.4 The distinguishing claim is narrower and more interesting than it sounds — that an agent here can say it does not know.2
Three paths reach that catalogue and they do not overlap. The host dispatches the twenty-three team agents by matching their description.5 A shell script assembles the ninety-seven reasoning patterns, and it is worth knowing exactly what it does: it prints the agent file. No model is called.6 Only the spawn script launches a process, and it is the one place where a guarantee has teeth.9
What actually blocks, and under what condition
Four guarantees hold without a way around them. A delegation cannot start without a validated contract.9 A definition with a live caller cannot be deleted, and the gate fails closed on “cannot determine” rather than open.21 A call that would surface secrets exits with a blocking code.20 And an acceptance gate passes only if its command exits zero, with an empty gate set counting as a rejection.22
Four more are enforced until someone changes a setting, and one of those is stated wrongly in the README. It says unsourced absolutes are blocked at any profile;17 under the permissive profile the checker counts the errors, prints them, and exits zero.16 The rule names themselves cannot be switched off — that attempt exits 2 — but the severity can.18 Then comes the one that matters most: the zetetic spine, the four-step method the product is named after, is report-only by default and blocks only if you opt in, and its detection is a heuristic that a single web search satisfies.19
The contract is the only thing standing before a mutation
When an agent is launched into a repository with permissions bypassed, the question is what stands between the launch and the working tree. The answer here is one file, validated before anything is touched: no contract, no contract file, or no Python interpreter each refuse the launch outright.9 The schema requires owned paths, excluded paths, a push authority with no default, handback artefacts, and an acceptance oracle it defines as the external signal that decides completion, never a self-report.10
Two details raise this above a checklist. The lock re-runs the whole validation at registration time, explicitly to cover the race between the first check and the launch,12 and two delegations whose paths overlap are denied rather than queued, which the schema states plainly and admits is not serialisation.11 The agent's identity is set by the spawn site, on the stated grounds that a subagent cannot be allowed to forge its own.13
What is measured, and what drifted past the gate
The measurement work here is real and unusually disciplined. A paired benchmark compares an inline skill against a spawned subagent with the protocol frozen before any run existed, five repetitions, randomised order, two blind graders, raw and scored data committed.28 Its result is not the flattering one: spawning cost more on every run, and was slower where the difference reached significance.29 The routing scorecard goes further and publishes a ceiling: two annotators agree on the shape only 39 % of the time, so every routing score is reported against 39 rather than against 100.3031 A structural auditor passes over all 120 agent files, and a count check verifies thirty-seven documented claims against the tree.327
Then there is the sentence every user reads. The session banner prints four numbers, and three are wrong at this revision: 63 skills where there are 81, 14 hooks where there are 23 registrations, 17 tools where there are 57.26 The command index claims seventeen commands where the tree holds twenty-seven.27 The counting document contradicts itself four paragraphs apart, and disagrees with its own prescribed command about the number of tools.8 None of those files is in the registry the drift check reads.7 The gate works; it was simply given a list of five files, and the most visible sentence in the product is not on it.
Why this is the right thing to publish
A project whose first rule is “no source, say I don't know” invites exactly this treatment, and mostly survives it. The enforcement that matters — contract, deletion, secrets, acceptance — holds and fails closed. The benchmark reports against its own product's interest. The routing score is published against a measured ceiling instead of a round number. What does not hold is narrower and fixable: one README line that contradicts the checker,17 a documented context budget resting on a hook that was deliberately removed,34 and four self-descriptions that drifted because nobody added them to a list.26278 Naming them is not a reservation about the tool. It is the tool's own standard, applied to the tool.