# ratchet changelog

Releases are immutable: a published version directory is never rewritten.
`latest.json` always points at the newest cut.

## v0.6.12 — 2026-09-15 (atlas v0.1.11)

- **atlas v0.1.11 — packaging only; the code is byte-identical to v0.1.10.** Release tarballs are now built with `tar --no-xattrs`. `COPYFILE_DISABLE=1` never stopped libarchive from recording extended attributes, and macOS stamps every copied file with a per-process `com.apple.provenance` value, so identical source packaged from a different session produced a different checksum. Re-packaging atlas 0.1.10 was refused by the release gate for exactly that (18 differing bytes, all inside the provenance record). 0.1.10 stays published and unchanged; 0.1.11 exists so the fixed packaging has a version to live under. Nothing to do on your side beyond the usual `ratchet update`.

- **The SessionStart packet names each section's true total.** `(5 of 15)` under the working set meant "5 of the 15 the brief kept", not "5 of the 97 this project has": the packet took the length of a list `brief` had already capped (15 / 15 / 8 / 6). `brief` now carries `working_set_total`, `paging_files_total`, `traps_total` and `slow_commands_total` as fields (trap and slow totals are null on an unswept store, never 0), and the packet reads them. The guard that shipped green built its Brief by hand and never ran the cap; its replacement goes store → gather → packet.
- **MCP `ratchet_query_hotspots` defaults to `order: occurrences`, like the CLI.** The handler still defaulted to seconds, which ranks WAIT ahead of every untimed shape — REPEAT and PAGING (ranks 1 and 2 by count) fell out of any limited view on the surface agents actually call.
- **MCP `kind` is validated on hotspots, signatures and trends.** An unknown kind returned `rows: []` with no caveat — "no friction of that kind" — while the CLI refused it. It is now a tool error naming all eight kinds.
- **`doctor` / `ratchet_status` wiring faults name `ratchet wire`.** `MCP_UNREGISTERED`, `MCP_STALE_PATH` and both atlas twins said `Fix: ratchet install`, a subcommand that has never existed. A test now runs every verb a caveat names against the binary.
- **MCP `ratchet_findings` refuses a negative `limit`.** `-1` was cast to an unsigned size, which returned every row with no TRUNCATED caveat, while the CLI's `--limit` rejects negatives.
- `--kind` help on `trends`, `query hotspots|signatures|trends` lists TRAIN; `ratchet_context`'s description names BYPASS and TRAIN among the null-seconds kinds; RATCHET.md shows `--limit` on `query who|sessions|intents`.

## v0.6.11 — 2026-09-08 (atlas v0.1.10)

- **atlas 0.1.10: `--json` on the CLI lookup.** `atlas <word> --json` and `atlas lookup <word> --json` emit the MCP tool's exact record (meta with notices, hits, exact_indexes, neighbours_only, qualifier, live, footer); `--no-build` travels through so a hook fails open in milliseconds on a tree with no manifest, the state named in `meta.notices`. Fleet Interop clause 1, asked by praxis for the Grep→atlas reach hook (#ratchet 673).
- **`ratchet query skills` + `ratchet_query_skills` (15th tool).** The USED axis for praxis's skills plans: Skill invocations per (skill_name, project) over the last 30 days from tool_use rows; scope is null on purpose (SCOPE_NOT_CAPTURED), absent = unmeasured, never 0; SKILL_NAME_ABSENT counts uncaptured names. Second witness to the vendor's /stats, which reported 0× for skills the store had seen 5, 3 and 2 times.

## atlas v0.1.9 — 2026-09-05 (ratchet unchanged at v0.6.10)

- **atlas 0.1.9: a declared symbol answers with its callers.** `atlas <word>` / `atlas_lookup` now run the one live whole-word search on every lookup, so a declared symbol returns its definition AND its references (declaration sites not counted) in one call; the footer lists code bodies as searched. Before, the live pass ran only when nothing declared the term — who-calls on a declared symbol got the definition and none of the callers, and the footer pointed at a `refs` verb the MCP tool does not expose (STNG field report, 2026-09-05: fifteen grep batches to one atlas call, and that was why).

## v0.6.10 — 2026-09-05 (atlas unchanged at v0.1.8)

- **`doctor` faults `SCHEDULE_LAST_EXIT`.** It reads the scheduler's last exit code (launchd `last exit code`, via `launchctl print`) and faults on non-zero, naming the code and the fix; `None` stays not-measured, never 0. On 2026-09-05 a launchd job that exited 78 EX_CONFIG on every run for two hours read healthy because only "loaded" was asked. Also the `run` summary now says `findings N promoted (M new; T in tier)` — 120 beside a tier of 182 had read as a discrepancy.

- **`ratchet update` / `wire` re-bootstrap the scheduler when the binary's bytes change.** `schedule` records the bootstrapped binary's sha256 in `~/.ratchet/schedule.sha256` and re-runs bootout+bootstrap when the installed bytes differ (absent sidecar = one harmless re-bootstrap). Incident 2026-09-05: the 0.6.9 in-place swap left launchd refusing the binary (78 EX_CONFIG, no output) while `wire` reported loaded — loaded is not runnable.

- **`ratchet update` hands the wire step to the NEW binary.** After a successful swap it runs `<installed binary> wire` instead of wiring in-process, so the re-bootstrap logic that runs is the just-installed version's, not the one being replaced. Consequence: the 0.6.9→next update is still wired by 0.6.9 — follow it once with `ratchet wire`; every update after that is self-healing.

## v0.6.9 — 2026-09-05 (atlas unchanged at v0.1.8)

- **MCP tool targets: a second probe tier.** For `mcp__*` tools only, `action` / `channel` / `cwd` / `intent_id` become `tool_target` when no path key applies (measured: the four keys cover every high-volume target-less MCP tool; no native tool carries them). Existing claude-code rows are filled once at open by `backfill_mcp_target_v1` (a rescan cannot reach rows that already produced an action — measured: 0 filled, 6,891 still NULL).

- **`ratchet normalize --rescan`.** Resets both normalize cursors and re-examines every raw record (idempotent). Retires the hand-SQL recipe — a DELETE against the store was the one documented step that looked exactly like the BYPASS the meter counts.

- **Two absence-as-zero sites fixed.** An unreadable install dir says so above the uninstall prompt instead of "0 files"; a missing scheduled binary makes the stale-session census say "unavailable" instead of "no stale sessions".

- **doctor escalates NEW_BUT_RAN.** A store with zero records after two or more ingest runs is a fault with the fix named; one run or none stays NEW and exit 0. New meta keys last_ingest_at / last_ingest_files_seen / ingest_runs, absent = never ran, never 0.

- **`--json` on `query <verb>` is a no-op, never a refusal.** Measured: a grok
  executor that learned `--json` from `trends` / `hotspots` / `grade` / `findings`
  lost a turn on `query hotspots --json` (2026-09-01, intent-78). Query output is
  always JSON; the flag is accepted for symmetry (`global = true` on the parent).
  Nothing breaks — the "wrong" spelling now routes.

- **BYPASS: a new count-only kind for the times an agent hand-rolls SQL against a substrate
  a registered tool already serves.** Asked by the `arc-mba` seat (#ratchet 613), which had
  measured 154 `sqlite3` calls on `.arc/state.db` against 39 query-shaped MCP calls on its
  own machine. Its premise that ratchet stores a digest but not the command was refuted
  first: `action.tool_target` carries the full Bash command on **40,794 of 40,794** rows
  here (avg 540 chars, max 50,470), so this is a sweep-time detector over an existing
  column and needs no new capture. Like-for-like the ratio is **worse** than reported: in
  project `arc`, 9 query-shaped MCP calls against 108 hand-rolled `sqlite3` — 12:1, not
  4:1. A first pass compared ALL `mcp__arc__*` (2,481, of which 2,250 are declare/merge/sync
  lifecycle calls) and would have published a false inversion. The corrected split says the
  bypass is specific to the QUERY surface: an intent merge cannot be hand-rolled.
- **The discriminator is the finding, so it survives ranking.** Of the 22 sessions here that
  hand-rolled against a served store, **18 had also called that tool's own surface in the
  same session** — a coverage gap, not a reach gap. Two shapes carry it, and
  `hotspots::grouping_for` threads the signature's own `evidence.shape` into the fourth slot
  of the existing `(kind, tool_name, verb)` grouping so hotspots AND trends rank by it
  through one helper. A trend cell and a hotspot row still cannot name the same friction
  differently. No new grouping key, no schema change.
- **A second arm was built and killed by its own first measurement.** It counted a tool's
  `--json` piped into `python3`/`jq`. On this corpus it fired 258 times, and inspection
  showed most hits were `--json | jq`, which is the INTENDED way to consume a machine
  surface, plus compound commands where `--json` and `| python3` sat in unrelated segments;
  only 14 were the reshaping heredoc it was aimed at. Firing on the known-good path is the
  false-alarm failure this project has already ruled on, so the arm is gone and the negative
  is pinned by a test (`a_json_pipe_is_not_a_bypass`).
- **First measured ranking (2026-09-04, this machine, sweep-001625):** 522 signatures —
  285 "the session also used that tool's surface", 237 "never used". By project: ratchet
  219, arc 109, mcp-server 92, kyber 33, wire 23. Note the semantics, because it is not the
  zero-check the ask assumed: a project with no `.arc` can still score, because the kind is
  about hand-rolling against a SERVED STORE (ratchet's or kyber's included), not about the
  project's own substrate.
- `query::KINDS` is 7; `--kind BYPASS` is accepted on `query hotspots`, `query signatures`
  and `trends`; `rule_version_for("BYPASS")` is `0.6.9`. Counts are a lower bound: SQL run
  through a python `sqlite3` import, a non-Bash shell tool, or a store reached by a path the
  const does not name are all missed.
- **`target_of` no longer treats every `name` key as a target.** Intent 84 put
  `skill` then `name` on the generic probe so Skill calls store the skill; the
  disclosed side effect was that ANY tool whose input carries `name` stored it
  as `tool_target` — 1 Workflow row on this store (`input.name = deep-research`).
  The Skill arm is now gated on `tool_name == "Skill"` (claude-code and rift);
  a non-Skill tool with only a `name` key stores NULL. The one-shot
  `backfill_skill_target` was already Skill-only and is unchanged. The leaked
  Workflow row is not repaired: it feeds no detector.

- **Findings JSON is bounded by default (council #wire 1463 C5).**
  `ratchet_findings` returned **119,000 characters** in one MCP result —
  unreadable, and the caller re-aggregated by hand, so the machine surface
  was less usable than the CLI piped to python. `findings --json` and MCP
  `ratchet_findings` now return the top 50 rows (highest `evidence_n`
  first) with `meta.floor = {limit, returned, total}` and a TRUNCATED
  caveat when cut; `--limit N` / `limit` lifts the cap; `--all` / `all:true`
  returns every row. `meta.by_kind` and `meta.contested` are computed from
  ALL rows regardless of the cut, so an aggregate never needs the full set.
  **Breaking for callers who parsed the unbounded list** — they must pass
  `--all` / `all:true`.

- **BREAKING: a missing store is a refusal, not a create.** Measured: `RATCHET_DB=/tmp/zz.db ratchet query hotspots` refused `no_sweep` (exit 1) but first minted a 4,096-byte empty store at that path, so a typo'd `RATCHET_DB` yielded a second store the next `run` filled while both read healthy. Read-only verbs now use `db::open_existing` and refuse with JSON `{"error":"no_store","detail":"...run ratchet init..."}` (exit 3). Only `init`, `run`, and `ingest` may create. `mcp-server` still STARTS with no store (a start-time refusal would crash-loop Claude Code) and every `tools/call` answers the same envelope until the store exists — re-checked per call, so `ratchet init` later recovers the same server without a restart. Scripts that relied on a read-only verb to create the store will now see exit 3 instead.
- **Atlas manifest counts are Option.** Unparseable manifest meta renders as unreadable (`?` counts), never as "0 files, 0 symbols"; an existing manifest that fails to open is reported as unreadable in the orient packet, `ratchet status` and `ratchet_orient`, instead of as no manifest. `BRIEF_MAP_UNREADABLE` and `ORIENT_MAP_ABSENT(UNREADABLE_MANIFEST)` name the case.

- **Backfill: `tool_target = ''` → NULL, once per store.** 7,104 pre-August rows on the reference machine carried '' where the normalizer now writes NULL; '' grouped every target-less call as one identical target. Only column with '' rows in any table (censused 2026-09-05).

- **TRAIN: the eighth kind.** Three or more consecutive same-tool same-stem calls in one session (Bash verb or MCP tool) — one call per turn that had one shape; the fix is a loop, a script, or one batched call. Count-only, one signature per (session, tool, stem), shape identical / varying / target_not_captured threaded into ranking. Read/Edit trains are excluded (PAGING/REPEAT own them). Measured 2026-09-05: 3,404 trains / 15,191 calls, 16.8% of tool use; only 2% of Bash trains repeat an identical command. Spec: docs/plans/2026-09-05-train-detector.md.

## v0.6.8 — 2026-09-02 (atlas v0.1.8)

- **atlas 0.1.8 — the Linux port's change finally carries a version.** `manifest_dir` became
  platform-aware on 2026-08-25 (intent 67: `~/Library/Caches/ratchet-map` on macOS,
  `~/.cache/ratchet-map` elsewhere; `XDG_CACHE_HOME` deliberately not consulted so fixture
  homes cannot clobber the live cache) but shipped in tree under `__version__ = "0.1.7"`.
  The release guard caught it as "0.1.7 is already published with different bytes" and
  refused — which is the guard doing its job, not an obstacle. Bumped, no force.

**Nine intents in one cut: the 2026-09-01 review fan-out (five), three verbs asked for on the wire the same night, the Skill-name fix, and a price retraction — every one merged behind a full suite on the integrated base, executed by rift/grok-4.6-build under arc intents 74–86. Breaking for machine consumers: findings document v3 (v2 refused), `delta_gate` null when no blanket gate applies, four new MCP tools (10 → 14).**

- **Retracted the future-dated Claude Sonnet 5 $3/$15 row.** Anthropic's
  pricing page (platform.claude.com/docs/en/about-claude/pricing, read
  2026-09-02): "The $2/$10 pricing for Claude Sonnet 5 … is now the standard
  price. The previously scheduled increase to $3/$15 on September 1, 2026 will
  not occur." Removed the `effective_from` 2026-09-01 row from
  `prices/2026-07-31.json`. Loaded price tables are immutable (`price::load`
  refuses conflicting re-loads), so a JSON edit alone would leave the row in
  every existing store: a one-shot migration (`retract_sonnet5_future_row_v1`)
  deletes it from `price_table` and NULLs `cost_usd`/`price_resolution` on
  `claude-sonnet-5` assistant actions dated ≥ 2026-09-01T00:00:00Z so the next
  `ratchet price` re-prices them at $2/$10. Measured on this machine's live
  store: 388 assistant actions across 11 sessions and 3 projects, $30.81
  under the retracted row, should be ~$20.54.

- **`ratchet query intents [--project P] [--since YYYY-MM-DD]`** — cost per
  completed task, not per call (commerce-agents post A27; gap analysis
  action 2). The task is the arc intent already stamped on
  `action.intent_id`. JSON-only (like every query verb); MCP
  `ratchet_query_intents` (13 → 14 tools). One row per `(project, intent_id)`
  plus a per-project outside-any-intent row so the share of work with no
  declared intent is visible, never hidden. Token sums and `cost_usd` are
  NULL when never measured, never 0. `INTENT_COVERAGE` names both counts
  (actions carrying an intent_id / all actions in scope) so a percentage
  cannot hide the uncovered share. `UNPRICED_PRESENT` fires when any row
  has unpriced turns (`cost_usd` on those rows is a floor, not a total).

- **`ratchet query sessions [--project P] [--since YYYY-MM-DD]`** — per-session
  cost envelope asked by kyber (#ratchet 569) after a consumer joined the
  store by raw SQL over `action`/`raw_record`/`source_file`. JSON-only (like
  every query verb); MCP `ratchet_query_sessions` (12 → 13 tools). One row
  per `session_id`: token sums, cost, unpriced turns, cache hit ratio. Token
  sums and `cost_usd` are NULL when never measured, never 0. `UNPRICED_PRESENT`
  fires when any session has unpriced turns (`cost_usd` on those rows is a
  floor, not a total).

- **`ratchet query who --path P --at T [--window 5m]`** — which sessions/commands
  touched or named a path in a window on this box. Asked by arc-mba (#wire 1565)
  after a forensic hunt that nine envelopes chased a write one transcript grep
  answered. JSON-only (like every query verb); MCP `ratchet_query_who`
  (11 → 12 tools). Two bounds travel on every envelope: `STORE_LAG` (ingest is
  30-minutely — forensics, never now; the caveat names the newest ingested
  `action.ts`) and `WHO_LOWER_BOUND` (only what the tool was handed; a path a
  script wrote without naming it is invisible). Empty is empty-with-caveat
  (`WHO_NO_MATCH`), never a bare empty.

- **`claude-fable-5-1` was unpriced.** Doctor reported UNPRICED on **206 assistant
  rows**, all of them `claude-fable-5-1`, first seen 2026-09-01T19:20Z —
  `price_resolution=unknown_model`, `cost_usd` NULL (ruling 2 held: unknown is
  never 0). Added the exact-id row to `prices/2026-07-31.json` at $10 in / $50
  out / $0.25 cache read / $12.50 cache write, `effective_from` 2026-06-24. No
  prefix matching, no fallback onto `claude-fable-5`. The table is `include_str!`
  at build (`src/main.rs:501`), so installed binaries see the row only after the
  next cut.

- **Bash verb extractor: block closers and branch keywords are no longer
  attributed as work.** Sibling of the 0.6.1 opener-keyword fix: KEYWORDS
  held `for`/`while`/`if`/`do`/`then`/`{` and not the closers, so the latest
  sweep's WAIT rows ranked `Bash done` 14, `Bash else` 8, `Bash break` 6
  (findings from 2026-08-20 also fossilised `Bash exit` / `unset` /
  `continue` / `fi`). Closers (`done` `fi` `else` `elif` `esac` `!`) skip
  the token and keep walking the segment (`else echo FAIL` is echo);
  control-flow (`break` `continue` `return` `exit`) and `unset` skip the
  whole segment like `cd`. Python HEREDOC bodies
  (`python3 - <<'PY' … PY`) are dropped from the original target before
  flattening, so `r = json.loads(…)` is not walked as shell. Measured on
  a store copy, 1,871 distinct Bash WAIT targets of sweep-001494: done
  13→0, else 8→0, break 5→0, continue 1→0, exit 1→0, unset 1→0, fi 1→0
  (return was already 0). Left standing with a sample: `r` 9 (`$R` folds
  to `r` under the 0.2.1 variable rule), `anonymous` 3 (quote-blind `|`
  inside a grep -E character class). Shapes regroup on the next sweep;
  fossilised findings supersede naturally.

- **Background-task spans are now a measurement, not an absence wearing a
  plausible small value.** A `run_in_background` Bash / Agent / SendMessage
  launch was priced at its spawn gap (1–6 s) while the work lived in the
  task. Completion arrives later as a `<task-notification>` whose
  `<tool-use-id>` is an exact key to the launching `tool_use`. Two delivery
  shapes: idle-session `user` (already an action) and busy-session
  `queue-operation` (never an action — 9,278 rows / 0 actions on the
  reference store). New table `background_span`, populated by a dedicated
  normalize pass with its own cursor (the action cursor's NOT EXISTS guard
  skips already-actioned `user` notifications and advances past
  queue-operation forever). `span_ms` is launch → first notification wall
  clock, a **ceiling** of the same class as `derived` — never written into
  `action.duration_ms`, never a `duration_source` WAIT could sum. Coverage
  travels inside the value (recovered / launches, both counts + pct).
  Surfaces: `ratchet query background` (JSON-only, like every query verb), a Background section on
  `ratchet report`, a `background` key on `ratchet_context` that is
  **ABSENT** (caveat `BACKGROUND_NOT_MEASURED`) until a notification has
  been seen so a fresh install cannot read as "no background work", and MCP
  tool `ratchet_query_background` (10 → 11 tools). `queue-operation` stays
  **out** of `MODELLED_KIND_SQL` — putting it in scope with 0 actions is
  `UNMEASURED_HARNESS` forever; `meta.counts` remains the action-production
  accounting. Spec: `docs/plans/2026-09-01-background-spans.md`. The finding
  is "the spawn-gap duration is WRONG", never "these hours are waste".

**A newly-visible harness no longer silences every trend on the machine, and a count
delta now travels with the rate it was computed from.** Measured 2026-09-01:
`ratchet trends --json` returned 111 deltas, all `comparable: false`, under one
machine-global `HARNESS_ABSENT_IN_PRIOR` ("rift has 6098 in-scope records in
2026-08 but none in 2026-07"). Claude-code-only cells — REPEAT 1,103 → 133,
PAGING 770 → 199 — were refused along with the rift ones until October. Underneath
the gate, claude-code in-scope fell 152,182 → 87,265 (−43%) while REPEAT fell 88%,
and no delta was normalised by volume: when the gate lifted, a month with half
the work would have read as a fix that held.

- **Per-cell harness gate.** Each (kind, tool, shape) delta is gated against the
  coverage rows of the harnesses that contribute signatures to it. Same reason
  codes (`HARNESS_ABSENT_IN_PRIOR`, `PRIOR_CAPTURE_SPARSE`, `COVERAGE_SHIFT`)
  live on each delta's `reason`. The machine-global fact is now the
  `HARNESS_NEWLY_VISIBLE` info caveat ("rift is newly visible in 2026-08; cells
  it contributes to are gated individually"). Top-level `delta_gate` stays as
  the remaining blanket (`NO_COMPLETE_PERIOD` / `NO_PRIOR_DATA`) and is **null
  when no blanket applies**. **Breaking for callers that parsed a present
  `delta_gate` as the only reason deltas were absent** — `ratchet_context`'s
  `trend_gate_or_deltas` is the XOR, and a null gate is the deltas arm, never
  "no trends".
- **Rate beside the count.** Every cell carries `per_1k_in_scope`
  (`occurrences * 1000 / in_scope` of that period, restricted to the cell's
  contributing harnesses); every delta carries `current_per_1k_in_scope` /
  `prior_per_1k_in_scope`, `pct_change_per_1k` beside `pct_change`, and
  `volume_change_pct`. NULL when that in-scope is 0 — never divide by zero,
  never 0.0. `VOLUME_SHIFT` (warning) fires when any delta has
  `|volume_change_pct| ≥ 25`, names the cells, and points at `pct_change_per_1k`.
- Coverage array stays machine-global; `--project` still scopes cells only.

- **Findings decay, coverage counts, and a detector-rule stamp — and they
  travel.** Measured 2026-09-01: **69 of 177 findings matched no hotspot in
  the latest sweep**, and **22 carried `confidence` NULL** (promoted before
  0.6.4 and never restamped). The portable tier had no decay, no
  denominator beside `coverage_pct`, and no way to tell a rule change from
  a behaviour change. Five nullable columns on `finding`
  (`last_promoted_at`, `dormant_since`, `coverage_measured`,
  `coverage_in_scope`, `detector_rule`): promote stamps the counts from
  `hotspots::corpus_counts` (never re-derived) and `detector_rule` from
  `sweep::rule_version_for`, and marks locally-promoted rows whose shape
  is absent from the sweep **dormant rather than deleting them** (contest
  verdicts survive). The export document carries `dormant` as a bool —
  never a timestamp — plus the two counts and the rule stamp. Import
  stores them; a dormant row's `dormant_since` is the import time (when
  *this* machine learned the origin saw it dormant — local knowledge, so a
  local timestamp is honest). NULL means not-measured. **Breaking for
  callers parsing the findings document: v3, v2 refused** — a v2 row
  lacks the fields and accepting one would launder absent-as-present.

- **Skill `tool_use` stored an empty `tool_target`.** 488 nameless Skill
  rows (claude-code 460, rift 28) even though the raw record carries the
  name: claude-code `input.skill`, rift `args.name`. `target_of` now
  probes those keys (NULL when absent, never `''`); a one-shot
  `backfill_skill_target` (meta-gated, from `raw_record.body`; pruned
  bodies stay NULL) repairs existing rows. That is what makes the
  "≥1/3 of sessions → hoist to CLAUDE.md, else leave as a skill" rule
  (commerce-agents post A3, council #wire 1463 X3) a query.

## v0.6.7 — 2026-08-17 (atlas unchanged at v0.1.7)

**0.6.6 taught REPHRASE to see how people search. It did not teach it to see the tools
that REPLACE searching — and the omission was deleting evidence, not just missing it.**
Caught by `agent:wire-mba` hours after 0.6.6 shipped.

- **`atlas` was not counted as a search.** Measured on the reference corpus:

  | atlas invoked via | calls |
  |---|---|
  | Bash (`atlas <word>`) | **166** |
  | MCP (`atlas_lookup`) | 9 |

  The review predicted this as an MCP gap; it is overwhelmingly a **shell** gap, in the
  code path 0.6.6 had just rewritten.

- **It suppressed signatures rather than merely missing them.** REPHRASE is
  adjacency-based, so a non-search **breaks the run**. An `atlas` call between two greps
  split one hunt into two sub-`MIN_RUN` fragments and the signature vanished. Atlas usage
  was actively erasing evidence of the refine loops around it — and, worse for the
  question everyone actually wants answered, it made the detector **structurally unable to
  observe an atlas-for-grep substitution**, which is precisely the by-construction
  blindness 0.6.6 existed to remove.

- **Fixed:** `atlas` added to the search binaries, plus a `SEARCH_TOOLS` list carrying
  `Grep`, `Glob`, `grep_search`, `mcp__atlas__atlas_lookup`, `mcp__arc__arc_search`.

- **`ToolSearch` and `WebSearch` are excluded, and pinned excluded by test.** They search
  tool schemas and the internet, not this tree. `ToolSearch` alone was **728 calls at cut
  time** on the reference corpus and would have swamped the metric — the same inflation
  the pipeline-head rule exists to prevent. (A live-corpus count only grows; re-derive
  before quoting it onward rather than carrying this figure.)

- **Corpus delta: 884 → 904 (+2.3%), and the count is not the point.** The value is that
  REPHRASE can now *see* an atlas-for-grep substitution at all.

**The generalisation, and it is the reusable part:** when you widen a detector's reach,
ask what **SUBSTITUTES** for the thing you just started counting — not only what spells it
differently. The 0.6.6 sibling sweep enumerated the *codebase* and found one harmless hit;
the real sibling was conceptual, a different tool doing the same job.

**`MIN_RUN = 3` and the accumulation surface are INDEPENDENT defects** (also wire-mba's
framing, adopted): fixing the surface without lowering the floor still misses a two-call
refine loop; lowering the floor without the surface still sees a fraction of it. Recorded
as two rows so neither is closed by the other's fix. **The floor is unchanged** —
deliberately, so the 884 → 904 delta stays attributable to one cause.

**Known weakness, stated rather than discovered later:** `SEARCH_TOOLS` is a hardcoded
list, so the next harness with a native search tool is invisible until somebody notices —
which is the defect being fixed, one layer up. A principled version would classify by tool
semantics rather than by name; no mechanism for that exists here yet.

## v0.6.6 — 2026-08-17 (atlas unchanged at v0.1.7)

**A detector that cannot fire on the path people use reports zero, and zero reads as
health.** Two fixes and a measurement, all one law: check what the instrument can
actually see before believing what it says.

### REPHRASE was watching ~2% of the search surface. 31 → 884.

`sweep/rephrase.rs` accumulated runs of the **`Grep` and `Glob` tools only**. On the
reference corpus that is **181 `Grep` + 67 `Glob` calls against 11,562 Bash `grep`/`rg`
invocations** — so the detector observed about **2%** of searching and reported **31**
signatures all-time. That number reads as *"this agent barely refines its searches"*,
a conclusion the instrument was structurally incapable of supporting.

- **Searches are now classified on the PIPELINE HEAD.** A search binary at the head of
  any command segment hunts the filesystem; the same binary downstream of a `|` filters a
  stream. `rg foo src/`, `grep -rn x .`, `cd web && rg X` and `find . -name '*.rs'`
  count. `ps -ax | grep node`, `cat f | grep x` and `cargo test | grep FAILED` do not.

  **This discriminator is the point.** Counting every command containing "grep" would
  have been worse than the bug — it folds ordinary shell plumbing into a friction metric
  and produces an alarm people learn to skip.

- **Measured delta, both rules over the same store: 31 → 884, a 28.5× undercount.**

- **Counts are a LOWER BOUND, said out loud.** `cat list | xargs rg x`, searches spelled
  through a script or alias, and `awk`/`sed` pattern work are not counted. The bias is
  toward undercounting, which is the correct direction for a detector whose failure mode
  is crying wolf.

- **The composition label was one arm from repeating its own recorded defect.** It
  hardcoded three outcomes with `Grep` as the fallback `else`; once Bash entered a run,
  every Bash search episode would have been filed as `Grep`-tool friction — hiding the
  exact path this change exists to reveal. The comment directly above it describes that
  same defect being committed once for `Glob`. The label is now built from what the run
  actually contained.

**Compatibility, stated because the numbers move:** REPHRASE counts from before 0.6.6 are
**not comparable** with later ones — the definition changed, not the behaviour. **Trends
are unaffected**: `trends` is single-sweep and every sweep is a full re-detection over all
history, so all periods inside one sweep share one rule. A comparability gate was
considered and deliberately **not built**, for a hazard that does not exist.

**Known gap, disclosed rather than fixed:** findings exchanged between a 0.6.5 and a
0.6.6 machine **understate REPHRASE** on the older side, and nothing in the document says
so. The fix is a detector-rule stamp per finding. A blanket `EXPORT_DOC_VERSION` bump was
rejected as disproportionate — five of six kinds are unchanged, and refusing every v2
document over one kind costs more than it saves.

> **The undercount is CORPUS-DEPENDENT and must not be used as a correction factor.**
> Measured on two machines within hours of each other: **28.5×** here (31 → 884) and
> **11.4×** on the MBA (52 → 594, consecutive sweeps with 34 records of corpus growth
> between them — an unusually clean control). **2.5× apart.**
>
> The factor is the ratio of shell-invoked to tool-invoked searching in that machine's
> work, which is a **fingerprint of how the agent was driven, not a property of ratchet**.
> Any single multiplier is therefore wrong somewhere else in the fleet. **Do not
> back-correct old numbers — re-sweep.**
>
> This *strengthens* the case for the per-finding rule stamp rather than weakening it: a
> blind spot that varies 2.5× by machine cannot be corrected for at read time at all, so
> a version-scoped disclosure can never close the gap. (Second-corpus measurement and
> this correction: `agent:orc`, MBA.)

### A percentage now ships with the denominator it was computed from

`meta` carried bare `in_scope_pct`/`coverage_pct` while the **rendered coverage table has
printed `captured / in_scope / measured` all along** — the machine surface was less
auditable than the human one, backwards for a tool whose consumer is an agent.

- **`meta.counts` = `{captured, in_scope, measured}`**, derived from a single
  `hotspots::corpus_counts` that both percentages now read. A second derivation is exactly
  what fabricated `confidence` in v0.6.4.
- **Third instance of one law:** when the misreading happens on the **healthy** path, no
  caveat can fire and the fix is a **field beside the number** — after
  `stated_occurrences`/`derived_occurrences` beside `confidence` (0.6.4) and
  `dominant_excluded_shape` beside `in_scope_pct` (0.6.5).

### ratchet now states what it does NOT measure

Raised by the rift seat, verified against `src/sweep/*` before adoption, and now carried
in the MCP `INSTRUCTIONS`, `llms.txt` and `RATCHET.md`:

> **ratchet measures time lost, never coverage of the work.** Every detector reads the
> shape of tool use and none inspects what was produced. A session that did nothing and a
> session that did everything correctly both emit zero signatures. A clean sweep licenses
> exactly one claim: **this build found no friction it can model in what it could read** —
> not completeness, and not even low friction, since a missing signature may be
> not-measured rather than measured-and-none. **Never a quality gate or a done-check.**

`coverage_pct: 100.0` in particular means normalization succeeded on what this build can
model. It says nothing about whether anything was missed.

### atlas v0.1.7 — documentation only, no artifact change

`ATLAS.md` gained a measured cost section, from a grok session on the MBA: **atlas is
slower per lookup** (16 `rg` ≈ 175 ms vs 16 atlas ≈ 3 s) and **costs more tokens than an
empty `rg` when the answer is "nothing"** — deliberately, since an empty result with a
`searched`/`not_modelled` footer is evidence where an empty grep is a shrug. It pays by
deleting the refine loop, not by being faster. The magnitudes (~15–25 k tokens, 1–3 min)
are marked as **not portable** — grok, batching, one tree, one machine — and any
six-figure token-saving figure for atlas is unsupported by anything measured.

## v0.6.5 — 2026-08-14 (atlas unchanged at v0.1.7)

**A coverage row now names what is deflating it.** `ratchet query coverage` reported
grok at **2.2% in-scope** and rift at **3.3%**, and four separate readers — two of them
agents on two machines — took those numbers to mean ratchet was nearly blind to those
harnesses. It is not. The denominator counts every captured record, including streams
that are structurally incapable of being in scope: grok emits 167,335 `phase_changed`
rows, rift 44,568 `assistant_text` rows. Neither is a tool call, so neither can ever
enter the numerator.

- **`CoverageRow` gained `dominant_excluded_shape` and `dominant_excluded_count`** — the
  single largest record shape the in-scope predicate rejected, per harness. The reference
  corpus at cut time (these grow; re-derive with `ratchet query coverage` before quoting
  them onward):

  | harness | in-scope | largest out-of-scope shape |
  |---|---|---|
  | claude-code | 74.6% | `attachment` ×10,266 |
  | antigravity | 89.1% | `CONVERSATION_HISTORY` ×52 |
  | codex | 44.3% | `response_item` ×58 |
  | rift | 3.3% | `assistant_text` ×44,568 |
  | grok | 2.2% | `phase_changed` ×167,335 |
  | gemini | 0.0% | `(unclassified)` ×8 |

  `(unclassified)` is the value that still means *blind* — gemini is captured raw with no
  normalizer, and the row says so instead of hiding inside the same 0% that a
  well-understood exclusion would produce.

- **It is derived from the NEGATED in-scope predicate**, sharing the one SQL constant the
  numerator uses. A second hand-written "what counts as out of scope" clause is exactly
  how a disclosure drifts away from the number it is disclosing — the same second
  derivation that fabricated the confidence field in v0.6.4.

- **It is a FIELD, not a caveat.** Four readers took the percentage; none took the
  surrounding prose. Anything a reader must read to avoid a wrong conclusion belongs in
  the row.

**Three bugs found in the disclosure query itself, none by its author:**

1. **SQL three-valued logic dropped the very rows being counted.** `json_extract(...) IN
   (...)` is NULL — not false — when the key is absent, so `NOT (...)` was NULL too and
   the `WHERE` excluded exactly the untyped records that dominate the denominator. Caught
   by verifying against the corpus, *not* by a test: the fixture used explicit types and
   never reached the NULL path.
2. **`json_extract` raises on a pruned record's empty body.** Caught by a pre-existing
   test whose stated job is guarding that short-circuit — the guard held.
3. **`COALESCE` then defeated that short-circuit**, so `json_valid(r.body)` had to move
   into the `WHERE` clause ahead of it.

**Retraction, recorded because the corrected belief is worth more than the fix.** This
work was ranked #1 in the backlog on the claim that the in-scope predicate was "spelled
in Claude Code's vocabulary" and therefore harness-blind. That claim was carried for five
days and quoted twice without anyone reading the code. The predicate has a dedicated arm
per harness and the comments already stated that the excluded streams "correctly stay out
of scope." The defect was one level up, in the denominator.

## v0.6.4 — 2026-08-13 (atlas unchanged at v0.1.7)

**The portable tier stops asserting a confidence nobody computed.** One fix, and it is
the most serious thing found this week.

- **`findings export` was FABRICATING its confidence field.** It computed
  `magnitude.is_some() → "measured"`, and `magnitude` is the **count** rank — a number
  with no relationship whatsoever to duration measurement. Measured on the reference
  store before the fix: **118 of 118 exported findings, including all 89 WAIT rows,
  shipped `confidence: "measured"`** — the strongest value in the vocabulary, in the one
  field whose entire job is carrying the caveat, on the one tier that crosses machines.
  A receiving machine was told every claim was fully measured; on this corpus essentially
  none of them are.

  Now the value is **derived at promotion from the hotspot rows the claim was built
  from**, folded to the **weakest** contributor (rank folds to the best, because a shape
  is as prominent as its strongest mount; confidence folds to the worst, because a claim
  is only as trustworthy as its weakest evidence). Same store, after: **0 measured, 69
  `derived_upper_bound`, 49 `not_measured`.**

  Findings also carry `stated_occurrences` / `derived_occurrences` so a receiver can
  **audit** the label rather than trust it. That was the whole failure mode: nothing sat
  beside the label to disagree with it, so a fabricated value was indistinguishable from
  a measured one for as long as it existed.

  `hotspots::confidence_for` and `DurationProvenance::from_counts` are now shared rather
  than duplicated — a second derivation is exactly what produced the fabrication.

- **`EXPORT_DOC_VERSION` 1 → 2, and v1 documents are REFUSED.** This is a *meaning*
  change, not a shape change, which is the more dangerous kind: a v1 document's
  `confidence` claims `measured` on every row regardless of its durations, so accepting
  one would import a fabricated all-clear and render it beside honestly-measured rows,
  indistinguishable. A refusal names the reason; a lenient import would launder it.
  **Fleet note:** machines exchange findings only once both sides are on 0.6.4.

## v0.6.3 — 2026-08-13 (atlas unchanged at v0.1.7)

Three fixes, all one shape: **a surface discarding what the data already knew.** Every
one arrived from the peer seat on the other machine, and every one was verified against
its own falsifier before being written.

- **Capped lists in the orientation packet now say what they dropped.** Only `traps` had
  a total, and it disclosed only when ZERO traps fit — a partial cut (2 shown of 8) said
  nothing, which is the common case. `working_set`, `paged` and `slow` had no total at
  all, so a project with 15 working-set files and one with 5 rendered **identically** —
  in a packet whose job is telling a cold agent how broad the project is, for a diagnosis
  (`WORKING_SET`) that is *about* breadth. Every list now discloses `(5 of 15)` when cut
  and stays silent when whole. `HEAT_FORMAT` 1→2; existing bakes go silent until the next
  `ratchet run`, by design.

- **A trap's WORKED command renders whole, or is withheld with its length named.** Never
  a prefix. `hook.rs` already ruled that "a truncated command is not a shorter trap, it is
  a wrong one — and the agent would run it", and dropped pairs that could not fit; `brief`
  applied no such rule and folded both halves, producing **15 `[truncated for display]`
  markers in a single brief** on the one surface guaranteed to be in an agent's context.
  The asymmetry is the fix: the FAILED half is an identifier and still folds; the WORKED
  half is the deliverable. Now 8 markers, all of them failed-halves.

- **Slow-command rows carry their duration provenance.** A row reading `177.3s` with one
  flagged call could not be told from a human away from a permission prompt — the
  retracted "8.9 h grep" mistake, reintroduced downstream of its own fix. The field was
  already sitting in the evidence JSON the duration is parsed from; the gather simply
  never read it. Rows now read `harness-stated`, `derived — includes approval wait`,
  `mixed (n/m)` or `unrecorded`, and n=1 rows are named as single observations rather
  than ranked silently beside real medians. On the reference corpus **every row reads
  derived**.

## v0.6.2 — 2026-08-11 (atlas unchanged at v0.1.7)

Four fixes, three of them one defect wearing different clothes: **an answer that
was correct and unreachable**. Every one arrived from a peer agent over the
wire, and two independent reviewers converged on the same mechanism before any
of it was written.

- **`ratchet query hotspots` now defaults to `--order occurrences`.** The old
  `seconds` default was not an ordering, it was a silent FILTER: only WAIT
  carries a duration, so the default top-50 on the reference corpus was 50 WAIT
  rows and every untimed shape was ejected — including the two LARGEST, REPEAT
  (rank 1, 1,285 occurrences) and PAGING (rank 2, 940). A reader asking for "the
  hotspots" got a view that structurally could not contain them, while the
  envelope's own prose already said to rank by occurrences. Confirmed
  independently on two machines. `--order seconds` is unchanged and still
  opt-in; the default is now a named constant so a test can assert it.

- **A row with no seconds now carries NO seconds rank.** `rank_seconds` was
  computed with `COALESCE(total_cost_seconds, 0)`, which handed every untimed
  row the same number — `count(timed rows) + 1`, observed as 69 on one machine
  and 66 on another. That is a sentinel, and it did not render as one: it
  rendered as a rank, beside a null `mean_seconds`, and read as "measured, and
  small". It is now `null`, ordered NULLS LAST so absence parks behind the timed
  rows rather than floating to the top. **Breaking for any caller that parsed
  `rank_seconds` as a non-null integer.** Same law as the NULL seconds beside
  it: absence is a third state, never a value — and disclosure has to be
  attached to the number, because a caveat sitting beside a value loses to the
  value every time.

- **`ratchet_context` no longer aims the reader away from its own verdict.**
  The description told callers to "read these first" of the hotspot rows, which
  outranked the computed `brief.diagnosis` sitting in the same payload. It cost
  a real misread: a reviewer holding a PROCESS diagnosis — *"5 of 6 REPEAT
  signatures re-read notes, not code; a code index will not help"* —
  recommended a code index anyway, citing paging cost that is `null` under an
  explicit `TIME_NOT_MEASURED` caveat. A structured refusal was present and was
  inverted by a competent reader. The description now states that a present
  diagnosis supersedes hotspot-derived actions, and that the rows exist to
  falsify it.

- **The below-floor orientation packet says what to do instead of only what it
  refuses.** `brief::diagnose` already emitted the sentence — *"the working set
  and traps below are still the best available prediction … read them as a
  prediction, not a verdict"* — and the packet's 260-character budget cut it,
  leaving the refusal alone. On the reference machine **23 of 35 tracked
  projects sit below the 5-session floor**, so that was the default orientation
  experience rather than an edge case. `INSUFFICIENT` details now travel whole;
  every generated diagnosis is still capped.

## v0.6.1 — 2026-08-07 (atlas unchanged at v0.1.7)

The first release whose entire content arrived over the wire: every fix below
answers a peer agent's field report, two of them filed channel-to-channel with
no human relay.

- **Bash verb attribution: four junk families closed** (a peer retested a
  0.2.0-era defect on 0.6.0 and was right — 1,244 of 23,702 distinct corpus
  targets attributed junk; now 58, all in the documented quote-blindness
  class). Lowercase shell locals no longer rank as verbs (`pass=0` — the old
  UPPERCASE-only assignment rule's stated rationale was positionally
  impossible); subshell parens are transparent (`(cd x && make)` is make);
  line-continuation backslashes are noise; and wrapper launchers unwrap to
  the work — `npx tsc` is tsc (614 distinct targets), `timeout 100 npm` is
  npm, `sigil run … -- sh -c` is shell-script while `sigil list` stays sigil.
  Command substitutions attribute their INNER command (`n=$(grep …)` runs
  grep) — but never arithmetic `$((…))` and never a single-quoted literal,
  which executes nothing. An external adversarial review (grok) caught three
  regressions in the first cut — bare launchers, value-taking flags,
  `sigil run` without `--` — all fixed and pinned. Hotspot shapes regroup on
  the next sweep; findings keyed on old wrapper shapes supersede naturally.

- **`ratchet brief` learned three things from a cold reader** (a peer's
  first-contact report, filed over the wire): trap pairs are now
  divergence-anchored — two distinct commands can never render as identical
  visible strings (the old head-truncation cut away exactly the
  distinguishing span; a truly identical pair now says so instead of drawing
  a fake distinction); the positional PROJECT is optional (cwd resolves via
  the same marker walk doctor uses; arc worktrees resolve to base); and a
  missing atlas manifest renders as `map absent — run: atlas build` instead
  of a bare dash — absence with its remedy, never an auto-build.

- **`ratchet update` ends with a stale-session census** when either binary
  changed: every live process still holding the replaced image is named (pid
  + argv), because an MCP server spawned before the update keeps running old
  code until its session restarts — measured on this fleet at eight live
  processes, two of them three days stale. A failed census reports as failed,
  never as all-fresh. 702 tests (681 at v0.6.0).

## v0.6.0 — 2026-08-04  ·  atlas v0.1.7

- **Antigravity is the fifth modelled harness.** New normalizer for the
  `~/.gemini/antigravity-cli` brain transcript stream (USER_INPUT /
  PLANNER_RESPONSE with tool_calls / DONE tool-result records). 91.4% of the
  1,223 already-captured records enter scope on the first `ratchet run` after
  this update — the cursor-reset contract fires once, automatically. The
  notable property: **every tool result carries harness-stated
  `Created At`/`Completed At` stamps, so antigravity's durations are 100%
  `explicit`** — the first WAIT tier with zero approval-latency
  contamination, a control group for every `derived_upper_bound` ranking.
  Absence stays honest throughout: the harness states no token counts (NULL,
  never 0), its status vocabulary states no failure (is_error NULL unless the
  typed exit_code says so — the 19 prose-only "command failed" records are
  deliberately unread), and `cost_basis='subscription'` marks the dollar
  question inapplicable rather than free.

- **Absent-vs-zero: the read side is closed.** The four sites recorded in the
  state doc were verified against the tree — three were already fixed by
  earlier intents (the ledger had gone stale in the optimistic direction) —
  and the two real gaps are gone: token sums over an unpriced store now
  render "not measured" instead of "0 tokens", and an empty caveat list says
  "not computed" instead of "no conditions currently hold" (that wording was
  reachable ONLY when nothing was computed). 681 tests (663 before).

**atlas v0.1.7**

- **Worktree detection is marker-first.** 0.1.5/0.1.6 resolved arc worktrees
  to their base project, but DETECTED them by the `<project>.arc-<id>-<slug>`
  naming scheme — if arc ever renamed its layout, atlas would never look for
  the contract file at all. `.arc-workspace` now counts as root evidence in
  the walk itself and the marker read no longer requires a conforming name;
  the naming scheme gates only the sibling fallback. An arc field session's
  observation ("the marker is the contract") made true within the hour.
  77 invariants.

## atlas v0.1.6 — 2026-08-04 (ratchet unchanged at v0.5.1)

- **The `.arc-workspace` rung of worktree resolution actually fires now.**
  Field verification of 0.1.5 (arc's agent, same day) showed the resolution
  succeeding through the name-derived sibling FALLBACK — the authoritative
  `.arc-workspace` rung was dead on arrival, because arc's live contract
  points `arc_root` at the `.arc` DIRECTORY rather than the project root and
  atlas required the declared path to itself carry a marker. Atlas now reads
  the field as written and as intended (a declared path whose basename is a
  root marker means its parent), and the invariant bakes arc's live spelling
  in, so a regression cannot hide behind the fallback. The case this rescues
  is the renamed base — the one layout the name fallback cannot save.
  77 invariants.

## v0.5.1 — 2026-08-04  ·  atlas v0.1.5

Both fixes in this cut came from OTHER agents using the tools under load —
an external adversarial review and a field report from another project's
session — which in two days has found more defects than reading the code ever
did.

- **Portability gate: home shorthand no longer hides behind punctuation.**
  `(~alice)` and `file=~bob` passed every `looks_like_path` heuristic (no
  slash, token does not START with `~`), so a machine-local home directory
  shorthand could ride a claim into the portable tier. The rule now: a `~`
  opening a token or following any non-alphanumeric byte is location-shaped;
  a mid-word `~` (`approx~5`) stays prose. `=D:` was the same bypass one
  heuristic over (the drive-letter check read raw byte positions) and is
  covered by the same stripped-prefix pass. Zero existing shapes trip the
  tightened rule — the over-refusal side measured free. Found by a headless
  Antigravity review. 663 tests (662 before).

**atlas v0.1.5**

- **An arc worktree never keys its own manifest.** Worktrees are disposable
  sibling clones that vanish at merge; a manifest keyed on the worktree path
  is a cold index built at exactly the moment work starts and thrown away
  days later — a field session measured four intents that would have meant
  four throwaway 459-file manifests. Resolution now lands on the BASE
  project: by arc's own `.arc-workspace` contract first (authoritative,
  survives a renamed base), by the name-declared sibling second, standing
  pat only when both are gone so a worktree whose base vanished still
  answers. The worktree diverges from base only in the files the intent
  itself edits, and pointers are verified against disk at answer time.
  77 invariants (76 before).

## v0.5.0 — 2026-08-03  ·  atlas v0.1.4

- **NEW — the findings tier gains a transport: `ratchet findings export [--out]` /
  `ratchet findings import <file>`.** The file primitive for moving portable
  findings between machines. What travels: identity (`kind`, `tool_name`,
  `shape`), occurrences, ranks, confidence, contest verdicts WITH their
  reasons, and provenance (`origin` — the machine the claim is about). What
  never travels: wall-clock time — findings carry no seconds by construction,
  the document's `meta.time_policy` states the exclusion, and a test walks
  every key in the document against banned time-substrings. Import is
  idempotent, never overwrites a local contest verdict or a locally-promoted
  row, applies the same portability gate `promote` uses on the way in
  (row-wise, counted), and refuses malformed or wrong-version documents with a
  JSON refusal and non-zero exit, writing nothing. Imported rows carry
  `origin` in `findings --json` and are annotated in the CLI listing and
  report. How documents MOVE between machines (file copy, relay, nothing) is
  deliberately still an open design question — this is the primitive any
  answer needs.

- **Absent-is-never-zero, write side.** v0.4.0-era work swept the READ side;
  this release sweeps the writers. Ten masked absences fixed, two of them live
  bugs: a failed sweep-id query could dress itself as "no sweeps yet" and mint
  `sweep-000001` over an existing store, and the v1 provenance backfill had
  labeled rift's harness-stated durations `derived` in live stores (corrected
  by a one-shot migration that reads the truth back out of the raw bodies —
  pruned bodies stay NULL, never guessed). Also: `tool_target` and a
  no-cwd record's `project` no longer write `''` for absence; rift and codex
  inserts now stamp `duration_source`; grok's `request_id` is no longer
  fabricated from an absent session id; hotspot and finding rows with no tool
  dimension store NULL — an **observable JSON change**: `tool_name` on SERIAL
  rows serializes as `null` where it was `""`. One residual is documented on
  the column rather than half-fixed (`is_sidechain` on pre-field claude-code
  history). 662 tests (639 before).

- **`--kind` docs and MCP schemas now name PAGING.** The sixth signature was
  accepted by every filter but named in none of the nine places that list the
  vocabulary — including the three MCP tool schemas agents read.

**atlas v0.1.4**

- **Vendored-by-declaration pruning.** A directory is pruned iff its name
  matches a dependency the project itself declares (root `package.json`, all
  four dependency sections) AND at least two siblings also match — the
  evidence is the project's own manifest, not name-guessing, and a lone match
  never prunes (measured: first-party `redux/` and `bull/` dirs wrap the dep
  they are named for). Pruned paths are recorded in manifest meta, disclosed
  on `refs` misses, and filtered identically in the live grep. Measured on the
  motivating project: manifest build 10.48 s → 0.56 s. The old `"packages"`
  name-guess is REMOVED — it was hiding first-party monorepo code in two
  measured projects, which now index.

- **The canonical source now lives in the ratchet tree** (`atlas/atlas.py`
  plus its 76-check invariant suite) — atlas was a 3,748-line file tracked by
  nothing. Release staging reads the in-tree path; the published artifact
  contract is unchanged.

## v0.4.0 — 2026-08-03  ·  atlas v0.1.3

**The MCP surface was a strict subset of the CLI.** Every surface in this tool
is built for agents, and in two places the agent could reach less than a human
could. Both are closed here.

- **BREAKING-ISH — `ratchet_status` now returns the WHOLE of `doctor --json`.**
  It called `doctor::render_json`, which reports the pipeline only, so the
  surface agents read could never emit `BINARY_MISSING`, `MCP_UNREGISTERED`,
  `MCP_STALE_PATH`, `LAUNCHD_UNLOADED`, `LAUNCHD_STALE_PATH`,
  `STORE_UNMIGRATED`, `STALE_BINARY` or any `ATLAS_*` code. A human running
  `ratchet doctor` on a machine whose MCP registration was missing learned the
  install was broken; an agent calling `ratchet_status` on the same machine at
  the same moment was told the pipeline was fine. That is the exact drift class
  this project spent two days fixing, sitting inside the tool that diagnoses
  it.

  **What changes for callers.** The response gains `install` and `atlas`
  sections, and `faults` becomes the flat union of all three fault lists.
  **Anything branching on `faults` being empty will now see codes it has never
  seen**, on machines that were always broken and never said so. An empty
  `faults` array is still the healthy state; the array is just no longer
  lying by omission. The tool's description was rewritten to say all of this,
  because the description is the only part of the contract an agent reliably
  reads.

  Each wiring fault also arrives as a structured `meta.caveat` carrying the
  OBSERVED value and the fix — `MCP_STALE_PATH` names the path the
  registration actually points at, not merely that it is wrong. A bare code in
  a string array is something an agent can branch on but not act on.
  `STALE_BINARY` is the one `info` among them: an installed build older than
  the store's last writer is what a completed upgrade looks like for one
  scheduled tick, and an alarm that fires on a healthy upgrade is one people
  learn to ignore (the `never_ingested` ruling, applied again). It still
  faults — the exit-code contract is unchanged.

  Cost, measured against an empty store so nothing else is in the number:
  **+0.20 s**, flat, for three subprocess probes (`launchctl list`, two
  `--version` execs). Independent of store size.

- **NEW `ratchet_orient`.** The fused orientation packet — ratchet's measured
  heat for the project composed with atlas's structural map for the tree — had
  no MCP door at all. It is the richest artifact this tool produces and the
  only ways to it were the `SessionStart` hook (Claude Code only, and only if
  installed) or shelling out to `ratchet hook orient`. On the reference machine
  that shell-out is what agents actually did.

  Same bytes, same renderer, same byte ceiling as the hook's packet — a test
  pins byte-equality, so a packet quoted from MCP and a packet seen in a
  transcript can never differ. It **never builds**: a primary-key `meta` lookup
  of the pre-baked heat row plus a read-only open of an EXISTING atlas
  manifest, with a test asserting no manifest is conjured and no heat row baked
  as a side effect.

  A separate tool rather than a key on `ratchet_context`, because
  `ratchet_context` refuses with `no_sweep` before it computes anything — it
  would have buried a packet that does not need a sweep behind a refusal that
  does not apply to it.

  Absence is named, never implied. Each missing half reports the exact key or
  path that was searched (`orient/<project>` in the store; the manifest path
  under `~/Library/Caches/ratchet-map`), and the heat half distinguishes THREE
  conditions that all render as the same empty section and mean entirely
  different things: `ORIENT_HEAT_ABSENT` (never baked — nothing measured here
  yet), `ORIENT_HEAT_STALE` (baked past the 7-day window, so the scheduled job
  has been DOWN, which does not self-clear — `warning`), and
  `ORIENT_HEAT_FORMAT` (a bake from a superseded packet format, which is what
  an upgrade looks like for one tick — `info`). Collapsing those into one code
  would have been the `BLIND_HARNESS`/`PRE_INSTRUMENTATION` mistake again.

- **The tool count was left alone, on the evidence.** atlas collapsed 8 verbs
  to 1 after measuring a routing tax; the question was whether ratchet's tools
  carry the same one. They do not, and the store says why: on the reference
  machine ratchet's MCP surface has been called **7 times in its life**, 6 of
  them `ratchet_context` — while `ratchet <subcommand>` was shelled out of Bash
  **277 times**. There is no routing tax in that data because there is barely
  any routing. `ratchet_context` already IS the one-door form, and the tools
  behind it are distinct DATASETS reached after it has said which one holds
  your answer — not competing shapes of one answer, which is what made
  guessing an atlas verb expensive. The measured gap is reach, not routing,
  and reach is what the two fixes above address.

- **`ratchet remove --purge` now reaches atlas's manifest cache.** The cache
  at `~/Library/Caches/ratchet-map` is the one part of the install that lives
  outside `~/.ratchet`, and `remove --purge` never touched it — a purge left
  every derived manifest behind. It is now listed and sized in the plan
  before the confirmation gate like every other part, and removed with the
  store; without `--purge` it stays, because reaching into
  `~/Library/Caches` costs a flag.

**atlas v0.1.3 — the exactness gate, enum variants as declarations, and a
server that discloses its own staleness.**

- **A thin answer is gated like an empty one.** Third in a family: `where
  admin` let an exact hit suppress the substring sweep (fixed 0.1.1), `wire
  /cmi5` let one index answer while the sibling table held the real answer
  (fixed 0.1.1) — this is the same defect one level up, a SUBSTRING hit
  suppressing the LIVE pass. Measured in the field: `idle_timeout` returned
  two neighbours (`idle_timeout_triggers`,
  `idle_timeout_orphan_from_trigger`) while one file held THIRTEEN
  occurrences of the term including the `sleep(idle_timeout)` the whole
  investigation turned on — a local `let`, which the index structurally
  cannot hold; two neighbours were enough to make `empty` false, so the live
  fallback never ran. The gate now asks "did we find the thing you NAMED?",
  never "did anything match?" — a count threshold was refused deliberately,
  because "fall through under N hits" is a guess that is wrong at N+1. Every
  answer carries `exact_indexes` (the doors that held the term itself;
  `name substrings` absent by construction — a name containing the term is
  the definition of a neighbour), `neighbours_only`, and a one-sentence
  `qualifier` printed as the second line, because a two-row answer does not
  prompt the footer read an empty one does. A query the index really does
  answer still costs zero greps.
- **Rust enum variants are declarations.** Measured on a sibling project's
  manifest: function 3891 · type 484 · impl 173 · value 138 — and no variant
  kind at all; `DaemonCommand::CaptureNow` was declared and simply not held.
  Command, state and error enums are how this fleet's Rust is written, so
  the variant is exactly the handle a reader arrives with. Variants index as
  `Enum.Variant`, kind `variant`; the bare `CaptureNow` resolves by name and
  the qualified `DaemonCommand::CaptureNow` through its parent — one
  `declaration_rows` definition serving `where` and the one-door lookup
  alike, because two answers to "what declares this" from one manifest is
  how a tool contradicts itself. The variant scanner is a hand-rolled scan,
  not a regex: the regex attempt BACKTRACKED on
  `Sha(String), // explicit start point` and reported a variant named `t` —
  three of that tree's 520 variant pointers confidently wrong, the one thing
  rule 3 forbids outright.
- **`atlas file` matches on a path boundary, and a tie is disclosed.**
  `file text.rs` was a suffix match and returned the record for a file named
  `context.rs`; `file mod.rs` picked ONE of nine same-named files from a
  query with no ORDER BY and suppressed every other index on the strength of
  it. Now a `/` boundary bounds the match, the shortest path wins
  deterministically, and the others are listed (`also_named`) — ambiguity is
  information.
- **`STALE_BINARY` — the server discloses its own staleness.** An MCP server
  is spawned once and lives for the session, so replacing atlas on disk
  leaves the RUNNING image executing the old inode — and nothing in any
  answer said so. Measured cost: a tester spent real effort inferring from
  byte-identical results and a `ps` call that they were testing 0.1.1 while
  0.1.2 sat on disk. The server now re-stats its own file on every call and
  attaches a `STALE_BINARY` warning naming the fix precisely: restart the
  SESSION, not the server — MCP registrations are fixed when the client
  session spawns, so respawning the child re-runs the same image.

639 tests (619 before), 0 failures.

## v0.3.9 — 2026-08-03  ·  atlas v0.1.2

**The enablement gap.** Every teaching surface this tool shipped targeted
`AGENTS.md`. Measured on the reference machine:

```
42 arc/git-tracked projects
22 have AGENTS.md          <- what `agents-block --write` targets
 8 have CLAUDE.md          <- what Claude Code actually reads
 1 has the tool blocks, and it is in AGENTS.md
```

Claude Code — the dominant runtime here — reads `CLAUDE.md`, auto-memory, MCP
server instructions and skills. It does not read `AGENTS.md`. Verified from
inside a live session working in `ratchet`, which has a 10.7 KB `AGENTS.md`:
not one byte of it reached the context. **For that runtime, enablement was 0 of
42.** The block's own header was technically honest — "Codex / Cursor / Gemini
/ Antigravity / any runtime that reads AGENTS.md" — and named only the
inclusion, which is how the gap stayed invisible.

- **NEW `ratchet skill [--write] [--remove] [--dry-run] [--skills-dir DIR]`.**
  Installs one USER-scope Claude Code skill per tool into
  `~/.claude/skills/<name>/SKILL.md` — `ratchet-telemetry` and `atlas-lookup`.
  User scope is the point: one write covers all 42 projects, including the 34
  with no `CLAUDE.md` to edit, and it creates nothing inside any repository.

  **Two skills, not one.** The tools version independently (atlas moved three
  times on 2026-08-03 while ratchet moved once), so a combined skill goes stale
  whenever either moves. The stronger reason is that a skill's `description` is
  an INVOCATION TRIGGER rather than documentation: "where does this symbol
  live" and "has this shape cost time before" are different questions, and one
  description covering both mis-fires in both directions. Each description says
  when to reach for the skill and names the neighbours it must not be confused
  with — `ratchet-advisor` and `recon-cache` already exist and REASON OVER
  ratchet's output; neither teaches that the tools exist.

  `~/.claude/skills` holds the user's own work, so every write is
  marker-guarded: a file at our path without ratchet's marker is refused, never
  overwritten and never removed. `--remove` takes the file and its emptied
  directory and leaves nothing — a thing that cannot be cleanly removed does
  not get installed.

- **NEW `agents-block --write --all`.** 24 projects had real session history
  and no blocks; that was 24 manual invocations nobody was going to perform.
  `--all` enumerates the projects the store has measured an editing session for
  (the same list `hook bake` uses, so the two cannot disagree), re-derives each
  project's root from its session cwds, and **re-checks for a `.arc`/`.git`/
  `.hg` marker immediately before the write**. A project whose tree moved or
  was deleted is reported and skipped, never guessed at. Already-current
  projects write nothing at all. One project's unbalanced marker pair does not
  stop the others — it is reported, and the exit code carries it. `--dry-run`
  reports the whole plan; with `--all` the argument is a bare file NAME joined
  to each root (`--write CLAUDE.md --all`), and a path is refused.

- **NEW `ratchet enabled [--projects] [--json]`.** Enablement you cannot see
  decays silently, and this audit had to be hand-written as a shell script to
  answer "which of my projects are wired" — a question that recurs every time a
  project is added. Machine-wide: both binaries and versions, both MCP
  registrations, the SessionStart hook, both skills. Per project: editing
  sessions, whether the orient heat is baked and still servable, whether atlas
  has a manifest, and which blocks are current/stale/absent/refusing in
  `AGENTS.md` **and** in `CLAUDE.md`. Deliberately not `doctor`: that owns
  health and an exit code, this never exits non-zero, because an un-enabled
  project is a normal state and not a fault.

- **`install.sh` gains an opt-in menu.** MCP / skills / SessionStart hook /
  `AGENTS.md` blocks, each independently choosable and independently
  idempotent. The components that touch only ratchet's own install keep their
  `RATCHET_NO_*` opt-outs; the three that write into files the installer did
  not create are OPT-IN — `RATCHET_SKILL=1`, `RATCHET_HOOK=1`,
  `RATCHET_BLOCKS=1` (`RATCHET_ALL=1` for all three, `RATCHET_BLOCKS_FILE=` to
  choose the file). A `curl | sh` may register its own MCP server; it may not
  silently edit somebody's repositories. **No prompts** — the script is piped
  from stdin and has no TTY, so the menu is printed with each component's state
  and the variable that flips it, and every step probes its subcommand before
  invoking it so a pinned older `VERSION=` reads as "this build cannot" rather
  than a failed install.

- **One canonical body per tool, rendered into three surfaces.** This release
  would otherwise have created a FOURTH hand-maintained copy of "here is what
  atlas does" — the drift that produced `handle_ms` vs `duration_source` and
  `install.sh`'s 64 mentions of atlas against `update.rs`'s zero. Instead
  `agents_block::Tool` carries the `include_str!`'d published block, and
  `Tool::body()` (the text strictly between the delimiter lines, so the "paste
  this into your AGENTS.md" instruction never travels) is what both the block
  writer and the skill renderer consume. The parts that genuinely cannot be
  shared — the skill's `name`/`description` frontmatter — are DERIVED from
  fields on `Tool`, never forked from the body. The MCP `initialize`
  instructions remain a separate rendering on purpose (they name MCP tools, not
  CLI commands, and arrive as one prose paragraph), and are now DRIFT-GUARDED:
  a test asserts every claim both surfaces make — `derived_upper_bound`,
  `unverified`, `comparable`, `occurrences`, `permission prompt`, `never zero`,
  `caveat` — is made in both. A guard is weaker than a shared source and is
  named as such rather than counted as one.

- `ratchet remove` now lists and removes the two skills, each by its own path
  ("skills: 2" would not say which two directories under `~/.claude/skills` a
  removal is about to reach into). A foreign file at one of those paths plans
  as absent, because that is what the removal does to it: nothing.

**Also in this cut: the `absent is never zero` sweep.** The rule this project
invented — `not_measured`, `derived_upper_bound`, `PRE_INSTRUMENTATION`, a null
cost that means not-measured-never-zero — had never been swept for as a
standard. Four sites fixed. The sharpest: **`total spend: $0.00` on a
subscription-only machine**, where `cost_usd` is NULL *by design*; a claim about
money from a tool that had measured none, in the report that displays the very
actions the NULL rule was written for. Also `pct_of`, whose correct form
(`naive_ratio_suffix`) sat one function below it; doctor's `last_*` counters
reading `0` errors from zero runs; and reads-per-file — where the real defect
was not the constant but that **the division had two homes**, with `brief`
recomputing what `orientation` already knew. Roughly fourteen further sites were
examined and **left alone with reasons**, including the promotion gate where
`0.0` is the value that *refuses* to promote and failing closed is the point.

**atlas v0.1.2 — the furniture rule, and Rust import resolution.**

- **atlas indexed the agent's own transcripts as code.** In a real audit,
  `.rift/sessions/*.jsonl` ranked **third and fourth** on `refs write_queue`,
  above genuine callers, because an agent had discussed the module eleven times
  in one conversation. The failure profile is the point: it bites when a term is
  **rare in code and heavy in chatter** — the exact profile of a term you are
  investigating *because you do not yet understand it*. On a code-dense term the
  pollution is invisible. ratchet already had the principle ("that is the
  agent's furniture, not a codebase"); atlas had one of its six `AGENT_DIRS` and
  that one for VCS reasons. Now ported name-for-name, with a single predicate
  feeding both the file walk and the live-grep post-filter so the index and
  `refs` cannot disagree. `refs write_queue`: **116 refs / 22 files → 94 / 20**,
  both transcripts gone, the real callers back at 3rd and 4th.
- **A project's own committed `.claude/` is excluded, diverging from ratchet
  deliberately.** The tools ask different questions of one path: ratchet asks
  *whose spend was this* (editing a repo's `.claude/` is work on that repo),
  atlas asks *is this the codebase*. Being committed does not make a file code —
  and `SKILL.md` files are the densest source of a project's own vocabulary that
  is not its source (arc's three say `intent` **114 times**), which is the same
  defect one level up. Measured cost, stated rather than hand-waved: on arc that
  directory contributed 10 file rows and **zero** symbols, config keys, routes
  or wire rows.
- **Rust `use` paths now resolve to files.** They never had: **0 of 190 on
  ratchet, 0 of 1,278 on arc**, while Python and JS both resolved. So `hubs` was
  not merely unavailable on Rust trees — on arc it was **confidently wrong**,
  naming a TypeScript file from a sub-service that is ~4% of the tree as the top
  hub of a 202-file Rust codebase, in the packet designed to orient an agent at
  session start. Wrong is worse than absent. Now: **arc 546 resolved, ratchet
  40**, and `crate::` — internal by construction — is **328/328 and 26/26**.
  What stays unresolved is external crates and the 143 `use super::*` inside
  `#[cfg(test)] mod tests`, where `super` means the file itself: the resolver
  reads the statement's indentation to detect the inline module and **refuses**,
  avoiding 133 confident wrong pointers.
- Invariants **46 → 56**, all ten mutation-tested across eleven mutations —
  including two *over*-exclusion mutations proving `notclaude/` and an ordinary
  `tmp/` still survive. That is the boundary proved in both directions.
- **Known, unfixed, and pre-existing:** STNG builds in **10.6 s**, not the
  sub-second the rest of the fleet manages. 6.8 s of it is eight vendored
  three.js/draco blobs under `lib/three/` — 401 of 519 JS files. `VENDOR_TREES`
  knows `node_modules` and `vendor` but not `lib/`. Detecting vendored code by
  content is a design decision, not a bug fix, so it was left rather than
  muddying this audit.

619 tests (562 before), 0 failures.

## v0.3.8 — 2026-08-03  ·  atlas v0.1.1 (unchanged)

Two first-impression defects. Both were things the tool said to a user in
their first minute with it, and both were wrong.

- **`ratchet update` silently ignored atlas.** Measured on the live install:
  `src/update.rs` contained **0** mentions of atlas; `scripts/install.sh`
  contained **64**. The installer was taught about the second artifact the day
  it shipped; `update` never was. So:

  ```
  $ ratchet update
  ratchet updated: 0.3.6 -> 0.3.7
  wiring already current

  published atlas: 0.1.1     installed atlas: 0.1.0    <- untouched, unreported
  ```

  A user who upgraded rather than re-running the installer could **never**
  receive an atlas fix, and no surface in the system said so — `doctor`'s
  atlas section has no published-version axis at all, only
  present/runnable/registered. The first atlas fix was already published
  (v0.1.1 removes a `SyntaxWarning` that hit stderr on every invocation) and
  was not reaching anyone who updated.

  `update` now reads the same `atlas_version` / `atlas_artifact` /
  `atlas_checksums` keys `install.sh` reads, through **one** verified
  downloader shared by both artifacts rather than a second copy of the
  fetch-and-verify dance — a second copy being how these two drifted apart in
  the first place. Absent from the manifest, current, behind, or missing
  entirely: each is a distinct sentence, and both tools are named on one line
  in every branch, including the branches where nothing happened.

  ```
  ratchet updated: 0.3.7 -> 0.3.8; atlas updated: 0.1.0 -> 0.1.1
  ratchet 0.3.8 is current; atlas installed: 0.1.1
  ```

  The downgrade refusal applies **per tool, independently**, because the
  versions are independent: atlas moved three times on 2026-08-03 while
  ratchet moved once, so one can legitimately be newer while the other is not,
  and a shared verdict would either block a real ratchet upgrade or perform a
  real atlas downgrade. A missing checksum entry stays a hard failure for
  either artifact — "nothing was published" and "something was published and
  we cannot verify it" still never share a branch. `--check` reports both and
  writes nothing.

  **If you are on v0.3.7 or earlier, you have a partial upgrade right now.**
  `update` moved ratchet alone and printed nothing about atlas, so a machine
  that installed once and has only run `ratchet update` since is still on
  whatever atlas was current on its install day. Running `ratchet update` from
  v0.3.8 catches it up in one go; so does re-running `install.sh`, which is
  safe to re-run and has always handled both tools. Nothing warns you about
  this state on 0.3.7 — that is the defect, and it is why it is named here
  rather than quietly working from now on.

  The upgrade path is `ratchet update`, and `doctor`'s atlas remedies now say
  so; every one of them used to point at `install.sh` and nowhere else, which
  was the codebase admitting the gap without anyone reading it that way.

- **A brand-new install failed its own health check.** Found by installing
  into a virgin `HOME` against the real published site. Ten seconds after
  `curl | sh` printed success:

  ```
  ratchet: STALE — no capture within 48h. Ingest has stopped and the source
    corpus self-deletes on a rolling window, so evidence is being lost
    permanently.
  ratchet: coverage 0.0% — some captured records are unmeasured
  ```

  Both derivable, both **false as claims**, both alarming, and together a
  non-zero exit. Nothing is being lost permanently when nothing was ever
  there; no captured record is unmeasured when there are no captured records.

  There are three states, not two. Captured recently: fine. Captured, then
  stopped: `STALE`, and evidence really is being destroyed on the rolling
  window. **Never captured: NEW** — `doctor` says so, names `ratchet run`, and
  does not fault. `stale` had already been wrong in *both* directions here (an
  earlier `map_or(false, ..)` read an empty store as healthy, and its
  correction to `true` produced the alarm above), which is the signal that two
  states were being squeezed into one boolean.

  NEW is `info`-shaped rather than a fault, by the project's own severity
  rule, and the reasoning is on the record because "run `ratchet run`" **is**
  an action: the condition clears itself on the next scheduled tick 30 minutes
  later; the exit code is read immediately after an install, where red means
  "the install failed"; and the genuinely actionable failure on a fresh
  machine — nothing scheduled, no binary, no MCP registration — is already
  owned by the install faults, which do fault. A second alarm for a condition
  another check owns better is how people learn to skip alarms. Same reasoning
  that earned `PRE_INSTRUMENTATION` its `info` severity.

- **Coverage over an empty corpus is `null`, not `0.0`.** `in_scope_pct` and
  `coverage_pct` are `Option<f64>` and report absent when nothing has been
  captured — in `doctor`, in `report`'s Health block, and in every JSON
  envelope's `meta`. `0.0` and absent are opposite claims: the first says this
  build modelled none of what it caught, which an agent would act on. The
  existing `in_scope == 0 → 0.0` reading is untouched and still correct — that
  one is a real statement about real records, and it is what
  `PRE_INSTRUMENTATION` and `BLIND_HARNESS` are built on. Only the empty-store
  case became absent. One stored `0.0` survives, in `hotspot.coverage_pct`,
  where the column is a promotion **gate** rather than a display value and
  failing closed is the point — nothing reads it as a measurement.

- **Tested against a stub origin, never the live site.** `update` now honours
  `RATCHET_BASE_URL` — the same knob `install.sh` already had — so the whole
  atlas round (manifest parse, download, checksum verify, tar extract, atomic
  install, version probe) runs against a tempdir served over `file://`. That
  absence is why this path had no end-to-end test before now: an update path
  testable only by deploying is one that gets tested after release.
  **562 tests, 0 failures** (541 before; +21 net, all in `tests/update.rs`).

- Also: `update::scheduled_binary_path` derives from `Layout` instead of
  hand-joining `~/.ratchet/bin/ratchet` a second time, and `llms.txt` no
  longer claims atlas v0.1.0 in two places while its own status line says
  v0.1.1.

**atlas is unchanged at v0.1.1.** It versions independently, and bumping it to
match a ratchet release would publish an artifact with no changes in it —
exactly the coupling the split versions exist to prevent.

## v0.3.7 — 2026-08-03  ·  atlas v0.1.1

The guard was asking the wrong question, and the docs it blocked.

- **The immutability guard tested a LOCAL directory, not a published one.**
  It refused `--deploy` because the script's own dry run had created
  `public/v0.3.6/` — while the origin returned **404** for that path and
  `latest.json` still read 0.3.4. So the documented two-step workflow
  (`release.sh` to stage, then `--deploy`) was self-blocking, and the only
  escape it offered was `FORCE_REPUBLISH=1`, the flag whose absence
  permanently destroyed v0.1.1. Pushing a user toward that to get past a
  false positive is the worst affordance in the script.

  The question is now the immutable path's own `checksums.txt` **at the
  origin** — deliberately not `latest.json`, which names only the newest cut
  and would read a published-but-superseded version as absent. Four verdicts:
  **404 → publish** (overwriting dry-run leftovers); **200 and the sums match
  → identical**, leave it alone; **200 and the sums differ → refuse**, naming
  both sums; **anything else → refuse, failing closed.** Local staging is once
  again re-runnable without consequence, which is what a dry run means.

- **And the `identical` branch RESTORES a missing local directory.** This is
  the mechanism that destroyed v0.1.1, stated precisely for the first time:
  `wrangler deploy` uploads `public/` wholesale, so a version directory absent
  **locally** is a version directory deleted **at the edge**. It was never
  "a re-cut overwrote it" — it was an upload that didn't contain it.

- **`--docs-only [--deploy]`.** The nine root docs (`llms.txt`, `RATCHET.md`,
  `ATLAS.md`, `CHANGELOG.md`, the two blocks, `RESULTS.md`, `install.sh`,
  `latest.json`) are unversioned and correcting them is not a re-cut. Refuses
  unless local `latest.json` is byte-identical to live and every artifact it
  names hashes to what the origin serves; what it cannot check — older version
  directories, since the edge cannot be enumerated — it says out loud.

- **`ATLAS.md`, atlas's full agent reference.** It shipped with a paste block
  and no peer to `RATCHET.md`. Carries the one-door form, the
  `SEARCHED`/`NOT SEARCHED`/`NOT MODELLED` footer as the thing that bounds a
  negative, `NOT-INDEXED` as a verdict distinct from empty, the 2,000-file
  autobuild ceiling, and **the CWD trap**: a user-scope MCP server inherits the
  CWD of the client process that spawned it — the directory `claude` was
  launched from, NOT the project. Measured live against every running MCP
  child; a session launched from `$HOME` returns `no_root` for every call,
  which is why `atlas_lookup` takes an optional `path`.

- **`RATCHET.md` and `llms.txt` gained the lifecycle surface** — `wire`,
  `doctor`, `remove --purge`, `agents-block --write`, `update` and its
  downgrade refusal, and `hook orient|install|bake`. The SessionStart hook was
  **entirely undocumented** on the published site: 0 mentions in either file.

- **atlas v0.1.1 — the `SCHEMA` literal is now a raw string.** Its SQL comments
  quote route regexes containing `\/`, which Python does not recognise as an
  escape, so 3.12+ emitted a `SyntaxWarning` **to stderr on every invocation** —
  and a directly-executed script gets no `.pyc` caching, so it re-warned every
  single call (102 bytes, measured on the shipped artifact, identical on runs 1
  and 2). Harmless to the MCP protocol, whose frames are on stdout — but it
  trains a caller to ignore atlas's stderr, which is exactly where a genuine
  refusal appears. The resulting string is byte-identical; `\/` is left as-is
  either way.

- 541 tests (+51, all driving the guard as a command against stub origins:
  every classify input, every decide cell, both refusals checked for naming the
  tripped condition **and** for never again saying "staged artifacts").

**Known, unfixed:** a brand-new install fails its own health check. `doctor` on
a store seconds old reports `STALE — no capture within 48h … evidence is being
lost permanently` and `coverage 0.0%`. Both are technically true and both are
wrong as the first thing a new user sees; it needs a never-ingested branch.
Found by installing into a virgin `HOME` against the real published site.

## v0.3.6 — 2026-08-02

Provenance, an un-blinded harness, and a second tool on the installer.
v0.3.5 made the durations honest; this makes every number say where it came
from — which tool actually stated it, which project actually owned it,
which harness it belongs to — and makes two surfaces that had been quietly
disagreeing report the same measure.

- **`curl | sh` now installs TWO tools: ratchet, and atlas v0.1.0.** atlas
  is a stdlib-only single-file manifest of pointers for a codebase
  (declarations, name substrings, paths, routes, endpoints, config keys),
  landing at `~/.ratchet/bin/atlas` with its own MCP server registered
  beside ratchet's. Previously it lived at a source path an agent had to
  know by heart and nothing put it on PATH.

  **Independently versioned, and that is the whole design.** atlas changed
  three times on 2026-08-03 while ratchet changed once; one shared number
  would have forced a ratchet release per atlas fix, and — worse — made a
  ratchet release silently claim an atlas change. So `latest.json` names
  both (`version` beside `atlas_version`/`atlas_artifact`/
  `atlas_checksums`, flat and prefixed because a nested `"version"` would
  match the installer's own sed and hand it two versions concatenated), and
  each tool gets its own write-once directory: `/v<version>/` and
  `/atlas/<version>/`, each with its own checksums file. `VERSION=` and
  `ATLAS_VERSION=` pin either side alone; `RATCHET_NO_ATLAS=1` opts out.

  atlas's artifact carries every guarantee ratchet's already had, all of
  them paid for in blood: a **reproducible** tarball (`gzip -n`, fixed
  mtime, explicit mode — plain `tar -czf` produced two different checksums
  for identical source and read as a tampered artifact against a cached
  CDN), a **banner-vs-bytes** check that extracts the staged tarball and
  runs the extracted file's `--version` rather than trusting the source it
  was read from, and **immutability**. The immutability rule is *same
  version, same bytes* rather than a blunt "this version exists": a
  ratchet-only cut must not be blocked by an unchanged atlas, so the
  reproducible hash is compared against the published one — identical means
  reuse and write nothing, DIFFERENT is refused.

  A missing checksum ENTRY is a hard failure in the installer, never a
  skip. "Nothing was published" (absent, reportable, fine) and "something
  was published and we cannot verify it" (refuse) never share a branch.

- **`wire`, `doctor`, `remove` and `agents-block` all cover atlas.** `wire`
  registers it through the same name-taking, raw-text-splice helper —
  verified to parse AND deep-equal the original with only that entry
  changed, backup first, restore on failure — and reports `absent` rather
  than registering a server that cannot spawn. `doctor` gains an atlas
  section reusing the existing fault vocabulary with an `ATLAS_` subject
  prefix (`ATLAS_BINARY_MISSING`, `ATLAS_MCP_UNREGISTERED`,
  `ATLAS_MCP_STALE_PATH`) in the one flat `faults` array, still printing
  every fault before exiting once. `remove` takes both registrations and
  lists everything with sizes first — the bin directory is sized ONCE with
  both binaries named, because the number above a confirmation prompt has
  to be the number that gets deleted. `agents-block` emits a second block
  between its own `atlas:start`/`atlas:end` delimiters, so an atlas wording
  fix never rewrites ratchet's advice in a project's `AGENTS.md`.

- **A fourth `confidence` value: `derived_upper_bound`.** Hotspot rows now
  carry `duration_provenance {harness_stated, derived, unknown}`, and
  `confidence` distinguishes a harness-stated measurement from an elapsed
  call→result gap. Precedence, worst-first: `not_measured` >
  `unverified` > `derived_upper_bound` > `measured`. **`measured` now
  means harness-stated only — 113 of 51,954 timed calls** (51,805 derived,
  36 with no recorded provenance at all) — and it stays reachable
  (`WebFetch`, `Glob`), so it can still read GOOD. v0.3.5 removed the
  arithmetically *impossible* durations; what survived was the merely
  implausible. `Edit` still ranks **#1 by time at a 32.66 s mean** for an
  operation that completes in milliseconds — it just no longer claims that
  figure is machine time. Not a filter and not a re-rank: no row is
  dropped, no stored rank silently reordered.
- **`ratchet_context` leads with count-true churn.** New
  `top_hotspots_by_count` beside the existing `top_hotspots_by_time` —
  the same rows, two orders. By count this store leads **REPEAT 1,303 and
  PAGING 777**, and both carry NULL `cost_seconds` *by construction*, so
  neither can appear in a by-time ranking at all: orienting an agent with
  the time table alone hid the two largest shapes. The by-time table was
  kept deliberately — deleting it would make absence read as "no time cost
  anywhere".
- **Three caveats, each with a tested silent direction.**
  `DERIVED_DURATION` moved from the human report into the envelope: the
  report was the only surface carrying it while `hotspots --json`, `query
  hotspots` and `ratchet_context` published the same seconds with **no
  provenance statement at all**. Its arithmetic was wrong too —
  `derived ÷ (derived + stated)` hid the 36 rows that have a duration and
  no `duration_source` from *both* halves of the fraction. New
  `TIME_RANK_DERIVED` (warning) says whether the rows in front of you
  right now are derived and names the top one; new `TIME_NOT_MEASURED`
  (info) names which kinds in a count-led view have null seconds — absent,
  never zero.
- **Cross-project attribution was wrong, and wrong in the direction that
  flatters.** `parse_project` took the **basename** of a session cwd, so
  every action inherited that string regardless of which path it touched.
  Live store: `public` merged **six unrelated projects**, `src` six,
  `test` four, and `client` was **1,526 actions of beacon**. **152 of
  ratchet's 319 Read targets — 48% — were not under any ratchet
  directory**, and `brief ratchet` named another project's gateway
  `index.js` its heaviest pager. Now: a marker walk to a *repository* root
  (`.arc`/`.git`/`.hg` — package manifests rejected, because
  `beacon/packages/client` carries a package.json and would have preserved
  the exact fragmentation the walk exists to remove), per-action
  attribution by the path actually touched (gated to `PATH_TOOLS`, since a
  Bash target is a COMMAND and a Grep target a PATTERN and both can start
  with `/`), agent furniture judged positionally so a repo's own
  `.claude/` stays with it, and unresolvable → NULL, never the session's
  project.
- **Paths are stored repo-relative** — found by running the fixed tool,
  not predicted. Absolute paths fragmented ONE file across the base
  checkout and every worktree: `src/main.rs` held three rows at 33%
  coverage when its true coverage was 100%. Worse than the cross-project
  leak, because a foreign path looks foreign and a split file looks fine.
- **grok is modelled — the harness is no longer BLIND.** New
  `normalize_grok.rs` over the ACP `session/update` stream in
  `updates.jsonl`, the only one of grok's five streams carrying timestamp,
  session, tool call id, tool NAME, arguments and turn tokens on one line
  (`chat_history.jsonl` has the arguments but **no `ts` on any record**;
  `events.jsonl` has explicit durations but no `tool_call_id` on 711 of
  711 July calls). **2,225 of 112,386 captured records are in scope
  (1.98%)**, and that small number is the honest one: **89% of everything
  captured under grok is `phase_changed` UI breadcrumbs**, so "90,784
  records captured, 0% in scope" always overstated the missing work by
  ~50×. Verified end to end against the real corpus in a throwaway store:
  100% coverage, `doctor` exit 0.
- **`action.cost_basis`, so a bundled harness reads NULL-with-a-reason.**
  grok runs on a subscription: `cost_usd` stays NULL and `cost_basis` is
  `subscription`. A 0 would read as measured-and-free and make every
  dollar-ranked routing comparison favour grok by construction no matter
  how much rework it caused; NULL alone would read as "we failed to price
  it".
- **The cursor-recovery contract is automatic instead of ritual.** A new
  normalizer reaches nothing on its own — its records sit below the
  normalize watermark — so activating one has always required a manual
  `DELETE` of the cursor. Nobody performs a manual step that ships inside
  a binary upgrade. `db::migrate` now does a one-shot reset keyed by
  *reason* and recorded in `meta`, so it runs exactly once and a store
  created today never replays it.
- **Latent bug: `MODELLED_KIND_SQL`'s arms were not harness-bound**,
  despite a doc comment asserting they were. The enclosing gate only says
  "this harness has *some* normalizer"; it never bound an arm to the
  harness it was written for. It went load-bearing the moment grok gained
  one: **594 records** in grok's `chat_history.jsonl` carry a top-level
  `type` of `assistant`/`user` and matched the claude-code arm verbatim —
  in scope, producing no action, leaving `doctor` raising a permanently
  unclearable `UNMEASURED_HARNESS`. Every arm is now gated by its own
  harness.
- **`brief` refuses instead of hedging below the floor.** Under
  `MIN_SESSIONS` it rendered a confident diagnosis — a named heaviest
  pager, a prescribed medicine — with a caveat underneath saying none of
  it was a measured verdict. Observed emitting a verdict on three
  sessions. The classification is now `INSUFFICIENT`; the lists stay, as a
  prediction.
- **grok's per-response key is `streamStartMs`, not `promptId`.**
  `promptId` is a TURN id; keying SERIAL on it made grok read as batches of
  up to 44 tools — a false cross-model claim, caught before it shipped by a
  second witness that agreed 318/318. With that key, the first cross-model
  batching comparison this tool has been able to make, **re-derived from the
  shipped build against the live store before this release was cut**:

  | harness | responses | 2+ tools | max |
  |---|---:|---:|---:|
  | grok | 316 | **82.9%** | 12 |
  | claude-code | 42,177 | **14.2%** | 24 |

  The development figure for claude-code was 16.6%; re-deriving after the
  duplicate guard and the harness-binding fix moved it to 14.2%, which is
  why it was re-derived rather than transcribed. **Read grok's number with
  its n**: 316 responses against claude-code's 42,177 is a 133x sample
  difference, and grok's 2,225 in-scope actions are 2.0% of what is captured
  for it. The direction is large and the precision is not.
- **Upgrade note:** grok's records enter scope the moment this build lands,
  but they become actions only after a normalize pass runs. In between,
  `doctor` correctly reports grok as `UNMEASURED_HARNESS` (2,225 in scope,
  0 measured). That is the fault working, not a regression — the automatic
  cursor reset above clears it on the next `ratchet run`.
- **`grade` and `brief` no longer disagree.** They reported different
  numbers for the same measure, because `orientation::gather`'s working-set
  medians were unfiltered while `brief` recomputed them from root-filtered
  rows. On a clone of the live store **10 of 11 projects disagreed** —
  tender at 37 files against 18 — and now 0 of 11 do. Every `grade` number
  moved onto `brief`'s rather than the two meeting in the middle: one
  extracted attribution pass now serves both surfaces, so a future
  divergence is structurally unavailable rather than merely tested against.
- **And it got 9.4× faster doing it:** `grade` 2.83 s → **0.30 s**,
  `brief tender` 5.90 s → 2.46 s. The old per-session correlated subselects
  *were* the cost. The trap is worth recording: the first implementation
  was correct and **27× slower** (81.5 s), and it looked fine when timed in
  the `sqlite3` CLI — because that timing ran the query *without the bound
  parameter*, which is the version that does not exist at runtime. Only
  `/usr/bin/time` against the real binary on a real-sized store caught it.
- **`BLIND_HARNESS` split in two, because it was asserting something
  false.** `captured > 0 AND in_scope_pct == 0` cannot tell "no normalizer
  exists" from "a normalizer exists and these records predate its gate" —
  and it claimed the first for both. Of rift's **10,186 records**, every one
  is `v = 1` or has no `v` at all, which `normalize_rift` refuses as
  pre-instrumentation *by design*; the caveat nonetheless read "this build
  cannot read any of them", sending readers off to fix a gap that was not
  there. New **`PRE_INSTRUMENTATION` (info)** covers that case;
  `BLIND_HARNESS` keeps `warning`, now names its own remedy ("the fix is a
  normalizer, not a re-run"), and applies only to antigravity (1,157) and
  gemini (16). grok raises neither.
- **Severity became a rule instead of a habit: `warning` = an action
  exists; `info` = a bounded claim with none.** Nothing anyone does to the
  store can move a pre-instrumentation record into scope, so filing it as a
  warning would be an alarm that can never clear — the class this project
  has already been burnt by twice. Documented publicly, not just observed.
- **`BRIEF_FOREIGN_READS_EXCLUDED` → `FOREIGN_READS_EXCLUDED` and
  `BRIEF_AGENT_FURNITURE_EXCLUDED` → `AGENT_FURNITURE_EXCLUDED`.** Both
  surfaces raise them now, so a prefix naming one surface was the same
  class of lie as the numbers that disagreed. `OrientationData` also gains
  `foreign_reads`, `furniture_reads` and `scope_project`, so what was
  excluded is a field and not only a sentence.
- **One asymmetry is deliberate and stated rather than hidden:** the
  working-set columns exclude reads of other projects' files, while
  `median_pre_edit_readonly_calls` still counts them. The primary endpoint
  measures navigation CALLS — a call spent opening another project's file
  was still spent — but the file it opened tells you nothing about this
  project's working set. Two attributions, on purpose, in one row.

- **A `SessionStart` orientation hook, and the `<ratchet-orient>` packet it
  emits.** `ratchet hook install` writes a user-scope entry into
  `~/.claude/settings.json`; `ratchet hook orient` composes a ~2 KB packet
  from two artifacts something else already produced — the heat row `ratchet
  run` now bakes on every tick, and the atlas manifest — and `ratchet hook
  bake` is the manual refresh. It creates neither artifact: a missing one is
  SILENCE, not a build.

  Why a hook rather than another tool: placement is the only delivery lever
  this fleet has measured working. `arc_context` gets called unprompted
  because it arrives as a protocol-level instruction, while `Grep` — bounded
  and pre-installed — sits at **0.8% against 8,961 shelled-out greps**
  because it is merely an entry in a tool list.

  Bounded (3,200 bytes, ~800 tokens, every section capped individually, whole
  lines dropped cheapest-first if the ceiling ever fires), **silent when there
  is nothing measured to say** — no root, no manifest, no baked heat means no
  output and exit 0 — and fail-silent on a 1,500 ms deadline, because a
  session start is a human waiting. Where a quantity is genuinely underivable
  (import fan-in on a tree whose `use` paths do not resolve to files) it says
  UNMEASURED rather than printing an empty list. A baked row older than 7 days
  is declined: the pipeline re-bakes every 30 minutes, so a week-old row means
  it has been down for a week.

  The heat is baked rather than computed live because `brief::gather` costs
  2.4 s on this machine's 1.2 GB store — two orientation passes scan the whole
  `action` table before the project filter applies. The hook does one
  primary-key lookup instead.

- **The release immutability gate now asks what is PUBLISHED, not what is on
  local disk.** The old gate refused when a local staging directory for the
  version existed — and `release.sh` with no `--deploy` writes exactly that
  directory, because staging IS writing it. So the documented two-step
  workflow poisoned its own second step: on 2026-08-03 it refused to deploy
  v0.3.6 while `/v0.3.6/checksums.txt` was a 404 and `latest.json` still read
  0.3.4. Nothing was published; the gate had noticed the script's own dry run.
  Worse, the only escape it offered was `FORCE_REPUBLISH=1` — the flag whose
  ABSENCE is why v0.1.1 is a cautionary comment rather than a release.

  The replacement asks the immutable path's own `checksums.txt` at the origin
  (not `latest.json`, which names only the NEWEST cut and would read a
  published-but-superseded version as absent), and decides four ways:
  **PUBLISH** (404 — nothing to protect), **IDENTICAL** (published, and the
  rebuilt bytes hash the same — nothing to re-cut and nothing to refuse),
  **DIFFER** (the real violation, and the only thing `FORCE_REPUBLISH` is
  for), **UNKNOWN** (undetermined — fails closed, because "not published" and
  "could not check" are different facts and must never share a branch). This
  is the *same version, same bytes* rule the atlas half already used,
  extracted into `scripts/release-guard.sh` so both halves share ONE
  implementation — and so every branch has a test. The old gate had none,
  which is how it shipped describing a condition it was not testing.

  **Root docs are unversioned and redeploy freely at the same version.**
  `llms.txt`, `RATCHET.md`, `ATLAS.md`, `CHANGELOG.md`, `AGENTS.block.md`,
  `ATLAS.block.md`, `RESULTS.md`, `install.sh` and `latest.json` are not
  content-addressed and no checksum names them; correcting documentation for a
  shipped release is not a re-cut, and it must never require the flag that
  exists for deliberately overwriting published binary bytes. The DIFFER
  refusal now says so in the refusal itself.

  One consequence is load-bearing rather than tidy: on IDENTICAL, a version
  directory missing LOCALLY is now RESTORED from the reproducible build
  (provably the published bytes, because the hash matched). `wrangler deploy`
  uploads `public/` wholesale, so a locally-absent version directory is a
  version directory deleted at the edge — which is exactly how v0.1.1 stopped
  resolving for people who already had it.

  And a new `scripts/release.sh --docs-only [--deploy]` publishes the root docs
  and nothing else — no build, no packaging, no version directory, no manifest
  rewrite. Once the source moves past a published version, the gate is RIGHT to
  refuse a re-cut, and a doc correction must not be held hostage to a version
  bump it has nothing to do with. It is the most cautious path in the file
  because of that wholesale upload: it refuses unless the local `latest.json`
  is byte-identical to the live one (changing which version is current is a
  release, not a doc fix), and unless every artifact that manifest NAMES is
  present locally and hashes to what the origin serves. What it cannot check —
  version directories `latest.json` no longer names, since the edge cannot be
  enumerated — it says out loud rather than implying.

- **`ATLAS.md` — atlas's full agent reference, peer to `RATCHET.md`.** atlas
  shipped a paste block and no reference. It documents the one-door form, the
  verbs and when knowing the shape is worth it, the
  `SEARCHED`/`NOT SEARCHED`/`NOT MODELLED` footer that is what lets a negative
  be trusted, `NOT-INDEXED` as a distinct verdict, the self-building
  out-of-tree manifest and its 2,000-file automatic ceiling, `--db` as an
  override atlas never auto-writes, the per-pointer staleness verdicts, and
  the argparse quirk that a query starting with `-` cannot be passed bare.

  And **the CWD trap**, measured live: a user-scope MCP server inherits the
  working directory of the CLIENT process that spawned it — the directory
  `claude` was launched from, not the project. A session launched from `$HOME`
  returned `no_root` for every `atlas_lookup` call, correctly reporting that
  it had searched nothing while the real project sat two directories away.
  `path` exists for exactly that case.

- **`RATCHET.md` and `llms.txt` gain the whole lifecycle surface** — `wire`,
  `doctor` (both halves of its fault vocabulary, install and pipeline),
  `remove [--purge]`, `agents-block [--write]`, `update` (including that it
  now refuses to downgrade), `hook orient|install|bake`, and `brief` with its
  `INSUFFICIENT` refusal. The `<ratchet-orient>` packet is now discoverable
  from the reference docs rather than only from `AGENTS.block.md`. Measured
  before this change: `hook` appeared 0 times in either document.
- 541 tests.

## v0.3.5 — 2026-08-02

Honest durations. The meter was measuring the wrong thing, and it was
measuring it at the top of every ranking.

- **Duration provenance.** `action` gains `duration_source`. Of 51,104
  timed claude-code `tool_result` rows, **113 are harness-stated and
  50,991 were derived** from the `tool_use → tool_result` gap — and that
  gap includes however long a permission prompt sat unanswered. The 113
  are all `WebFetch` and `Glob`; `Bash`, `Read` and `Edit` never state a
  duration, so every tool the WAIT tier ranked was timed by approval
  latency. Proof was arithmetic: the top three "grep" waits were ~9,000 s
  each against the Bash tool's 600 s execution ceiling, and `Edit` — which
  completes in milliseconds — carried a 32.7 s mean.
- **The correction, same corpus, old binary vs new:** `Bash grep` #1 at
  8.86h → **#4 at 0.62h**. `Bash find` 3.71h → 0.46h. `node`/`npm` left
  the top 8 entirely. **Median grep is 197 ms.** grep was never slow. If
  you carried "prefer targeted tools, grep is expensive" as a standing
  rule, it was sourced from this artifact — drop it.
- **SERIAL's premise was itself an artifact.** Its doc comment claimed
  zero batched turns, "verified three independent ways (uuid, parent_uuid,
  raw_record_id)". Claude Code writes one `tool_use` block per JSONL line,
  so all three keys are blind to a parallel batch **by construction** —
  one witness wearing three hats. The real key is `requestId`, which
  parallel calls share. Actual rate: **83.4% of model responses issue one
  tool, 16.6% batch 2 or more** (max 24). SERIAL is still real; its
  magnitude was not. *(Superseded by v0.3.6: re-derived after the duplicate
  guard and the harness-binding fix, the rate is 85.8% / **14.2%**. The
  16.6% here is what this release measured, kept as the record of what it
  claimed.)*
- **Cross-file duplicate ingestion.** The dedup index keyed on
  `raw_record_id`, so one session line reachable by two paths survived
  twice: 5,937 duplicate groups, 5,936 spanning different source files,
  **~2.9% inflation on every count**. Fixed with a `WHERE NOT EXISTS`
  identity guard — not a UNIQUE index (uncreatable over data already
  violating it) and not a DELETE (a destructive fix for a counting bug).
- Migration backfills `duration_source` over existing rows. This was
  required, not optional: new columns alone leave 208k rows NULL forever,
  because normalize's cursor skips seen records and the new duplicate
  guard makes a forced rescan a no-op.
- `docs/handoff/STATE.md` squared with what measurement showed, with
  inline supersession markers at three overturned claims so a reader
  landing mid-file cannot mistake a retracted finding for a live one.
- 298 tests.

## v0.3.4 — 2026-08-01

The split-proof measure.

- `ratchet grade` gains a **working set** section: median distinct files
  read and total reads per editing session, with the derived reads-per-file
  ratio. It counts the knowledge a unit of work required, deliberately
  independent of how that knowledge is laid out on disk.
- Why it had to exist: every layout metric is gameable. Splitting a
  monolith silences PAGING while the agent still loads the same knowledge —
  and now has to find which of the twenty files it lives in. Measured on
  the reference corpus, a well-decomposed project loaded **53 distinct
  files per editing session** against a monolithic project's 25, and
  carried the worse orientation tax of the two. Fragmentation moves the
  cost from the ratio into the file count; merging does the reverse; only
  cohesion — a decomposition matching task boundaries — lowers both. That
  is the one movement this measure rewards, and no refactor can fake it.
- 277 tests.

## v0.3.3 — 2026-08-01

REPEAT splits: redundancy and paging are different diseases.

- The re-read detector now compares the stored input digest of consecutive
  reads. Same region twice is **REPEAT** — true redundancy, agent
  behaviour. Different regions is **PAGING** — a file too large to read in
  one call, and the shape says plainly that it measures the *file*, not the
  agent. Ties go to PAGING deliberately: mislabelling structural cost as
  redundancy prescribes "remember harder" to agents who forgot nothing.
- Why it mattered: on the reference corpus, **37% of what the top signature
  reported was paging** — and on the four largest hot files, 96%. Four
  projects were about to receive orientation maps for a problem that was
  never orientation. The prescription is now an index or a split.
- Evidence rows carry both pair counts, so a caller can see a mixed streak
  rather than trusting the majority label.
- 275 tests.

## v0.3.2 — 2026-08-01

Codex modelled; sweep retention; the action loop closed.

- **Codex normalizer**: rollout records become actions — measured token
  identities (cached input subtracted so cache reads are never
  double-billed; reasoning not folded), model tracked per turn (mid-session
  switches respected), exec command targets extracted for verb
  disaggregation, OpenAI models priced `no_rate_table` (visible, never a
  fault). First activation used the tested recovery contract: cursor reset,
  full idempotent rescan, zero duplicates — and moved the store's time
  floor from June back to **March 2026**.
- **Sweep retention**: superseded re-detection snapshots are released each
  run (keeping the latest 4 sweeps + each day's last as historical
  anchors). First firing reclaimed 118k signature rows; monthly growth
  drops ~48×. Amends the signatures-permanent ruling on its own logic:
  actions are evidence, sweeps are cache.
- `ratchet grade`'s data now has fleet companions: the `ratchet-advisor`
  and `recon-cache` skills close the measure → reason → adjust → re-measure
  loop (first instance live, carrying a falsifiable orientation thesis).
- 271 tests.

## v0.3.1 — 2026-08-01

Capture-first harness expansion.

- Four new transcript sources captured raw before their rotation windows
  destroy history: Codex (`~/.codex/sessions` rollouts — recovering session
  history back to March 2026), Grok (`~/.grok/sessions`), Antigravity
  (`transcript_full.jsonl` only — the subset sibling is excluded by name to
  prevent double-counting), and Gemini CLI (`chats/session-*.jsonl`).
  Normalizers come later; the raw tier preserves verbatim evidence now, and
  capture-only harnesses read BLIND_HARNESS (captured, 0% in scope) — the
  designed truthful state, never a fault.
- `normalize` gained a known-harness gate with a tested cursor-recovery
  contract: when a harness gains a normalizer, list it and reset the cursor
  once — the rescan is idempotent by the dedup index.
- The in-scope predicate gained the third clause of its sync duty: a kind
  match on a harness with no normalizer is not a modelling claim. (Grok's
  54k captured records use overlapping kind vocabulary; without the gate
  they false-faulted the doctor within minutes of capture — caught by the
  instrument's own honesty machinery, fixed the same hour.)
- Legacy pre-rename store fallbacks retired.
- 255 tests.

## v0.3.0 — 2026-08-01

The orientation instrument.

- `ratchet grade` / `ratchet query orientation` / MCP `ratchet_query_orientation`
  — per-project medians of what agents spend BEFORE their first edit. Primary
  endpoint is navigation intensity (pre-edit read-only calls, pre-edit
  re-search density); tokens and dollars are secondary and labeled
  model-conditioned. Editing sessions only, censoring disclosed (no-edit
  session rate), ≥5-session floor with dropped projects counted, post-edit
  rework medians bundled on every row (anti-Goodhart), alphabetical rows with
  a test-asserted guarantee that no field is a rank — this is an instrument
  for measuring a project against itself over time, never a league table.
  Structural caveats: TASK_MIX (always), MODEL_MIX, CENSORED_SESSIONS,
  ORIENTATION_NO_SWEEP.
- Field results page: [RESULTS.md](/RESULTS.md) — what the meter has caught,
  stated in portable-tier vocabulary.
- Legacy pre-rename store fallbacks retired (both machines migrated); the
  file-wise migration recipe survives as a history note.
- Release guard extended by v0.2.x's doc-rev rule: this entry exists because
  the release script refuses to cut without it.
- 245 tests.

## v0.2.3 — 2026-08-01

- REPEAT is compaction-aware: a re-read that follows a context-compaction
  boundary in the same session is rehydration, not friction. Boundaries are
  read from the raw transcript record (`type=system`,
  `subtype=compact_boundary`); pre-boundary repeats stand, post-boundary
  re-re-reads flag fresh. Measured on the reference corpus: compaction
  accounted for ~5 of 2,058 REPEAT streaks — the re-read churn is genuine.
- 234 tests.

## v0.2.2 — 2026-08-01

- The in-scope predicate (`MODELLED_KIND_SQL`) now mirrors the Rift
  normalizer's `v >= 2` gate: pre-instrumentation Rift records are
  captured-but-out-of-scope instead of permanently "unmeasured", so the
  UNMEASURED_HARNESS fault introduced in v0.2.1 cannot false-fire forever on
  stores carrying pre-v2 Rift history. Coverage numbers become honest as a
  side effect.
- 230 tests.

## v0.2.1 — 2026-08-01

Field-test round: fixes from a self-dogfood and an independent agent test on
a second machine.

- `command_verb` treats leading `echo`/`printf` as banner-preamble: an
  echo-led compound attributes to the downstream verb; terminal
  `echo x > file` remains real work. ("Bash echo" had ranked #3 by time on
  borrowed cost.)
- The `Agent` tool is excluded from WAIT — a subagent call's duration is an
  entire delegated session, not recoverable wait.
- Every agent-surface JSON refusal (`no_sweep`, `unknown_sweep`) exits
  non-zero; exit codes are API for headless callers.
- `--kind` is validated against the closed five-detector vocabulary — a
  bogus kind is refused loudly instead of returning an empty envelope that
  reads as "no friction".
- UNMEASURED_HARNESS: a harness with in-scope records and zero measured is a
  shared caveat on every surface AND a `doctor` fault.
- Shell-variable commands (`"$ARC"`, `${RATCHET}`) fold onto the lowercased
  variable name, ending split verb shapes.
- `ratchet run` reports `findings N (M new)` — created and refreshed counts
  split.
- 229 tests.

## v0.2.0 — 2026-08-01

The agent surface. "SQLite is the query surface" retired.

- Envelope contract: every JSON output is `{meta, ...}` — both coverage
  percentages, the occurrence floor, and caveats as structured
  `{code, severity, detail}` objects on every machine-readable response.
- `--json` on `report`, `doctor` (with a structural `faults` array matching
  the exit code), and `hotspots` (full table, per-row `confidence`).
- `ratchet query hotspots|signatures|coverage|trends` — filtered, ranked,
  truncation-announcing structured queries; signature rows carry parsed
  evidence and anchor timestamps.
- `ratchet trends` — calendar-period friction on the anchor-action time
  axis, per-period per-harness coverage beside every bucket, and deltas
  gated by `comparable` + a structured reason (a month where ingest was down
  must not read as a month friction dropped).
- The `finding` table is live: shapes persisting across consecutive sweeps
  promote into fleet-portable claims, privacy-gated at write time (no paths,
  no machine ids, MCP server segments reduced to `mcp:<tool>`, fail-closed
  refusal counter). `ratchet findings contest <id> --reason` freezes a
  disproved claim.
- `ratchet mcp-server` — hand-rolled stdio MCP (zero new dependencies):
  `ratchet_context`, `ratchet_query_*`, `ratchet_findings`,
  `ratchet_status`.
- Report gains Trends and Findings sections rendered from the same structs
  the agent surfaces emit.
- NULL-never-0 extended to durations: unmeasured `total_cost_seconds` is
  null with `confidence: "not_measured"`, never a zero that reads as
  "instant".
- 219 tests.

## v0.1.2 — 2026-07-31

- Release pipeline hardening: immutable published versions (re-cutting a
  published version is refused before any destructive step), reproducible
  tarballs (`gzip -n` + fixed mtime), missing checksum entries are hard
  failures in both `install.sh` and `ratchet update`, and the deploy gate
  downloads the served artifact and verifies it against the served checksum.
- Store migration notes corrected (file-wise `mv`, WAL siblings travel with
  the db).
- 138 tests.

## v0.1.1 — withdrawn

Destroyed by re-running the release script before the immutability guard
existed; permanently 404. The guard in v0.1.2 is this version's legacy.

## v0.1.0 — 2026-07-31

Initial public release: durable incremental ingest of Claude Code and Rift
transcripts into SQLite, date-keyed immutable pricing with cache-split
accounting, five friction detectors (WAIT / REPEAT / REPHRASE / SERIAL /
CEREMONY), ranked hotspots, `doctor` health contract, launchd scheduling,
`curl | sh` installer with checksum verification. 124 tests.
