# Field results — what the meter has caught

Real findings from ratchet observing live agent work. Everything here is
stated the way the portable tier states it: counts, ratios, and shape
vocabulary — no file paths, no machine identifiers, no third-party project
names. Each claim carries its own evidence, and where a claim was later
corrected, the correction is part of the story: the meter is only worth
trusting because it distrusts itself first.

## The meter indicts its own clock (v0.3.5–0.3.6)

The largest correction the tool has made is to its own headline number. Of
~51,000 timed calls in the reference corpus, **113 carried a harness-stated
duration**; every other "duration" was derived from the elapsed call→result
gap — which includes however long a permission prompt sat unanswered. The
proof was arithmetic, not statistical: the top three "slow search" waits
measured ~9,000 seconds each against the harness's own 600-second execution
ceiling for that tool, and an edit operation that completes in milliseconds
carried a 32.7-second mean.

The correction, same corpus, old binary against new: the #1 "slow tool" fell
from 8.9 hours to 37 minutes (rank #4), the runner-up from 3.7 hours to 28
minutes, and **the median search call was 197 ms all along**. Nothing was
dropped or silently re-ranked: every surviving figure now carries
`confidence: "derived_upper_bound"` — a ceiling on machine time, never an
execution cost — plus a store-wide caveat stating the derived fraction. If an
agent on this fleet carried "search is expensive, prefer targeted lookups" as
a standing rule, it was sourced from this artifact, and it is retracted.

The same audit overturned a detector's own premise: "zero batched tool
calls, verified three independent ways" turned out to be one witness wearing
three hats — all three keys were blind to parallel batches by construction.
The real key showed ~14% of model responses batching 2 or more calls. The
signature (SERIAL) is still real; its magnitude was not, and the record now
says which.

## Newest: the loop runs end to end (v0.3.x)

Three capabilities landed in a single day, each proven live before release:

- **The orientation instrument** (`ratchet grade`): what agents spend
  *before their first edit*, per project — navigation intensity as the
  primary endpoint, censoring disclosed, rework bundled, deliberately no
  rank fields (an instrument for measuring a project against itself, never
  a league table). Verified non-hardcoded by independent recomputation.
- **The recon cache**: an orientation doc generated from measured desire
  paths (the files agents re-read, the commands they mis-fire), every
  claim verified against real files and keyed to the paths it rests on. The
  first instance survived an adversarial review that forced five rounds of
  edits — recipes over directory tours, failed→succeeded trap pairs — and
  now carries a falsifiable thesis the next orientation read will judge.
- **History recovery + honest expansion**: six harnesses captured, three
  modelled. One harness's archive survived where another's 30-day rotation
  had already destroyed spring — first ingest moved the store's time floor
  back three and a half months. Capture-only harnesses read as BLIND
  (captured, honestly unreadable) — and when 54k newly-captured records
  used lookalike vocabulary, the false "unmeasured" alarm they triggered
  was caught by the meter's own machinery and gated the same hour.
- **Retention that respects evidence**: superseded re-detection snapshots
  (fully re-derivable from permanent actions) are now released each run —
  first firing reclaimed 118k rows, cutting derived-data growth ~48× —
  while raw evidence remains untouchable by design.

## The month's biggest catch: two diseases wearing one name

The top signature by volume — files read again with no edit between — had a
behavioural name and a behavioural prescription: batch your reads, remember
what you loaded. A forensic pass asked a question the count could not
answer: *were the re-reads of the same region, or different ones?* The
store already held an input digest for every call, so it was one query.

**37% of them were different regions** — agents paging through files too
large to read in one call. On the four largest hot files, 96%. Those agents
had forgotten nothing; the files simply did not fit. Four projects were
queued to receive orientation maps for a problem that was never
orientation.

The detector now emits two shapes. Same region twice is redundancy, and the
advice stands. Different regions is *paging*, and the shape says plainly
that it measures the file, not the agent — because the fix is an index or a
split, and telling a correctly-behaving agent to "remember harder" is how
an instrument turns into a scold. Ties break toward paging on purpose:
between blaming a file and blaming a worker, the honest default is the one
that cannot shame someone for doing it right.

## Worked example: one advisor pass over one month

What a single "find my problem areas" run produced, verbatim shapes and
counts (projects anonymized):

- **The volume king was process, not code.** Of ~2,000 repeat-read
  signatures, the most re-read individual files turned out to be *task
  brief documents* — 115 no-edit re-reads in a month. Agents were
  re-consulting their own specs every few turns instead of holding them.
  Adjustment: one behavioral line where every session reads it ("read the
  brief once, extract the criteria"). Judged by: repeat-read count on
  brief-shaped files next period.
- **Four sibling projects held ~45% of all repeat-reads**, centered on
  three source files — a shared-structure legibility gap, queued for the
  recon-cache treatment. Judged by: those projects' orientation medians.
- **Search waits topped time again (8.9h) but flagged themselves
  `unverified`** — the mean sits at the plausibility threshold, so the
  surface itself says "verify before acting." The honest instrument
  declined to oversell its own headline number — and the follow-up
  measurement killed it: that 8.9h was approval latency, not execution
  (see "The meter indicts its own clock" above). **Retracted as machine
  time.** The flag existed so the number would not be acted on before it
  was verified; it wasn't, and it did not survive verification.
- **A previously-filed tool latency held its ranking** (3.0h across 114
  calls) — already instrumented upstream; the next complete period's
  comparable-gated delta renders the verdict.
- Search *churn* (REPHRASE), meanwhile, measured at 6 occurrences —
  effectively solved, and the meter is allowed to say something is fine.

## The tool caught its own detector lying

On its first self-audit, ratchet's #3 hotspot by time was "Bash echo" —
4.3 hours across 103 occurrences. Pulling the underlying evidence rows
showed 13 of 15 echo-led WAIT findings were compound commands where `echo`
printed a banner and the real cost ran downstream (`echo "=== hdr ===" &&
<45-minute probe loop>`). The verb extractor was attributing borrowed time.
One fix later, "Bash echo" left the rankings and its hours redistributed to
the verbs that actually earned them. A ranking that can indict an innocent
verb needs exactly this kind of cross-examination — the evidence rows made
it a ten-minute investigation instead of a debate.

## A detector learned what "waiting" means

WAIT flags calls far above their own tool's median duration. Its top
finding was the subagent-delegation tool at 9.3 hours across 30 calls —
correctly flagged `unverified` by the plausibility check, and still wrong
to count: a delegated subagent session is a machine *working*, not a
machine being waited on. The detector's contract was sharpened to "time a
machine could have given back," and the phantom hours vanished. The
plausibility flag did its job: it marked the number untrustworthy before
anyone acted on it.

## An 876-second write, measured from outside

The single largest per-call waits in the corpus were intent-completion
writes to a companion versioning tool — 110 to 876 seconds per call,
totaling hours across a working day. Ratchet measures end-to-end from the
caller's side; the finding was filed upstream, the tool's maintainers
instrumented their write path within a day (decomposing queue wait,
execution, and pre-hook serialization), and shipped fixes. The whole-call
number — the one the agent actually waits on — gets its verdict from the
next full month's trend comparison, gated for comparability. Outside-in
measurement forced the question; inside-out instrumentation answered it;
neither can quietly agree with the other's mistake.

## A design debate ended by a number

Are re-reads after context compaction legitimate rehydration or real
friction? Reasonable people could argue either way. The detector was made
compaction-aware and the corpus answered: 5 of 2,058 repeat-read streaks
spanned a compaction boundary — about 0.24%. The re-read churn is genuine
behavior, and the remaining 2,053 are a real target. The debate took a day;
the measurement took one sweep.

## An alarm that caught a gap, then caught itself

A field test on a second machine found a harness whose records were
captured and modellable but never measured — invisible to the global
coverage number (a small harness drowns in a large one), reported clean by
the health check. A per-harness fault was added; it immediately fired on
the reference machine too. Then it revealed its own defect: hundreds of
those records predated the harness's instrumentation format and were
unmeasurable *by design* — an alarm that could never clear. The in-scope
predicate was taught the version gate, the false fault went quiet, and the
real one still fires. An alarm that cannot read both good and bad is not
an alarm; it is a decoration.

## Orientation tax: a pilot worth repeating

Across eleven projects on one machine with the same models, the median
spend *before an agent's first edit* — the cost of walking into a codebase
cold — spanned roughly half an order of magnitude, with the extremes
separating cleanly (tens of exploratory calls and tens of thousands of
reasoning tokens versus a handful of each). The ratio of exploration calls
to reasoning tokens separates "hard to find" from "hard problem." Pilot
data, hypothesis-generating, deliberately not a league table: the
defensible instrument is a project measured against itself over time,
after an orientation intervention, behind comparability gates. What the
pilot does establish: on identical models, the cost of entering a project
varies enormously — which is worth remembering the next time the model
gets blamed.

## An independent test on a second machine, same day

An agent with no involvement in the build was pointed at the tool cold. It
verified the pipeline honestly (and said so), then found four real defects
— refusals that exited zero, an unvalidated filter vocabulary that returned
silent empties, the unmeasured-harness blindness above, and verb shapes
fragmented by shell variables. All four were fixed and released the same
day. The tool's own telemetry recorded the session that tested it.

---

## How to read these honestly

Every number above came with structured caveats attached — plausibility
flags on implausible durations, floors reporting what was dropped,
comparability gates refusing deltas across unequal capture. That is the
product: not the findings, but the discipline that keeps findings honest.
NULL never becomes zero; absence never reads as health; an unverifiable
number says so on its face.

ratchet is a machine-local observer — transcripts never leave the machine.
Install: `curl -fsSL https://ratchet.daystra.com/install.sh | sh` · Full
reference: [RATCHET.md](/RATCHET.md) · Agent quick-start:
[llms.txt](/llms.txt) · History: [CHANGELOG.md](/CHANGELOG.md)
