# ATLAS.md — the full agent reference

atlas is a manifest of pointers for the codebase you are standing in:
declarations, name substrings, file paths, path routes, endpoints, config
keys, module registry. One stdlib-only Python file, no install step of its
own, no service, no index checked into your tree.

It exists to change the KIND of answer a lookup gives you, not the speed of
one. Three properties are the whole pitch, and each is something you can
check:

- **A hit carries `file:line`, verified against the file on disk at answer
  time.** A pointer that moved since indexing says so in its own `note`
  (`[moved since index, was :3]`). A confident wrong line is the one failure
  mode that costs more than no answer at all, so it is disclosed rather than
  believed.
- **An empty answer is EMPTY, not "this tool did not find it" — and so is a
  thin one.** The call has already fallen back to a live whole-word search of
  the tree whenever no index holds the term ITSELF (0.1.3), so a neighbour
  like `idle_timeout_triggers` never stands in for `idle_timeout`. There is
  no grep left for you to run.
- **Every answer ends with a footer bounding the negative** — what was
  searched, what was not, and what this build does not model at all. Absence
  inside a searched index is absence. Absence inside an unmodelled one is
  nothing.

### What it costs, measured — because a pitch without its cost is marketing

First field measurement of the claim above: **a grok session on the MBA,
2026-08-17, during rift work** — a route/SEO scan. It is reported here
**against** atlas wherever it goes that way, which is why it is worth
believing:

- **atlas is SLOWER per lookup, not faster.** 16 `rg` calls ≈ **175 ms**;
  16 atlas calls ≈ **3 s** of work, ~1.5–2 s wall when batched. The index is
  not where the win is, and anyone selling atlas on lookup speed is selling
  the wrong thing.
- **On absences it costs MORE tokens than an empty grep** — roughly
  **1.6–2.3 KB** of envelope and footer to say "nothing", against a few bytes
  for an empty `rg`. **This is the trade working as designed, not a
  regression:** an empty result with a `searched` / `not_modelled` footer is
  *evidence*; an empty grep is a shrug. Paying two kilobytes to turn "I found
  nothing" into "nothing is there, and here is what I looked in" is the whole
  point.
- **On well-posed names** (`ROUTE_META`, `sitemap`, `vlab`) atlas and a
  careful grep land within a few thousand tokens of each other. No advantage
  claimed.
- **Where it actually pays: the hunt, not the read.** ~**15–25 k tokens**
  avoided in that session — the wrong-guess search explosions that never
  happened, plus ~15–20 neighbour files listed and correctly skipped. Roughly
  **1–3 minutes off a ~6-minute scan**. The saving is **not doing 2–4 refine
  greps and 8–15 extra file reads in series** — the loop that eats the first
  half of a session.
- **It does not replace reading the files that matter.** The same session
  still opened `app-routes.js`, `prerender.js`, `route-schema.js` and should
  have. atlas replaces the *hunt* for them.

**If you have seen a six-figure token-saving figure for atlas, it is not
supported by anything measured here** — the honest order of magnitude is tens
of thousands per session, and the mechanism is a deleted refine loop rather
than a cheaper lookup.

**Caveats on the magnitudes, and they are not small.** One session, one tree,
one machine, **and one harness that is not the one most readers of this file
are running.**

- **Cross-machine.** Measured on the MBA, not the reference machine. Time is
  not portable across machines here — measured divergences run 3×–17× — so
  the seconds above describe that box.
- **Cross-harness, which is the caveat I would stress.** These are *grok*
  numbers. Batching is harness-specific and the wall-clock figure explicitly
  depends on it ("~1.5–2 s wall, they were batched"), so a harness that issues
  calls serially would see a different wall time from the identical work.
  ratchet's own corpus shows how far apart harness record-shapes are — grok
  emits 167,335 `phase_changed` records against 4,206 in-scope ones — and
  token accounting differs alongside.
- **Version: v0.1.7, confirmed.** Every atlas payload in that session carried
  `"atlas_version": "0.1.7"` — **17 of 17**, the same fleet cut. It was a
  *reporting* miss in the savings recap, not an unknown. Cite the build with
  the numbers next time; a measurement whose subject is unnamed is a
  measurement nobody can reproduce.

### What travels and what does not

Settled with the measuring agent rather than asserted here:

| claim | travels? | why |
|---|---|---|
| slower per call than `rg`; pays by deleting the refine loop | **yes** | structural |
| dearer on absences (the envelope tax) | **yes** | atlas payload shape |
| the wrong-guess explosions are where the tokens go | **yes** | tree + whole-word vs substring |
| **15–25 k tokens · 175 ms · 1–3 minutes** | **no** | grok + batching + that agent's judgment about what it would otherwise have dumped |

**The payload ratio is a tree-and-version fact, not a harness one:** ~**55 kB**
of atlas output against ~**1 MB** of raw `rg` and ~**336 kB** capped. That is
the mechanism behind the token claim, and it is reproducible by a script with
no agent in the loop.

**How NOT to re-measure this.** Replaying that session's query list under
another harness would copy one agent's questions and understate the other's own
grep habits — agents only run the searches they already know how to phrase, so
a staged replay structurally undercounts. If you want the harness axis,
instrument **one real hunt** as it happens and compare **refine-loop count**,
not a re-run of someone else's term list.

Ships with ratchet, independently versioned. Current: see
https://ratchet.daystra.com/CHANGELOG.md · paste block at
https://ratchet.daystra.com/ATLAS.block.md · ratchet's own reference at
https://ratchet.daystra.com/RATCHET.md

## Install

```
curl -fsSL https://ratchet.daystra.com/install.sh | sh
```

That installs TWO tools. atlas lands at `~/.ratchet/bin/atlas` and is
registered as its own MCP server (`atlas`, one tool). Pin it alone with
`ATLAS_VERSION=`; opt out with `RATCHET_NO_ATLAS=1`. `ratchet wire` registers
it, `ratchet doctor` checks it (`ATLAS_BINARY_MISSING`,
`ATLAS_MCP_UNREGISTERED`, `ATLAS_MCP_STALE_PATH`), `ratchet remove` takes it
away with everything else.

The command is **safe to re-run**: the PATH line is marker-guarded and skipped
outright when `$PATH` already contains the bin directory, the MCP registration
is a verified splice that writes nothing when it already matches, and launchd
is only reloaded when the plist actually changed.

atlas is versioned separately from ratchet on purpose: atlas moved three times
on 2026-08-03 while ratchet moved once. A shared number would force a ratchet
release per atlas fix and — worse — let a ratchet release silently claim an
atlas change.

### Upgrading

```
ratchet update
```

Since **ratchet v0.3.8** that carries atlas too, on its own version line, from
the same `latest.json` the installer reads — installing it if absent,
upgrading it if behind, saying so if current, and refusing a downgrade
independently of ratchet's. One line reports both:

```
ratchet updated: 0.3.7 -> 0.3.8; atlas updated: 0.1.0 -> 0.1.1
```

**On ratchet v0.3.7 and earlier it did not.** `update` moved ratchet alone and
printed nothing about atlas — no error, no warning, no mention — so a machine
that installed once and has only run `ratchet update` since is still on
whatever atlas was current on its install day. That was measured live on
2026-08-03: published atlas 0.1.1, installed atlas 0.1.0, and `ratchet update`
reporting complete success. If you are coming from 0.3.7, run `ratchet update`
once more after it lands, or re-run `install.sh` — which is safe to re-run and
has always handled both tools.

## Day to day

The whole working loop, in four lines:

1. **`atlas <word>` from anywhere inside the project.** No flags, no setup, no
   verb. The project root is walked up from your cwd; you never point it at
   anything.
2. **The manifest looks after itself.** It is built on your first query in a
   tree and rebuilt when the tree has moved on. There is no index in your repo
   and nothing to schedule.
3. **Give it a topic word when you do not know the name; give it the symbol
   when you do.** Both go through the same door — `atlas checkout`,
   `atlas LaunchData`, `atlas /api/v1/session`, `atlas DATABASE_URL` are all
   the same call.
4. **Read the footer before you conclude something is absent.** Every answer
   ends with `SEARCHED` / `NOT SEARCHED` / `NOT MODELLED`. An empty result
   under `SEARCHED` is real absence; an empty result where the footer says
   `NOT SEARCHED code bodies` just means the next call is `atlas refs <name>`.

Everything below is the reference for when one of those four is not enough.

## One door

```
atlas <word>
```

Any word. A symbol, a topic, an endpoint, a config key, a path or path
fragment, a CSS literal. **There is no verb to choose and nothing to aim.**
The project root is walked up from the cwd, the manifest is found or built,
every index is consulted, and the results come back kind-labelled.

Qualified names resolve too (0.1.3): `atlas DaemonCommand::CaptureNow` finds
the arm through its parent, and the bare `CaptureNow` finds it by name —
Rust enum variants are indexed as declarations (kind `variant`), because
command, state and error enums are how this fleet's Rust is written and the
variant is the handle a reader arrives with. The same rule answers
`Widget.render` for Python.

The verbless form is not sugar over the verbs — it is the measured default.
Two agents audited the same cmi5 implementation on the same day, one blind and
one with atlas, and issued the same number of tool calls (34 each). The atlas
arm reached orientation in 4 calls against 9 — and then spent roughly 3 of its
13 atlas calls learning only that it had guessed the wrong verb:

```
where LMS.LaunchData  -> NOT-INDEXED   (refs held the answer)
wire /cmi5            -> "no endpoint" (a 7-resource surface existed)
deps <path>           -> byte-identical to `file`
```

A verb menu makes the caller commit to the SHAPE of the answer before it has
the information that determines the shape. **n = 2 agents, 1 codebase, 1 day**
— a shape worth acting on, not a rate to quote.

## The verbs, and when knowing the shape is worth it

The specific verbs stay for when you already know what shape you want:
scriptable, and no routing to pay for. Reach for one only when the shape is
genuinely known in advance.

| verb | the question | when it beats the one door |
|---|---|---|
| `atlas where <name>` | where is this DEFINED | you have an exact declared name and want only definitions |
| `atlas refs <name>` | where is it USED | the one door said `NOT SEARCHED code bodies` — this greps the tree live |
| `atlas file <path>` | one file's summary, both edge directions | you already have the path and want its imports and importers |
| `atlas config <KEY>` | resolve an env/config key across layers | you know it is a config key and want every layer that sets it |
| `atlas wire <endpoint\|/path>` | client → handler → policy, or the route that serves a path | you know it is an endpoint, not a symbol |
| `atlas stale` | which pointers drifted since the build | you are auditing the manifest itself (see below) |
| `atlas stats` | what is in the manifest | you want to know whether a negative came from a thin index |
| `atlas build [root]` | index a tree by hand | a directory with NO repository marker; no file ceiling applies |
| `atlas mcp-server` | speak MCP over stdio | how the MCP server is launched |

`--also <manifest,manifest>` searches sibling manifests alongside this one —
the answer to "I need a dependency's symbols" that does not involve indexing
`node_modules`.

`atlas file <path>` matches on a path boundary (0.1.3): `file text.rs` no
longer returns a record for `context.rs`, and a tie — nine files named
`mod.rs` — is disclosed with the alternatives listed rather than answered
with an arbitrary one. Ambiguity is information.

## The footer — why a negative here is worth acting on

Every answer ends with three lists. They are the half of the answer that
bounds it:

```
SEARCHED      declarations · name substrings · file paths · path routes ·
              endpoints · config keys · module registry
NOT SEARCHED  code bodies — `atlas refs <name>` greps the tree live
NOT MODELLED  SQL predicates · localStorage keys ·
              response-field -> client-state · dynamic dispatch
```

- **`SEARCHED`** — absence here is absence. This is the only line that lets a
  negative mean anything.
- **`NOT SEARCHED`** — a real index that this call skipped, and it names its
  own fix. The common one is `code bodies`; the fix is `atlas refs`.
- **`NOT MODELLED`** — this build has no representation of these at all.
  Absence here is not evidence of anything. A symbol reached only through
  dynamic dispatch, or a key that exists only inside a SQL string, will not
  appear and its non-appearance means nothing.

**Absent is never zero.** A thin answer over an empty manifest is not the same
fact as a thin answer over 1,095 indexed symbols, which is why the MCP
envelope carries `meta.indexed {files, symbols, built_at}` and `atlas stats`
exists on the CLI.

## `NOT-INDEXED` is its own verdict

```
NOT-INDEXED (absence here is not absence in the code)
```

This is not an empty result. It means the question fell outside what was
indexed, and the two must never be read as the same thing. An empty result has
already been checked against the live tree; `NOT-INDEXED` has not been checked
at all.

## A thin answer is gated like an empty one (0.1.3)

The live tree search runs whenever no index holds the term ITSELF — not only
when everything came back empty. Measured in the field: an agent asked for
`idle_timeout` and got two rows, `idle_timeout_triggers` and
`idle_timeout_orphan_from_trigger`, while one file held thirteen occurrences
of the term including the `sleep(idle_timeout)` the whole investigation
turned on — a local `let`, which the index structurally cannot hold. Two
neighbours were enough to suppress the live pass, and the question was never
answered. The gate now asks "did we find the thing you NAMED?", never "did
anything match?" — and a count threshold ("fall through under N hits") was
refused deliberately, because exactness is the property the caller asked
about, not a tuning knob.

The verdict travels as fields on every answer: `exact_indexes` names the
doors that held the exact term (`name substrings` is absent by construction —
a name that contains the term is the definition of a neighbour); when it is
empty, `neighbours_only` is true and a one-sentence `qualifier`, printed as
the second line rather than buried in the footer, states that everything
above merely CONTAINS what you asked for — plus what the live pass actually
did. A query the index really does answer still costs zero greps.

## MCP

`atlas mcp-server` speaks MCP over stdio. `ratchet wire` registers it; by hand
it is:

```
claude mcp add -s user atlas -- ~/.ratchet/bin/atlas mcp-server
```

**One tool: `atlas_lookup(query, path?)`.** Not eight. An eight-verb MCP
surface would rebuild inside the tool list the exact routing tax the one-door
form deletes. The envelope is `{meta, query, answer, empty, footer}` with
`footer.searched` / `footer.not_searched` / `footer.not_modelled` as the same
three lists, structured.

### The CWD trap — read this before your first call

**A user-scope MCP server inherits the working directory of the CLIENT PROCESS
that spawned it — the directory `claude` was launched from — NOT the project
you are working in.** Nothing about a user-scope registration makes it
project-aware, and nothing in a normal answer tells you the server is looking
somewhere else.

Measured live: a session launched from `$HOME` returned `no_root` for every
`atlas_lookup` call in that session. `$HOME` contains no `.arc`/`.git`/`.hg`
and the root walk stops there by design, so every call correctly reported that
it had searched nothing — while the project the agent was actually editing sat
two directories away with a perfectly good manifest.

**`path` exists for exactly this.** Pass any absolute file or directory path
inside the project you mean:

```json
{"query": "suspend_data", "path": "/Users/you/projects/thing/src"}
```

Rules:

- Pass `path` **whenever the client's launch directory and the project might
  differ** — which is most of the time for a user-scope server.
- Pass it **always** after a `NO_ROOT` notice. That notice means nothing was
  searched; it is not an empty result and the footer says so
  (`not_searched` lists every index, `not_searched_why` says "no manifest was
  opened; this answer covers nothing at all").
- `meta.cwd` and `meta.cwd_source` (`"path argument"` or `"server working
  directory"`) are on every response, so a surprising root is one read to
  diagnose rather than a bisect.

Every call resolves its own project. A long-lived server pinned to one
manifest would answer questions about the wrong tree for the rest of a
session, so `mcp-server` deliberately takes no `--db` and no root.

### Notices are fields, not prose

`meta.notices` is a list of `{code, severity, detail}` — the same contract as
ratchet's `meta.caveats`.

| code | severity | meaning |
|---|---|---|
| `BUILT` | info | THIS call created the manifest. It is derived, out of tree and self-maintaining; you never build it by hand |
| `REBUILT` | info | this call refreshed a manifest the tree had moved past |
| `STALE_BUT_OVER_CEILING` | warning | you are reading an OLD manifest — the tree is past the automatic-build ceiling, so the answer may be behind the code |
| `NO_ROOT` | error | **nothing was searched at all.** Not an empty result. Pass `path` |
| `NO_MANIFEST` | error | an explicitly named `--db` does not exist. atlas will not create it |
| `TOO_LARGE` | error | this tree exceeds the automatic-build ceiling; `atlas build <root>` has no ceiling |
| `STALE_BINARY` | warning | atlas ITSELF was replaced on disk after this long-lived server started — every answer, including this one, comes from the old code. Restart the SESSION, not the server: MCP registrations are fixed when the client session spawns, so respawning the child re-runs the same image (0.1.3) |

Over MCP there is no stderr you read and no progress channel, so a call that
spent time building a manifest is indistinguishable from one that did not —
unless the build travels in the notices. That is why they are in the envelope
rather than printed.

## The manifest builds and refreshes itself

You never maintain it.

- **Where.** `~/Library/Caches/ratchet-map/<label>-<digest>.db` — out of the
  tree, always. Regenerable bytes inside a tracked tree get snapshotted once
  and live in the store forever; out-of-tree makes that impossible rather than
  merely ignored, and a cache directory is disposable by design.
- **Identity is the ABSOLUTE ROOT, not the basename.** The 48-bit digest is
  the correctness; the label is only there to make the directory readable.
  This is measured, not cautious: on ratchet's own store, `public` named six
  unrelated projects merged into one row, `src` six, `worker` four — all from
  taking a basename. For a MAP, that failure answers a question about one
  project with pointers into another. Worktrees keep their own manifests: a
  worktree is a different tree with different contents.
- **Root resolution.** Walk up from the cwd looking for `.arc`, `.git` or
  `.hg`, stopping at `$HOME` — the walk never crosses it.
- **Freshness.** Before every answer, a stat-only walk compares the file count
  and newest mtime against what the build recorded. Nothing is read, nothing
  is hashed. On the three trees this was verified against it costs 3–8 ms
  (345 / 100 / 450 files) against builds of 0.45 / 0.12 / 0.66 s, which is the
  whole argument for not simply rebuilding every time. One shape escapes it —
  a file deleted and another created between two calls, leaving the count
  equal — which is why `atlas stale` reads and hashes every file and is
  deliberately never automatic.
- **The automatic build has a ceiling: 2,000 indexable files.** Not a
  performance tuning knob — a bound on the worst failure mode for a tool that
  answers in one call, which is appearing to hang. Measured: this fleet's
  largest tree is 610 files; admitting one scope of `node_modules` took a
  build from 214 files to 3,203 and from 0.65 s to 47.7 s. The ceiling bounds
  the AUTOMATIC build only. `atlas build <root>` has none, because there a
  human asked and can watch it run.
- **Agent furniture is never indexed.** Transcript and harness directories
  (`.claude`, `.codex`, `.rift`, `.ratchet`, `.cursor`, `.arc`) and agent
  scratch directories are excluded from the file walk AND from `refs`' live
  grep by one shared predicate, so the index and the fallback cannot
  disagree. Measured before the rule existed: an agent's own session
  transcripts ranked 3rd and 4th on a `refs` query for a module it had
  merely discussed. A project's own committed `.claude/` is excluded too —
  being committed does not make a file code.
- **Installed dependencies and retired trees are never indexed, and there is
  no flag to re-admit them.** The moment such an override exists, someone
  points it at `node_modules` and the manifest stops being about code anyone
  edits. `--also` is the supported answer.

## `--db` is an override atlas never writes

```
atlas <word> --db PATH      # query a manifest YOU name
atlas build <root>          # print the manifest path it wrote
atlas <word> --no-build     # a missing manifest stays an error, never a side effect
atlas <word> --json         # the MCP tool's exact record (meta, hits, exact_indexes, footer) — v0.1.10
```

`$ATLAS_DB` is the same promise made once per shell.

**A manifest you named belongs to you: atlas will not create it and will not
overwrite it.** A frozen artifact stays frozen — an experiment that passes
`--db` to a fixed manifest must not have it rebuilt underneath, and a drift
check that proves disclosure by drifting a tree would pass vacuously if the
drift were silently erased. Automatic build and rebuild are properties of the
DERIVED path only: the door that takes no flags at all.

## Staleness: per-pointer, not per-file

"This file changed" is not actionable — it cannot say whether the line you are
about to open is still the definition. Each symbol carries an anchor (a hash
of its normalised declaration line), so `atlas stale` gives every pointer one
of three verdicts:

| verdict | meaning |
|---|---|
| `VALID` | the anchor is still at the stored line |
| `MOVED` | the anchor is elsewhere in the file — re-anchorable; `atlas stale --fix` rewrites the line number without a rebuild |
| `INVALID` | the anchor is gone; the declaration was edited or removed |

Staleness is disclosed on stderr, never guessed at. **stdout is always the
answer alone**, so piping is safe.

## CLI quirks

- **A query beginning with `-` cannot be passed bare.** argparse claims it as
  a flag and you get a usage error, not a result:
  ```
  atlas lookup -- -webkit-mask       # works; the `--` must FOLLOW the verb
  atlas -- -webkit-mask              # does NOT work
  ```
  Over MCP the query is a string field and no quoting rule applies at all.
- `--db` and `--no-build` are accepted on both sides of the subcommand.
- `--json` (v0.1.10) emits the same record `atlas_lookup` returns, so a hook or a
  script reads one shape from either door; `--no-build --json` on a tree with no
  manifest exits 0 in milliseconds with `NO_ROOT` in `meta.notices`. Put global
  flags AFTER the bare word — a flag before it defeats the verbless form.
  Position-sensitive flags are the ceremony this tool exists to delete.
- `atlas --version` prints atlas's own version, which is not ratchet's.

## What it does not model

`SQL predicates` · `localStorage keys` · `response-field → client-state` ·
`dynamic dispatch`.

These are printed on every answer for a reason. They are the areas where a
negative from atlas is worth nothing, and a tool that let you forget which
those were would be selling a bounded answer as a complete one.
