thinker / docs

thinker documentation

Thinker is a persistent knowledge cache for coding agents. It prevents agents from re-grepping and re-reading your codebase every turn by pre-seeding verified architecture notes, conventions, and invariants.

The Problem: When you start a prompt in a complex repository, coding agents spend 50–70% of their early turns running grep, reading package.json, tracing import trees, and hitting dead ends. This wastes minutes of wall time and burns tens of thousands of context tokens before any code is generated.

The Mental Model

Thinker is not a vector database or an ungrounded RAG system. It is a dependency-keyed knowledge cache stored locally under .thinker/ in your repository.

Thinker maintains concise 3–12 line notes answering recurring architectural questions (callpaths, locations, invariants, gotchas). Each note is explicitly tied to the exact files and symbols it describes via SHA-256 hashes.

1 Mine & Distill

During setup, Thinker mines git co-change history, merged PRs, and subsystem areas, synthesizing structured notes with explicit code pointers.

2 Inject Context

Before your agent begins a turn, Thinker runs BM25 ranking against your prompt with coverage floors and injects the 2–3 most relevant notes via native agent hooks or MCP tools.

3 Invalidate & Verify

Every note tracks symbol-level content hashes against your working tree. When code changes, notes are flagged stale and verified with a fast model—never serving outdated claims.

Supported Agents

Thinker hooks into your existing coding tools without changing your daily command workflow.

Claude Code

Thinker configures .claude/settings.local.json:

  • Prompt hook: UserPromptSubmit automatically injects the orientation notes into the prompt bundle.
  • Session learning: Stop hook condenses turn activity and distills new notes incrementally in the background.
  • Late file notes: Serves relevant context when Claude opens specific files (--late).

Cursor

Thinker writes .cursor/hooks.json and approves the MCP server:

  • Cursor's prompt hook computes candidate notes and delivers them on the first tool invocation.
  • Cursor's agent can also explicitly call the orient and remember MCP tools.

Codex CLI

Thinker registers hooks in .codex/hooks.json and marks the project and hook hashes as trusted in ~/.codex/config.toml, avoiding repetitive approval prompts.

Gemini CLI / Antigravity

Wires prompt-time context injection directly into .gemini/settings.json.

How It Works

What a Note Is

A Thinker note answers a recurring developer question, not a redundant restatement of "what this file does":

Kind What It Answers
callpath How control/data flows across files for an operation
location Where a recurring concern is handled
cochange What files or modules must change together
howto How to build, test, or lint with non-obvious flags
convention Local codebase rules an agent would otherwise violate
gotcha A trap (similar names, hidden ordering, cache invalidation)

Dependency-Keyed Invalidation

Traditional documentation goes stale the moment code changes. Thinker avoids this using fine-grained symbol hashing:

  • A dependency is tracked as {path, symbol?}. With a symbol, Thinker hashes just that definition block (e.g. AuthManager.verifyToken).
  • A note about AuthManager.verifyToken does not go stale just because an unrelated method in the same file was modified.
  • Every prompt serving re-hashes candidate deps against your working tree, catching uncommitted edits immediately.
  • Stale notes are deprioritized, served with a ⚠ STALE banner, and scheduled for background verification.

Verifying in Your Repo

You don't need to take benchmark claims on faith. Thinker includes built-in commands to measure the impact directly on your repository:

# 1. Benchmark agent efficiency on a recent merged PR:
thinker benchmark pr [number]

# 2. Or test a specific question with vs. without Thinker context:
thinker benchmark run "explain how permissions are validated on project exports"

# 3. View the latest side-by-side comparison report:
thinker benchmark report

How it works: Thinker makes two read-only agent calls—one without Thinker context and one with the relevant notes—and compares wall time, agent turns, tool calls, and token usage side-by-side. Both answers are saved under .thinker/benchmarks/ for human inspection.

What Outcomes to Expect

Evaluated on real tasks from merged pull requests across frontier and fast coding models (Claude Fable via Claude Code, Gemini 3.8 Flash via Antigravity CLI, and GPT-6 Astra via Codex CLI), independently graded against calibrated acceptance criteria:

  • Tool & Exploration Efficiency (Consistent across all models): Reduces total tool calls by 17% to 22% and cuts file reads by 14% to 27%. Pre-seeded architecture notes prevent agents from burning early turns on repetitive grep and find loops.
  • Context & Token Savings (Consistent across all models): Cuts input context re-reads by 18% to 21%, directly lowering cost per task by 10% to 16% across metered models.
  • Wall Clock Time (Consistent across all models): Completes tasks 7% to 14% faster (saving between 30 and 100+ seconds on complex tasks).
  • Correctness & Accuracy (Model-dependent): On frontier models (Claude Fable), pre-seeding conventions and invariants raised acceptance criteria accuracy by +10% and lifted complete task passes from 45% to 60% (+33% more tasks fully solved). Fast and reasoning models (Gemini Flash, GPT-6 Astra) maintained strict correctness parity (within single-run noise) while achieving significant token and tool reductions.

Note: The cache accelerates code navigation and eliminates redundant file exploration; it does not turn a weak model into a strong one. Run thinker benchmark report in your repository to measure your own baseline.

CLI Reference

Command Description
thinker onboard Interactive 3-stage onboarding (detect harnesses, build cache, benchmark)
thinker benchmark pr [n] Run paired agent benchmark on a recent PR change
thinker benchmark run <query> Run paired agent benchmark on an architectural query
thinker benchmark report Reprint the latest benchmark comparison table
thinker check Check all notes against the current working tree and report stale deps
thinker verify Verify stale notes against changed diffs with a fast model
thinker update Update Thinker CLI to the latest version
thinker uninstall [--purge] Remove agent hooks and MCP registrations (--purge removes notes)

Privacy & Telemetry

Thinker collects anonymous daily effectiveness metrics (cache hit rate, total note count, estimated token savings) to monitor cache quality.

No code snippets, note bodies, prompt text, file paths, or repository URLs are ever collected or transmitted.

To opt out completely at any time:

# In your shell environment:
export THINKER_TELEMETRY=off

# Or in .thinker/config.json:
{ "telemetry": false }