Shire
.,:lccc:,.
.,codxkkOOOOkkxdoc,.
.;ldkkOOOOOOOOOOOOOOOkkdl;.
.:oxOOkxdollccccccccllodxkOOkxo:.
,lkOOxl;.. ..,lxOOkl,
.ckOOd:. .:dOOkc.
;xOOo, .,clllc,. ,oOOx;
lOOk; .:dkOOOOOOkd:. ;kOOl
oOOx, .ckOOOOOOOOOOOOkc. ,xOOo
lOOk, ;xOOOkdl:;;:ldkOOOx; ,kOOl
;OOO; lOOOd;. .;dOOOl ;OOO;
dOOd :OOOl lOOO: dOOd
kOOl oOOx .;;. xOOo lOOk
kOOl oOOx .xOOx. xOOo lOOk
dOOd :OOOl .oOOo. lOOO: dOOd
;OOO; lOOOd;. .,,. .;dOOOl ;OOO;
lOOk, ;xOOOkdl:,:ldkOOOx; ,kOOl
oOOx, .ckOOOOOOOOOOOOkc. ,xOOo
lOOk; .:dkOOOOOOkd:. ;kOOl
;xOOo, .,clllc,. ,oOOx;
.ckOOd:. .:dOOkc.
,lkOOxl;.. ..,lxOOkl,
.:oxOOkxdollccccccccllodxkOOkxo:.
.;ldkkOOOOOOOOOOOOOOOkkdl;.
.,codxkkOOOOkkxdoc,.
.,:lccc:,.
One index to rule them all.
Search, Hierarchy, Index, Repo Explorer — a monorepo package indexer that builds a dependency graph in SQLite and serves it over Model Context Protocol.
Point it at a monorepo. It discovers every package, maps their dependency relationships, extracts symbols from source code, and gives your AI tools structured access to the result.
Get started in 30 seconds
brew install justinjdev/shire/shire
shire init --global
shire build
That’s it. Claude Code can now search your packages, symbols, files, and dependency graph. See Setup for details.
Installation
Homebrew (macOS, Linux)
brew tap justinjdev/shire
brew install shire
From prebuilt binary
Download the latest release from GitHub Releases and add the binary to your PATH.
Nix
# Install into your profile
nix profile install github:justinjdev/shire
# Or run without installing
nix run github:justinjdev/shire
From source
Requires Rust toolchain.
cargo install --path .
Setup
Claude Code
One command configures shire globally for all projects:
shire init --global
This creates:
~/.claude/shire.toml— shared config withdb_path = "~/.claude/shire/{repo}/{worktree}/index.db"(auto-namespaced per repo and worktree)mcpServers.shirein~/.claude.json— serves the index viashire servePostToolUsehook in~/.claude/settings.json— auto-rebuilds the index after file edits (Edit,Write,NotebookEdit,Bash)~/.claude/rules/shire.md— rules file guiding Claude Code to prefer Shire tools
The {repo} placeholder is replaced with the repository directory name at runtime, and {worktree} with the worktree name (or _primary for the main checkout), so each repo and worktree gets its own index file automatically.
After running shire init --global, open any repo and run:
shire build
The index is ready. Claude Code will automatically use it via the MCP server.
Interactive setup
Run in a terminal, shire init asks its questions in sections:
- Index: the install scope, rebuild strategy, database path, whether to gitignore the index directory, and extra directories to exclude.
- Extras: one checklist (space toggles an item, enter confirms):
- the cross-reference tools
- the rules file
- the
~/.claude/CLAUDE.mdguidance - the Claude Code status mod
- Review: lists every file it is about to write, then asks Apply these changes?
Answering no writes nothing.
- If
shire.tomlalready exists, you are asked first whether to overwrite it. Answer no to keep it and still set up the rest.
- If
--yes (or running without a terminal) skips the questions and uses the defaults.
Rules file
shire init creates ~/.claude/rules/shire.md with guidance on when to use Shire tools vs Grep/Glob. This helps Claude Code default to Shire for codebase searches, so you spend fewer tool calls on broad exploration.
If it already exists, shire init leaves it untouched — with one exception: turning on the cross-reference index (symbols.references_enabled) for a repo that already has a rules file appends the extra reference-tools guidance in place, so your other customizations are preserved.
CLAUDE.md integration
In interactive setup this is the Search guidance in ~/.claude/CLAUDE.md item in the
Extras checklist, checked by default. If selected, it appends a one-liner to ~/.claude/CLAUDE.md directing Claude Code to prefer Shire MCP tools over Grep/Glob for code search. The line is idempotent — running init again won’t duplicate it. If ~/.claude/CLAUDE.md doesn’t exist yet, it creates the file.
Claude Code status mod (experimental)
In interactive setup this is the Claude Code status mod item in the Extras checklist,
unchecked by default. Checking it, or passing --mod, installs it (--mod and --no-mod both answer the question, so the item is left out of the checklist). It is a Claude Code
mod
that polls shire status --json and shows index health in Claude Code’s status
line (for example shire ● 412 pkgs · 38.2k syms · 4m ago), shows a toast when something
changes (new build failures, an interrupted build, the watch daemon stopping), and adds a
/shire pane with rebuild buttons.
The mod is always installed for your user, even from a project-level shire init. Claude
Code loads extra plugin folders only from CLAUDE_CODE_PLUGIN_DIRS in ~/.claude/settings.json,
never from a project’s settings, so a cloned repository cannot turn it on. shire init writes
the mod’s files, which are compiled into the shire binary, to ~/.claude/shire-mod/shire-status/
and adds that folder to env.CLAUDE_CODE_PLUGIN_DIRS, keeping any folders already listed there.
New Claude Code sessions pick it up.
shire installrefreshes an installed mod’s files, so they keep matching the binary after an upgrade. It never installs the mod, and never touches the settings: if you removed the folder fromCLAUDE_CODE_PLUGIN_DIRSto switch the mod off, it stays off.shire uninstallremoves the folder and its entry inCLAUDE_CODE_PLUGIN_DIRS.
The mod uses Claude Code’s early-access mod API, which may change between Claude Code releases.
If a Claude Code update breaks it, Claude Code names the mod in the transcript, and
shire uninstall, or deleting the folder, turns it off.
Terminal output
shire init uses styled terminal output to show what it does:
- ✓ (green) — a file or config entry was created or updated
- – (dimmed) — a file or config entry already exists, skipped
- Section headers appear in cyan
Most file writes (.gitignore, CLAUDE.md, settings.json, .mcp.json, ~/.claude.json) use atomic writes — content is written to a temporary file first, then renamed into place. This prevents partial writes if the process is interrupted.
Project-level setup
To create a shire.toml in the current repo instead of globally:
shire init
This generates a minimal shire.toml (just db_path; everything else uses built-in defaults — see Configuration), writes the MCP server config to .mcp.json, and (in hook mode) a PostToolUse hook to .claude/settings.json plus .claude/rules/shire.md. If the db_path points to a local directory (e.g., .shire/index.db), it offers to add that directory to .gitignore.
Manual setup
If you prefer manual configuration, add to ~/.claude.json (global) or .mcp.json (project-level):
{
"mcpServers": {
"shire": {
"command": "shire",
"args": ["serve"]
}
}
}
To keep the index fresh during a session, add a PostToolUse hook to ~/.claude/settings.json:
{
"hooks": {
"PostToolUse": [
{
"matcher": "Edit|Write|NotebookEdit|Bash",
"hooks": [{ "type": "command", "command": "shire rebuild --stdin" }]
}
]
}
}
Claude Desktop
Add Shire to your claude_desktop_config.json:
{
"mcpServers": {
"shire": {
"command": "shire",
"args": ["serve", "--db", "/path/to/repo/.shire/index.db"]
}
}
}
Other MCP clients
Shire speaks standard MCP over stdio. Any client that supports MCP can connect:
shire serve --db /path/to/repo/.shire/index.db
Use --root to enable on-demand reindexing. Before answering a query the
server re-checks the working tree — at most once per serve.debounce_s
window (default 5 seconds) — by running an incremental build, so unstaged
edits are picked up without any hook:
shire serve --root /path/to/repo
Editor registration
For editors and CLIs other than Claude Code, shire install registers the built shire
binary as an MCP server with every supported tool it finds on the machine: Claude Code
(via the claude CLI, falling back to file patching), Codex CLI, Cursor, Windsurf,
Gemini CLI, VS Code, and Zed. It writes the current binary’s absolute path (via
std::env::current_exe) rather than a bare shire, so registrations keep working even
if the tool that launches them doesn’t inherit your shell’s PATH.
shire install # register with every detected tool
shire install --dry-run # show what would change without writing anything
shire install --force # overwrite existing registrations (e.g. after moving the binary)
shire uninstall # remove shire's registration from every detected tool
shire uninstall --dry-run
install/uninstall only touch each tool’s own MCP config file (or the tool’s own CLI,
for Claude Code and Codex); they do not create or modify shire.toml, PostToolUse hooks,
or the rules file — use shire init for those.
CLI reference
Build an index
shire build --root /path/to/repo
Rebuild from scratch
Ignore cached hashes and re-parse everything:
shire build --root /path/to/repo --force
Custom database location
shire build --root /path/to/repo --db /tmp/my-index.db
The index defaults to .shire/index.db inside the repo root. Override with --db or db_path in shire.toml (see Configuration).
Clean up
Remove the index database, WAL/SHM files, the .shire directory, and stop the watch daemon:
shire clean
Index status
Show the index’s state without rebuilding or writing anything: whether a build is running,
when the index was built and at which commit (and whether HEAD has moved since), counts,
packages still owed a source re-check, the last build’s failures, whether the file walk saw
the whole tree, and the watch daemon’s liveness:
shire status # human-readable
shire status --json # one JSON object, for scripts and editor integrations
state is one of missing, refused (symlinked db_path), unreadable, building,
interrupted (the last build died part-way; the next build repairs it) or ok. With no
--root, the repo is found by walking up from the current directory. The command always
exits 0, even when shire.toml cannot be read (state is then unreadable, db_path is null and
error says why); read state to decide.
For Claude Code, shire init --mod installs a mod that polls shire status --json and shows
index health in the status line (see Claude Code status mod).
Watch daemon status
Check whether the watch daemon is running for a repo (PID, socket path, and whether it’s actually reachable — see Watch Daemon):
shire watch --root /path/to/repo --status
Incremental builds
Subsequent builds are incremental — only manifests whose content has changed (by SHA-256 hash) are re-parsed. Source files are tracked at per-file granularity: if individual source files change without a manifest change, only those files have their symbols re-extracted. An mtime pre-check skips hash computation entirely for packages whose source files haven’t been touched since the last build.
File indexing is also incremental — a file-tree hash detects structural changes, skipping the file indexing phase entirely when no files have been added, removed, or resized.
Symbol extraction and source hashing are parallelized across packages and within packages using rayon for multi-core throughput. Files are read once per build (single-pass hash + extraction). All database writes use batched multi-row INSERTs within explicit transactions, with FTS5 triggers temporarily disabled during bulk operations for maximum SQLite throughput.
Build progress
shire build shows real-time progress for each build phase:
- Spinners for quick phases (discovering manifests, workspace context, recomputing internals, indexing files)
- Progress bars with ETAs for longer phases (parsing manifests, extracting symbols)
Progress bars persist after completion so you can see the full build history in your terminal. Quiet mode (used internally by the MCP server for on-demand rebuilds) hides all progress output.
Configuration
Drop a shire.toml in the repo root to customize behavior:
# Custom database location (default: .shire/index.db)
db_path = "/path/to/custom/index.db"
[discovery]
manifests = ["package.json", "go.mod", "go.work", "Cargo.toml", "pyproject.toml", "pom.xml", "build.gradle", "build.gradle.kts", "settings.gradle", "settings.gradle.kts", "cpanfile", "Gemfile", "flake.nix"]
exclude = ["node_modules", "vendor", "dist", ".build", "target", "third_party", ".shire", ".gradle", "build"]
# Symbol extraction
[symbols]
exclude_extensions = [".proto", ".pl"]
exclude_patterns = [] # file name patterns to skip (suffix match, e.g. "_generated.go"; or prefix, e.g. "zz_generated.")
references_enabled = false # EXPERIMENTAL, default false — see below
max_file_size = 0 # 0 = disabled (default); set to e.g. 2097152 for 2 MiB cap
max_references_per_file = 10000 # 0 = unlimited; default 10000 — caps cross-references per file
include_private = true # default true — index private/unexported symbols; see below
# Documentation indexing
[docs]
extensions = [".md", ".rst", ".txt", ".adoc"]
max_file_size = 262144 # 256 KB — files larger than this are truncated
# MCP server on-demand rebuild
[serve]
debounce_s = 5 # `serve --root` re-checks the working tree at most this often
# Override package descriptions
[[packages]]
name = "legacy-auth"
description = "Deprecated auth service — do not add new dependencies"
File discovery and .gitignore
The manifest walk, the file index, and symbol extraction all skip:
- Directories named in
discovery.exclude - Anything matched by a committed
.gitignoreor.ignorefile, at the repo root or in any nested directory
They deliberately do not honor a developer’s personal global gitignore (git config core.excludesFile) or the untracked, per-clone .git/info/exclude — only files checked into the repo affect what gets indexed, so the index is the same for every collaborator and in CI regardless of local git configuration.
A malformed pattern in a .gitignore (one git accepts but the ignore crate’s glob compiler rejects — e.g. brace alternation, a trailing backslash) is logged as a warning and otherwise ignored; it never aborts a build.
Config precedence
Config is resolved in this order, with no merging — the first one found is used whole, and none of the others are read:
--config <PATH>— explicit path, must exist./shire.toml— repo-root config~/.claude/shire.toml— global config (created byshire init --global)- Built-in defaults
Because the fallback is whole-file replacement rather than a merge, a local shire.toml containing only db_path discards every other setting in ~/.claude/shire.toml (excludes, custom discovery rules, etc.) rather than layering on top of it.
What gets walked
Every shire walk — manifests, files, and source files for symbol extraction —
skips hidden entries, the directories in discovery.exclude, and anything
matched by a committed .gitignore (the repo root’s and any nested ones).
Ignore files that are not part of the repository are deliberately not
consulted: neither your personal global gitignore (core.excludesFile,
usually ~/.gitignore_global) nor the per-clone .git/info/exclude. Both are
machine-local, so honouring them would make the index — and what
search_symbols can find — depend on which machine built it, with no
diagnostic and nothing in the repo to explain the difference.
To keep a path out of the index for everyone, add it to the repo’s
.gitignore or to discovery.exclude.
Watch daemon
[watch]
debounce_ms = 2000 # milliseconds to wait after last change before rebuilding
Logging
[log]
level = "warn" # error, warn, info, debug, trace
dir = ".shire/logs" # log directory (relative to repo root). Set to "" to disable file logging
max_days = 30 # automatically delete log files older than this
The SHIRE_LOG environment variable overrides the config level (e.g., SHIRE_LOG=debug shire build). Log files are daily-rotated with filenames like shire.log.2026-03-26. Each session includes a unique session ID for correlation across concurrent processes.
All fields are optional. Defaults are shown above. The --db CLI flag takes precedence over db_path in config.
Cross-reference index (experimental)
symbols.references_enabled (default false) populates the symbol_refs
table so the symbol_references, symbol_callers, and symbol_callees
MCP tools can answer “where is this used?” / “who calls this?” questions.
Reference extraction is supported for 8 tier-1 languages: Go, Python,
Java, TypeScript, JavaScript, Perl, Ruby, Scala.
Opt-in: shire init asks whether to enable this (prompt labelled
experimental), and writes references_enabled = true to shire.toml
when you say yes. You can also add it manually:
[symbols]
references_enabled = true
Cost: DB grows substantially — roughly +30% on TS/JS repos to +150% on Go-heavy repos (benchmarks on shire-bench: turborepo +29%, grafana +152%, kubernetes +104% vs main baseline). Build time grows ~5-7%.
Toggling the flag takes effect on the next build. Disabling wipes
symbol_refs at the start of the build; re-enabling repopulates it on
the next full rebuild (shire build --force).
This feature is marked experimental: its schema and coverage may change in minor versions as language support broadens and edge cases surface.
Private symbols
symbols.include_private (default true) controls whether private and
unexported symbols are indexed. They are tagged visibility = "private", by
each language’s own convention — a lowercase Go name, a leading _ in
Python, a non-pub Rust item, a private Java member (see the Visibility
column in Supported Ecosystems) — and search_symbols ranks them after
the public ones, so they are there when you look for a helper by name without
crowding out the API.
[symbols]
include_private = false # index only what other code can use
With false, symbols whose visibility is private are dropped at extraction
time. internal and protected symbols are kept either way, and
cross-references (references_enabled) are unaffected — calls made from
inside a private function are still recorded.
Size: private code is often most of a codebase. As a guide, Shire’s own
(Rust) source indexes roughly 3.4x as many symbols with the default as with
include_private = false, and symbols and symbols_fts grow by about that
much. Expect a smaller jump in
code that is mostly exported, and a larger one in application code full of
helpers.
Toggling the option takes effect on the next build, which re-extracts every
source file once (no --force needed). The same happens once after upgrading
to a Shire version whose extractor output changed.
Custom package discovery
For codebases where packages aren’t defined by standard manifest files — Go single-module monorepos, repos that use ownership.yml + build files, or any non-standard convention — you can define custom discovery rules:
# Discover Go apps: directories containing both main.go and ownership.yml
[[discovery.custom]]
name = "go-apps"
kind = "go"
requires = ["main.go", "ownership.yml"]
paths = ["services/", "cmd/"]
exclude = ["testdata", "examples"]
max_depth = 3
name_prefix = "go:"
# Discover proto packages: directories containing *.proto and buf.yaml
[[discovery.custom]]
name = "proto-packages"
kind = "proto"
requires = ["*.proto", "buf.yaml"]
paths = ["proto/", "services/"]
max_depth = 4
| Field | Required | Description |
|---|---|---|
name | yes | Rule identifier |
kind | yes | Package kind for symbol extraction (go, proto, npm, etc.) |
requires | yes | File patterns that must ALL exist in a directory (supports globs like *.proto) |
paths | no | Limit search to specific subtrees (default: repo root) |
exclude | no | Rule-specific directory exclusions (on top of global excludes) |
max_depth | no | Maximum depth to search from each paths entry |
name_prefix | no | Prefix prepended to directory-derived package name (e.g., go:services/auth) |
extensions | no | Override which file extensions get symbol extraction |
Custom discovery runs alongside manifest-based discovery. Directories already found by manifest parsers are skipped. Subdirectories of matched directories are also skipped to prevent nested matches. A manifest package nested under a custom package still owns its own subtree: its files are indexed for it, not for the custom package. Custom packages are re-checked on every incremental build like manifest packages, but one whose directory stops matching its rule is not removed (not even by shire build --force); run shire clean and rebuild to drop it.
MCP Tools & Prompts
Tools
Shire exposes the following tools over the Model Context Protocol:
| Tool | Description |
|---|---|
search_packages | Search packages by name or description. Use instead of Grep for finding packages. |
list_packages | List all indexed packages, optionally filtered by kind |
package_dependencies | List a package’s dependencies. Set depth>1 for transitive graph (returns edge list with different schema; limit caps the edge list too). |
package_dependents | Find all packages that depend on this package |
search_symbols | Find functions, classes, types, methods by identifier or identifier prefix (not regex or substring). handle matches handleRequest; verify jwt matches verifyJwtToken. Matches the symbol name and its sub-tokens only, never signatures or file paths. Omit query with a package filter to list that package’s symbols, capped at limit: non-private symbols in (file, line) order, then private ones. Private symbols are indexed and returned with their visibility; search ranks them after the public ones (an exactly-named symbol still comes first). |
get_file_symbols | List all symbols defined in a specific file. Use instead of reading the file to understand its exports. |
search_files | Find files by path or name. Use instead of Glob/find for locating files. Useful for “middleware”, “proto files”, or files in a specific directory. |
search_docs | Search documentation files by content, title, or path — returns matching docs with text snippets |
list_package_files | List all files in a package, optionally filtered by extension. Use instead of Glob for listing package contents. |
explore | Explore a concept across the codebase — searches packages, symbols, files, and documentation semantically. Use as the first tool when investigating unfamiliar code or broad topics like “authentication” or “error handling”. Returns a structured context map organized by package. |
index_status | Index build metadata: timestamp, git commit, counts |
symbol_references | Find all references to a symbol by name. Returns [{name, kind, file_path, line, package, enclosing_symbol}]. Accepts optional kind and package filters. Requires symbols.references_enabled = true (experimental, opt-in). Note: matching is name-based. enclosing_symbol is dot-qualified (AuthService.login); a qualified name passed as name is resolved through symbols.parent_symbol, and references written in packages that define their own symbol of that name are left out — see Qualified names. |
symbol_callers | List all callers of a symbol (call-site references). Returns [{caller_name, caller_file, caller_line, caller_package, call_sites}], where caller_name is the dot-qualified enclosing path (AuthService.login) and can be fed straight back in as name — a qualified name is resolved through the type that defines the method (see Qualified names). Accepts optional package filter. Requires symbols.references_enabled = true. Same name-based-match caveat as symbol_references. |
symbol_callees | List what a function calls (outbound call graph). Returns [{callee_name, first_file, first_line, call_sites}]. Accepts a bare method name (login, which matches every qualified form such as AuthService.login) or a qualified one (AuthService.login, which matches only that method), plus an optional package filter. Requires symbols.references_enabled = true. |
change_impact | Analyze the blast radius of changing a symbol. Combines cross-references with the dependency graph to return {direct_impact, cross_package_impact, transitive_impact, summary}. Use before renaming, changing a signature, or deleting a symbol. Accepts optional package (home package hint, for disambiguation), transitive_depth (default 2), and limit. Requires symbols.references_enabled = true. A dot-qualified name sets home_package from the type that defines it and reports excluded_packages; when the name is defined in more than one package the response also carries defined_in and a home_package_note saying the home package was a choice among them. Same name-based-match caveat as symbol_references. |
schema_consumers | Find all files generated from a schema file (e.g. .proto). Returns generated file paths and their packages. Use to understand the blast radius of a schema change. |
generated_from | Find the source schema file that generated a given file. Use to trace a generated file (e.g. user.pb.go) back to its source proto. |
How matching works
All four search tools (search_symbols, search_packages, search_files,
search_docs) run the same FTS5 query builder:
- The query is split on whitespace and every token must match (implicit AND).
- Each token matches by prefix:
handlematcheshandleRequestandhandle_request. Tokens of one character are matched exactly instead —packages_ftsanddocs_ftsindex 2- and 3-character prefixes, whilesymbols_ftsandfiles_ftsdeliberately carry no prefix index (a prefix query there walks a term range instead, measured as equally fast and ~28% smaller on disk), so a single-character prefix would have to scan every term. - Each tool searches only the columns it is about.
search_symbolsmatches the symbol name and its sub-tokens — not signatures, file paths or kinds (filter by kind withkind, find paths withsearch_files, and use Grep for text inside a signature).search_filesmatches the path,search_packagesthe package name, description and path, andsearch_docsthe doc title, body and path. search_symbolsorders exact name matches first, so searchinghandle— orhandle*, or a pastedhandle.— never buries a symbol actually calledhandleunder its own prefixes. After that, private symbols rank after the others: the results FTS returns are reordered sopublic,protectedandinternalsymbols lead, keeping their relative rank. A private symbol is never left out for being private — only ranked later.- Symbol names are additionally indexed by their sub-tokens:
verifyJwtTokenis indexed asverify,jwt,token, soverify jwt,jwtandtokenall find it. This applies to symbol names only, not to file paths or doc bodies. - Matching is by identifier, not regex or substring:
andleRequfinds nothing. - Operators in a query (
OR,NEAR,*,-,column:) are treated as literal text, not as FTS5 syntax.
Qualified names
symbol_references, symbol_callers and change_impact take a symbol name.
The reference index stores bare names (run), while enclosing_symbol /
caller_name come back dot-qualified (AuthService.run) and are meant to be
fed straight back in. A qualified name is resolved in three steps:
- Literally. Some refs really are dot-named — an
importofos.path. If the name matches refs as given, that is the answer. - Through the qualifier. Otherwise the last segment before the dot is
looked up in
symbols.parent_symbol:A.runfinds the symbols namedrunwhose parent isA, which is where the symbol lives. A reference row records the name and the package it was written in, never the type it resolves to, so the qualifier cannot filter references directly — what it can do is attribute them. Arunwritten inside a package that defines its ownrunon some other type belongs to that package’s method, so those packages are left out; every other package is kept, because a cross-package call site is exactly what these tools exist to find. - Bare, and flagged. If no indexed symbol carries that qualifier, the qualifier is dropped and the bare name is matched on its own.
Whenever the rows were matched on a name other than the one passed, the result
carries matched_name (the name actually matched), matched_note (what that
means) and, for step 2, defined_in and excluded_packages. Because those
fields need somewhere to live, a rewritten name always returns the single
object form described under Result limits — results plus
the match fields — even when nothing was truncated. change_impact already
returns an object, and gains matched_name, qualifier_dropped and
excluded_packages; it takes its home_package from the resolved symbol.
excluded_packages is worth reading before acting on a change_impact
answer: those packages were left out of direct_impact,
cross_package_impact, summary.affected_packages and the reverse-dep walk
seeded from it, so a call site in one of them is blast radius the qualifier
chose to attribute elsewhere. The list names every package that defines a
symbol of that name, not only the ones that turned out to reference it — most
entries will have had nothing to drop. summary.excluded_ref_count is the
number that matters: how many references those packages actually held. When it
is not zero, re-run with the bare name to see them.
A package filter is applied on top of the resolution, so asking for a
qualified name and a package that step 2 excluded is a contradiction and
returns nothing — excluded_packages in the response is what says why. In
change_impact, where package is a home-package hint rather than a filter,
the same combination empties direct_impact by construction (that package’s
references were attributed to its own symbol), and home_package_note says
so.
Which package is home_package
change_impact splits references into direct_impact and
cross_package_impact by comparing each reference’s package against
home_package, so that one value decides the whole answer. It is the
package argument when given, and otherwise the first — by name — of the
packages defining the resolved symbol. When there is more than one, that
first is a tiebreak, not a fact, and the response says so:
defined_inlists every package defining the symbol (present only when there is more than one, and narrowed to the definitions under the qualifier for a qualified name).home_package_noteexplains what was chosen and how to choose differently — passpackageto make one of the others the home package.
What this cannot do is separate two same-named methods inside one package:
with only the bare name recorded, A.run and B.run in the same package still
merge, and both are reported. Pass package to narrow the answer; use Grep
when the distinction has to be exact.
Result limits
Tool output is pasted verbatim into a model’s context, so every list-returning tool is bounded:
| Tool | limit default | Maximum |
|---|---|---|
search_symbols, search_packages, search_files, search_docs | 20 | 200 |
get_file_symbols, list_package_files, list_packages, package_dependencies, package_dependents, schema_consumers, generated_from | 100 | 200 |
symbol_references, symbol_callers, symbol_callees, change_impact | 100 | 200 |
limit: 0 means “use the default”, not “one row”. The limit is applied in
SQL, and one row beyond it is fetched to tell a page that was cut from a list
that merely ends there.
A complete result is the bare JSON array. A truncated one — or, for the reference tools, one whose name was rewritten (see Qualified names) — is a single JSON object instead:
{"results": [...], "truncated": true, "limit": 20, "max": 200, "note": "showing the first 20 results …"}
so a capped list is never presented as a complete one, and a client that
concatenates the result’s text blocks still gets parseable JSON. (At
limit = 200 the extra row cannot be fetched, so a result that fills the
ceiling is always flagged, with a note saying more rows may exist rather
than that they do — and the note then asks for a narrower request rather
than a bigger limit, which is already clamped at 200.)
change_impact returns an object rather than a list; when a bucket is capped
it gains the same truncated / limit / max / note fields.
summary.direct_count and summary.cross_package_count count every reference
scanned rather than only the rows returned — but the scan itself stops at
10 000 references, so they are totals only while summary.counts_capped is
false; when it is true they are floors. summary.transitive_package_count is
never a total: the reverse-dep walk stops at limit, so a capped result
reports a floor.
Index freshness under serve --root
With --root the server reindexes on demand. Before answering a tool call it
checks how long ago the index was last built: inside the serve.debounce_s
window (default 5 seconds) it answers straight from the current index, and
outside it, it runs an incremental build first and answers from the result.
That build is the only freshness oracle — it compares the repo’s file tree,
per-package mtimes and per-file content hashes itself, and costs on the order
of 60-200 ms when nothing has changed. So an ordinary working-tree edit is
picked up on the first tool call more than serve.debounce_s after it, with
no need to stage anything: git add and .git/index play no part.
Raise serve.debounce_s to trade freshness for fewer rebuilds during bursts
of tool calls; lower it for a repo where builds are cheap and edits frequent.
Without --root (plain shire serve --db …) the server is strictly
read-only and never rebuilds — refresh the index with shire build, the
watch daemon, or the PostToolUse hook.
When to use Shire vs Grep/Glob
| Task | Use | Not |
|---|---|---|
| Find a function, class, or type by name | search_symbols | Grep |
| Find a file by name or path | search_files | Glob / find |
| List files in a package | list_package_files | Glob |
| Find a package | search_packages | Grep |
| Explore an unfamiliar area | explore | multiple Grep calls |
| Search for a literal string or log message | Grep | Shire |
| Search inside function bodies | Grep | Shire |
| Pattern match on file contents | Grep | Shire |
Prompts
Prompts are pre-built templates that compose multiple queries into structured context. They give your AI a map of where concepts live in the codebase.
| Prompt | Args | Description |
|---|---|---|
explore | query | Search packages, symbols, files, and documentation for a concept — returns a structured context map organized by package |
reference_audit | name | Guides refactor-safety analysis for a symbol: classifies refs by kind, traces the call graph via symbol_callers, identifies cross-package impact, and assesses rename/change risk. Requires symbols.references_enabled = true (experimental). |
Watch Daemon
shire watch starts a background daemon that rebuilds the index whenever a rebuild
signal arrives — from the Claude Code PostToolUse hook (shire rebuild --stdin) or a
manual shire rebuild. It does not watch the filesystem itself (no inotify/FSEvents):
an edit made outside Claude Code — by another editor, git checkout, a codegen script run
outside the hook — is never picked up until something explicitly signals a rebuild. It
uses Unix domain socket IPC with configurable debounce (default 2s).
Start the daemon
Idempotent — safe to call multiple times:
shire watch --root /path/to/repo
Check whether it’s running
shire watch --root /path/to/repo --status
Prints the daemon’s PID, socket path, and whether it’s actually reachable (a stale PID
file left behind by a crash, kill -9, or a bind failure reads as not running, not as a
false “yes”).
Signal a rebuild manually
shire rebuild --root /path/to/repo
If no daemon is listening at that root, this prints a warning to stderr and still exits 0 (so it’s safe to call from a hook) — the index is simply not updated.
Signal a rebuild from a Claude Code hook
Reads JSON from stdin. The repo root is resolved by walking up from the hook’s cwd to
the nearest ancestor containing shire.toml or .git — so a session launched in a
package subdirectory of a monorepo still reaches the daemon’s socket at the repo root. A
bare .shire/ directory does not count as a marker: shire creates .shire/logs
under any directory it is pointed at, so treating it as a marker would let a stray
.shire left behind by a one-off shire build --root <subdir> silently divert future
lookups to that subdirectory instead of the real repo root:
shire rebuild --stdin
Stop the daemon
shire watch --root /path/to/repo --stop
Sends SIGTERM and waits up to 5s for the process to actually exit before removing its
PID/socket files. The daemon only handles SIGTERM between rebuilds, so a slow or
uninterruptible rebuild can outlast that wait — if it does, --stop exits non-zero and
prints the daemon’s PID rather than reporting success while it is still running; its
PID/socket files are left in place, and retrying (or checking --status) is safe.
The daemon is identified as shire’s own by its executable — an exact basename match, a
basename starting with shire- or shire. (a versioned install like shire-v0.7, a
renamed download, or a manual mv shire shire.old && cp new shire-style in-place
upgrade), or the exact same file as the shire binary currently invoking
--stop/--status (including after an in-place upgrade replaces that file while the
daemon is still running, so long as it’s still at the same install path) — so a renamed
or versioned binary is recognized correctly, while an unrelated binary whose name merely
starts with “shire” with no separator (e.g. shireling) is not. This is combined with a
cmdline check requiring argv[0] to itself look like shire’s own binary, a --root
argument naming this exact repository, and the literal tokens “watch” and “–foreground”,
before anything is signalled.
Executable identity can only be confirmed when --stop/--status/clean run from the
exact same binary file that started the daemon (or one satisfying the basename rule
above) — a different shire binary checking on a renamed install it didn’t start cannot
positively verify it. In that case shire does not guess: if the daemon’s socket is still
answering, its pid/socket files are left alone and nothing is signalled, rather than being
treated as stale and deleted out from under a process that is demonstrably still running.
shire clean inherits the same caution and refuses (non-zero exit, nothing removed)
rather than removing .shire while such a daemon is alive.
One residual case has no automatic recovery: if the repository directory itself is
renamed while the daemon is running, the --root recorded in its own argv no longer
matches the (now different) path passed to --stop, so shire refuses to signal it —
stop it directly instead with kill $(cat .shire/watch.pid).
Smart filtering
The watch daemon avoids unnecessary rebuilds:
- Edit/Write/NotebookEdit tools — the changed file is checked for relevance against
the same configuration the indexer itself uses: manifest filenames
(
discovery.manifests), source file extensions, doc extensions (docs.extensions), anddiscovery.customrules. Paths under an excluded directory (discovery.exclude, e.g.node_modules) are skipped even if the extension would otherwise match. - Bash commands — filtered against an allowlist of known read-only commands (
ls,git status,cargo test, etc.) that are skipped; unknown commands default to triggering a rebuild.
Troubleshooting
.shire/watch-stderr.log— the daemon’s stderr, including a bind failure (e.g. the socket path exceeding the platform’sSUN_LEN, which is more likely on deeply nested repo paths). Ifshire watchfails to start, this file has the reason..shire/logs/shire.log.<date>— the daemon’s regular tracing output (setSHIRE_LOG=debugfor verbose per-rebuild logging, including which files were skipped as irrelevant).
Git Worktrees
Shire automatically detects git worktrees and maintains separate indexes for each one. This works out of the box with no configuration required.
How it works
When you run shire build or shire serve inside a linked worktree, shire:
- Detects the worktree by inspecting
.git— a directory means primary working tree, a file means linked worktree - Resolves a per-worktree DB path using the
{worktree}placeholder indb_path - Seeds from the primary worktree’s DB on first build, so you don’t start from scratch — only changed packages need reindexing
The primary worktree uses the reserved name _primary. Linked worktrees use Git’s stable worktree ID (the directory name under .git/worktrees/<id>).
Configuration
The default global config (generated by shire init --global) already includes worktree support:
db_path = "~/.claude/shire/{repo}/{worktree}/index.db"
This produces separate databases like:
~/.claude/shire/my-project/_primary/index.db
~/.claude/shire/my-project/feat-auth/index.db
~/.claude/shire/my-project/bugfix-123/index.db
Placeholders
| Placeholder | Description |
|---|---|
{repo} | Name of the main repository (directory basename) |
{worktree} | Worktree identifier — _primary for the main working tree, or Git’s worktree ID for linked worktrees |
Shared vs separate databases
If your db_path does not include {worktree}, all worktrees share the same database. This is fine for read-only use; concurrent builds from different worktrees serialize on the database (each connection waits up to 5 seconds for the other writer), and a build that still cannot get the lock fails rather than corrupting anything.
Including {worktree} in the path gives each worktree its own database, which is the recommended setup.
DB seeding
When shire builds in a linked worktree for the first time and no database exists yet, it checks whether the primary worktree has an existing database. If so, it copies that database as a seed — giving you a fully populated index immediately. Only packages that differ between the worktrees need reindexing.
$ cd ~/worktrees/feat-auth
$ shire build
Seeded DB from /Users/you/.claude/shire/my-project/_primary/index.db
Building index...
Local config
A local shire.toml at the repo root can use a relative db_path (e.g., .shire/index.db). Since each worktree has its own root directory, relative paths naturally resolve to separate databases per worktree. Seeding still applies in this case.
Supported Ecosystems
| Manifest | Kind | Workspace support |
|---|---|---|
package.json | npm | workspace: protocol versions normalized |
go.mod | go | go.work member metadata |
go.work | go | use directives parsed for workspace context |
Cargo.toml | cargo | workspace = true deps resolved from root |
pyproject.toml | python | — |
pom.xml | maven | Parent POM inheritance (groupId, version) |
build.gradle / build.gradle.kts | gradle | settings.gradle project inclusion |
cpanfile | perl | requires / on 'test' blocks |
Gemfile | ruby | gem / group :test blocks |
flake.nix | nix | inputs attrset (dotted and block forms) |
Package naming
A package’s name is the join key everything else carries (symbols.package,
dependencies.package, the package filter on every MCP tool), so it is
never empty:
| Manifest | Name |
|---|---|
package.json, pyproject.toml, Cargo.toml | the declared name |
go.mod | the last segment of the module path |
pom.xml, build.gradle | group:artifact (artifact alone when there is no group) |
cpanfile, Gemfile, flake.nix | no name field exists — see below |
When a manifest declares no name (a Gemfile, a tooling-only root
pyproject.toml, a private package.json), the name is derived from its
location: a nested manifest takes its directory path with / replaced by -
(services/api/Gemfile → services-api), and a manifest at the repo root
takes the repository directory’s own name.
Two Gradle subprojects can compute the same group:projectName (two
directories both called app). The one indexed first keeps that name; the
colliding one falls back to its path-derived name, with -2, -3… appended
if that name is taken as well. A warning naming both directories is logged,
and neither package is dropped.
Symbol extraction
Shire extracts symbols (functions, classes, types, methods, interfaces) from source files using tree-sitter, with full signatures, parameters, and return types.
Private and unexported symbols are indexed too, not skipped. Every symbol
carries a visibility — public, protected, internal or private —
derived from the language’s own convention (the Visibility column below). A
member is narrowed by its enclosing type: a public method of a private class
is private. Search ranks private symbols after the others (see
MCP Tools), and symbols.include_private = false drops them
(see Configuration) — only private ones: internal
(e.g. Java package-private, Rust pub(crate)) and protected symbols are
always kept. Where a language has no visibility rule that Shire reads, every
symbol is public.
| Language | Extractor | Visibility |
|---|---|---|
| TypeScript / JavaScript | tree-sitter | Module-level declarations: exported (export ... or named in an export { ... } clause) is public, anything else private. Methods: private / #name / protected modifiers. CommonJS module.exports is not recognised. |
| Go | tree-sitter | Capitalised name public, otherwise private |
| Rust | tree-sitter | pub → public; pub(crate) / pub(super) / pub(in …) → internal; no modifier or pub(self) → private. Trait-impl methods are public. |
| Python | tree-sitter | Leading _ → private; dunder names (__init__) are public |
| Java | tree-sitter | public / protected / private; package-private (no modifier) → internal. Interface members and enum constants are implicitly public. |
| Kotlin | tree-sitter | private / protected / internal; no modifier → public |
| Dart | tree-sitter | Leading _ → private (including named constructors such as Foo._internal) |
| Protobuf | tree-sitter | all public |
| C | tree-sitter | static → private, otherwise public |
| C++ | tree-sitter | Class members from the nearest public: / protected: / private: label (default private in a class, public in a struct); non-member static → private |
| C# | tree-sitter | public / protected / internal / private; with no modifier a class member is private, an interface member public, a top-level type internal |
| Swift | tree-sitter | private / fileprivate → private; explicit internal → internal; no modifier → public |
| PHP | tree-sitter | private / protected; no modifier → public |
| Scala | tree-sitter | private / private[this] → private; private[pkg] → internal; protected |
| Zig | tree-sitter | pub → public, otherwise private |
| Bash / Shell | tree-sitter | all public |
| R | tree-sitter | all public |
| Haskell | tree-sitter | all public (export lists are not read) |
| YAML | tree-sitter | all public |
| SQL | tree-sitter | all public |
| HCL / Terraform | tree-sitter | all public |
| TOML | tree-sitter | all public |
| Perl | tree-sitter | Leading _ → private |
| Ruby | tree-sitter | Methods after a bare private / protected line, or written private def …; private :name is not tracked |
| OCaml | tree-sitter | all public (.mli signatures are not read) |
| Lua | tree-sitter | local function / local f = function → private |
| Elixir | tree-sitter | defp / defmacrop / defguardp / @typep → private |
| Clojure | tree-sitter | defn- and ^:private metadata → private |
| Erlang | tree-sitter | all public (-export lists are not read) |
| Julia | tree-sitter | all public (export statements are not read) |
| Gleam | tree-sitter | pub → public, otherwise private |
| Odin | tree-sitter | all public |
| Nix | tree-sitter | all public |
| Nim | tree-sitter | * export marker → public, otherwise private |
| COBOL | regex-based | all public |
An index built by an older Shire picks the private symbols up on its first build after upgrading: the extractor version is stored in the index, and a mismatch re-extracts every source file once.
Reference extraction
Shire extracts cross-references (calls, type references, imports, and interface implementations) for a subset of languages. These are stored in the symbol_refs table and exposed via the symbol_references, symbol_callers, and symbol_callees MCP tools.
| Language | Call | Type | Import | Impl |
|---|---|---|---|---|
| Go | yes | yes | yes | — (implicit interfaces) |
| Python | yes | yes | yes | yes |
| Java | yes | yes | yes | yes |
| TypeScript | yes | yes | yes | yes |
| JavaScript | yes | — | yes | yes |
| Perl | yes | — | yes | — |
| Ruby | yes | yes | yes | yes |
| Scala | yes | yes | yes | yes |
All other languages: symbol definitions only; references are not extracted.
Performance
Shire is designed to index large monorepos quickly and answer queries instantly. This page documents benchmark methodology, results, and how to reproduce them.
Test repos
Benchmarks run against three real-world open-source monorepos covering a range of sizes and ecosystems:
| Repo | Size | Packages | Symbols | Files | Primary languages |
|---|---|---|---|---|---|
| turborepo | small | 400 | 10,686 | 5,451 | Rust, TypeScript, Go |
| grafana | medium | 28 | 35,104 | 14,054 | Go, TypeScript |
| kubernetes | large | 34 | 78,458 | 18,275 | Go |
Build performance
Full rebuild (no incremental cache), median of 4 iterations after a warmup run:
| Repo | Median | Min | P95 | Std dev |
|---|---|---|---|---|
| turborepo | 571ms | 525ms | 678ms | 58ms |
| grafana | 1,150ms | 1,025ms | 1,172ms | 58ms |
| kubernetes | 1,703ms | 1,607ms | 1,897ms | 108ms |
Build time scales roughly linearly with symbol count. The pipeline is parallelized with rayon across packages and files, with batched multi-row SQLite inserts within explicit transactions.
Incremental builds are significantly faster – only packages with changed source files are re-extracted, and an mtime pre-check skips SHA-256 computation entirely for untouched packages.
2026-09-07: The table above predates two behavior changes not yet re-benchmarked end to end. (1) Incremental builds now always scan the working tree directly instead of relying on
.git/index, so unstaged edits are indexed on the next build — this adds a fixed per-build walk cost. Measured on turborepo: a no-op incremental build went from 33ms to 173ms, a one-file incremental build from 298ms to 170ms, and a full rebuild from 1,769ms to 1,558ms. (2)search_symbolsnow matches by prefix and against identifier sub-tokens (e.g.verify jwtmatchesverifyJwtToken) rather than exact substring, which changes result sets but was not separately benchmarked for latency.
Query performance
Median latency over 100 iterations per query:
| Query | Small | Medium | Large |
|---|---|---|---|
search_symbols("parse") | 0.09ms | 0.09ms | 0.04ms |
search_symbols("Config") | 0.20ms | 0.29ms | 1.01ms |
search_files("mod") | 0.05ms | 0.03ms | 0.04ms |
search_files("test") | 0.07ms | 0.60ms | 1.99ms |
list_packages(None) | 0.11ms | 0.01ms | 0.01ms |
All queries use SQLite FTS5 full-text search with the unicode61 tokenizer. packages_fts and docs_fts carry a prefix='2,3' index; symbols_fts and files_fts deliberately do not (a prefix query there walks a term range instead). Query latency depends primarily on result set size, not total index size.
Reproducing benchmarks
Shire includes an autoresearch binary for reproducible benchmarking. It is a
development tool, not part of the shipped CLI: it lives behind the non-default
bench cargo feature, so every invocation needs --features bench.
The lifecycle and quality phases rewrite source files in the target repo.
They refuse to run against a repo with uncommitted changes to tracked files
(git status --porcelain --untracked-files=no must be empty, and an
indeterminable state — no git, not a repo — is treated as dirty). Untracked
paths are ignored, since the benchmark and setup-bench-repo.sh leave their
own artifacts (shire.toml, .shire/bench.db) in the repo. Each phase
restores exactly the bytes it captured before its first write, including when
it aborts part way through a run. If the guard turns away every repo, the
phase reports the skipped repos and exits non-zero rather than printing an
empty result with a success status.
Setup
Run the benchmark repo setup script to clone and prepare the test repos:
scripts/setup-bench-repo.sh
This clones the three repos into ~/.cache/shire-bench/ and creates a shire.toml in each.
Running benchmarks
# Build the benchmark binary
cargo build --release --features bench --bin autoresearch
# Run build benchmarks (all repos)
cargo run --release --features bench --bin autoresearch -- --phase build
# Run query benchmarks (all repos)
cargo run --release --features bench --bin autoresearch -- --phase query
# Filter by repo size
cargo run --release --features bench --bin autoresearch -- --phase build --size small
# Point at a specific repo (must have a clean worktree for lifecycle/quality)
cargo run --release --features bench --bin autoresearch -- --phase build --repo /path/to/repo
Build benchmarks run 5 iterations (1 warmup + 4 measured) per repo. Query benchmarks run 100 iterations per query. Results are printed as JSON to stdout.
Environment notes
- Results vary by machine (CPU, disk speed, available memory)
- Close other applications for more stable measurements
- The warmup iteration primes filesystem caches and SQLite page cache
- Numbers in this document were captured on an Apple M-series Mac
Architecture
src/
├── main.rs # CLI (clap): build, serve, watch, rebuild, init, install, uninstall, clean, status subcommands
├── lib.rs # Library re-exports for embedding shire as a crate
├── claude_mod.rs # Optional Claude Code status mod: embedded files, install/refresh/uninstall
├── config.rs # shire.toml parsing
├── git.rs # Git worktree detection and repo root resolution
├── init.rs # `shire init` setup (config, MCP server, hooks, rules)
├── install.rs # `shire install`/`uninstall` — registers/removes shire as an MCP server across detected AI tools
├── logging.rs # Rotating file logging (tracing-appender), per-session IDs
├── status.rs # `shire status` — read-only index/build/daemon snapshot (text or JSON)
├── db/
│ ├── mod.rs # SQLite schema, open/create
│ └── queries.rs # FTS search, dependency graph BFS, listing
├── index/
│ ├── mod.rs # Walk + incremental index orchestrator
│ ├── custom_discovery.rs # Config-driven custom package discovery
│ ├── manifest.rs # ManifestParser trait
│ ├── hash.rs # SHA-256 content hashing for incremental builds
│ ├── lock.rs # Cross-process build lock (flock on <db_path>.lock)
│ ├── ref_writer.rs # Cross-reference write strategy threaded through the build phases
│ ├── npm.rs # package.json parser (workspace: protocol)
│ ├── go.rs # go.mod parser
│ ├── go_work.rs # go.work parser (workspace use directives)
│ ├── cargo.rs # Cargo.toml parser (workspace dep resolution)
│ ├── python.rs # pyproject.toml parser
│ ├── maven.rs # pom.xml parser (parent POM inheritance)
│ ├── gradle.rs # build.gradle / build.gradle.kts parser
│ ├── gradle_settings.rs # settings.gradle parser (project inclusion)
│ ├── perl.rs # cpanfile parser (requires, on 'test')
│ ├── ruby.rs # Gemfile parser (gem, group blocks)
│ └── nix.rs # flake.nix parser (inputs attrset, dotted and block forms)
├── symbols/
│ ├── mod.rs # Symbol types, kind-agnostic extraction orchestrator
│ ├── walker.rs # Source file discovery (extension filtering, excludes)
│ ├── registry.rs # Language registry: maps extensions to tree-sitter grammars + hooks
│ ├── query_extract.rs # Generic tree-sitter query executor with hook callbacks
│ ├── queries/ # Tree-sitter .scm query files (one per language)
│ ├── hooks/ # Language-specific hooks (visibility, signatures, params, post-processing)
│ └── cobol.rs # COBOL extractor (regex-based — the only non-tree-sitter language)
│ # Cross-reference extraction (call, type, import, impl) is supported
│ # for 8 tier-1 languages: Go, Python, Java, TypeScript, JavaScript,
│ # Perl, Ruby, Scala. References are captured via @reference.* captures
│ # in the language's .scm query and written to the symbol_refs table.
│ # Coverage is asymmetric per language: JavaScript omits Type refs
│ # (no type system), and Go/Perl omit Impl refs (no extends/implements).
├── mcp/
│ ├── mod.rs # MCP server setup (rmcp, stdio transport)
│ ├── tools.rs # 17 tool handlers
│ └── prompts.rs # 2 prompt templates (explore, reference_audit)
├── watch/
│ ├── mod.rs # Daemon event loop (UDS listener, debounce, rebuild)
│ ├── daemon.rs # Process management (start/stop/is_running via PID)
│ └── protocol.rs # Hook input parsing, Bash read-only allowlist
└── bin/
└── autoresearch.rs # Benchmark harness, gated behind the non-default `bench` feature
What an incremental build re-reads
A build re-hashes a package’s source files when any of them changes mtime, size, or path. Between them those cover the ordinary cases: an editor save moves the mtime, a rename or a new/deleted file moves the path set, and a restore that rewinds an mtime is still caught if the size differs.
One case is not covered: an edit that keeps the file’s size and its mtime.
That is what cp -p/install -p out of a build cache, rsync -a, a tar -x
restore and some patch tools do. Nothing on disk distinguishes it from an
untouched file without reading every byte of the repo on every build, so
shire does not try — run shire build --force after one.
Which package owns a file
A source file’s symbols and references belong to its nearest enclosing
package only — the same longest-prefix rule the files table uses. Each
package’s source walk stops at the directories of packages nested inside it,
so a file under pkgs/a/sub/b/ is extracted once, for b, not once more for
a and again for the root package.
Packages found by discovery.custom rules take part in incremental builds like
any other: their edits are picked up, and a nested manifest package appearing
or disappearing under one moves the affected files out of it or back into it.
A custom package whose directory stops matching its rule is not removed,
though, even by shire build --force; run shire clean and rebuild after
reorganising those.
Indexes written by shire 0.7.0 and earlier extracted nested files once per
ancestor package, and could leave a nested file’s references under an ancestor
alone. The first build against such an index re-extracts every package once
(comparable in cost to a full rebuild) and records the current
extractor_state in shire_meta; later builds are incremental again. The
same one-time pass runs whenever extractor_state — the extractor version
plus symbols.include_private — changes: after an upgrade whose extractor
emits different symbols, or when that option is toggled. If the build is
interrupted, the next one repeats the pass; a package whose walk failed
during it is recorded in pending_source_reextract and retried on the next
build.
When a walk cannot see the whole tree
Both walks a build makes — the repo-wide file walk and the per-package source walk — distinguish “this is not there” from “I could not look”. The difference decides whether rows may be deleted.
An unreadable path (a directory whose permissions changed, a network share
that blipped, a container volume mid-remount) is recorded as a blind spot. Rows
under it are kept rather than deleted, no file-tree hash is stored (so the next
build re-checks the tree instead of short-circuiting), and a package whose
source walk failed keeps every symbol, reference and file hash it already had.
shire build reports that package and exits non-zero; the watch daemon and the
MCP server’s on-demand rebuild only log it, because the build itself committed.
Hitting the 500,000-file cap is the same thing: the walk stopped partway
through a nondeterministic traversal, so it is treated as a blind spot over the
whole tree — nothing is deleted, no file-tree hash is stored, and the build
prints a warning naming the cap. Exclude directories in discovery.exclude to
bring a repository back under it.
A malformed pattern in a committed .gitignore is not a blind spot: the
ignore crate reports it as an error but still walks the whole tree, so it is
warned about once and skipped.
The build lock
Every build holds an advisory flock on <db_path>.lock for its whole run, so
two builders cannot both take the insert-only full-build path and double every
symbol. shire build and the watch daemon wait for a competing build (up to
index::lock::LOCK_TIMEOUT, minutes rather than seconds) and then report the
conflict; the MCP server’s per-tool-call rebuild skips instead and logs it,
since its trigger comes round again.
shire clean takes the same lock before removing anything, and refuses with
“a shire build is running” rather than deleting a database out from under a
build. Neither command creates the lock file until it has established that
db_path is shire’s to use: db_path comes from the repository’s own
shire.toml, so before the lock file is created build checks what is already
at that path. It refuses a symlink (the open that creates the index would
follow it, and a symlinked index cannot be auto-repaired either — db_path
must name a regular file), a file that is not a SQLite database, and a SQLite
database holding tables of its own and no shire_meta, the mark every shire
index carries. It will not write its schema into a database it did not create.
A database that cannot be inspected at all — damaged, or held by another
process mid-write — is not decided before the lock: a shire build already
running against the same db_path looks exactly like that, since builds run
under journal_mode=MEMORY, whose write transactions block readers. Taking the
lock waits that build out, and the check runs again under it. If the database
still cannot be identified it is refused, unless it sits in <repo>/.shire/ or
~/.claude/shire/, where a damaged index is the only thing it can be and is
rebuilt as before. The cost of asking twice is that a foreign database held by
its own writer gets an empty <db_path>.lock sidecar beside it before the
refusal; the database itself is never touched.
The schema is created in a single transaction, so an interrupted first build
leaves either a complete index or an empty file — never a half-schema that the
checks above could only read as someone else’s database. The <db_path>.lock file itself is left in place — it is an empty
sidecar like -wal/-shm, and unlinking it while a builder holds a lock on
that inode is what would let a second builder take the lock at the same path.
symbol_refs table
The symbol_refs table stores cross-reference records extracted alongside symbol definitions. Each row captures a reference to a named symbol:
| Column | Type | Description |
|---|---|---|
name | TEXT | The name being referenced (function, type, module, etc.) |
kind | TEXT | One of: call, type, import, impl |
file_id | INTEGER | REFERENCES files(id) ON DELETE CASCADE — the file containing the reference |
line | INTEGER | Line number of the reference |
package | TEXT | Package the referencing file belongs to (nullable) |
enclosing_symbol | TEXT | Dot-qualified path of the enclosing scopes, innermost last (AuthService.login, Outer.outer.Inner.run); nullable |
file_id stores a compact reference into files(id) rather than a duplicated path string; phase_index_files runs before symbol extraction so files is already populated when refs are inserted, and read queries JOIN files to resolve file_path. B-tree indexes on name, file_id, and enclosing_symbol (plus composite and partial covering indexes for the callers/callees/package-scoped queries) support the exact-match lookups used by the symbol_references, symbol_callers, and symbol_callees MCP tools. No FTS5 table — reference queries are exact-name only.
Incremental behavior mirrors symbol extraction: references for a file are dropped and re-extracted whenever the file’s SHA-256 hash changes. No separate pass is needed — references are extracted in the same tree-sitter walk as symbol definitions.