Bread builds a workspace-local codebase index for agent retrieval. Index state,
optional semantic assets, and watcher checkpoints remain in .bread/ and are
never uploaded by the indexer.
Language recognition
bread::engines::language_catalog::detect(path, source) returns a stable
language identifier. The native registry recognizes 115 identifiers through
conventional filename, compound suffix, extension, Emacs modeline, shebang, or
document signature.
The hot path is deliberately small:
- Exact conventional filename.
- Direct compound suffix or extension match.
- At most the first 512 source bytes for modelines, shebangs, HTML/XML, or a unified-diff signature.
plaintextwhen no safe classification exists.
Filename and extension recognition uses static direct dispatch. It does not load metadata, construct glob rules, scan a registry, or allocate. Source inspection is skipped for a known filename or extension. An unknown or non-text-like input remains indexable through the lexical tier; classification failure never blocks indexing.
SUPPORTED_LANGUAGE_IDS is the public inventory used by regression tests.
Adding a language requires an explicit native rule and a representative test
sample. Classifications are intentionally deterministic rather than
classifier-based: ambiguous extensions use a narrow content discriminator only
where it is necessary (.m and .pp), otherwise they follow the documented
native rule.
Indexing tiers
| Tier | Scope | Result |
|---|---|---|
| Native structural | Rust, Python, JavaScript, TypeScript/TSX | Tree-sitter declaration ranges, syntax status, and tested local imports |
| Bounded structural | Other recognized languages with a conservative declaration form | At most 256 named declarations and bounded source chunks |
| Lexical | Unknown or declaration-free text | Exact and lexical chunks without guessed structure |
| Excluded | Denied, ignored, .bread/, and .git paths | Never indexed |
A language label does not imply compiler-accurate definitions, references, or type information. Generic structural extraction only recognizes conservative declaration patterns. It does not claim AST precision, and no identifier is discarded when a declaration is not recognized: lexical retrieval remains available.
Dependency graph and freshness
Local graph edges are intentionally narrower than language recognition. Bread resolves only tested repository-local forms: Rust, Python, JavaScript and TypeScript imports; quoted C/C++ includes; and relative shell, Ruby, and PHP imports. It never guesses package, URL, absolute-path, standard-library, or bare-import edges.
Generations are published atomically. A query observes either the prior complete generation or its replacement, never an in-progress graph. Deleted and stat-mismatched chunks are excluded. While a replacement is pending, a stale candidate receives a bounded direct search of that path.
Safety boundaries
- Ignore and deny rules run before indexing; restricted files never enter index records or prompt context.
- Malformed source, missing optional semantic assets, parser errors, and unavailable daemons fail open to the strongest safe lower tier.
- Prompt context remains provenance-marked, session-deduplicated, score-gated, and budget-bounded. Vague or instruction-like prompts receive no injected repository context.
- No hook runs compilers, language servers, network downloads, or external language classifiers.
Validation contract
Every registry identifier has a native detection sample. The deterministic Codex matrix contains exactly 240 cases:
- 40
SessionStartstatic inventory-cue responses; - 40 restricted-path
PreToolUsedenials; - 40 safe
PreToolUsepass-throughs; - 20 concise and 20 oversized Bash
PostToolUseresults; - 40 eligible
UserPromptSubmitcontext injections; and - 40 vague or instruction-like prompt no-injection checks.
The matrix is a local adapter regression suite, not 200 paid remote calls. A separate low-cost live Codex sandbox smoke validates the installed-hook path. The current live target is Codex CLI only, using the shared isolated sandbox and its configured low-cost model.
Performance methodology
The default build includes no additional parser grammars or registry runtime dependency. The indexing benchmark uses a deterministic fixture of 2,000 one-line declarations, five warmed release runs, and median values. Release size and clean release build time are recorded separately. It does not govern the README's live Codex sandbox rows, which identify their own sample sizes, or the scorecard's 128-event pass-through percentile. Measurements are hardware- and corpus-specific regression evidence, not universal performance claims. The current observed A/B values are maintained in the README.