BreadDocs
Browse documentation

Technical contract

Codebase indexing

Index tiers, language recognition, freshness, and graph limits.

Open source

Bread builds a workspace-local codebase index for agent retrieval. Index state, optional semantic assets, and watcher checkpoints remain in .bread/ and are never uploaded by the indexer.

Language recognition

bread::engines::language_catalog::detect(path, source) returns a stable language identifier. The native registry recognizes 115 identifiers through conventional filename, compound suffix, extension, Emacs modeline, shebang, or document signature.

The hot path is deliberately small:

  1. Exact conventional filename.
  2. Direct compound suffix or extension match.
  3. At most the first 512 source bytes for modelines, shebangs, HTML/XML, or a unified-diff signature.
  4. plaintext when no safe classification exists.

Filename and extension recognition uses static direct dispatch. It does not load metadata, construct glob rules, scan a registry, or allocate. Source inspection is skipped for a known filename or extension. An unknown or non-text-like input remains indexable through the lexical tier; classification failure never blocks indexing.

SUPPORTED_LANGUAGE_IDS is the public inventory used by regression tests. Adding a language requires an explicit native rule and a representative test sample. Classifications are intentionally deterministic rather than classifier-based: ambiguous extensions use a narrow content discriminator only where it is necessary (.m and .pp), otherwise they follow the documented native rule.

Indexing tiers

TierScopeResult
Native structuralRust, Python, JavaScript, TypeScript/TSXTree-sitter declaration ranges, syntax status, and tested local imports
Bounded structuralOther recognized languages with a conservative declaration formAt most 256 named declarations and bounded source chunks
LexicalUnknown or declaration-free textExact and lexical chunks without guessed structure
ExcludedDenied, ignored, .bread/, and .git pathsNever indexed

A language label does not imply compiler-accurate definitions, references, or type information. Generic structural extraction only recognizes conservative declaration patterns. It does not claim AST precision, and no identifier is discarded when a declaration is not recognized: lexical retrieval remains available.

Dependency graph and freshness

Local graph edges are intentionally narrower than language recognition. Bread resolves only tested repository-local forms: Rust, Python, JavaScript and TypeScript imports; quoted C/C++ includes; and relative shell, Ruby, and PHP imports. It never guesses package, URL, absolute-path, standard-library, or bare-import edges.

Generations are published atomically. A query observes either the prior complete generation or its replacement, never an in-progress graph. Deleted and stat-mismatched chunks are excluded. While a replacement is pending, a stale candidate receives a bounded direct search of that path.

Safety boundaries

  • Ignore and deny rules run before indexing; restricted files never enter index records or prompt context.
  • Malformed source, missing optional semantic assets, parser errors, and unavailable daemons fail open to the strongest safe lower tier.
  • Prompt context remains provenance-marked, session-deduplicated, score-gated, and budget-bounded. Vague or instruction-like prompts receive no injected repository context.
  • No hook runs compilers, language servers, network downloads, or external language classifiers.

Validation contract

Every registry identifier has a native detection sample. The deterministic Codex matrix contains exactly 240 cases:

  • 40 SessionStart static inventory-cue responses;
  • 40 restricted-path PreToolUse denials;
  • 40 safe PreToolUse pass-throughs;
  • 20 concise and 20 oversized Bash PostToolUse results;
  • 40 eligible UserPromptSubmit context injections; and
  • 40 vague or instruction-like prompt no-injection checks.

The matrix is a local adapter regression suite, not 200 paid remote calls. A separate low-cost live Codex sandbox smoke validates the installed-hook path. The current live target is Codex CLI only, using the shared isolated sandbox and its configured low-cost model.

Performance methodology

The default build includes no additional parser grammars or registry runtime dependency. The indexing benchmark uses a deterministic fixture of 2,000 one-line declarations, five warmed release runs, and median values. Release size and clean release build time are recorded separately. It does not govern the README's live Codex sandbox rows, which identify their own sample sizes, or the scorecard's 128-event pass-through percentile. Measurements are hardware- and corpus-specific regression evidence, not universal performance claims. The current observed A/B values are maintained in the README.