Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering SHA-1 content hashing, manifest diffing algorithms, SQL OR-chained row invalidation, and real-world benchmark metrics (1.49s no-op vs multi-minute full rebuild).
Opening & Scope:
"In this subsystem (indexer.py:112 (build_index)), I engineered a content-hash incremental re-indexing pipeline that reduces re-indexing times on large codebases from minutes to sub-second runs by caching per-file SHA-1 signatures."
Step 1: The Full-Rebuild Bottleneck:
"In the naive implementation, re-indexing required re-parsing and re-embedding every single file in the repository on every invocation. When testing against a production repository containing 495 files and 5,231 chunks, a full rebuild took several minutes and exceeded execution timeouts. For a CLI tool to be usable during active development, re-indexing must be near-instantaneous."
Step 2: The SHA-1 Manifest Architecture:
"To track file state across runs, I introduced a lightweight manifest at .sherlock/manifest.json (indexer.py:21). Stored as a simple JSON dictionary mapping relative file paths to their 40-character hex SHA-1 digests ({rel_path: sha1_hex}), this manifest persists the exact cryptographic state of every indexed file."
Step 3: Fast Diffing Algorithm (Changed vs Removed):
"When build_index() runs, it loads the prior manifest via _load_manifest().
if old_manifest.get(rel) != h: changed_paths.append(path).removed_rels = set(old_manifest) - set(new_manifest).Step 4: Atomic Row Invalidation via SQL OR Chains:
"For changed or deleted files, old chunks in LanceDB must be purged to prevent duplicate or phantom results. In _sql_or() (indexer.py:80), I built a query generator that escapes single quotes and constructs an atomic SQL predicate: file = 'path1' OR file = 'path2'. In indexer.py:134, table.delete(to_delete) purges all stale rows in a single batch before newly embedded chunks are appended via table.add(rows)."
Step 5: Verified Benchmark Results:
"I validated this pipeline on the 495-file Conductor repository:
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. Manifest Check | Load cached SHA-1 manifest | _load_manifest() reading JSON |
If missing or --full passed, trigger _build_full() |
indexer.py:113-115 |
| 2. Content Hashing | Compute SHA-1 per discovered file | hashlib.sha1(data).hexdigest() |
Cryptographic hash avoids mtime false-positives (e.g. git checkout touch) |
indexer.py:122-126 |
| 3. Set Diffing | Identify changed and deleted sets | set(old) - set(new) for deletes; hash diff for changes |
Isolates exactly which files require LanceDB row updates | indexer.py:128-129 |
| 4. Row Invalidation | Delete stale rows from Lance table | _sql_or() building SQL predicate |
Atomic deletion of obsolete rows before inserting new chunks | indexer.py:132-134 |
| 5. Selective Append | Chunk & embed only changed files | chunk_file() + table.add() |
Limits expensive embedding inference strictly to modified files | indexer.py:137-141 |
The incremental indexing engine was benchmarked against the real Conductor codebase (495 files, 5,231 chunks, including C, Python, Shell, and JSON):
| Scenario | Files Scanned | Files Re-indexed | Chunks Added / Deleted | Wall-Clock Time |
|---|---|---|---|---|
| Full Rebuild (Cold Start) | 495 | 495 | +5,231 / -0 | ~2+ minutes (background run) |
| Incremental No-Op (Clean Repo) | 495 | 0 | 0 / 0 | 1.49 seconds |
| Single-File Modification (1 line added) | 495 | 1 | +1 / -1 | 0.984 seconds |
| Revert Change (git checkout) | 495 | 1 | +1 / -1 | ~1.0 second |
The incremental update logic in indexer.py guarantees transactional consistency:
# Set difference for removed files
removed_rels = set(old_manifest) - set(new_manifest)
changed_rels = {str(p.relative_to(repo_root)) for p in changed_paths}
# Atomic row deletion via OR predicate
to_delete = changed_rels | removed_rels
if to_delete:
table.delete(_sql_or("file", to_delete))
# Re-chunk and embed only changed files
new_chunks = []
for path in changed_paths:
new_chunks.extend(chunk_file(path, repo_root))
rows = _rows(new_chunks)
if rows:
table.add(rows)
# Atomically update disk manifest
_save_manifest(repo_root, new_manifest)
mtime)?"Answer: mtime timestamps are notoriously unreliable in software development workflows. Running git checkout, switching git branches, running build scripts, or touching files updates mtime without altering file contents, which would trigger expensive, unnecessary re-embeddings across hundreds of files. Cryptographic SHA-1 hashing ensures that a file is only re-embedded when its actual bytes change, guaranteeing 100% precision.
Answer: Sherlock uses file-level diffing: modifying one line in a 50-function file causes all 50 functions in that file to be re-embedded. A chunk-level diffing engine would hash individual AST nodes and only re-embed the mutated function. While chunk-level diffing would save embedding cycles on massive files, it adds significant complexity (tracking AST node hashes, handling line shifts for downstream functions in the same file). For codebases up to 1,000 files, file-level diffing provides sub-second re-indexing with minimal architectural complexity.