Topic 03: Content-Hash Incremental Indexing & Manifest Diffing

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering SHA-1 content hashing, manifest diffing algorithms, SQL OR-chained row invalidation, and real-world benchmark metrics (1.49s no-op vs multi-minute full rebuild).


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. Manifest Check Load cached SHA-1 manifest _load_manifest() reading JSON If missing or --full passed, trigger _build_full() indexer.py:113-115
2. Content Hashing Compute SHA-1 per discovered file hashlib.sha1(data).hexdigest() Cryptographic hash avoids mtime false-positives (e.g. git checkout touch) indexer.py:122-126
3. Set Diffing Identify changed and deleted sets set(old) - set(new) for deletes; hash diff for changes Isolates exactly which files require LanceDB row updates indexer.py:128-129
4. Row Invalidation Delete stale rows from Lance table _sql_or() building SQL predicate Atomic deletion of obsolete rows before inserting new chunks indexer.py:132-134
5. Selective Append Chunk & embed only changed files chunk_file() + table.add() Limits expensive embedding inference strictly to modified files indexer.py:137-141

📊 Performance Benchmarks & Empirical Results

The incremental indexing engine was benchmarked against the real Conductor codebase (495 files, 5,231 chunks, including C, Python, Shell, and JSON):

Scenario Files Scanned Files Re-indexed Chunks Added / Deleted Wall-Clock Time
Full Rebuild (Cold Start) 495 495 +5,231 / -0 ~2+ minutes (background run)
Incremental No-Op (Clean Repo) 495 0 0 / 0 1.49 seconds
Single-File Modification (1 line added) 495 1 +1 / -1 0.984 seconds
Revert Change (git checkout) 495 1 +1 / -1 ~1.0 second

🔍 Under-the-Hood Deep Dive: The Diff Algorithm

The incremental update logic in indexer.py guarantees transactional consistency:

# Set difference for removed files
removed_rels = set(old_manifest) - set(new_manifest)
changed_rels = {str(p.relative_to(repo_root)) for p in changed_paths}

# Atomic row deletion via OR predicate
to_delete = changed_rels | removed_rels
if to_delete:
    table.delete(_sql_or("file", to_delete))

# Re-chunk and embed only changed files
new_chunks = []
for path in changed_paths:
    new_chunks.extend(chunk_file(path, repo_root))
rows = _rows(new_chunks)
if rows:
    table.add(rows)

# Atomically update disk manifest
_save_manifest(repo_root, new_manifest)

💡 Tough Interview Questions & Detailed Answers

Q1: "Why use SHA-1 content hashing rather than file modification time (mtime)?"

Answer: mtime timestamps are notoriously unreliable in software development workflows. Running git checkout, switching git branches, running build scripts, or touching files updates mtime without altering file contents, which would trigger expensive, unnecessary re-embeddings across hundreds of files. Cryptographic SHA-1 hashing ensures that a file is only re-embedded when its actual bytes change, guaranteeing 100% precision.

Q2: "What is the trade-off of file-level vs chunk-level incremental diffing?"

Answer: Sherlock uses file-level diffing: modifying one line in a 50-function file causes all 50 functions in that file to be re-embedded. A chunk-level diffing engine would hash individual AST nodes and only re-embed the mutated function. While chunk-level diffing would save embedding cycles on massive files, it adds significant complexity (tracking AST node hashes, handling line shifts for downstream functions in the same file). For codebases up to 1,000 files, file-level diffing provides sub-second re-indexing with minimal architectural complexity.