4–5 Minute Detailed Overview: Sherlock (Local Codebase RAG & Debugger)

Target Duration: 4–5 minutes (~650–800 spoken words)
Goal: Deliver an end-to-end architectural explanation covering AST parsing, embedded vector tables, incremental manifest diffing, multi-modal query routing, traceback interval matching, and offline LLM synthesis. Plant deep-dive hooks along the way.


🎙️ Spoken Script & Section Walkthrough

1. High-Level Architectural Thesis (30s)

"To explore how modern developer assistants can achieve privacy, sub-second latency, and zero API costs, I designed Sherlock—a local-first codebase RAG and runtime debugging system built in Python.

Rather than relying on remote hosted services or naive line-by-line grep, Sherlock combines compile-time syntactic analysis with in-process vector mathematics. It breaks code down into semantic AST units, embeds them locally, stores them in an embedded columnar vector table, routes queries intelligently through rule-based heuristics, and synthesizes answers using a local Ollama model.

I structured the system into five core functional pillars:

  1. Syntax-Aware AST Chunking & Whitelisting Engine
  2. In-Process Embedded Vector Database & Local Embeddings
  3. Content-Hash Incremental Indexing & Manifest Diffing
  4. Multi-Modal Query Routing & Exact Symbol Prefiltering
  5. Traceback Frame Extraction & Local Offline Synthesis
"


2. Pillar 1: Syntax-Aware AST Chunking & Grammar Mapping (60s)

"Standard RAG systems divide documents using fixed character lengths or line sliding windows. On source code, this breaks functions in half and ruins variable scope.

In chunker.py:98 (chunk_file), I implemented AST-aware chunking using Tree-sitter.

(Hook: Deep dive into topic_01_ast_chunking_and_tree_sitter)


3. Pillar 2: In-Process Columnar Vector Table & Local Embeddings (60s)

"Rather than deploying an external database cluster like Milvus or Pinecone, I built the storage layer around LanceDB—an embedded, serverless vector database operating directly on local NVMe disk via Lance columnar files (.sherlock/db).

Key architectural properties include:

(Hook: Deep dive into topic_02_vector_storage_and_fastembed)


4. Pillar 3: Content-Hash Incremental Indexing & Manifest Diffing (60s)

"In production repositories, developers modify a handful of files at a time. Re-embedding thousands of unchanged files on every run causes intolerable latency bottlenecks.

In indexer.py:112 (build_index), I engineered an incremental indexing pipeline driven by a SHA-1 content manifest (.sherlock/manifest.json):

(Hook: Deep dive into topic_03_incremental_indexing_and_manifest)


5. Pillar 4: Multi-Modal Query Routing & Symbol Resolution (45s)

"User queries vary drastically—from raw error tracebacks and function names to free-form architectural questions. Relying on an LLM to classify query intent introduces 500ms of unnecessary latency.

In router.py:60 (route), I implemented a zero-latency heuristic cascade:

(Hook: Deep dive into topic_04_query_routing_and_symbol_resolution)


6. Pillar 5: Traceback Frame Extraction & Offline LLM Synthesis (60s)

"For error diagnosis and code explanation, Sherlock bridges retrieval into grounded synthesis without cloud dependencies:

(Hook: Deep dive into topic_05_traceback_frame_parsing_and_diagnosis or topic_06_local_llm_synthesis_and_cli)


7. Hands-On Verification & Usage

Sherlock is fully verifiable using the command line:

# 1. Build or incrementally update the local index for any target repository
uv run sherlock index --path /path/to/repo

# 2. Ask architectural questions or look up specific symbols with code citations
uv run sherlock ask --path /path/to/repo "what does conductor.sh do?"
uv run sherlock ask --path /path/to/repo "divide" --chunks --show-code

# 3. Diagnose runtime errors directly from piped stack traces
python failing_script.py 2>&1 | uv run sherlock diagnose --path /path/to/repo

# 4. Launch the interactive multi-turn terminal REPL
uv run sherlock chat --path /path/to/repo

# 5. Run the end-to-end integration smoke test
uv run python test_smoke.py

8. Summary & Wrap-up (30s)

"In summary, Sherlock gave me deep, practical experience building local AI systems: navigating compiler AST structures with Tree-sitter, managing embedded vector tables with LanceDB, amortizing re-indexing through content-addressable hashing, and orchestrating offline LLM synthesis with zero cloud telemetry."


🧭 Interviewer Follow-Up Mapping

graph LR
    DetailedSpeech[4-5 Min Detailed Speech] --> T1[topic_01: AST Chunking & Tree-sitter]
    DetailedSpeech --> T2[topic_02: Vector Storage & FastEmbed]
    DetailedSpeech --> T3[topic_03: Incremental Indexing & Manifest]
    DetailedSpeech --> T4[topic_04: Query Routing & Symbol Resolution]
    DetailedSpeech --> T5[topic_05: Traceback Parsing & Diagnosis]
    DetailedSpeech --> T6[topic_06: Local LLM Synthesis & CLI/REPL]