1–2 Minute Brief Pitch: Sherlock (Local Codebase RAG & Debugger)

Target Duration: 60–90 seconds (~180–220 spoken words)
Goal: Deliver a clear, plain-English overview of why the project was built, its full scope, and key mental models learned, avoiding low-level implementation minutiae while dropping strategic follow-up hooks for the interviewer.


🎙️ Spoken Script (Plain English & Conversational)

"In this project, I built Sherlock—a fully local, offline codebase intelligence and error diagnosis assistant in Python, combining syntax-aware Tree-sitter AST chunking, local FastEmbed dense vector embeddings, embedded LanceDB vector storage, content-hash incremental indexing, three-way query routing, and offline Ollama LLM answer synthesis.

Why I Built It:
Exploring large codebases or diagnosing runtime stack traces usually forces developers to either rely on crude grep searches or upload proprietary code to cloud AI services with per-token billing and privacy risks. My goal was to build a fast, zero-cost, completely offline developer assistant that indexes an entire repository locally, performs hybrid vector and symbol lookups in-process, and synthesizes grounded explanations and traceback diagnoses without a single byte of code leaving the local machine.

What I Learned & Engineering Takeaways:

  1. Syntax-Aware Parsing over Blind Windowing:
    I learned why arbitrary character-count sliding windows fail on source code, mastering concrete syntax tree traversal to slice files along semantic boundaries like functions and classes, preserving precise line ranges for runtime traceback matching.

  2. Zero-Overhead In-Process Vector Architecture:
    I gained hands-on experience pairing embedded columnar storage with local embedding inference, discovering how serverless vector tables can simultaneously serve exact SQL-filtered lookups and approximate nearest-neighbor similarity searches in-process.

  3. Content-Addressable Incremental Amortization:
    I internalized that developer tooling must be near-instant on re-runs. By engineering content-hash manifest diffing, I learned how to isolate only changed files, cutting re-indexing times from minutes to sub-second updates on production-sized repositories.

  4. Deterministic Multi-Modal Routing:
    I mastered query classification using fast rule-based cascades, routing tracebacks, exact symbols, and free-form questions to dedicated retrieval strategies without incurring extra LLM latency.

Key Takeaway:
This project shifted my understanding from treating AI as a cloud API wrapper to engineering an end-to-end, privacy-preserving systems pipeline—bridging AST compilers, vector mathematics, and local model orchestration."


🪝 What the Interviewer Gets Hooked Into (Natural Follow-Ups)

What You Mentioned Why It Was Done (The Motivation) Problems Faced & How Solved (The Reality) Target Deep-Dive Document
"Syntax-aware Tree-sitter AST chunking" "Why parse full ASTs with tree-sitter instead of using standard text splitters or regexes?" "What happened with non-code files like shell scripts and huge vendored C libraries, and how did whole-file fallbacks and byte limits solve it?" topic_01_ast_chunking_and_tree_sitter
"Embedded LanceDB storage & local FastEmbed" "Why choose an embedded serverless vector DB over external vector stores like Pinecone, Milvus, or Chroma?" "What broke when LanceDB dropped built-in FastEmbed registry integration, and how did a standalone singleton solve it?" topic_02_vector_storage_and_fastembed
"Content-hash incremental indexing manifest" "Why use SHA-1 content hashing instead of file mtime timestamps or chunk-level diffing?" "How did self-indexing `.sherlock` metadata cause infinite re-index loops, and what performance was measured on a 495-file repo?" topic_03_incremental_indexing_and_manifest
"Three-way rule-based query routing" "Why use rule-based regex cascades instead of an LLM query classifier?" "How did you handle ambiguous single-word queries like 'divide' that could be either an exact function name or a general concept?" topic_04_query_routing_and_symbol_resolution
"Traceback parsing & error diagnosis" "Why match traceback frames against chunk line ranges rather than just vector searching the raw error message?" "What bug caused frame matching to fail completely across relative paths, and how did reversing the suffix match fix it?" topic_05_traceback_frame_parsing_and_diagnosis
"Offline Ollama synthesis & CLI architecture" "Why pivot away from cloud APIs (Claude) to a local Ollama model, and how is context budgeted?" "How did you resolve the argparse `--path` flag positioning bug, and how does the Rich interactive REPL handle connection dropouts?" topic_06_local_llm_synthesis_and_cli

🧭 Next Step Transitions