Target Duration: 4–5 minutes (~650–800 spoken words)
Goal: Deliver an end-to-end architectural explanation covering AST parsing, embedded vector tables, incremental manifest diffing, multi-modal query routing, traceback interval matching, and offline LLM synthesis. Plant deep-dive hooks along the way.
"To explore how modern developer assistants can achieve privacy, sub-second latency, and zero API costs, I designed Sherlock—a local-first codebase RAG and runtime debugging system built in Python.
Rather than relying on remote hosted services or naive line-by-line grep, Sherlock combines compile-time syntactic analysis with in-process vector mathematics. It breaks code down into semantic AST units, embeds them locally, stores them in an embedded columnar vector table, routes queries intelligently through rule-based heuristics, and synthesizes answers using a local Ollama model.
I structured the system into five core functional pillars:
"Standard RAG systems divide documents using fixed character lengths or line sliding windows. On source code, this breaks functions in half and ruins variable scope.
In chunker.py:98 (chunk_file), I implemented AST-aware chunking using Tree-sitter.
EXT_TO_LANG)._walk()) isolates meaningful syntax boundaries—such as function declarations, method definitions, and class blocks defined in CHUNK_NODE_TYPES—capturing exact start_line and end_line boundaries alongside parent namespace context.TEXT_EXTENSIONS and NO_EXT_FILES), ensuring files like Dockerfile or Makefile remain fully discoverable.MAX_FILE_BYTES) are automatically pruned, preventing huge vendored C binaries or database blobs from polluting the index."(Hook: Deep dive into topic_01_ast_chunking_and_tree_sitter)
"Rather than deploying an external database cluster like Milvus or Pinecone, I built the storage layer around LanceDB—an embedded, serverless vector database operating directly on local NVMe disk via Lance columnar files (.sherlock/db).
Key architectural properties include:
ChunkSchema) stores all chunk metadata alongside a 384-dimensional dense vector. This enables instantaneous exact-match SQL prefiltering (e.g. WHERE name = 'add') alongside approximate nearest neighbor (ANN) vector searches on the same dataset.BAAI/bge-small-en-v1.5 ONNX model directly through FastEmbed, managing the model as a process singleton (_embedder() via lru_cache). This avoided heavy PyTorch dependencies and eliminated inter-process RPC latency during batch embedding."(Hook: Deep dive into topic_02_vector_storage_and_fastembed)
"In production repositories, developers modify a handful of files at a time. Re-embedding thousands of unchanged files on every run causes intolerable latency bottlenecks.
In indexer.py:112 (build_index), I engineered an incremental indexing pipeline driven by a SHA-1 content manifest (.sherlock/manifest.json):
OR predicate (_sql_or()) to purge stale chunk rows from LanceDB before selectively chunking, embedding, and appending only the newly mutated files.(Hook: Deep dive into topic_03_incremental_indexing_and_manifest)
"User queries vary drastically—from raw error tracebacks and function names to free-form architectural questions. Relying on an LLM to classify query intent introduces 500ms of unnecessary latency.
In router.py:60 (route), I implemented a zero-latency heuristic cascade:
IDENTIFIER_RE), it executes an exact SQL prefiltered lookup (symbol_lookup()) for instantaneous $O(1)$ symbol resolution.semantic_search())."(Hook: Deep dive into topic_04_query_routing_and_symbol_resolution)
"For error diagnosis and code explanation, Sherlock bridges retrieval into grounded synthesis without cloud dependencies:
router.py:36 (diagnose), stack frames are parsed via regex (PY_FRAME_RE and GENERIC_FRAME_RE). The system matches each frame's execution line against indexed chunk ranges (start_line <= line <= end_line), retrieving the exact function executing at the moment of failure. It simultaneously vector-searches the exception message to capture semantically related callers.build_context()) and submitted to a local Ollama instance (http://localhost:11434/api/chat). A strict system prompt enforces grounded responses with mandatory file:line citations and root-cause analysis.cli.py) provides interactive multi-turn chat (cmd_chat()) rendered directly in the terminal via Rich markdown."(Hook: Deep dive into topic_05_traceback_frame_parsing_and_diagnosis or topic_06_local_llm_synthesis_and_cli)
Sherlock is fully verifiable using the command line:
# 1. Build or incrementally update the local index for any target repository
uv run sherlock index --path /path/to/repo
# 2. Ask architectural questions or look up specific symbols with code citations
uv run sherlock ask --path /path/to/repo "what does conductor.sh do?"
uv run sherlock ask --path /path/to/repo "divide" --chunks --show-code
# 3. Diagnose runtime errors directly from piped stack traces
python failing_script.py 2>&1 | uv run sherlock diagnose --path /path/to/repo
# 4. Launch the interactive multi-turn terminal REPL
uv run sherlock chat --path /path/to/repo
# 5. Run the end-to-end integration smoke test
uv run python test_smoke.py
"In summary, Sherlock gave me deep, practical experience building local AI systems: navigating compiler AST structures with Tree-sitter, managing embedded vector tables with LanceDB, amortizing re-indexing through content-addressable hashing, and orchestrating offline LLM synthesis with zero cloud telemetry."
graph LR
DetailedSpeech[4-5 Min Detailed Speech] --> T1[topic_01: AST Chunking & Tree-sitter]
DetailedSpeech --> T2[topic_02: Vector Storage & FastEmbed]
DetailedSpeech --> T3[topic_03: Incremental Indexing & Manifest]
DetailedSpeech --> T4[topic_04: Query Routing & Symbol Resolution]
DetailedSpeech --> T5[topic_05: Traceback Parsing & Diagnosis]
DetailedSpeech --> T6[topic_06: Local LLM Synthesis & CLI/REPL]