Sherlock — Local Codebase RAG & Debugging Assistant: Master Index

A structured navigation hub for explaining Sherlock—a fully local, privacy-first codebase search, explanation, and error diagnosis system built in Python using Tree-sitter AST chunking, embedded LanceDB vector storage, local FastEmbed embeddings, and offline Ollama LLM synthesis.


📂 File Directory

File Type Duration Primary Focus
intro Speech 1–2 mins Crisp elevator pitch on motivation, mental models learned, and full architectural scope without low-level code minutiae.
detailed Speech 4–5 mins End-to-end architectural blueprint touching all 6 subsystems with embedded breadcrumb hooks to topic deep dives.
topic_01_ast_chunking_and_tree_sitter Deep Dive 2–3 mins Tree-sitter recursive AST walk, multi-language grammar mappings (11 languages), whole-file fallback, and file size limits.
topic_02_vector_storage_and_fastembed Deep Dive 2–3 mins Embedded serverless LanceDB table, 384-dim BGE-small FastEmbed singleton, hybrid exact prefilter + cosine ANN search.
topic_03_incremental_indexing_and_manifest Deep Dive 2–3 mins SHA-1 content-hash manifest diffing, SQL OR row invalidation, and sub-second incremental re-indexing benchmarks (1.49s no-op vs multi-minute full rebuild).
topic_04_query_routing_and_symbol_resolution Deep Dive 2–3 mins Three-way rule-based heuristic routing cascade: traceback detection, identifier regex symbol lookup, and semantic vector fallback.
topic_05_traceback_frame_parsing_and_diagnosis Deep Dive 2–3 mins Regex stack frame parsing, interval containment line matching, solving the backwards .endswith() bug, and root-cause prompt engineering.
topic_06_local_llm_synthesis_and_cli Deep Dive 2–3 mins Offline Ollama REST API integration, 12-chunk context budgeting, interactive REPL, Rich markdown formatting, and the argparse parent-parser pattern.
code.html (Source Code Browser) Interactive Tool Interactive 7 Sherlock source and config files, 58 symbol definitions, syntax highlighting, click-to-highlight, and quick-open navigation.

🏗️ Architecture Overview

graph TD
    subgraph Ingestion ["1. AST-Aware Ingestion & Indexing Pipeline"]
        Repo[Target Codebase] --> Disc[discover_files: Extensions & Size Filter]
        Disc --> Hash[Compute SHA-1 Content Hash]
        Hash --> Diff{Manifest Diff: Changed?}
        Diff -->|Unchanged| Skip[Skip Embedding]
        Diff -->|Changed / Added| Parse[Tree-sitter Parser: get_parser]
        Parse --> Walk[Recursive AST Walk: CHUNK_NODE_TYPES]
        Walk --> Chunks[Function/Class Chunk Dataclasses]
        Chunks --> Embed[FastEmbed Singleton: BAAI/bge-small-en-v1.5]
        Embed --> Lance[LanceDB: .sherlock/db Chunks Table]
        Hash --> Manifest[.sherlock/manifest.json]
    end

    subgraph Retrieval ["2. Query Routing & Multi-Modal Retrieval"]
        UserQ[User CLI Query / Pasted Traceback] --> Router[router.route]
        Router -->|Looks like traceback| Diag[parse_frames: PY_FRAME_RE & GENERIC_FRAME_RE]
        Diag --> FrameMatch[Interval Match: start_line <= line <= end_line]
        Router -->|Identifier regex match| Sym[Exact SQL Prefilter: name = '...']
        Router -->|Free-form text| Sem[Vector ANN Search: table.search]
        Sym -->|Miss| Sem
    end

    subgraph Synthesis ["3. Offline LLM Synthesis & Terminal UI"]
        FrameMatch --> TopK[Top Chunks: MAX_CONTEXT_CHUNKS = 12]
        Sem --> TopK
        TopK --> Ctx[answer.build_context: file:line code headers]
        Ctx --> Prompt[SYSTEM_PROMPT: Grounded & Concise]
        Prompt --> Ollama[Local Ollama REST API: /api/chat]
        Ollama --> Reply[Markdown Response]
        Reply --> RichUI[rich.markdown.Markdown Console Render]
    end