Topic 01: Syntax-Aware AST Chunking & Tree-Sitter Parser Pipeline

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering Tree-sitter AST parsing, multi-language grammar registries, AST node extraction, whole-file fallback strategies, and defensive file-size filtering.


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. Discovery & Filter Walk repo, prune junk directories & blobs rglob("*") checking SKIP_DIRS, size < 500KB Avoid indexing .git, .venv, node_modules, and massive vendor files chunker.py:129 (discover_files)
2. Grammar Resolution Map extension to Tree-sitter parser Inspect path.suffix against EXT_TO_LANG Dynamic parser dispatch across 11 languages with pre-compiled grammars chunker.py:101
3. AST Parsing Build CST and recursively walk nodes get_parser(lang).parse(source); recursive _walk() Guarantees complete syntax boundary isolation; no arbitrary cutoffs chunker.py:115-118
4. Symbol Extraction Isolate identifier names and line ranges _node_name() via child_by_field_name("name") Stores exact 1-indexed lines and symbol names for fast prefiltering and traceback lookup chunker.py:82-92
5. Fallback Chunking Single-chunk emit for scripts/configs UTF-8 decode into kind="file" chunk Ensures non-code assets (shell scripts, Dockerfiles) remain queryable chunker.py:105-113

🔍 Under-the-Hood Deep Dive: AST Chunking Mechanics

The Chunk Data Structure

Each chunk produced by the chunker adheres to the following dataclass:

@dataclass
class Chunk:
    file: str          # Relative path from repository root (e.g. "sherlock/router.py")
    language: str      # Language identifier (e.g. "python", "c", "text")
    kind: str          # AST node type (e.g. "function_definition", "class_definition", "file")
    name: str          # Symbol identifier (e.g. "chunk_file", "ChunkSchema", "")
    parent: str | None # Enclosing class or scope name (e.g. "Calculator")
    start_line: int    # 1-indexed start line in source file
    end_line: int      # 1-indexed end line in source file
    code: str          # Raw source code slice

AST Traversal & Child Recursion

When _walk() encounters a node whose type matches CHUNK_NODE_TYPES[lang], it performs two actions:

  1. Emits a Chunk representing that node, setting parent=parent_name.
  2. Recursively invokes _walk(child, source, lang, name, out, file_str) with the new name as the parent for any nested functions or methods.
If the node is not a chunk-worthy type (e.g., an if_statement or block), it passes through without changing parent_name. This produces a flat list of semantically rich chunks with explicit hierarchy links.


⚠️ Debugging Stories & Critical Edge Cases

Bug Story 1: The Missing Shell Script Context

Symptom: When asking sherlock ask "what does conductor.sh do?" against a container project, the assistant answered "No relevant code found in the index", despite the file existing on disk.
Root Cause: discover_files() initially checked if path.suffix not in EXT_TO_LANG: continue. Because .sh, .md, and Dockerfile lacked Tree-sitter AST entries, they were silently discarded during file discovery before chunking ever ran.
Resolution: Broadened the discovery filter in chunker.py:136 to include TEXT_EXTENSIONS and NO_EXT_FILES, pairing it with the whole-file fallback in chunker.py:105.

Bug Story 2: The Self-Indexing Loop

Symptom: After adding .json to TEXT_EXTENSIONS, the automated smoke test failed: a clean no-op re-index reported files_changed=1 on every single execution.
Root Cause: The incremental indexing engine saves its hash manifest at .sherlock/manifest.json. Because .json was now whitelisted and .sherlock was not in SKIP_DIRS, Sherlock indexed its own manifest, computed a new hash for it, updated the manifest, and perpetually detected its own state file as mutated!
Resolution: Added .sherlock explicitly to SKIP_DIRS (chunker.py:43), enforcing that internal runtime state directories are excluded at the walk root.


💡 Tough Interview Questions & Detailed Answers

Q1: "Why use Tree-sitter for chunking instead of Python's built-in ast module or regex?"

Answer: Python's native ast module only parses Python. Tree-sitter provides a unified C-based Concrete Syntax Tree (CST) parser capable of handling 11+ languages with high performance and error-recovery capabilities (it can parse incomplete or syntactically invalid code during active development without throwing exceptions). Furthermore, regex-based parsers fail on nested classes, multi-line decorators, and inner closures, while Tree-sitter guarantees exact byte offsets and line boundaries.

Q2: "What happens if a developer edits code inside an anonymous lambda or top-level script?"

Answer: If an AST walk finds zero chunk-worthy nodes (such as a top-level script or procedural block), chunk_file() falls back to emitting a single whole-file chunk spanning lines 1 to N. For lambdas inside functions, the lambda's code is encapsulated within the enclosing function's chunk, ensuring that semantic vector searches for the lambda logic still retrieve the containing function.