Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering Tree-sitter AST parsing, multi-language grammar registries, AST node extraction, whole-file fallback strategies, and defensive file-size filtering.
Opening & Scope:
"In this subsystem (sherlock/chunker.py), I designed an AST-aware code chunking pipeline using Tree-sitter to parse arbitrary codebases into semantically self-contained units of code—specifically at function, method, and class granularity."
Step 1: The Problem with Fixed-Size Text Splitting:
"Standard RAG chunkers slice text into fixed 500-token or 1000-character windows with arbitrary overlap. When applied to source code, this is catastrophic: functions get sliced mid-expression, class contexts are lost, and line numbers no longer map cleanly to physical source files. This breaks exact symbol search and ruins traceback diagnosis."
Step 2: Grammar Mapping & Target Node Registries:
"To parse code syntactically, I mapped 16 source extensions across 11 programming languages to their corresponding Tree-sitter grammars in EXT_TO_LANG (chunker.py:9). For each language, I defined a target AST node whitelist in CHUNK_NODE_TYPES (chunker.py:29)—for example, targeting function_definition and class_definition in Python, and function_item, struct_item, and impl_item in Rust."
Step 3: Recursive AST Traversal & Chunk Extraction:
"When indexing a file in chunk_file() (chunker.py:98), Tree-sitter builds a Concrete Syntax Tree. I implemented a recursive walker _walk() (chunker.py:78) that traverses the tree. Whenever it hits a whitelisted node, it resolves the symbol name via _node_name() (chunker.py:67), captures the exact 1-indexed start_line and end_line, extracts the raw slice, tracks the parent enclosing class name, and emits a structured Chunk (chunker.py:56) dataclass."
Step 4: Dual Fallback & Non-Code Whitelisting:
"Real-world repositories contain more than just compiled functions. To ensure infrastructure scripts, configurations, and documentation remain searchable, I established two fallback layers:
kind="file" chunk.TEXT_EXTENSIONS and NO_EXT_FILES to index shell scripts (.sh), markdown (.md), configs (.yaml, .json), and Dockerfiles."Step 5: Defensive Memory Pruning (File Size Guards):
"Finally, in chunker.py:99, I enforce a strict 500 KB limit (MAX_FILE_BYTES = 500_000) alongside directory exclusions in SKIP_DIRS. When testing against a real production repository, this prevented multi-megabyte vendored C files—like a 5.5 MB sqlite3.c blob—from causing out-of-memory crashes or embedding timeouts."
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. Discovery & Filter | Walk repo, prune junk directories & blobs | rglob("*") checking SKIP_DIRS, size < 500KB |
Avoid indexing .git, .venv, node_modules, and massive vendor files |
chunker.py:129 (discover_files) |
| 2. Grammar Resolution | Map extension to Tree-sitter parser | Inspect path.suffix against EXT_TO_LANG |
Dynamic parser dispatch across 11 languages with pre-compiled grammars | chunker.py:101 |
| 3. AST Parsing | Build CST and recursively walk nodes | get_parser(lang).parse(source); recursive _walk() |
Guarantees complete syntax boundary isolation; no arbitrary cutoffs | chunker.py:115-118 |
| 4. Symbol Extraction | Isolate identifier names and line ranges | _node_name() via child_by_field_name("name") |
Stores exact 1-indexed lines and symbol names for fast prefiltering and traceback lookup | chunker.py:82-92 |
| 5. Fallback Chunking | Single-chunk emit for scripts/configs | UTF-8 decode into kind="file" chunk |
Ensures non-code assets (shell scripts, Dockerfiles) remain queryable | chunker.py:105-113 |
Chunk Data StructureEach chunk produced by the chunker adheres to the following dataclass:
@dataclass
class Chunk:
file: str # Relative path from repository root (e.g. "sherlock/router.py")
language: str # Language identifier (e.g. "python", "c", "text")
kind: str # AST node type (e.g. "function_definition", "class_definition", "file")
name: str # Symbol identifier (e.g. "chunk_file", "ChunkSchema", "")
parent: str | None # Enclosing class or scope name (e.g. "Calculator")
start_line: int # 1-indexed start line in source file
end_line: int # 1-indexed end line in source file
code: str # Raw source code slice
When _walk() encounters a node whose type matches CHUNK_NODE_TYPES[lang], it performs two actions:
Chunk representing that node, setting parent=parent_name._walk(child, source, lang, name, out, file_str) with the new name as the parent for any nested functions or methods.if_statement or block), it passes through without changing parent_name. This produces a flat list of semantically rich chunks with explicit hierarchy links.
Symptom: When asking sherlock ask "what does conductor.sh do?" against a container project, the assistant answered "No relevant code found in the index", despite the file existing on disk.
Root Cause: discover_files() initially checked if path.suffix not in EXT_TO_LANG: continue. Because .sh, .md, and Dockerfile lacked Tree-sitter AST entries, they were silently discarded during file discovery before chunking ever ran.
Resolution: Broadened the discovery filter in chunker.py:136 to include TEXT_EXTENSIONS and NO_EXT_FILES, pairing it with the whole-file fallback in chunker.py:105.
Symptom: After adding .json to TEXT_EXTENSIONS, the automated smoke test failed: a clean no-op re-index reported files_changed=1 on every single execution.
Root Cause: The incremental indexing engine saves its hash manifest at .sherlock/manifest.json. Because .json was now whitelisted and .sherlock was not in SKIP_DIRS, Sherlock indexed its own manifest, computed a new hash for it, updated the manifest, and perpetually detected its own state file as mutated!
Resolution: Added .sherlock explicitly to SKIP_DIRS (chunker.py:43), enforcing that internal runtime state directories are excluded at the walk root.
ast module or regex?"Answer: Python's native ast module only parses Python. Tree-sitter provides a unified C-based Concrete Syntax Tree (CST) parser capable of handling 11+ languages with high performance and error-recovery capabilities (it can parse incomplete or syntactically invalid code during active development without throwing exceptions). Furthermore, regex-based parsers fail on nested classes, multi-line decorators, and inner closures, while Tree-sitter guarantees exact byte offsets and line boundaries.
Answer: If an AST walk finds zero chunk-worthy nodes (such as a top-level script or procedural block), chunk_file() falls back to emitting a single whole-file chunk spanning lines 1 to N. For lambdas inside functions, the lambda's code is encapsulated within the enclosing function's chunk, ensuring that semantic vector searches for the lambda logic still retrieve the containing function.