Topic 04: Three-Way Query Routing & Symbol Resolution

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering zero-overhead rule-based query classification, regex cascade heuristics, exact symbol matching vs semantic vector fallback, and resolving identifier ambiguity.


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. Traceback Check Detect error logs & stack traces Substring check + GENERIC_FRAME_RE Diverts stack traces immediately to frame-range extraction router.py:17-18
2. Identifier Regex Detect bare function/class names re.match(r"^[A-Za-z_][A-Za-z0-9_]*$", query) Isolates exact programming symbols without natural-language tokens router.py:9
3. SQL Symbol Prefilter Query LanceDB name column table.search().where(f"name = '{safe}'", prefilter=True) Sub-millisecond exact return; avoids vector cosine inaccuracies indexer.py:163-166
4. Semantic Fallback Vector search if symbol not found semantic_search(repo_root, query, k=8) Catches natural language queries or fuzzy concepts seamlessly router.py:71-72

🔍 Under-the-Hood Deep Dive: The Routing Cascade Logic

# The complete router decision engine in router.py
def route(repo_root: Path, query: str, k: int = 8) -> dict:
    query = query.strip()

    # 1. Error traceback detection
    if looks_like_traceback(query):
        return diagnose(repo_root, query, k=k)

    # 2. Bare programming identifier match
    if IDENTIFIER_RE.match(query):
        exact = indexer.symbol_lookup(repo_root, query)
        if exact:
            return {"type": "symbol", "query": query, "chunks": exact}

    # 3. Semantic vector search (natural language or symbol miss)
    results = indexer.semantic_search(repo_root, query, k=k)
    return {"type": "semantic", "query": query, "chunks": results}

💡 Tough Interview Questions & Detailed Answers

Q1: "Why prefer rule-based routing over an LLM function-calling agent?"

Answer: An LLM agent (like OpenAI Tools or LangChain Router) introduces three major problems for a local CLI: (1) Latency: an LLM round-trip takes 500–1500ms just to decide the search mode; (2) Cost & Hardware: it doubles the local inference load on Ollama; and (3) Non-Determinism: an LLM might misclassify an exact function name as a conversational prompt. Regex heuristics execute in under 0.1ms, are 100% deterministic, and cleanly handle 99% of developer workflows.

Q2: "How does the router prevent SQL injection in LanceDB queries?"

Answer: In indexer.py:165, symbol names are sanitized before being interpolated into LanceDB's SQL filter string: safe = name.replace("'", "''"). This escapes single quotes, ensuring that malicious or odd symbol names (e.g. operators or syntax strings) cannot escape the SQL string literal.