Topic 05: Traceback Frame Extraction & Runtime Error Diagnosis

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering regex stack frame extraction, interval containment matching, the backwards `.endswith()` path resolution bug, hybrid frame-plus-semantic context gathering, and root-cause synthesis.


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. Frame Extraction Parse file and line numbers from text Regex match with PY_FRAME_RE and GENERIC_FRAME_RE Extracts ordered call stack from pasted terminal output router.py:21-33
2. Suffix Path Match Normalize absolute vs relative paths file.endswith(r["file"]) Matches long runtime paths (/var/log/.../app.py) against relative repo paths (app.py) router.py:45
3. Interval Matching Locate chunk enclosing the failure line r["start_line"] <= line <= r["end_line"] Retrieves the entire crashing function/class context, not just an isolated line router.py:45
4. Error Line Semantic Search Vector search on exception message semantic_search(repo_root, error_line, k=5) Surfaces callers or configuration files related to the specific error message router.py:49-50
5. Grounded Diagnosis Synthesize fix via local LLM answer.synthesize() with root-cause prompt Enforces diagnosis discipline: pinpoint root cause before proposing code changes cli.py:68

🔍 Under-the-Hood Deep Dive: Frame Matching Engine

# Exact stack frame resolution in router.py
def diagnose(repo_root: Path, traceback_text: str, k: int = 5) -> dict:
    frames = parse_frames(traceback_text)
    frame_chunks = []
    
    # 1. Deterministic stack frame interval match
    for file, line in frames:
        table = indexer.open_table(repo_root)
        rows = table.search().to_list()
        matches = [
            r for r in rows
            if file.endswith(r["file"]) and r["start_line"] <= line <= r["end_line"]
        ]
        frame_chunks.extend(matches)

    # 2. Semantic vector search on exception description
    error_line = traceback_text.strip().splitlines()[-1] if traceback_text.strip() else ""
    semantic = indexer.semantic_search(repo_root, error_line, k=k) if error_line else []

    return {
        "type": "diagnose",
        "frames": frames,
        "frame_chunks": frame_chunks,
        "semantic_chunks": semantic,
    }

⚠️ Debugging Stories: The Backwards Path Suffix Trap

The Failure

In early testing of sherlock diagnose, feeding valid Python tracebacks produced zero matched stack frames. The CLI repeatedly reported:

-- matched stack frames --
(empty)
-- semantically related --
...

Diagnosis & Root Cause

Inspecting the frame matching logic revealed:

# BUGGY ORIGINAL LINE:
r["file"].endswith(file.lstrip("./"))
In this comparison: Calling "app.py".endswith("/Users/.../app.py") is mathematically impossible because the string being checked is shorter than the suffix argument! The comparison was directionally inverted.

The Resolution

Reversed the comparison argument in router.py:45 to:

file.endswith(r["file"])
Now, the absolute traceback path "/Users/.../app.py" correctly matches when ending with the indexed relative path "app.py". All stack frames resolved instantly.


💡 Tough Interview Questions & Detailed Answers

Q1: "Why combine stack frame matching with semantic search on the exception line?"

Answer: The stack trace identifies where the program crashed, but often the root cause originated elsewhere (e.g. an invalid config loaded at startup, or a caller passing unvalidated None). Frame interval matching guarantees you retrieve the exact crashed function, while semantic search on the exception message (e.g., "KeyError: 'timeout'") retrieves configuration loaders, dictionary definitions, or upstream callers that handle timeouts.

Q2: "What happens if a traceback refers to third-party library files in site-packages?"

Answer: Because .venv, site-packages, and Python standard libraries are excluded from the repository index by SKIP_DIRS, file.endswith(r["file"]) simply yields no match for library frames. Only the stack frames corresponding to first-party application code are matched, filtering out external framework noise and focusing the LLM's context window entirely on the user's code.