Topic 06: Grounded Local LLM Synthesis & CLI/REPL Architecture

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering local Ollama REST API integration, strict citation prompt engineering, context budgeting, interactive REPL session management, and solving the argparse sub-parser inheritance bug.


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. Context Assembly Format retrieved chunks with metadata headers build_context() capping at 12 chunks Prevents prompt context overflow while preserving file:line coordinates answer.py:18-23
2. Grounded Prompting Enforce strict citation rules SYSTEM_PROMPT forbidding external knowledge Eliminates hallucinated functions or third-party methods answer.py:10-15
3. Ollama REST API Send HTTP POST to local Ollama daemon requests.post(OLLAMA_URL, json={...}) Local offline inference without cloud token billing or external API keys answer.py:26-36
4. Stateful Chat Memory Maintain conversation history in-memory chat_turn() mutating message list Enables multi-turn conversational code reasoning in REPL answer.py:56-63
5. Parent Parser Architecture Share --path across root and subparsers path_parent with parents=[path_parent] Allows --path to appear before or after subcommands without syntax errors cli.py:99-105

🔍 Under-the-Hood Deep Dive: Context Formatting & Prompt Engineering

# Formatting chunks with grounded headers in answer.py
def build_context(chunks: list[dict]) -> str:
    parts = []
    for c in chunks[:MAX_CONTEXT_CHUNKS]:
        header = f"# {c['file']}:{c['start_line']}-{c['end_line']} ({c['kind']} {c['name']})"
        parts.append(f"{header}\n{c['code']}")
    return "\n\n".join(parts)

# Strict grounding system prompt
SYSTEM_PROMPT = (
    "You are a codebase assistant. Answer using ONLY the provided code context — "
    "don't invent functions or files that aren't shown. Cite locations as file:line. "
    "When diagnosing an error, identify the root cause first, then suggest a concrete fix. "
    "Be concise."
)

⚠️ Debugging Stories: The Argparse Sub-Parser Trap

The Failure

When running commands like:

sherlock index --path /Users/angshuman/git/Conductor
Argparse threw an immediate syntax error:
error: unrecognized arguments: --path /Users/angshuman/git/Conductor
However, running sherlock --path /Users/angshuman/git/Conductor index worked. Users naturally place flags after the subcommand.

The Resolution

In standard argparse, options declared on the root parser are not inherited by subparsers by default. In cli.py:99, I created a parent parser without help:

path_parent = argparse.ArgumentParser(add_help=False)
path_parent.add_argument("--path", default=".", help="repo root (default: cwd)")

# Inherited across root and all subcommands
p = argparse.ArgumentParser(prog="sherlock", parents=[path_parent])
p_index = sub.add_parser("index", parents=[path_parent])
p_ask = sub.add_parser("ask", parents=[path_parent])
p_diag = sub.add_parser("diagnose", parents=[path_parent])
p_chat = sub.add_parser("chat", parents=[path_parent])
This makes --path fully valid anywhere in the command string without duplicating argument definitions.


💡 Tough Interview Questions & Detailed Answers

Q1: "Why use Ollama REST API over embedding an ONNX or llama.cpp LLM directly in Python?"

Answer: While FastEmbed runs small 30MB embedding models efficiently inside the Python process, generative LLMs require 4 GB to 16 GB of RAM, GPU acceleration (Metal on macOS, CUDA on Linux), and continuous KV-cache management. Embedding a full 7B or 14B parameter generative model directly inside Python bloats process startup time and complicates memory management. Delegating generative synthesis to a native daemon like Ollama keeps the Python CLI lightweight (~50ms startup), while leveraging Ollama's highly optimized native C++ runtime for token generation.

Q2: "What packaging configuration was required to make sherlock executable via uv run?"

Answer: In pyproject.toml, defining [project.scripts] sherlock = "sherlock.cli:main" was initially insufficient: uv run sherlock threw "Failed to spawn: sherlock — No such file or directory". Because uv defaults to virtualenv virtual mode, it skips console-script entrypoint installation unless explicitly declared as a package. Adding [tool.uv] package = true instructed uv to build the editable wheel, properly installing the sherlock binary into .venv/bin/sherlock.