Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering local Ollama REST API integration, strict citation prompt engineering, context budgeting, interactive REPL session management, and solving the argparse sub-parser inheritance bug.
Opening & Scope:
"In this subsystem (sherlock/answer.py and sherlock/cli.py), I built the local LLM synthesis engine and terminal CLI interface, connecting retrieved code chunks to an offline Ollama model with zero cloud token billing and rich in-terminal rendering."
Step 1: The Design Pivot from Cloud APIs to Ollama:
"Initially, the synthesis engine was implemented with the Anthropic Claude API. However, sending proprietary codebase files to an external cloud API violates offline privacy constraints and incurs recurring token costs. I pivoted completely to an offline architecture, purging the anthropic dependency (pyproject.toml) and pointing synthesis to a local Ollama daemon at http://localhost:11434/api/chat (answer.py:6)."
Step 2: Context Window Budgeting & Citation Formatting:
"To assemble LLM prompts without exceeding local context windows, build_context() (answer.py:18) caps retrieved chunks at MAX_CONTEXT_CHUNKS = 12. Each chunk is formatted as a structured code block prefixed with `# file:start_line-end_line (kind name)`. Combined with SYSTEM_PROMPT, this strictly forces the model to answer only from the provided code and cite exact physical lines, preventing hallucinated APIs."
Step 3: Stateful Multi-Turn Interactive REPL:
"For conversational exploration, cmd_chat() (cli.py:71) establishes an interactive REPL loop. In chat_turn() (answer.py:56), the conversation history is maintained as an in-memory message list mutated in place. On every prompt, fresh chunks are retrieved and appended to the context, allowing developers to ask follow-up questions about complex logic."
Step 4: Graceful Degradation & Rich Terminal UI:
"Rather than allowing connection drops to crash the CLI with raw Python tracebacks, _chat_request() intercepts requests.ConnectionError and outputs a clean, actionable prompt: 'Couldn't reach Ollama — is `ollama serve` running?'. Output responses are rendered using Rich (console.print(Markdown(text))), formatting markdown tables, syntax highlights, and bullet points natively in the terminal."
Step 5: Solving the Argparse Sub-Parser Bug:
"A subtle CLI bug occurred where running sherlock index --path /path/to/repo failed with unrecognized arguments: --path because argparse by default restricts options defined on the root parser from appearing after subcommands. I solved this by creating a reusable path_parent sub-parser and inheriting it via parents=[path_parent] across the root and all subcommands."
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. Context Assembly | Format retrieved chunks with metadata headers | build_context() capping at 12 chunks |
Prevents prompt context overflow while preserving file:line coordinates | answer.py:18-23 |
| 2. Grounded Prompting | Enforce strict citation rules | SYSTEM_PROMPT forbidding external knowledge |
Eliminates hallucinated functions or third-party methods | answer.py:10-15 |
| 3. Ollama REST API | Send HTTP POST to local Ollama daemon | requests.post(OLLAMA_URL, json={...}) |
Local offline inference without cloud token billing or external API keys | answer.py:26-36 |
| 4. Stateful Chat Memory | Maintain conversation history in-memory | chat_turn() mutating message list |
Enables multi-turn conversational code reasoning in REPL | answer.py:56-63 |
| 5. Parent Parser Architecture | Share --path across root and subparsers |
path_parent with parents=[path_parent] |
Allows --path to appear before or after subcommands without syntax errors |
cli.py:99-105 |
# Formatting chunks with grounded headers in answer.py
def build_context(chunks: list[dict]) -> str:
parts = []
for c in chunks[:MAX_CONTEXT_CHUNKS]:
header = f"# {c['file']}:{c['start_line']}-{c['end_line']} ({c['kind']} {c['name']})"
parts.append(f"{header}\n{c['code']}")
return "\n\n".join(parts)
# Strict grounding system prompt
SYSTEM_PROMPT = (
"You are a codebase assistant. Answer using ONLY the provided code context — "
"don't invent functions or files that aren't shown. Cite locations as file:line. "
"When diagnosing an error, identify the root cause first, then suggest a concrete fix. "
"Be concise."
)
When running commands like:
sherlock index --path /Users/angshuman/git/Conductor
Argparse threw an immediate syntax error:
error: unrecognized arguments: --path /Users/angshuman/git/Conductor
However, running sherlock --path /Users/angshuman/git/Conductor index worked. Users naturally place flags after the subcommand.
In standard argparse, options declared on the root parser are not inherited by subparsers by default. In cli.py:99, I created a parent parser without help:
path_parent = argparse.ArgumentParser(add_help=False)
path_parent.add_argument("--path", default=".", help="repo root (default: cwd)")
# Inherited across root and all subcommands
p = argparse.ArgumentParser(prog="sherlock", parents=[path_parent])
p_index = sub.add_parser("index", parents=[path_parent])
p_ask = sub.add_parser("ask", parents=[path_parent])
p_diag = sub.add_parser("diagnose", parents=[path_parent])
p_chat = sub.add_parser("chat", parents=[path_parent])
This makes --path fully valid anywhere in the command string without duplicating argument definitions.
Answer: While FastEmbed runs small 30MB embedding models efficiently inside the Python process, generative LLMs require 4 GB to 16 GB of RAM, GPU acceleration (Metal on macOS, CUDA on Linux), and continuous KV-cache management. Embedding a full 7B or 14B parameter generative model directly inside Python bloats process startup time and complicates memory management. Delegating generative synthesis to a native daemon like Ollama keeps the Python CLI lightweight (~50ms startup), while leveraging Ollama's highly optimized native C++ runtime for token generation.
sherlock executable via uv run?"Answer: In pyproject.toml, defining [project.scripts] sherlock = "sherlock.cli:main" was initially insufficient: uv run sherlock threw "Failed to spawn: sherlock — No such file or directory". Because uv defaults to virtualenv virtual mode, it skips console-script entrypoint installation unless explicitly declared as a package. Adding [tool.uv] package = true instructed uv to build the editable wheel, properly installing the sherlock binary into .venv/bin/sherlock.