1–2 Minute Brief Pitch: Linux Kernel Modules & Internals

Target Duration: 60–90 seconds (~180–220 spoken words)
Goal: Deliver a clear, plain-English overview of why the project was built, its full scope, and key mental models learned, avoiding low-level implementation minutiae while dropping strategic follow-up hooks for the interviewer.


🎙️ Spoken Script (Plain English & Conversational)

"In this project, I built a set of Linux kernel modules and character device drivers to explore different interfaces between userspace and the Linux kernel.

The project had three main parts:

First, I wrote Linux kernel modules to inspect and measure a process's memory usage. I implemented manual page-table walks to directly translate a process's virtual address to its physical address. I also compared the kernel's fast RSS-based estimate of a process's physical memory usage with a precise method that scans each memory page and checks whether it is present in RAM, including detection of 2 MiB Transparent Huge Pages.

Second, I developed custom character device drivers using ioctl to allow userspace applications to trigger privileged kernel operations. One driver takes a virtual address from userspace, translates it to a physical address, and writes data directly to physical memory through the kernel's direct mapping. Another driver modifies the process hierarchy maintained by the kernel to reparent existing processes at runtime.

Third, I implemented /proc and /sysfs interfaces for exposing kernel information to userspace. The /proc interface streams system-wide page-fault statistics, while the /sysfs interface accepts a PID and allows userspace to query its virtual and physical memory usage in different units.

Overall, the project gave me hands-on experience with Linux kernel modules, page-table walks, character drivers, ioctl, /proc, /sysfs, kernel memory, and process structures, and helped me understand how userspace interacts with privileged kernel functionality."


🪝 What the Interviewer Gets Hooked Into (Strategic Follow-Ups)

What You Mentioned Why It Was Done (The Motivation) Problems Faced & How Solved (The Reality) Target Deep-Dive Document
"Kernel modules & lockless process inspection" Safely inspect task_struct and children list in Ring 0 without stalling the kernel scheduler. Solved using rcu_read_lock() and list_for_each_entry_rcu(), preventing race conditions during process fork/exit. Topic 01: LKM Lifecycle & Tasks
"5-level manual page table walks (VA → PA)" Understand hardware MMU translation directly from software. Walked PGD → P4D → PUD → PMD → PTE with null/bad entry checking at each level, extracting PFN via pte_pfn(). Topic 02: MMU Page Table Walks
"Memory accounting & 2 MiB Transparent Huge Pages" Compare $O(1)$ cached RSS against exact $O(N)$ VMA frame traversal. Traversed VMAs using Linux 6.1 Maple Tree API; detected THPs via pmd_trans_huge() and advanced by 2 MiB instead of checking 512 individual 4 KiB pages. Topic 03: Memory Accounting & THP
"Character drivers & physical memory writing" Allow userspace to trigger privileged memory writes without bypassing MMU paging. Translated VA to PA, converted PA to virtual address via kernel direct mapping (phys_to_virt()), and wrote data directly. Topic 04: Character Devices & ioctl
"Live process tree surgery & reparenting" Allow dynamic runtime adoption of running child processes without calling fork(). Manipulated kernel children and sibling lists and updated real_parent/parent pointers under tasklist_lock. Topic 05: Process Reparenting
"Kernel telemetry via /proc and sysfs" Contrast procedural stream interfaces with structured object-oriented sysfs attributes. Implemented struct proc_ops for lockless page-fault streaming and a sysfs kobject with container_of for per-PID memory queries. Topic 06: Kernel Telemetry

❓ Anticipated Quick-Fire Follow-Up Questions (Direct from the 90s Pitch)

Q1: Why does the kernel's fast RSS estimate differ from your precise page-by-page scan?

Answer: The kernel's RSS counter (in task_struct->mm->rss_stat) is an $O(1)$ heuristic updated asynchronously during page faults and unmaps. It can lag behind actual physical page state, especially with file-backed mappings and lazy unmapping. In contrast, our page-by-page scan traverses the actual page tables down to the PTEs and checks the hardware _PAGE_PRESENT bit directly, providing the ground truth of resident RAM frames. Furthermore, our traversal accounts for 2 MiB Transparent Huge Pages (THP) where a single PMD entry represents 512 physical frames.

Q2: Why did you write an ioctl character driver to write to physical memory instead of using /dev/mem?

Answer: Modern production Linux kernels enable CONFIG_STRICT_DEVMEM, which restricts user access to /dev/mem to protect physical system RAM and hardware registers from arbitrary user-space tampering. Building a custom character driver allowed us to create a safe, controlled kernel interface where a user process submits a virtual address, and the driver translates it through the process's page table before accessing physical memory via the kernel's direct mapping (phys_to_virt()).

Q3: What practical problem does dynamic process reparenting without fork() solve?

Answer: In standard Linux, process parentage is immutable after fork(). If a worker process loses its supervisor daemon, it gets orphaned and reparented to init (PID 1) or a subreaper. If a new manager process wants to take over custody, monitor execution, and receive death notifications via waitpid() without restarting the child, standard Linux has no system call. Our character driver performs dynamic process tree surgery by atomically rewriting the parent/real_parent pointers and splicing the children/sibling lists under tasklist_lock.

Q4: Why did you implement both /proc and /sys interfaces for kernel telemetry?

Answer: They serve fundamentally different design philosophies in Linux. /proc is procedural and sequential—well-suited for streaming multi-line event metrics like our system-wide page fault statistics via proc_ops. In contrast, sysfs is strictly object-oriented adhering to the "one value per file" rule—ideal for exposing per-PID memory metrics as individual attributes (pid, virtmem, physmem, unit) linked to custom kernel objects via kobject and sysfs_emit.