4–5 Minute Detailed Architecture: Linux Kernel Modules & Drivers

Target Duration: ~3.5 to 5 minutes (~650–850 spoken words)
Goal: Deliver a comprehensive end-to-end architectural narrative covering kernel execution contexts, virtual memory mechanics, driver development, and kernel-userspace data flow.


🎙️ Spoken Script & Architectural Walkthrough

To explain the complete architecture, I divide the project into four parts: kernel modules and process inspection, memory translation and accounting, character device drivers with ioctl, and exposing kernel information through procfs and sysfs.


First, Linux kernel modules and process inspection

I wrote Linux kernel modules that can be loaded into the running kernel using insmod. I built these modules using Kbuild (1/Kbuild), the Linux kernel's build system, which produces a .ko kernel module.

When the module is loaded, module_init() runs the initialization function, and when it is removed using rmmod, module_exit() runs the cleanup function.

These modules run in kernel space, so they can access kernel data structures and kernel APIs.

I started by working with Linux's internal process representation. In the kernel, each process is represented by a task_struct, which contains information about the process.

Given a PID, I first use find_vpid() to get the corresponding struct pid. Then I use pid_task() to get the corresponding task_struct.

I then use the kernel's process-list functions (list_for_each_entry_rcu()) to traverse the child processes and inspect the parent-child relationships.

While traversing the process structures, I use Read-Copy-Update, or RCU, with rcu_read_lock() and rcu_read_unlock(). This makes the traversal safe when processes are being created or terminated concurrently, because it prevents the objects I am reading from being reclaimed while I am accessing them.

🎯 Follow-up: Why RCU?
If the interviewer asks "Why do you use RCU?", keep it simple:
"Because the process structures can change while I am traversing them. RCU ensures that the objects I am currently reading are not reclaimed while I am accessing them."

(Hook: Deep dive into Topic 01: Lifecycle & Tasks)


Second, memory translation and accounting

For memory inspection, I implemented a manual 5-level page-table walk (1/lkm3.c) to translate a process's virtual address into a physical address.

I started with the process's mm_struct (task->mm), which gives me access to the process's page-table hierarchy. From there, I walked through the five levels: PGD → P4D → PUD → PMD → PTE.

At each level, I checked whether the entry was valid before moving to the next level. At the final PTE level, I used pte_pfn() to get the Page Frame Number, and combined it with the offset within the page to calculate the physical address.

I also worked on memory accounting. I compared the kernel's fast, RSS-based estimate of a process's physical memory usage (get_mm_rss()) with a more detailed method (get_mapped_size()) that walks through the process's virtual memory pages and checks which pages are currently present.

For this, I iterated through the process's Virtual Memory Areas, or VMAs, using the Linux VMA iterator (VMA_ITERATOR). For each VMA, I scanned the virtual address range in 4 KiB page-sized steps (vaddr += PAGE_SIZE) and checked the page-table entries to determine whether the page is present.

I also handled 2 MiB Transparent Huge Pages. A huge page can be detected at the PMD level using pmd_trans_huge(). When I detect one, I count the full 2 MiB and advance the address by 2 MiB (HPAGE_PMD_SIZE) instead of checking the 512 individual 4 KiB pages.

🎯 Follow-Up Hook: What happens if a page table entry is missing or swapped out during a manual walk?
The defensive checks (_none() or missing _present()) flag the page as not resident in RAM, returning an error rather than dereferencing a garbage or swap identifier which would cause a fatal kernel Oops.

(Hook: Deep dive into Topic 02: MMU Walks and Topic 03: Memory Accounting)


Third, character device drivers and ioctl

To allow userspace programs to trigger these kernel operations, I developed custom character device drivers.

I first defined the operations supported by the driver using the struct file_operations structure (2.1/chardev.c:110). In my driver, the main operation is ioctl(), which is mapped to my device_ioctl() function through the .unlocked_ioctl field. So, when a userspace program calls the ioctl() system call on this character device, the kernel invokes my device_ioctl() function to handle that request.

During initialization, I dynamically allocate the major and minor numbers using alloc_chrdev_region(). I then initialize the cdev and associate it with my file operations using cdev_init(), and register the character device with the kernel using cdev_add().

Finally, I use class_create() and device_create() to create the corresponding device entry under /dev. This allows userspace programs to access the device and send commands through ioctl().

I developed two main character drivers.

The first is a physical memory driver (2.1/chardev.c). Userspace provides a virtual address and data. The driver performs the 5-level page-table walk (page_table_walk()) to find the corresponding physical address, converts that physical address to a kernel virtual address using phys_to_virt(), and writes the data to that physical memory location.

The second driver handles process reparenting (2.2/chardev.c). It allows an existing process to be reparented to another process at runtime without creating a new process using fork(). The driver updates the required parent information and moves the child between the corresponding parent-child lists (change_parent()).

🎯 Follow-up: What is a character device?
"A character device is a device that provides data as a stream of bytes and allows userspace to communicate with a kernel driver through operations such as read, write, or ioctl."
🎯 Follow-up: Why did you use a character device?
"I used a character device because my operations are command-based rather than continuous block-based data access. I needed a simple interface through which userspace could send commands and arguments to my kernel driver using ioctl()."
🎯 Follow-up: Why ioctl() instead of standard read() / write()?
"read() and write() only accept an unstructured byte stream. My operations required structured control commands with custom argument structs (such as passing a virtual address and byte payload). ioctl() lets userspace pass strongly typed request codes and pointer structs directly to the driver."
🎯 Follow-up: Why is write_lock_irq(&tasklist_lock) mandatory during process reparenting?
"Reparenting modifies parent and real_parent pointers and splices circular linked lists. If an interrupt occurs or another CPU calls do_exit concurrently while lists are half-modified, the kernel will crash or deadlock."

(Hook: Deep dive into Topic 04: Character Devices & ioctl and Topic 05: Process Reparenting)


Fourth, exposing kernel information through procfs and sysfs

Finally, I implemented two standard Linux interfaces to expose kernel information to userspace: procfs and sysfs.

For the /proc interface, I created a proc entry using proc_create() and defined its operations using struct proc_ops. When userspace reads the entry, my read function (procfile_read()) collects the system-wide page-fault statistics using all_vm_events() and returns them to userspace.

For the sysfs interface, I created a kobject under /sys/kernel/mem_stats using kobject_init_and_add() with four attributes: pid, virtmem, physmem, and unit. These attributes allow userspace to provide a PID and query the process's virtual and physical memory usage in different units.

So, /proc is used for system-wide page-fault statistics, while sysfs is used for process-specific memory information.

So overall, the project connects userspace, Linux kernel data structures, virtual memory, and physical memory. It gave me hands-on experience with kernel modules, process structures, page tables, VMAs, character drivers, ioctl, procfs, and sysfs.

🎯 Follow-up: When do you use procfs vs. sysfs?
"procfs (/proc) is intended for procedural streams, process metadata, and kernel runtime execution events (like page fault counters or /proc/cpuinfo). sysfs (/sys) is strictly an object-oriented hierarchy representing kernel objects, devices, and buses, following the Unix rule of 'one value per file' for reading or tuning configurable attributes."
🎯 Follow-up: Why use per-CPU counters for page fault monitoring?
"A shared global atomic integer would suffer severe cacheline bouncing across multi-core CPUs during heavy page fault workloads. Reading local per-CPU counters via all_vm_events() allows cores to increment their own counters without lock contention, scaling linearly with core count."
🎯 Follow-up: How does container_of() work in your sysfs attribute handlers?
"Sysfs callbacks only receive a pointer to the generic struct kobject. Using container_of(kobj, struct sysfs_entry, kobj), the kernel calculates the offset of the kobj member within my outer struct sysfs_entry and subtracts it from the pointer, safely recovering the parent struct containing the target PID and display unit."

(Hook: Deep dive into Topic 06: Kernel Telemetry)


❓ Anticipated Deep-Dive Architecture Questions & Crisp Answers

Q1: Walk me through the complete end-to-end execution flow when a userspace application triggers your character driver.

Answer: The userspace application opens /dev/mem_driver and executes an ioctl() call containing a request code and memory payload. Inside the kernel, the VFS routes the call to our driver's unlocked_ioctl handler. The driver uses copy_from_user() to safely import user data, looks up the target process's task_struct under RCU, and acquires mmap_read_lock() on its mm_struct. It then executes the 5-level page-table walk from PGD down to PTE, converts the physical page frame number to a direct-mapped kernel address via phys_to_virt(), writes the payload, releases locks, and returns 0 to userspace.

Q2: What is the locking strategy across the different modules in your architecture?

Answer: We applied three distinct synchronization mechanisms tailored to specific kernel constraints:
1. Lockless RCU (rcu_read_lock): Used for process enumeration and task hierarchy reads, ensuring readers never block writers or degrade scheduler throughput.
2. Reader/Writer Semaphore (mmap_read_lock): Acquired when reading memory maps and page tables to prevent race conditions against concurrent mmap()/munmap() calls.
3. Global IRQ-Disabled Spinlock (write_lock_irq(&tasklist_lock)): Strictly reserved for live process reparenting to guarantee atomic list surgery across all CPU cores and interrupt contexts.

Q3: What happens if an administrator unloads the module (rmmod) while a userspace application holds an open file descriptor?

Answer: The character driver assigns .owner = THIS_MODULE in its struct file_operations. When userspace opens the device file, the VFS automatically increments the module's reference counter (refcnt). If an administrator runs rmmod, the kernel checks this counter, detects an active consumer, and rejects the unload with -EBUSY ("Device or resource busy"). The module can only be unloaded after all userspace descriptors are closed, preventing dangling function pointers and kernel panics.