Target Duration: 2โ4 minutes (~300โ450 spoken words)
Focus: Pointwise verbal delivery covering Ring 0 out-of-tree execution,module_init/module_exit,module_param, circulartask_structtraversal, and RCU locking semantics.
What You Mentioned: "Kernel modules & lockless process inspection"
Why It Was Done (The Motivation): Safely inspect task_struct and children list in Ring 0 without stalling the kernel scheduler.
Problems Faced & How Solved (The Reality): Solved using rcu_read_lock() and list_for_each_entry_rcu(), preventing race conditions during process fork/exit.
Opening & Scope:
"In this part of the project, I developed foundational Linux Kernel Modules (1/lkm1.c and 1/lkm2.c) targeting Linux 6.1 to inspect operating system processes directly from Ring 0 without relying on user-space system calls."
Step 1: Out-of-Tree LKM Lifecycle (module_init vs module_exit):
"First, to execute code in kernel space, I structured each module using the standard out-of-tree Kbuild pattern. I used module_init() (1/lkm1.c:30) to designate the initialization function invoked upon insmod, where kernel data structures are initialized. Symmetrically, I implemented module_exit() (1/lkm1.c:31) for rmmod to clean up registered resources. Ensuring complete cleanup in module_exit is vital in Ring 0 because any leaked registration or lingering pointer will trigger a kernel panic on subsequent access."
Step 2: Dynamic Configuration via module_param:
"Next, to allow runtime parameter injection without recompiling, I used the module_param() and MODULE_PARM_DESC() (1/lkm2.c:11-12) macros. In lkm2, this allowed passing target PIDs at module load time (e.g., insmod lkm2.ko pid=1337), which the module parses directly into static kernel variables."
Step 3: Process List Iteration with for_each_process:
"Then, in lkm1, to discover running processes, I traversed the kernel's circular doubly linked list of all processes using the for_each_process() (1/lkm1.c:17) macro. Starting from init_task, this macro advances through task_struct pointers. For each task, I inspected task_is_running(task) (1/lkm1.c:18) to filter for tasks in the TASK_RUNNING state, logging their PIDs and command names to the kernel ring buffer using pr_info()."
Step 4: PID Resolution and Child Hierarchy Traversal:
"In lkm2, to inspect a process's children, I resolved the input PID into a struct task_struct*. In modern kernels, you cannot simply index an array; I used find_vpid() & pid_task() (1/lkm2.c:39) to acquire the task struct within the current PID namespace. From there, I iterated over children using list_for_each_entry_rcu() (1/lkm2.c:46) to inspect each child's process descriptors."
Step 5: Concurrency Safety with RCU (Read-Copy-Update):
"Finally, a crucial OS principle I applied was lockless synchronization using Read-Copy-Update (RCU). Because tasks can fork or exit concurrently while the module is reading the task list, holding a heavy global lock would harm scheduler throughput. Instead, I wrapped traversals inside rcu_read_lock() & rcu_read_unlock() (1/lkm2.c:37, 50). This disabled preemption and informed the kernel's RCU subsystem that a read-side critical section was active, guaranteeing that no task_struct in the list would be freed by the garbage collector until the read completed."
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. Lifecycle Setup | Implemented module_init & module_exit |
Compiles as .ko via Kbuild; hooks into insmod and rmmod. |
Ring 0 modules must cleanly release all allocated structures to prevent kernel panics. | 1/lkm1.c:30-311/lkm2.c:59-60 |
| 2. Parameter Binding | Configured module_param |
Defines static variables populated by insmod arguments. |
Allows dynamic target selection (e.g., target PID) without recompiling module source. | 1/lkm2.c:11-12 |
| 3. Task Iteration | Traversed all tasks via for_each_process |
Iterates circular doubly linked list of task_struct starting at init_task. |
Enables global process inspection directly from kernel memory. | 1/lkm1.c:17-21 |
| 4. Hierarchy Walk | Resolved PID & traversed children list |
Used find_vpid() $\rightarrow$ pid_task() and traversed task->children. |
Safely navigates the process tree to identify parent-child relationships. | 1/lkm2.c:39-48 |
| 5. RCU Synchronization | Guarded traversal with rcu_read_lock |
Enclosed reads in RCU read-side critical sections. | Prevents use-after-free without acquiring expensive blocking locks during concurrent forks/exits. | 1/lkm2.c:37, 50 |
/proc or running ps from user space?Answer: User-space utilities like
psparse thousands of pseudo-files under/proc/[pid]/stat, which incurs significant VFS overhead, context switching, and string parsing. Furthermore,/procsnapshots are not atomic across processesโby the time user space finishes reading, child processes may have exited or reparented. Inspecting processes directly in Ring 0 allows atomic, sub-microsecond traversal of kernel task structures (struct task_struct) directly from memory, resolving PID namespaces viafind_vpid()andpid_task()without file system latency.
Answer: The Linux task list is constantly modified as processes fork, terminate, and exit. If a module iterates the list using raw pointers, another CPU might free a terminating task's
task_struct, leading to a fatal use-after-free kernel crash. Holding a heavy global write lock liketasklist_lockwould halt scheduling across all CPUs. Instead, wrapping traversals inrcu_read_lock()/rcu_read_unlock()ensures lockless read access: even if a task is unlinked, its memory cannot be reclaimed until our CPU completes its critical section and passes a grace period.
module_exit()?Answer: In Linux kernel space, there is no automatic garbage collection. If
module_exit()fails to unregister a character device, procfs entry, or sysfs attribute, the kernel's core dispatch tables continue pointing to the memory where the module was mapped. Whenrmmodcompletes and unmaps the module's code pages, any subsequent read/write access to that dangling address triggers an immediate Ring 0 page fault and fatal kernel panic (Oops). All allocations and registrations must be strictly unwound in reverse order of initialization.