Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering MMU address translation, 5-level page table navigation (PGD, P4D, PUD, PMD, PTE), entry validation, cross-architecture PFN extraction, and user-space pagemap verification.
What You Mentioned: "5-level manual page table walks (VA → PA)"
Why It Was Done (The Motivation): Understand hardware MMU translation directly from software.
Problems Faced & How Solved (The Reality): Walked PGD → P4D → PUD → PMD → PTE with null/bad entry checking at each level, extracting PFN via pte_pfn().
Opening & Scope:
"In lkm3.c:23 and chardev.c:24, I implemented a manual 5-level page table walk to resolve virtual memory addresses to physical RAM frames, simulating the exact translation steps performed by the CPU's Memory Management Unit (MMU) on Linux 6.1."
Step 1: Accessing the Process Memory Descriptor (mm_struct):
"First, given a target task, I acquired its memory descriptor via task->mm (1/lkm3.c:45). In Linux, the root of the per-process page table is stored in mm->pgd (1/lkm3.c:48), which is loaded into the CPU's CR3 control register on x86_64 during a context switch. If task->mm is null, it indicates a kernel thread, which borrows memory contexts and has no private user address space."
Step 2: 5-Level Traversal Sequence:
"Next, I walked the five hierarchical page table levels supported in Linux 6.1 using the kernel's offset macros:
pgd = pgd_offset(mm, vaddr) (1/lkm3.c:48) to find the Page Global Directory entry.p4d = p4d_offset(pgd, vaddr) (1/lkm3.c:55) for the Page 4th Directory.pud = pud_offset(p4d, vaddr) (1/lkm3.c:62) for the Page Upper Directory.pmd = pmd_offset(pud, vaddr) (1/lkm3.c:69) for the Page Middle Directory.pte = pte_offset_kernel(pmd, vaddr) (1/lkm3.c:76) for the final Page Table Entry pointing to the physical frame."Step 3: Defending Against Unmapped and Bad Entries:
"Then, at every single stage of the walk, defensive validation was mandatory. If an intermediate directory pointer is empty or points to a non-existent table, dereferencing it causes an immediate kernel crash. I verified every level using pgd_none(), !pgd_present(), and pgd_bad() (1/lkm3.c:49-53) checks. If any entry failed validation, the module aborted cleanly, reporting that the virtual page was not backed by physical memory."
Step 4: Physical Address Calculation & Cross-Arch Portability:
"Once I reached a valid leaf pte, I computed the physical address using the formula:
$$\text{PhysAddr} = (\text{pte\_pfn}(*\text{pte}) \ll \text{PAGE\_SHIFT}) \mid (\text{vaddr} \ \& \ \sim\text{PAGE\_MASK})$$
To ensure cross-architecture portability across x86_64 and arm64, I deliberately avoided architecture-specific bitmasks like PTE_PFN_MASK or PTE_ADDR_MASK. Instead, I used the architecture-agnostic pte_pfn() (1/lkm3.c:83) macro to extract the Page Frame Number, guaranteeing consistent behavior regardless of hardware bit shifts (implemented identically in 2.1/chardev.c:64)."
Step 5: User-Space Verification via /proc/self/pagemap:
"Finally, to prove correctness, I built a user-space client (dev_user.c:39-58) that allocated dynamic heap memory, queried the module via ioctl for its physical address, and cross-referenced the result by reading /proc/self/pagemap. The physical frame address returned by my page table walk matched /proc/self/pagemap with 100% bitwise accuracy."
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. Memory Context | Acquired task->mm |
Checks if process has valid user memory context. | Kernel threads lack an mm_struct (mm == NULL) and must be skipped. |
1/lkm3.c:452.1/chardev.c:70 |
| 2. Table Navigation | Traversed 5 directory levels | PGD $\rightarrow$ P4D $\rightarrow$ PUD $\rightarrow$ PMD $\rightarrow$ PTE via *_offset macros. |
Accurately mirrors hardware MMU radix tree traversal on 5-level paging kernels. | 1/lkm3.c:48-762.1/chardev.c:34-58 |
| 3. Entry Validation | Verified entries with _none / _bad / _present |
Checked presence and integrity flags before descending. | Prevents kernel null-pointer dereferences and page faults in Ring 0. | 1/lkm3.c:49-812.1/chardev.c:35-62 |
| 4. Address Synthesis | Extracted PFN with pte_pfn() |
Shifted PFN by PAGE_SHIFT (12) and bitwise-ORed lower 12 offset bits. |
pte_pfn() abstracts architecture differences between x86_64 and arm64. |
1/lkm3.c:832.1/chardev.c:64 |
| 5. Verification | Cross-referenced /proc/self/pagemap |
Read 64-bit pagemap descriptor from user space and compared PFN. | Validates driver correctness against the kernel's authoritative page accounting. | 2.1/dev_user.c:39-58 |
get_user_pages() or /proc/[pid]/pagemap?Answer:
get_user_pages()(GUP) forcibly faults in unmapped pages, pins physical frames into memory, and alters process state. A memory inspection tool or translation driver must be non-intrusive—observing page residency without mutating state or triggering disk/swap I/O./proc/[pid]/pagemapis a user-space pseudo-file requiring VFS file reads, context switches, and string parsing. A manual walk in Ring 0 directly reads the MMU radix tree in real time with microsecond latency and zero side effects.
Answer: Two fatal failure modes occurred: 1) Kernel threads have
task->mm == NULL, so attempting to readmm->pgdcauses an immediate null pointer dereference panic. 2) Because Linux uses demand paging, memory allocated withmalloc()ormmap()does not have physical frames populated until the first byte write. Attempting to traverse intermediate directory pointers when entries are empty (pgd_none,p4d_none,pud_none,pmd_none,pte_none) or marked bad (*_bad()) triggers kernel page faults in Ring 0. We guarded every single level with defensive_noneand_badvalidation before advancing to the next level.
Answer: We wrote a dedicated user-space verification harness (
dev_user.c) that allocated a page, wrote a unique signature into it, and resolved its physical page frame number (PFN) via/proc/self/pagemapby seeking to(vaddr / PAGE_SIZE) * sizeof(uint64_t)and masking bit 63 (present) and bits 0-54 (PFN). We then sent that same virtual address to our kernel character driver viaioctl. The harness cross-referenced the driver's returned physical address against(pagemap_pfn << 12) | page_offset, proving bit-for-bit translation accuracy.