Target Duration: 2โ4 minutes (~300โ450 spoken words)
Focus: Pointwise verbal delivery covering $O(1)$ fast RSS counters vs $O(N)$ Maple tree VMA iteration, 2 MiB Transparent HugePage (THP) detection at the PMD level, and cursor stride optimizations.
What You Mentioned: "Memory accounting & 2 MiB Transparent Huge Pages"
Why It Was Done (The Motivation): Compare $O(1)$ fast RSS estimate against exact $O(N)$ VMA and page table walk.
Problems Faced & How Solved (The Reality): Traversed VMAs using Linux 6.1 Maple Tree API; detected THPs via pmd_trans_huge() and advanced by 2 MiB instead of checking 512 individual 4 KiB pages.
Opening & Scope:
"In 1/lkm4.c and 1/lkm5.c, I engineered process memory accounting subsystems to analyze virtual address allocations, resident physical memory (RSS), and 2 MiB Transparent HugePages (THP) under Linux 6.1."
Step 1: O(1) Fast RSS vs O(N) Deep Page Walks:
"First, I investigated how the kernel accounts for resident physical memory. The kernel provides an $O(1)$ fast lookup helperโget_mm_rss(mm) (1/lkm4.c:90)โwhich aggregates atomic counters (MM_FILEPAGES, MM_ANONPAGES, MM_SHMEMPAGES). This is what utilities like top and /proc/[pid]/statm read. However, to verify ground truth, I implemented an $O(N)$ deep traversal (get_mapped_size in 1/lkm4.c:19) that iterates through every Virtual Memory Area (VMA) and checks each individual page entry's presence."
Step 2: Modern VMA Iteration with the Linux 6.1 Maple Tree API:
"Next, in Linux 6.1, the historical red-black tree and linked list for VMAs were replaced by the Maple Tree data structure. To walk the VMAs safely and idiomatically, I used the Maple Tree iterator API:
VMA_ITERATOR(vma_iter, mm, 0) & for_each_vma() (1/lkm4.c:107-112) (also implemented in 1/lkm5.c:108-110).
For each VMA, I accumulated the virtual range (vma->vm_end - vma->vm_start) and walked the underlying page tables to verify physical residency."
Step 3: Detecting 2 MiB Transparent HugePages (THP):
"Then, in lkm5, I focused on Transparent HugePages. Modern x86_64 and arm64 architectures support 2 MiB huge pages, where the PMD points directly to a large physical frame rather than a sub-level PTE page table. To identify THPs, I walked the page tables down to the PMD stage and checked pmd_trans_huge(*pmd) (1/lkm5.c:72). If true, the memory region is backed by a 2 MiB compound huge page."
Step 4: The 2 MiB Cursor Stride Optimization:
"A critical algorithmic optimization I implemented was the stride skip. Normally, a page walk loop increments the virtual address by PAGE_SIZE (4 KiB). When encountering a 2 MiB THP, continuing with a 4 KiB step would erroneously attempt 512 redundant lookups on non-existent PTE tables, causing corrupt statistics and redundant CPU cycles. Instead, once pmd_trans_huge returned true, I incremented the virtual cursor directly by HPAGE_PMD_SIZE (2 MiB) (1/lkm5.c:76-78), guaranteeing exact $O(1)$ processing for the huge block."
Step 5: Testing with madvise and MADV_HUGEPAGE:
"Finally, to validate the module under controlled conditions, I wrote a user-space test harness (1/test2.c:8). It allocated aligned memory buffers and issued madvise(addr, size, MADV_HUGEPAGE) to instruct the kernel's khugepaged daemon to back the allocation with 2 MiB pages. Running lkm5 against this process accurately reported the exact number of 2 MiB THP blocks allocated."
| Step | What Was Done | How It Works | Why This Mechanism / Order | Code Reference |
|---|---|---|---|---|
| 1. RSS Comparison | Evaluated $O(1)$ vs $O(N)$ RSS | Compares atomic counter get_mm_rss() against full PTE walk. |
Atomic counters offer microsecond speed; full walks reveal exact physical residency per VMA. | 1/lkm4.c:901/lkm4.c:110-112 |
| 2. VMA Iteration | Iterated VMAs via Maple Tree | Utilized VMA_ITERATOR and for_each_vma macros on Linux 6.1. |
Replaces deprecated red-black tree traversal with modern RCU-safe B-tree variant. | 1/lkm4.c:107-1121/lkm5.c:108-110 |
| 3. THP Identification | Inspected PMDs via pmd_trans_huge |
Checks if PMD leaf entry maps directly to 2 MiB physical memory. | THPs bypass the 4th level (PTE); treating them as standard PMDs would fail. | 1/lkm5.c:72 |
| 4. Stride Advance | Advanced cursor by HPAGE_PMD_SIZE |
Jumped vaddr += 2\text{ MiB} upon detecting huge page. |
Prevents 512 redundant lookups and prevents double-counting identical pages. | 1/lkm5.c:76-78 |
| 5. Test Verification | Formed THP workload via madvise |
Triggered huge page allocation with MADV_HUGEPAGE. |
Provides repeatable, deterministic test data for kernel verification. | 1/test2.c:8-25 |
get_mm_rss())?Answer:
get_mm_rss()(which powersps,top, and/proc/[pid]/statm) is an $O(1)$ read of atomic counters (MM_FILEPAGES,MM_ANONPAGES,MM_SHMEMPAGES) updated lazily during page fault handling and unmapping. In multi-threaded systems or copy-on-write sharing, atomic counters can diverge or over-count shared pages. Furthermore, atomic counters cannot distinguish which specific VMA or memory segment (heap, stack, anonymous mmap) owns the physical frames. Our $O(N)$ VMA traversal inspects present PTE flags directly, providing the ground-truth physical footprint broken down by address region.
Answer: Standard page table traversals advance virtual addresses in 4 KiB increments (
PAGE_SIZE). But a THP is mapped at the PMD level as a single contiguous 2 MiB leaf entry (pmd_trans_huge(*pmd)is true). If a walker treats a huge PMD as a pointer to a PTE page or increments by 4 KiB, two things break: 1) callingpte_offset_kernel()corrupts kernel pointers because the PMD holds raw data, not a pointer table; and 2) stepping 4 KiB causes 512 redundant lookups of the exact same huge page, artificially inflating physical residency by 512x. Upon detectingpmd_trans_huge(), our loop immediately accounts for 2 MiB and strides the address forward byHPAGE_PMD_SIZE(2 MiB), avoiding 511 redundant iterations.
Answer: The Linux kernel's background
khugepageddaemon collapses pages opportunistically and non-deterministically. To reliably trigger THP allocation for our test harness, we allocated an aligned 10 MiB buffer withposix_memalign(), calledmadvise(buf, size, MADV_HUGEPAGE)to explicitly register the VMA for huge pages, and touched every 4 KiB page. Our test harness then compared our kernel module's reported huge page count against/proc/self/smaps(AnonHugePages: 10240 kB), verifying 100% detection accuracy.