Topic 03: Process Memory Accounting & Transparent HugePages (THP)

Target Duration: 2โ€“4 minutes (~300โ€“450 spoken words)
Focus: Pointwise verbal delivery covering $O(1)$ fast RSS counters vs $O(N)$ Maple tree VMA iteration, 2 MiB Transparent HugePage (THP) detection at the PMD level, and cursor stride optimizations.

๐ŸŽฏ Strategic Follow-Up Hook (From Elevator Pitch)

What You Mentioned: "Memory accounting & 2 MiB Transparent Huge Pages"

Why It Was Done (The Motivation): Compare $O(1)$ fast RSS estimate against exact $O(N)$ VMA and page table walk.

Problems Faced & How Solved (The Reality): Traversed VMAs using Linux 6.1 Maple Tree API; detected THPs via pmd_trans_huge() and advanced by 2 MiB instead of checking 512 individual 4 KiB pages.


๐ŸŽ™๏ธ Pointwise Spoken Speech (Word-for-Word Delivery)


๐Ÿ“‹ Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order Code Reference
1. RSS Comparison Evaluated $O(1)$ vs $O(N)$ RSS Compares atomic counter get_mm_rss() against full PTE walk. Atomic counters offer microsecond speed; full walks reveal exact physical residency per VMA. 1/lkm4.c:90
1/lkm4.c:110-112
2. VMA Iteration Iterated VMAs via Maple Tree Utilized VMA_ITERATOR and for_each_vma macros on Linux 6.1. Replaces deprecated red-black tree traversal with modern RCU-safe B-tree variant. 1/lkm4.c:107-112
1/lkm5.c:108-110
3. THP Identification Inspected PMDs via pmd_trans_huge Checks if PMD leaf entry maps directly to 2 MiB physical memory. THPs bypass the 4th level (PTE); treating them as standard PMDs would fail. 1/lkm5.c:72
4. Stride Advance Advanced cursor by HPAGE_PMD_SIZE Jumped vaddr += 2\text{ MiB} upon detecting huge page. Prevents 512 redundant lookups and prevents double-counting identical pages. 1/lkm5.c:76-78
5. Test Verification Formed THP workload via madvise Triggered huge page allocation with MADV_HUGEPAGE. Provides repeatable, deterministic test data for kernel verification. 1/test2.c:8-25

โ“ Anticipated Interview Questions & Crisp Answers

Q1: Why did you build an $O(N)$ VMA and page table walk when the kernel already provides fast atomic RSS counters (get_mm_rss())?

Answer: get_mm_rss() (which powers ps, top, and /proc/[pid]/statm) is an $O(1)$ read of atomic counters (MM_FILEPAGES, MM_ANONPAGES, MM_SHMEMPAGES) updated lazily during page fault handling and unmapping. In multi-threaded systems or copy-on-write sharing, atomic counters can diverge or over-count shared pages. Furthermore, atomic counters cannot distinguish which specific VMA or memory segment (heap, stack, anonymous mmap) owns the physical frames. Our $O(N)$ VMA traversal inspects present PTE flags directly, providing the ground-truth physical footprint broken down by address region.

Q2: What performance bug or accounting error occurred when traversing Transparent HugePages (THPs)?

Answer: Standard page table traversals advance virtual addresses in 4 KiB increments (PAGE_SIZE). But a THP is mapped at the PMD level as a single contiguous 2 MiB leaf entry (pmd_trans_huge(*pmd) is true). If a walker treats a huge PMD as a pointer to a PTE page or increments by 4 KiB, two things break: 1) calling pte_offset_kernel() corrupts kernel pointers because the PMD holds raw data, not a pointer table; and 2) stepping 4 KiB causes 512 redundant lookups of the exact same huge page, artificially inflating physical residency by 512x. Upon detecting pmd_trans_huge(), our loop immediately accounts for 2 MiB and strides the address forward by HPAGE_PMD_SIZE (2 MiB), avoiding 511 redundant iterations.

Q3: How did you create a reproducible test workload in user space to verify your THP detection logic?

Answer: The Linux kernel's background khugepaged daemon collapses pages opportunistically and non-deterministically. To reliably trigger THP allocation for our test harness, we allocated an aligned 10 MiB buffer with posix_memalign(), called madvise(buf, size, MADV_HUGEPAGE) to explicitly register the VMA for huge pages, and touched every 4 KiB page. Our test harness then compared our kernel module's reported huge page count against /proc/self/smaps (AnonHugePages: 10240 kB), verifying 100% detection accuracy.