Topic 02: Resource Control via cgroups v2 & CPU Quotas

Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering cgroups v2 controller hierarchy, CFS bandwidth quotas, and race-free bootstrap process attachment.


🎙️ Pointwise Spoken Speech (Word-for-Word Delivery)


📋 Step-by-Step Summary (What, How & Why)

Step What Was Done How It Works Why This Mechanism / Order
1. Subtree Delegation Enabled CPU controller down hierarchy Delegated +cpu controller in root cgroup.subtree_control. In cgroups v2, child groups cannot enforce limits unless explicitly enabled top-down.
2. CFS Quota Setup Enforced hard CPU bandwidth cap Wrote 50,000 µs quota per 100,000 µs period into cpu.max. Caps container at 50% core; scheduler throttles threads when budget is exhausted.
3. Policy Choice Chose CFS quota over shares Used non-work-conserving quota instead of proportional weights. Guarantees hard ceiling so noisy workloads cannot spike host latency.
4. Race-Free Attachment Attached shell PID before exec Added launcher PID to cgroup.procs prior to container bootstrap. Membership is inherited across fork/exec, eliminating unthrottled startup windows.

❓ Anticipated Interview Questions & Crisp Answers

Q1: What is the main difference between cgroups v1 and cgroups v2?

Answer: cgroups v1 had independent hierarchies for each resource controller, making it impossible to coordinate limits across controllers (e.g., memory limits and block I/O writebacks couldn't sync). cgroups v2 enforces a single unified hierarchy where every process belongs to exactly one cgroup, with explicit top-down controller delegation.

Q2: How does CFS bandwidth quota differ from CPU shares/weights?

Answer: CPU shares (cpu.weight) provide proportional, work-conserving sharing—a process can consume 100% of an idle CPU if no other process wants it. In contrast, CFS bandwidth quota (cpu.max) is a hard non-work-conserving limit: once the cgroup uses up its allocated quota within the current period, tasks are throttled and put to sleep until the next period, regardless of idle CPU availability.

Q3: Why does cgroup.procs membership persist across exec?

Answer: exec replaces the process address space and register state but preserves the underlying kernel task_struct and its pointer to the cgroup_subsys_state. Since the process itself is not destroyed, its cgroup association remains strictly intact.

Q4: What is CFS (Completely Fair Scheduler), and how do you know your kernel uses it?

Answer:
- What it is: CFS is the Linux kernel's default process scheduler for normal tasks (SCHED_NORMAL/SCHED_OTHER). It models an "ideal multi-tasking CPU" on real hardware by tracking each task's execution time via virtual runtime (vruntime) using a self-balancing Red-Black Tree. Tasks with the smallest vruntime are scheduled next to guarantee fair CPU proportion. When paired with cgroups bandwidth control (CONFIG_CFS_BANDWIDTH), it enforces hard runtime quotas per period.
- How you know your kernel uses it:
1. Check kernel compile configuration: grep CONFIG_CFS_BANDWIDTH /boot/config-$(uname -r) (or /proc/config.gz), which confirms CFS bandwidth controller support.
2. Inspect /proc/sched_debug or /proc/sys/kernel/sched_* to view active CFS scheduling parameters (like sched_latency_ns and sched_min_granularity_ns).
3. Inspect cgroup throttling stats: cat /sys/fs/cgroup/<group>/cpu.stat—seeing nr_throttled and throttled_usec metrics increment when a container hits its cpu.max limit proves the kernel's CFS bandwidth throttler is actively enforcing the quota.