Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering the 4 namespaces, syscall sequence, PID 1 inheritance rule, and supervisor process controls.
Opening & Scope:
"For process isolation, I isolated the container across four primary Linux namespaces: UTS for hostname, PID for process numbering, Mount for private filesystems, and Network for private interfaces."
Step 1: Choosing clone() vs unshare():
"First, to create these boundaries, I worked with two low-level syscalls: clone() and unshare(). clone() spawns a brand-new child process directly inside new namespaces with its own allocated stack, whereas unshare() detaches the current running process from parent namespaces in place. I used unshare for the container supervisor so the calling script can pivot into the container root before executing."
Step 2: Handling the PID Namespace Inheritance Rule:
"Next, a critical kernel detail I had to handle was PID namespace inheritance. When a process calls unshare for a PID namespace, its own PID does not change because the kernel cannot rewrite a running process's tracking structures. Instead, the new PID namespace only applies to subsequent children. So, I had the supervisor immediately fork a child process after entering the namespace—and this forked child became PID 1 inside the container."
Step 3: Private Mount Table and Remounting /proc:
"Then, for filesystem isolation, I created a private mount namespace and mounted a fresh /proc filesystem inside it. This was essential because if you don't remount /proc, tools like ps and top will read the host's /proc and expose all host processes, breaking process isolation."
Step 4: Preventing Orphan Zombie Processes:
"To handle process cleanup, I configured the supervisor to automatically send a SIGKILL to all child processes when it exits. Without this, if the container supervisor crashes, background container processes would become orphaned zombies reparenting to the host's PID 1."
Step 5: Attaching to Running Containers with pidfd_open and setns:
"Finally, to attach to an already-running container for commands like exec, I used pidfd_open to get a stable file descriptor to the container's PID 1 and called setns to join its UTS, PID, Mount, and Network namespaces. Using pidfd_open instead of file paths in /proc prevents race conditions if a PID gets recycled before attachment."
| Step | What Was Done | How It Works | Why This Mechanism / Order |
|---|---|---|---|
| 1. Syscall Selection | Chose clone() vs unshare() |
clone() creates a new child in new namespaces; unshare() detaches the current process. |
unshare allows the supervisor to prepare namespaces and pivot root before executing the target binary. |
| 2. PID 1 Creation | Forked child after PID unshare | unshare(CLONE_NEWPID) applies only to future children. |
Kernel cannot change a running process's PID dynamically without breaking scheduler structs. |
3. Mount & /proc |
Remounted /proc in new mount ns |
Pivoted apparent root and mounted fresh procfs instance. | Prevents container processes from reading host /proc and seeing host processes via ps. |
| 4. Process Cleanup | Configured automatic child killing | Supervisor signals all children in namespace on exit. | Prevents orphaned container processes from leaking to host PID 1 as zombies. |
| 5. Exec Attachment | Attached via pidfd_open + setns |
Acquired process FD to PID 1 and joined its namespaces. | Eliminates PID recycling race conditions when joining active containers. |
Answer: In Linux, a process's PID struct is allocated at process creation and deeply tied to scheduler task structs, signal handling tables, and credentials. Changing it dynamically would break thread group semantics and create kernel inconsistencies. Thus,
CLONE_NEWPIDonly sets the namespace for descendants created on subsequentfork()/clone()calls.
pidfd_open() instead of opening /proc/<pid>/ns/pid?Answer: Accessing
/proc/<pid>/ns/pidis vulnerable to a PID reuse race condition. If the target process exits and another process is assigned the recycled PID beforeopen()runs,setns()would attach to an unrelated process.pidfd_open()creates a file descriptor tied to that exact process instance, eliminating the race.
ps aux inside a container without remounting /proc?Answer:
psreads/procto display process state. If/procis not remounted in the new mount namespace, the container shares the host's/proc, sopswill list all host processes even though the container's PID namespace is isolated. Remounting procfs reflects only the container's private PID space.