Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering container lifecycle transitions (build,run,exec,stop),nsenterchild discovery, and reverse unmount teardown.
Opening & Scope:
"For the container lifecycle, I designed a complete state machine managing container creation, runtime execution, live command attachment, and clean resource teardown."
Step 1: The build Stage:
"First, the build command compiles the build recipe into an ordered list of immutable OverlayFS layer paths and registers this metadata under the image's name. This decouples the build artifact from container execution so any number of containers can be stamped out from the same image."
Step 2: The run Stage:
"Next, when run is invoked, the runtime creates dedicated container directories, mounts the OverlayFS union root, and bind-mounts essential /dev devices into the container tree. It then starts the container supervisor in new UTS, PID, Mount, and Network namespaces. The supervisor mounts private /proc, /sys, and /tmp filesystems, changes directory into the container root, and replaces itself with the container's main entrypoint process."
Step 3: The exec Stage (Attaching to Running Containers):
"Then, to implement exec—running a new command inside an already-active container without restarting it—I had to solve a process discovery problem. In our architecture, the unshare supervisor process runs in host namespaces, while its immediate child runs as PID 1 inside the container's PID namespace. So, I searched the process table for the supervisor, queried the process tree to find its child PID 1, and attached to that child using nsenter. This seamlessly joins the container's UTS, PID, Mount, and Network namespaces while setting the root and working directories to match."
Step 4: The stop Stage (Strict Reverse Teardown):
"Finally, when stopping a container, cleaning up resources in the correct order is critical. I sent termination signals to the container process tree, deleted the host-side virtual network endpoint, and unmounted filesystems in strict reverse hierarchical order: /proc, /sys, /dev, and /tmp before unmounting the root OverlayFS union. Unmounting leaf filesystems first is required because attempting to unmount the root overlay while sub-mounts are active will fail with an EBUSY error in the kernel. Once the last container exits, the runtime flushes global packet forwarding rules and resets state."
| Step | What Was Done | How It Works | Why This Mechanism / Order |
|---|---|---|---|
| 1. Build | Compiled image layer stack | Persisted ordered layer path list in .images/<name>/layers. |
Decouples image construction from container execution. |
| 2. Run | Initialized rootfs and supervisor | Mounted OverlayFS union, bind-mounted /dev, mounted procfs/sysfs. |
Prepares complete isolated runtime sandbox before launching entrypoint. |
| 3. Exec | Attached to active container | Located supervisor $\rightarrow$ found child PID 1 $\rightarrow$ attached via nsenter. |
Must target child PID (and not supervisor) to enter the container's PID namespace. |
| 4. Stop | Cleaned up process & mounts | Sent kill signal, deleted veth, unmounted leaf mounts before root overlay. | Reverse unmount prevents kernel EBUSY errors when detaching active sub-mounts. |
exec?Answer: The
unsharesupervisor process itself runs in the host PID namespace—its sole job is to supervise the container and hold its namespace descriptors. The child process created by the supervisor is the actual process running inside the container's private PID namespace. Attaching to the supervisor would attach the user to the host PID space instead of the container.
Answer: The supervisor process was configured with automatic child-killing (
kill-child). When the container's primary process exits, the supervisor immediately receives the exit signal and instructs the kernel to sendSIGKILLto all remaining child processes in that PID namespace, preventing dangling orphan processes on the host.
Answer: Mount points in Linux form a tree hierarchy. Leaf filesystems like
/procand/sysare mounted on top of directory inodes inside the rootmerged/overlay. Attempting to unmountmerged/first fails withEBUSYbecause active filesystem mounts are pinned to its child directories.