Target Duration: 2–4 minutes (~300–450 spoken words)
Focus: Pointwise verbal delivery covering image build mechanics, SHA-256 Merkle layer caching, ephemeral build sandboxes, and OverlayFS CoW / whiteouts.
Opening & Scope:
"For container filesystems and image building, I implemented a Docker-like layered storage engine using OverlayFS union mounts and content-addressable caching."
Step 1: Base Image Initialization (FROM):
"First, the build pipeline starts by bootstrapping a minimal Linux base root filesystem. I cached this base rootfs on disk as an immutable root layer so that multiple images can share the same OS foundation without re-downloading distribution packages."
Step 2: Content-Addressable Layer Hashing (Merkle Chain):
"Next, for every instruction in the build file—such as RUN or COPY—I computed a deterministic SHA-256 cryptographic layer ID. For a RUN instruction, I hashed the parent layer's hash together with the command string; for COPY, I hashed the parent hash together with the checksums of all source files. By chaining the parent's hash into each child layer like a Merkle DAG, any change to an early layer automatically alters the IDs of all downstream layers, giving us automatic transitive cache invalidation without needing complex dependency tracking."
Step 3: Cache Evaluation & Ephemeral Build Sandboxes:
"Then, before executing an instruction, the builder checks if that layer hash already exists in cache. On a cache hit, it skips execution completely. On a cache miss, it mounts a temporary OverlayFS where lowerdir is the parent layer stack and upperdir is a fresh diff directory. It launches a temporary container sandbox to run the command inside this merged view. When the command finishes and the overlay is unmounted, the upperdir contains strictly the new and modified files produced by that step, which are saved as the new layer."
Step 4: OverlayFS Runtime Mechanics (CoW & Whiteouts):
"Finally, at runtime, the container mounts all read-only image layers as lowerdir, a per-container writable directory as upperdir, and an internal scratchpad as workdir. When a container reads a file, OverlayFS searches top-down; when it modifies a file from an image layer, the kernel performs Copy-on-Write by copying the file up into upperdir before modifying it. When a container deletes an existing image file, OverlayFS creates a special 0/0 character device—known as a whiteout—in upperdir to mask the file without touching the underlying immutable layers."
| Step | What Was Done | How It Works | Why This Mechanism / Order |
|---|---|---|---|
| 1. Base Layer | Bootstrapped base rootfs | Downloaded minimal distro once and marked as immutable base. | Provides shared foundational OS files across all images without duplication. |
| 2. Layer Hashing | Computed chained SHA-256 IDs | hash = SHA256(type + parent_hash + content_hash). |
Merkle chaining guarantees automatic transitive cache invalidation. |
| 3. Ephemeral Sandbox | Ran cache-miss steps in overlay | Mounted parent layers as lower, ran command in sandbox, captured upper diff. |
Captures strictly the incremental filesystem delta created by that command. |
| 4. Runtime CoW | Applied Copy-on-Write at runtime | Copies lower files to upperdir on edit; uses 0/0 whiteout devices for deletes. |
Allows multiple containers to share the same image layers safely in parallel. |
workdir required in OverlayFS? Why can't OverlayFS write directly to upperdir?Answer: POSIX file operations like file creation, renaming, and truncating must appear atomic. If a crash or power cut occurs mid-operation, writing directly to
upperdircould leave corrupted files in the container root. OverlayFS prepares the metadata and files inworkdirfirst, then issues an atomic kernel rename intoupperdir.
Answer: For regular files, OverlayFS places a
0/0character device (whiteout) inupperdir. For directories, OverlayFS creates an opaque directory by setting the extended attributetrusted.overlay.opaque=yon the directory inupperdir, instructing the kernel to stop searching lower layers.
COPY?Answer: File modification timestamps (mtime) change during git checkouts, cloning, or file transfers without any actual code change. Hashing the concatenated SHA-256 digests of all source files ensures cache hits are purely content-deterministic across different machines and builds.
Answer:
- What it is: A Merkle DAG (Directed Acyclic Graph) is a graph data structure where every node is uniquely identified by a cryptographic hash of its own contents plus the cryptographic hashes of its parent nodes. It is acyclic (no circular references) and directed (child nodes explicitly point back to their parents).
- How it is used in container layer caching:
1. Content-Addressability: Each container layer's identity is derived asSHA256(instruction_type + parent_layer_hash + content_or_command_hash).
2. Transitive Invalidation: Because child node IDs embed their parent's hash, changing an intermediate layer $K$ instantly alters its hash. Consequently, all descendant layers $K+1 \dots N$ have different parent inputs, causing their hashes to change and automatically missing the cache without requiring a separate dependency tracking database.
3. Deduplication: Multiple different images that share the same base commands (e.g.FROM debianfollowed byRUN apt update) compute identical hashes for those shared prefixes and reuse the exact same layer directories on disk.