3.5–5 Minute Detailed Architecture: Bare-Metal KVM Hypervisor

Target Duration: ~3.5 to 5 minutes (~650–850 spoken words)
Goal: Deliver an end-to-end architectural narrative covering initialization, execution loops, trap-based communication, address translation, and multi-VM coordination.


🎙️ Spoken Script & Architectural Walkthrough

To explain the complete architecture of the hypervisor, I divide the project into four main stages: VM and memory initialization, vCPU configuration and execution, VM exits and guest-host communication, and dual-VM scheduling and synchronization.


Stage 1: VM & Memory Initialization

Everything begins in userspace by opening /dev/kvm and using the Linux KVM API to create and manage virtual machines.

  1. I open /dev/kvm to obtain the system-level file descriptor and verify the API version using KVM_GET_API_VERSION.
  2. I create a virtual machine instance via ioctl(dev_fd, KVM_CREATE_VM, 0), which yields vm_fd.
  3. For guest physical RAM, I allocate 2 MB using mmap() with MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE.
  4. I register this memory with KVM via ioctl(vm_fd, KVM_SET_USER_MEMORY_REGION, &region), establishing that Guest Physical Address (GPA) 0x0 corresponds to (uint64_t)vm->mem.
  5. I create the vCPU via ioctl(vm_fd, KVM_CREATE_VCPU, 0) and map the shared struct kvm_run structure using mmap() with MAP_SHARED.

(Hook: Deep dive into Topic 01: Core KVM API & Memory Layout)


Stage 2: vCPU Configuration & Guest Execution

Before running guest code, the hypervisor configures the initial CPU register state:

  1. I read and configure special registers via KVM_GET_SREGS and KVM_SET_SREGS (setting code segment base to 0, data segment base to 0).
  2. I set general registers via KVM_SET_REGS: setting the instruction pointer rip = 0x0000 and flags rflags = 0x0002 (x86 reserved bit).
  3. I load the compiled guest machine code payload directly into vm->mem.
  4. To execute the guest, the hypervisor enters an execution loop calling ioctl(vcpu_fd, KVM_RUN, 0).
  5. Under hardware virtualization (Intel VT-x / AMD-V), KVM executes the VMLAUNCH/VMRESUME instruction, transitioning the CPU into non-root guest execution mode. The guest executes natively at silicon speed until a privileged instruction forces a VM exit.

(Hook: Deep dive into Topic 02: VM Exits & Hypercall Traps)


Stage 3: VM Exits, I/O Traps & Address Translation

When a VM exit occurs, control returns from the CPU to KVM, and KVM_RUN returns to userspace. The hypervisor checks vcpu->kvm_run->exit_reason:

For buffer operations (like string printing), the guest passes a Guest Virtual Address (GVA). Because real mode maps addresses 1:1 (GVA == GPA), and guest physical RAM starts at vm->mem, the hypervisor computes:

$$\text{HVA} = \text{vm}\to\text{mem} + \text{GVA}$$

The hypervisor safely dereferences the host virtual address to access the guest's data buffer.

(Hook: Deep dive into Topic 03: GVA→HVA Address Translation)


Stage 4: Dual-VM Producer/Consumer Scheduling

In the second phase, I expanded the architecture to run two concurrent VMs: VM1 as a Producer and VM2 as a Consumer.

  1. Isolated Memory Spaces: Each VM has its own independent 2 MB RAM allocation and vCPU context. Neither VM can access the other's address space directly.
  2. Registration Phase: During boot, both VMs notify the hypervisor of their local 20-element circular buffer and synchronization pointers via I/O port exits. The hypervisor translates these GVAs to HVAs and maintains shadow pointers.
  3. Producer-Consumer Synchronization:
  1. Userspace vCPU Scheduling: The hypervisor implements a time-sliced round-robin scheduler driven by a trace file. For each step, it calls KVM_RUN on the designated vCPU, processes the exit, and moves to the next scheduled VM.

(Hook: Deep dive into Topic 04: Dual-VM Circular Buffer and Topic 05: Userspace vCPU Scheduling)


🎯 Architectural Summary

Dimension Implementation in Hypervisor
Virtualization Model Type-2 Hypervisor leveraging Linux /dev/kvm hardware-assisted virtualization
Execution Mode x86 Real Mode with identity paging (GVA == GPA)
Memory Allocation 2 MB per VM via anonymous mmap with MAP_NORESERVE
Guest-Host Hypercalls Synchronous x86 in/out port traps handled via KVM_EXIT_IO
Address Translation Direct offset translation: HVA = vm->mem + GVA
Concurrency & Sync Shadow circular ring buffer mediated entirely by userspace hypervisor