Interviewer: “How did you implement scheduling between the two VMs in userspace?”
In the second part of the project, I also implemented scheduling between the two VMs in userspace.
The basic idea was that I had a scheduling trace containing a sequence of VM IDs, where 1 represents the Producer VM and 2 represents the Consumer VM. The hypervisor reads this sequence and uses it to decide which VM should run next.
For each step, the hypervisor selects the corresponding VM and vCPU and calls KVM_RUN on that vCPU:
int next_vm = sched_trace[step];
if (next_vm == 1) {
ioctl(vm1->vcpu_fd, KVM_RUN, 0);
handle_exit(vm1);
} else {
ioctl(vm2->vcpu_fd, KVM_RUN, 0);
handle_exit(vm2);
}
The guest then starts or continues executing until it produces an exit, such as an I/O exit.
When KVM_RUN returns, the hypervisor handles that exit. For example, if the Producer has generated new data, the hypervisor synchronizes the buffer with the Consumer. It then moves to the next VM specified by the scheduling trace.
Because KVM maintains the register state of each vCPU, I don't have to manually save and restore all the guest registers when switching between the two VMs. I simply run the appropriate vCPU through its vCPU file descriptor, and KVM resumes that vCPU with its saved state.
I also used a delay between scheduling steps to make the execution sequence observable during the experiment.
So overall, the userspace hypervisor is doing two things together: it decides which vCPU runs next, and it acts as the broker that handles VM exits and synchronizes data between the two VMs.
In commercial hypervisors like QEMU/KVM, each vCPU typically runs inside an independent host POSIX thread (pthread_create), leaving time-slicing to the host Linux CFS scheduler.
In our specialized hypervisor, scheduling is implemented deterministically in userspace using a single thread. The hypervisor reads an interleaved scheduling trace (e.g. 1 1 2 1 2 2), issuing KVM_RUN calls sequentially. This allows precise observation of race conditions, buffer full/empty states, and synchronization timing.
/dev/kvm, allocating memory via mmap, trapping I/O port exits, and orchestrating vCPU time-slices directly.They explore compute abstraction and isolation at every layer of the modern systems hierarchy:
| Mechanism / Subsystem | Exact Function / Code Pattern | |||
|---|---|---|---|---|
| Open KVM Subsystem | dev_fd = open("/dev/kvm", O_RDWR);ioctl(dev_fd, KVM_GET_API_VERSION, 0); |
|||
| Create VM Instance | vm_fd = ioctl(dev_fd, KVM_CREATE_VM, 0); |
|||
| Allocate Guest Memory | `vm->mem = mmap(NULL, 210241024, PROT_READ\ | PROT_WRITE, MAP_PRIVATE\ | MAP_ANONYMOUS\ | MAP_NORESERVE, -1, 0);` |
| Register Guest RAM | struct kvm_userspace_memory_region region = { .slot = 0, .guest_phys_addr = 0, .memory_size = 2MB, .userspace_addr = (uint64_t)vm->mem };ioctl(vm_fd, KVM_SET_USER_MEMORY_REGION, ®ion); |
|||
| Create vCPU | vcpu_fd = ioctl(vm_fd, KVM_CREATE_VCPU, 0); |
|||
Map Shared kvm_run |
mmap_size = ioctl(dev_fd, KVM_GET_VCPU_MMAP_SIZE, 0);`vcpu->kvm_run = mmap(NULL, mmap_size, PROT_READ\ |
PROT_WRITE, MAP_SHARED, vcpu_fd, 0);` | ||
| Configure Initial RIP & Flags | struct kvm_regs regs = { .rip = 0, .rflags = 2 };ioctl(vcpu_fd, KVM_SET_REGS, ®s); |
|||
| Run vCPU Loop | ioctl(vcpu_fd, KVM_RUN, 0);if (vcpu->kvm_run->exit_reason == KVM_EXIT_IO) { ... } |
|||
| Guest I/O Out Assembly | asm("out %0, %1" : : "a"(value), "Nd"(port) : "memory"); |
|||
| GVA to HVA Formula | uint32_t *hva = (uint32_t *)((char *)vm->mem + gva); |
|||
| Inspect I/O Exit Payload | uint32_t val = *((uint32_t *)((char *)vcpu->kvm_run + vcpu->kvm_run->io.data_offset)); |