Topic 03: Guest-to-Host Virtual Address (GVA → HVA) Translation


🎙️ Deep-Dive Spoken Script

Interviewer: “How did you implement address translation between the guest virtual address and host virtual address?”

Address translation is an important concept in virtualization because the guest and host operate in completely different virtual address spaces.

A pointer inside the guest is a Guest Virtual Address (GVA). If the guest passes a pointer to a buffer or string to the hypervisor, the hypervisor cannot directly dereference that pointer because it belongs to the guest's address space.

When I registered guest memory using KVM_SET_USER_MEMORY_REGION, I mapped guest physical address 0 to the beginning of the userspace memory buffer allocated using mmap.

For the guest configuration used in this buffer-sharing path, the relevant guest virtual addresses map directly to guest physical addresses. Therefore, the hypervisor can obtain the host virtual address by adding the guest address to the base of its allocated memory:

$$\text{HVA} = \text{vm}\to\text{mem} + \text{GVA}$$

For example, when the guest wants the hypervisor to read a string or array:

  1. The guest passes the buffer address through an I/O port instruction (e.g. outl(0xE2, (uint32_t)str_ptr)).
  2. The hypervisor receives the resulting KVM_EXIT_IO.
  3. It reads the guest address from kvm_run->io.data_offset.
  4. It adds that address to vm->mem to obtain the corresponding host pointer:
   char *host_str = (char *)vm->mem + guest_addr;
  1. The hypervisor can then safely access and print the guest buffer through its own address space.

So, for this setup, the translation path is:

$$\mathbf{GVA} \longrightarrow \mathbf{GPA} \longrightarrow \mathbf{HVA}$$

with the GVA-to-GPA step being a direct mapping in our real-mode guest configuration, followed by the GPA-to-HVA mapping provided by my userspace memory layout.


🔬 In-Depth Q&A Database

Q10: Walk through the complete four-tier address translation model in virtualization.

  1. Guest Virtual Address (GVA): Virtual address generated by applications inside the guest OS.
  2. Guest Physical Address (GPA): The physical address perceived by the guest OS kernel. Translated from GVA via Guest Page Tables.
  3. Host Virtual Address (HVA): Userspace virtual address in the host hypervisor process (e.g. vm->mem).
  4. Host Physical Address (HPA): Actual physical RAM address on the physical hardware motherboard.

In hardware-assisted virtualization with EPT/NPT, the CPU's memory management unit translates GPA to HPA in hardware across two-dimensional page table walks.

Q11: Why is GVA to HVA translation trivial in our project (HVA = vm->mem + GVA)?

Because:

  1. Our guest runs in x86 real mode without paging enabled. In real mode, segmentation uses a base of 0, meaning $GVA \equiv GPA$.
  2. In our hypervisor, KVM_SET_USER_MEMORY_REGION maps GPA 0x00000000 directly to userspace_addr = (uint64_t)vm->mem.
  3. Therefore, any GPA is simply a linear offset from the start of the vm->mem buffer:

$$\text{HVA} = \text{vm}\to\text{mem} + \text{GPA} = \text{vm}\to\text{mem} + \text{GVA}$$