brintos

brintos / linux-shallow public Read only

0
0
Text · 14.2 KiB · c15d6fe Raw
261 lines · plain
1.. SPDX-License-Identifier: GPL-2.02 3Confidential Computing VMs4==========================5Hyper-V can create and run Linux guests that are Confidential Computing6(CoCo) VMs. Such VMs cooperate with the physical processor to better protect7the confidentiality and integrity of data in the VM's memory, even in the8face of a hypervisor/VMM that has been compromised and may behave maliciously.9CoCo VMs on Hyper-V share the generic CoCo VM threat model and security10objectives described in Documentation/security/snp-tdx-threat-model.rst. Note11that Hyper-V specific code in Linux refers to CoCo VMs as "isolated VMs" or12"isolation VMs".13 14A Linux CoCo VM on Hyper-V requires the cooperation and interaction of the15following:16 17* Physical hardware with a processor that supports CoCo VMs18 19* The hardware runs a version of Windows/Hyper-V with support for CoCo VMs20 21* The VM runs a version of Linux that supports being a CoCo VM22 23The physical hardware requirements are as follows:24 25* AMD processor with SEV-SNP. Hyper-V does not run guest VMs with AMD SME,26  SEV, or SEV-ES encryption, and such encryption is not sufficient for a CoCo27  VM on Hyper-V.28 29* Intel processor with TDX30 31To create a CoCo VM, the "Isolated VM" attribute must be specified to Hyper-V32when the VM is created. A VM cannot be changed from a CoCo VM to a normal VM,33or vice versa, after it is created.34 35Operational Modes36-----------------37Hyper-V CoCo VMs can run in two modes. The mode is selected when the VM is38created and cannot be changed during the life of the VM.39 40* Fully-enlightened mode. In this mode, the guest operating system is41  enlightened to understand and manage all aspects of running as a CoCo VM.42 43* Paravisor mode. In this mode, a paravisor layer between the guest and the44  host provides some operations needed to run as a CoCo VM. The guest operating45  system can have fewer CoCo enlightenments than is required in the46  fully-enlightened case.47 48Conceptually, fully-enlightened mode and paravisor mode may be treated as49points on a spectrum spanning the degree of guest enlightenment needed to run50as a CoCo VM. Fully-enlightened mode is one end of the spectrum. A full51implementation of paravisor mode is the other end of the spectrum, where all52aspects of running as a CoCo VM are handled by the paravisor, and a normal53guest OS with no knowledge of memory encryption or other aspects of CoCo VMs54can run successfully. However, the Hyper-V implementation of paravisor mode55does not go this far, and is somewhere in the middle of the spectrum. Some56aspects of CoCo VMs are handled by the Hyper-V paravisor while the guest OS57must be enlightened for other aspects. Unfortunately, there is no58standardized enumeration of feature/functions that might be provided in the59paravisor, and there is no standardized mechanism for a guest OS to query the60paravisor for the feature/functions it provides. The understanding of what61the paravisor provides is hard-coded in the guest OS.62 63Paravisor mode has similarities to the `Coconut project`_, which aims to provide64a limited paravisor to provide services to the guest such as a virtual TPM.65However, the Hyper-V paravisor generally handles more aspects of CoCo VMs66than is currently envisioned for Coconut, and so is further toward the "no67guest enlightenments required" end of the spectrum.68 69.. _Coconut project: https://github.com/coconut-svsm/svsm70 71In the CoCo VM threat model, the paravisor is in the guest security domain72and must be trusted by the guest OS. By implication, the hypervisor/VMM must73protect itself against a potentially malicious paravisor just like it74protects against a potentially malicious guest.75 76The hardware architectural approach to fully-enlightened vs. paravisor mode77varies depending on the underlying processor.78 79* With AMD SEV-SNP processors, in fully-enlightened mode the guest OS runs in80  VMPL 0 and has full control of the guest context. In paravisor mode, the81  guest OS runs in VMPL 2 and the paravisor runs in VMPL 0. The paravisor82  running in VMPL 0 has privileges that the guest OS in VMPL 2 does not have.83  Certain operations require the guest to invoke the paravisor. Furthermore, in84  paravisor mode the guest OS operates in "virtual Top Of Memory" (vTOM) mode85  as defined by the SEV-SNP architecture. This mode simplifies guest management86  of memory encryption when a paravisor is used.87 88* With Intel TDX processor, in fully-enlightened mode the guest OS runs in an89  L1 VM. In paravisor mode, TD partitioning is used. The paravisor runs in the90  L1 VM, and the guest OS runs in a nested L2 VM.91 92Hyper-V exposes a synthetic MSR to guests that describes the CoCo mode. This93MSR indicates if the underlying processor uses AMD SEV-SNP or Intel TDX, and94whether a paravisor is being used. It is straightforward to build a single95kernel image that can boot and run properly on either architecture, and in96either mode.97 98Paravisor Effects99-----------------100Running in paravisor mode affects the following areas of generic Linux kernel101CoCo VM functionality:102 103* Initial guest memory setup. When a new VM is created in paravisor mode, the104  paravisor runs first and sets up the guest physical memory as encrypted. The105  guest Linux does normal memory initialization, except for explicitly marking106  appropriate ranges as decrypted (shared). In paravisor mode, Linux does not107  perform the early boot memory setup steps that are particularly tricky with108  AMD SEV-SNP in fully-enlightened mode.109 110* #VC/#VE exception handling. In paravisor mode, Hyper-V configures the guest111  CoCo VM to route #VC and #VE exceptions to VMPL 0 and the L1 VM,112  respectively, and not the guest Linux. Consequently, these exception handlers113  do not run in the guest Linux and are not a required enlightenment for a114  Linux guest in paravisor mode.115 116* CPUID flags. Both AMD SEV-SNP and Intel TDX provide a CPUID flag in the117  guest indicating that the VM is operating with the respective hardware118  support. While these CPUID flags are visible in fully-enlightened CoCo VMs,119  the paravisor filters out these flags and the guest Linux does not see them.120  Throughout the Linux kernel, explicitly testing these flags has mostly been121  eliminated in favor of the cc_platform_has() function, with the goal of122  abstracting the differences between SEV-SNP and TDX. But the123  cc_platform_has() abstraction also allows the Hyper-V paravisor configuration124  to selectively enable aspects of CoCo VM functionality even when the CPUID125  flags are not set. The exception is early boot memory setup on SEV-SNP, which126  tests the CPUID SEV-SNP flag. But not having the flag in Hyper-V paravisor127  mode VM achieves the desired effect or not running SEV-SNP specific early128  boot memory setup.129 130* Device emulation. In paravisor mode, the Hyper-V paravisor provides131  emulation of devices such as the IO-APIC and TPM. Because the emulation132  happens in the paravisor in the guest context (instead of the hypervisor/VMM133  context), MMIO accesses to these devices must be encrypted references instead134  of the decrypted references that would be used in a fully-enlightened CoCo135  VM. The __ioremap_caller() function has been enhanced to make a callback to136  check whether a particular address range should be treated as encrypted137  (private). See the "is_private_mmio" callback.138 139* Encrypt/decrypt memory transitions. In a CoCo VM, transitioning guest140  memory between encrypted and decrypted requires coordinating with the141  hypervisor/VMM. This is done via callbacks invoked from142  __set_memory_enc_pgtable(). In fully-enlightened mode, the normal SEV-SNP and143  TDX implementations of these callbacks are used. In paravisor mode, a Hyper-V144  specific set of callbacks is used. These callbacks invoke the paravisor so145  that the paravisor can coordinate the transitions and inform the hypervisor146  as necessary. See hv_vtom_init() where these callback are set up.147 148* Interrupt injection. In fully enlightened mode, a malicious hypervisor149  could inject interrupts into the guest OS at times that violate x86/x64150  architectural rules. For full protection, the guest OS should include151  enlightenments that use the interrupt injection management features provided152  by CoCo-capable processors. In paravisor mode, the paravisor mediates153  interrupt injection into the guest OS, and ensures that the guest OS only154  sees interrupts that are "legal". The paravisor uses the interrupt injection155  management features provided by the CoCo-capable physical processor, thereby156  masking these complexities from the guest OS.157 158Hyper-V Hypercalls159------------------160When in fully-enlightened mode, hypercalls made by the Linux guest are routed161directly to the hypervisor, just as in a non-CoCo VM. But in paravisor mode,162normal hypercalls trap to the paravisor first, which may in turn invoke the163hypervisor. But the paravisor is idiosyncratic in this regard, and a few164hypercalls made by the Linux guest must always be routed directly to the165hypervisor. These hypercall sites test for a paravisor being present, and use166a special invocation sequence. See hv_post_message(), for example.167 168Guest communication with Hyper-V169--------------------------------170Separate from the generic Linux kernel handling of memory encryption in Linux171CoCo VMs, Hyper-V has VMBus and VMBus devices that communicate using memory172shared between the Linux guest and the host. This shared memory must be173marked decrypted to enable communication. Furthermore, since the threat model174includes a compromised and potentially malicious host, the guest must guard175against leaking any unintended data to the host through this shared memory.176 177These Hyper-V and VMBus memory pages are marked as decrypted:178 179* VMBus monitor pages180 181* Synthetic interrupt controller (synic) related pages (unless supplied by182  the paravisor)183 184* Per-cpu hypercall input and output pages (unless running with a paravisor)185 186* VMBus ring buffers. The direct mapping is marked decrypted in187  __vmbus_establish_gpadl(). The secondary mapping created in188  hv_ringbuffer_init() must also include the "decrypted" attribute.189 190When the guest writes data to memory that is shared with the host, it must191ensure that only the intended data is written. Padding or unused fields must192be initialized to zeros before copying into the shared memory so that random193kernel data is not inadvertently given to the host.194 195Similarly, when the guest reads memory that is shared with the host, it must196validate the data before acting on it so that a malicious host cannot induce197the guest to expose unintended data. Doing such validation can be tricky198because the host can modify the shared memory areas even while or after199validation is performed. For messages passed from the host to the guest in a200VMBus ring buffer, the length of the message is validated, and the message is201copied into a temporary (encrypted) buffer for further validation and202processing. The copying adds a small amount of overhead, but is the only way203to protect against a malicious host. See hv_pkt_iter_first().204 205Many drivers for VMBus devices have been "hardened" by adding code to fully206validate messages received over VMBus, instead of assuming that Hyper-V is207acting cooperatively. Such drivers are marked as "allowed_in_isolated" in the208vmbus_devs[] table. Other drivers for VMBus devices that are not needed in a209CoCo VM have not been hardened, and they are not allowed to load in a CoCo210VM. See vmbus_is_valid_offer() where such devices are excluded.211 212Two VMBus devices depend on the Hyper-V host to do DMA data transfers:213storvsc for disk I/O and netvsc for network I/O. storvsc uses the normal214Linux kernel DMA APIs, and so bounce buffering through decrypted swiotlb215memory is done implicitly. netvsc has two modes for data transfers. The first216mode goes through send and receive buffer space that is explicitly allocated217by the netvsc driver, and is used for most smaller packets. These send and218receive buffers are marked decrypted by __vmbus_establish_gpadl(). Because219the netvsc driver explicitly copies packets to/from these buffers, the220equivalent of bounce buffering between encrypted and decrypted memory is221already part of the data path. The second mode uses the normal Linux kernel222DMA APIs, and is bounce buffered through swiotlb memory implicitly like in223storvsc.224 225Finally, the VMBus virtual PCI driver needs special handling in a CoCo VM.226Linux PCI device drivers access PCI config space using standard APIs provided227by the Linux PCI subsystem. On Hyper-V, these functions directly access MMIO228space, and the access traps to Hyper-V for emulation. But in CoCo VMs, memory229encryption prevents Hyper-V from reading the guest instruction stream to230emulate the access. So in a CoCo VM, these functions must make a hypercall231with arguments explicitly describing the access. See232_hv_pcifront_read_config() and _hv_pcifront_write_config() and the233"use_calls" flag indicating to use hypercalls.234 235load_unaligned_zeropad()236------------------------237When transitioning memory between encrypted and decrypted, the caller of238set_memory_encrypted() or set_memory_decrypted() is responsible for ensuring239the memory isn't in use and isn't referenced while the transition is in240progress. The transition has multiple steps, and includes interaction with241the Hyper-V host. The memory is in an inconsistent state until all steps are242complete. A reference while the state is inconsistent could result in an243exception that can't be cleanly fixed up.244 245However, the kernel load_unaligned_zeropad() mechanism may make stray246references that can't be prevented by the caller of set_memory_encrypted() or247set_memory_decrypted(), so there's specific code in the #VC or #VE exception248handler to fixup this case. But a CoCo VM running on Hyper-V may be249configured to run with a paravisor, with the #VC or #VE exception routed to250the paravisor. There's no architectural way to forward the exceptions back to251the guest kernel, and in such a case, the load_unaligned_zeropad() fixup code252in the #VC/#VE handlers doesn't run.253 254To avoid this problem, the Hyper-V specific functions for notifying the255hypervisor of the transition mark pages as "not present" while a transition256is in progress. If load_unaligned_zeropad() causes a stray reference, a257normal page fault is generated instead of #VC or #VE, and the page-fault-258based handlers for load_unaligned_zeropad() fixup the reference. When the259encrypted/decrypted transition is complete, the pages are marked as "present"260again. See hv_vtom_clear_present() and hv_vtom_set_host_visibility().261