提交 · a205bc19f0d203f31a15bd2f9464ef05f918b77e · openeuler / raspberrypi-kernel

10 9月, 2009 40 次提交

KVM: MMU: Fix MMU_DEBUG compile breakage · a205bc19

由 Joerg Roedel 提交于 7月 09, 2009

Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

a205bc19

KVM: add ioeventfd support · d34e6b17

由 Gregory Haskins 提交于 7月 07, 2009

ioeventfd is a mechanism to register PIO/MMIO regions to trigger an eventfd
signal when written to by a guest.  Host userspace can register any
arbitrary IO address with a corresponding eventfd and then pass the eventfd
to a specific end-point of interest for handling.

Normal IO requires a blocking round-trip since the operation may cause
side-effects in the emulated model or may return data to the caller.
Therefore, an IO in KVM traps from the guest to the host, causes a VMX/SVM
"heavy-weight" exit back to userspace, and is ultimately serviced by qemu's
device model synchronously before returning control back to the vcpu.

However, there is a subclass of IO which acts purely as a trigger for
other IO (such as to kick off an out-of-band DMA request, etc).  For these
patterns, the synchronous call is particularly expensive since we really
only want to simply get our notification transmitted asychronously and
return as quickly as possible.  All the sychronous infrastructure to ensure
proper data-dependencies are met in the normal IO case are just unecessary
overhead for signalling.  This adds additional computational load on the
system, as well as latency to the signalling path.

Therefore, we provide a mechanism for registration of an in-kernel trigger
point that allows the VCPU to only require a very brief, lightweight
exit just long enough to signal an eventfd.  This also means that any
clients compatible with the eventfd interface (which includes userspace
and kernelspace equally well) can now register to be notified. The end
result should be a more flexible and higher performance notification API
for the backend KVM hypervisor and perhipheral components.

To test this theory, we built a test-harness called "doorbell".  This
module has a function called "doorbell_ring()" which simply increments a
counter for each time the doorbell is signaled.  It supports signalling
from either an eventfd, or an ioctl().

We then wired up two paths to the doorbell: One via QEMU via a registered
io region and through the doorbell ioctl().  The other is direct via
ioeventfd.

You can download this test harness here:

ftp://ftp.novell.com/dev/ghaskins/doorbell.tar.bz2

The measured results are as follows:

qemu-mmio:       110000 iops, 9.09us rtt
ioeventfd-mmio: 200100 iops, 5.00us rtt
ioeventfd-pio:  367300 iops, 2.72us rtt

I didn't measure qemu-pio, because I have to figure out how to register a
PIO region with qemu's device model, and I got lazy.  However, for now we
can extrapolate based on the data from the NULLIO runs of +2.56us for MMIO,
and -350ns for HC, we get:

qemu-pio:      153139 iops, 6.53us rtt
ioeventfd-hc: 412585 iops, 2.37us rtt

these are just for fun, for now, until I can gather more data.

Here is a graph for your convenience:

http://developer.novell.com/wiki/images/7/76/Iofd-chart.png

The conclusion to draw is that we save about 4us by skipping the userspace
hop.

--------------------
Signed-off-by: NGregory Haskins <ghaskins@novell.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d34e6b17

KVM: make io_bus interface more robust · 090b7aff

由 Gregory Haskins 提交于 7月 07, 2009

Today kvm_io_bus_regsiter_dev() returns void and will internally BUG_ON
if it fails.  We want to create dynamic MMIO/PIO entries driven from
userspace later in the series, so we need to enhance the code to be more
robust with the following changes:

   1) Add a return value to the registration function
   2) Fix up all the callsites to check the return code, handle any
      failures, and percolate the error up to the caller.
   3) Add an unregister function that collapses holes in the array
Signed-off-by: NGregory Haskins <ghaskins@novell.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

090b7aff

KVM: add module parameters documentation · fef07aae

由 Andre Przywara 提交于 7月 10, 2009

Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

fef07aae

KVM: PIT support for HPET legacy mode · e9f42757

由 Beth Kon 提交于 7月 07, 2009

When kvm is in hpet_legacy_mode, the hpet is providing the timer
interrupt and the pit should not be. So in legacy mode, the pit timer
is destroyed, but the *state* of the pit is maintained. So if kvm or
the guest tries to modify the state of the pit, this modification is
accepted, *except* that the timer isn't actually started. When we exit
hpet_legacy_mode, the current state of the pit (which is up to date
since we've been accepting modifications) is used to restart the pit
timer.

The saved_mode code in kvm_pit_load_count temporarily changes mode to
0xff in order to destroy the timer, but then restores the actual
value, again maintaining "current" state of the pit for possible later
reenablement.

[avi: add some reserved storage in the ioctl; make SET_PIT2 IOW]
[marcelo: fix memory corruption due to reserved storage]
Signed-off-by: NBeth Kon <eak@us.ibm.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

e9f42757

KVM: Always report x2apic as supported feature · 0d1de2d9

由 Gleb Natapov 提交于 7月 12, 2009

We emulate x2apic in software, so host support is not required.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

0d1de2d9

KVM: No need to kick cpu if not in a guest mode · c7f0f24b

由 Gleb Natapov 提交于 7月 07, 2009

This will save a couple of IPIs.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Acked-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

c7f0f24b

KVM: Add trace points in irqchip code · 1000ff8d

由 Gleb Natapov 提交于 7月 07, 2009

Add tracepoint in msi/ioapic/pic set_irq() functions,
in IPI sending and in the point where IRQ is placed into
apic's IRR.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

1000ff8d

KVM: ignore msi request if !level · 07fb8bb2

由 Michael S. Tsirkin 提交于 7月 05, 2009

Irqfd sets level for interrupt to 1 and then to 0.
For MSI, check level so that a single message is sent.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

07fb8bb2

KVM: fix MMIO_CONF_BASE MSR access · f7c6d140

由 Andre Przywara 提交于 7月 02, 2009

Some Windows versions check whether the BIOS has setup MMI/O for
config space accesses on AMD Fam10h CPUs, we say "no" by returning 0 on
reads and only allow disabling of MMI/O CfgSpace setup by igoring "0" writes.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

f7c6d140

A
KVM: Trace shadow page lifecycle · f691fe1d
由 Avi Kivity 提交于 7月 06, 2009
```
Create, sync, unsync, zap.
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
f691fe1d

KVM: Document basic API · 9c1b96e3

由 Avi Kivity 提交于 6月 09, 2009

Document the basic API corresponding to the 2.6.22 release.
Signed-off-by: NAvi Kivity <avi@redhat.com>

9c1b96e3

A
KVM: MMU: Trace guest pagetable walker · 07420171
由 Avi Kivity 提交于 7月 06, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
07420171

Revert "KVM: x86: check for cr3 validity in ioctl_set_sregs" · dc7e795e

由 Jan Kiszka 提交于 7月 01, 2009

This reverts commit 6c20e1442bb1c62914bb85b7f4a38973d2a423ba.

To my understanding, it became obsolete with the advent of the more
robust check in mmu_alloc_roots (89da4ff17f). Moreover, it prevents
the conceptually safe pattern

 1. set sregs
 2. register mem-slots
 3. run vcpu

by setting a sticky triple fault during step 1.
Signed-off-by: NJan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

dc7e795e

KVM: handle AMD microcode MSR · 6098ca93

由 Andre Przywara 提交于 7月 03, 2009

Windows 7 tries to update the CPU's microcode on some processors,
so we ignore the MSR write here. The patchlevel register is already handled
(returning 0), because the MSR number is the same as Intel's.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

6098ca93

KVM: Fix apic_mmio_write return for unaligned write · 756975bb

由 Sheng Yang 提交于 7月 06, 2009

Some in-famous OS do unaligned writing for APIC MMIO, and the return value
has been missed in recent change, then the OS hangs.
Signed-off-by: NSheng Yang <sheng@linux.intel.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

756975bb

KVM: Use temporary variable to shorten lines. · 70f93dae

由 Gleb Natapov 提交于 7月 05, 2009

Cosmetic only. No logic is changed by this patch.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

70f93dae

KVM: x2apic interface to lapic · 0105d1a5

由 Gleb Natapov 提交于 7月 05, 2009

This patch implements MSR interface to local apic as defines by x2apic
Intel specification.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

0105d1a5

KVM: Add Directed EOI support to APIC emulation · fc61b800

由 Gleb Natapov 提交于 7月 05, 2009

Directed EOI is specified by x2APIC, but is available even when lapic is
in xAPIC mode.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

fc61b800

A
KVM: Trace apic registers using their symbolic names · cb247721
由 Avi Kivity 提交于 7月 01, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
cb247721
A
KVM: Trace mmio · aec51dc4
由 Avi Kivity 提交于 7月 01, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
aec51dc4

KVM: Ignore PCI ECS I/O enablement · c323c0e5

由 Andre Przywara 提交于 6月 24, 2009

Linux guests will try to enable access to the extended PCI config space
via the I/O ports 0xCF8/0xCFC on AMD Fam10h CPU. Since we (currently?)
don't use ECS, simply ignore write and read attempts.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

c323c0e5

A
KVM: Trace irq level and source id · ae8c1c40
由 Avi Kivity 提交于 7月 01, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
ae8c1c40

KVM: fix lock imbalance · 27c4ba60

由 Jiri Slaby 提交于 6月 29, 2009

There is a missing unlock on one fail path in ioapic_mmio_write,
fix that.
Signed-off-by: NJiri Slaby <jirislaby@gmail.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

27c4ba60

KVM: document lock nesting rule · 22fc0294

由 Michael S. Tsirkin 提交于 6月 29, 2009

Document kvm->lock nesting within kvm->slots_lock
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

22fc0294

KVM: remove in_range from io devices · bda9020e

由 Michael S. Tsirkin 提交于 6月 29, 2009

This changes bus accesses to use high-level kvm_io_bus_read/kvm_io_bus_write
functions. in_range now becomes unused so it is removed from device ops in
favor of read/write callbacks performing range checks internally.

This allows aliasing (mostly for in-kernel virtio), as well as better error
handling by making it possible to pass errors up to userspace.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

bda9020e

KVM: convert bus to slots_lock · 6c474694

由 Michael S. Tsirkin 提交于 6月 29, 2009

Use slots_lock to protect device list on the bus.  slots_lock is already
taken for read everywhere, so we only need to take it for write when
registering devices.  This is in preparation to removing in_range and
kvm->lock around it.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

6c474694

KVM: switch pit creation to slots_lock · 108b5669

由 Michael S. Tsirkin 提交于 6月 29, 2009

switch pit creation to slots_lock. slots_lock is already taken for read
everywhere, so we only need to take it for write when creating pit.
This is in preparation to removing in_range and kvm->lock around it.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

108b5669

KVM: switch coalesced mmio changes to slots_lock · d5c2dcc3

由 Michael S. Tsirkin 提交于 6月 29, 2009

switch coalesced mmio slots_lock.  slots_lock is already taken for read
everywhere, so we only need to take it for write when changing zones.
This is in preparation to removing in_range and kvm->lock around it.

[avi: fix build]
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d5c2dcc3

KVM: document locking for kvm_io_device_ops · 69fa2d78

由 Michael S. Tsirkin 提交于 6月 29, 2009

slots_lock is taken everywhere when device ops are called.
Document this as we will use this to rework locking for io.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

69fa2d78

KVM: use vcpu_id instead of bsp_vcpu pointer in kvm_vcpu_is_bsp · d3efc8ef

由 Marcelo Tosatti 提交于 6月 17, 2009

Change kvm_vcpu_is_bsp to use vcpu_id instead of bsp_vcpu pointer, which
is only initialized at the end of kvm_vm_ioctl_create_vcpu.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d3efc8ef

KVM: remove old KVMTRACE support code · 2023a29c

由 Marcelo Tosatti 提交于 6月 18, 2009

Return EOPNOTSUPP for KVM_TRACE_ENABLE/PAUSE/DISABLE ioctls.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

2023a29c

KVM: powerpc: convert marker probes to event trace · 46f43c6e

由 Marcelo Tosatti 提交于 6月 18, 2009

[avi: make it build]
[avi: fold trace-arch.h into trace.h]

CC: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

46f43c6e

KVM: introduce module parameter for ignoring unknown MSRs accesses · ed85c068

由 Andre Przywara 提交于 6月 25, 2009

KVM will inject a #GP into the guest if that tries to access unhandled
MSRs. This will crash many guests. Although it would be the correct
way to actually handle these MSRs, we introduce a runtime switchable
module param called "ignore_msrs" (defaults to 0). If this is Y, unknown
MSR reads will return 0, while MSR writes are simply dropped. In both cases
we print a message to dmesg to inform the user about that.

You can change the behaviour at any time by saying:

 # echo 1 > /sys/modules/kvm/parameters/ignore_msrs
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

ed85c068

KVM: ignore reads from AMDs C1E enabled MSR · 1fdbd48c

由 Andre Przywara 提交于 6月 24, 2009

If the Linux kernel detects an C1E capable AMD processor (K8 RevF and
higher), it will access a certain MSR on every attempt to go to halt.
Explicitly handle this read and return 0 to let KVM run a Linux guest
with the native AMD host CPU propagated to the guest.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

1fdbd48c

KVM: ignore AMDs HWCR register access to set the FFDIS bit · 8f1589d9

由 Andre Przywara 提交于 6月 24, 2009

Linux tries to disable the flush filter on all AMD K8 CPUs. Since KVM
does not handle the needed MSR, the injected #GP will panic the Linux
kernel. Ignore setting of the HWCR.FFDIS bit in this MSR to let Linux
boot with an AMD K8 family guest CPU.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

8f1589d9

KVM: x86: missing locking in PIT/IRQCHIP/SET_BSP_CPU ioctl paths · 894a9c55

由 Marcelo Tosatti 提交于 6月 23, 2009

Correct missing locking in a few places in x86's vm_ioctl handling path.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

894a9c55

KVM: Prepare memslot data structures for multiple hugepage sizes · ec04b260

由 Joerg Roedel 提交于 6月 19, 2009

[avi: fix build on non-x86]
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

ec04b260

hugetlbfs: export vma_kernel_pagsize to modules · f340ca0f

由 Joerg Roedel 提交于 6月 19, 2009

This function is required by KVM.
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

f340ca0f

KVM: s390: Fix memslot initialization for userspace_addr != 0 · 3eea8437

由 Christian Borntraeger 提交于 6月 23, 2009

Since
commit 854b5338196b1175706e99d63be43a4f8d8ab607
Author: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
    KVM: s390: streamline memslot handling

s390 uses the values of the memslot instead of doing everything in the arch
ioctl handler of the KVM_SET_USER_MEMORY_REGION. Unfortunately we missed to
set the userspace_addr of our memslot due to our s390 ifdef in
__kvm_set_memory_region.
Old s390 userspace launchers did not notice, since they started the guest at
userspace address 0.
Because of CONFIG_DEFAULT_MMAP_MIN_ADDR we now put the guest at 1M userspace,
which does not work. This patch makes sure that new.userspace_addr is set
on s390.
This fix should go in quickly. Nevertheless, looking at the code we should
clean up that ifdef in the long term. Any kernel janitors?
Signed-off-by: NChristian Borntraeger <borntraeger@de.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

3eea8437