提交 · eab4b8aa34fc64e3a91358e1612e6d059396193b · openanolis / cloud-kernel

10 9月, 2009 40 次提交

KVM: VMX: Optimize vmx_get_cpl() · eab4b8aa

由 Avi Kivity 提交于 8月 04, 2009

Instead of calling vmx_get_segment() (which reads a whole bunch of
vmcs fields), read only the cs selector which contains the cpl.
Signed-off-by: NAvi Kivity <avi@redhat.com>

eab4b8aa

KVM: x86: Disallow hypercalls for guest callers in rings > 0 · 07708c4a

由 Jan Kiszka 提交于 8月 03, 2009

So far unprivileged guest callers running in ring 3 can issue, e.g., MMU
hypercalls. Normally, such callers cannot provide any hand-crafted MMU
command structure as it has to be passed by its physical address, but
they can still crash the guest kernel by passing random addresses.

To close the hole, this patch considers hypercalls valid only if issued
from guest ring 0. This may still be relaxed on a per-hypercall base in
the future once required.

Cc: stable@kernel.org
Signed-off-by: NJan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

07708c4a

KVM: MMU: fix bogus alloc_mmu_pages assignment · b90c062c

由 Marcelo Tosatti 提交于 7月 28, 2009

Remove the bogus n_free_mmu_pages assignment from alloc_mmu_pages.

It breaks accounting of mmu pages, since n_free_mmu_pages is modified
but the real number of pages remains the same.

Cc: stable@kernel.org
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

b90c062c

KVM: MMU: make __kvm_mmu_free_some_pages handle empty list · 3b80fffe

由 Izik Eidus 提交于 7月 28, 2009

First check if the list is empty before attempting to look at list
entries.

Cc: stable@kernel.org
Signed-off-by: NIzik Eidus <ieidus@redhat.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

3b80fffe

KVM: remove superfluous NULL pointer check in kvm_inject_pit_timer_irqs() · 95fb4eb6

由 Bartlomiej Zolnierkiewicz 提交于 7月 29, 2009

This takes care of the following entries from Dan's list:

arch/x86/kvm/i8254.c +714 kvm_inject_pit_timer_irqs(6) warning: variable derefenced in initializer 'vcpu'
arch/x86/kvm/i8254.c +714 kvm_inject_pit_timer_irqs(6) warning: variable derefenced before check 'vcpu'
Reported-by: NDan Carpenter <error27@gmail.com>
Cc: corbet@lwn.net
Cc: eteo@redhat.com
Cc: Julia Lawall <julia@diku.dk>
Signed-off-by: NBartlomiej Zolnierkiewicz <bzolnier@gmail.com>
Acked-by: NSheng Yang <sheng@linux.intel.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

95fb4eb6

KVM: report 1GB page support to userspace · 344f414f

由 Joerg Roedel 提交于 7月 27, 2009

If userspace knows that the kernel part supports 1GB pages it can enable
the corresponding cpuid bit so that guests actually use GB pages.
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

344f414f

KVM: MMU: shadow support for 1gb pages · 7e4e4056

由 Joerg Roedel 提交于 7月 27, 2009

This patch adds support for shadow paging to the 1gb page table code in KVM.
With this code the guest can use 1gb pages even if the host does not support
them.

[ Marcelo: fix shadow page collision on pmd level if a guest 1gb page is mapped
           with 4kb ptes on host level ]
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

7e4e4056

KVM: MMU: make page walker aware of mapping levels · e04da980

由 Joerg Roedel 提交于 7月 27, 2009

The page walker may be used with nested paging too when accessing mmio
areas.  Make it support the additional page-level too.

[ Marcelo: fix reserved bit check for 1gb pte ]
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

e04da980

J
KVM: MMU: make direct mapping paths aware of mapping levels · 852e3c19
由 Joerg Roedel 提交于 7月 27, 2009
```
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
852e3c19

KVM: MMU: rename is_largepage_backed to mapping_level · d25797b2

由 Joerg Roedel 提交于 7月 27, 2009

With the new name and the corresponding backend changes this function
can now support multiple hugepage sizes.
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d25797b2

KVM: MMU: make rmap code aware of mapping levels · 44ad9944

由 Joerg Roedel 提交于 7月 27, 2009

This patch removes the largepage parameter from the rmap_add function.
Together with rmap_remove this function now uses the role.level field to
find determine if the page is a huge page.
Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

44ad9944

KVM: limit lapic periodic timer frequency · 1444885a

由 Marcelo Tosatti 提交于 7月 27, 2009

Otherwise its possible to starve the host by programming lapic timer
with a very high frequency.

Cc: stable@kernel.org
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

1444885a

KVM: Align cr8 threshold when userspace changes cr8 · 5f0269f5

由 Mikhail Ershov 提交于 8月 03, 2009

Commit f0a3602c20 ("KVM: Move interrupt injection logic to x86.c") does not
update the cr8 intercept if the lapic is disabled, so when userspace updates
cr8, the cr8 threshold control is not updated and we are left with illegal
control fields.

Fix by explicitly resetting the cr8 threshold.
Signed-off-by: NAvi Kivity <avi@redhat.com>

5f0269f5

KVM: VMX: Avoid to return ENOTSUPP to userland · 7f582ab6

由 Jan Kiszka 提交于 7月 22, 2009

Choose some allowed error values for the cases VMX returned ENOTSUPP so
far as these values could be returned by the KVM_RUN IOCTL.
Signed-off-by: NJan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

7f582ab6

G
KVM: PIT: Unregister ack notifier callback when freeing · 84fde248
由 Gleb Natapov 提交于 7月 16, 2009
```
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
```
84fde248

KVM: VMX: Introduce KVM_SET_IDENTITY_MAP_ADDR ioctl · b927a3ce

由 Sheng Yang 提交于 7月 21, 2009

Now KVM allow guest to modify guest's physical address of EPT's identity mapping page.

(change from v1, discard unnecessary check, change ioctl to accept parameter
address rather than value)
Signed-off-by: NSheng Yang <sheng@linux.intel.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

b927a3ce

KVM: x86: use kvm_get_gdt() and kvm_read_ldt() · b792c344

由 Akinobu Mita 提交于 7月 19, 2009

Use kvm_get_gdt() and kvm_read_ldt() to reduce inline assembly code.

Cc: Avi Kivity <avi@redhat.com>
Cc: kvm@vger.kernel.org
Signed-off-by: NAkinobu Mita <akinobu.mita@gmail.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

b792c344

KVM: x86: use get_desc_base() and get_desc_limit() · 46a359e7

由 Akinobu Mita 提交于 7月 18, 2009

Use get_desc_base() and get_desc_limit() to get the base address and
limit in desc_struct.

Cc: Avi Kivity <avi@redhat.com>
Cc: kvm@vger.kernel.org
Signed-off-by: NAkinobu Mita <akinobu.mita@gmail.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

46a359e7

KVM: MMU: fix missing locking in alloc_mmu_pages · 6a1ac771

由 Marcelo Tosatti 提交于 7月 15, 2009

n_requested_mmu_pages/n_free_mmu_pages are used by
kvm_mmu_change_mmu_pages to calculate the number of pages to zap.

alloc_mmu_pages, called from the vcpu initialization path, modifies this
variables without proper locking, which can result in a negative value
in kvm_mmu_change_mmu_pages (say, with cpu hotplug).
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

6a1ac771

KVM: Discard unnecessary kvm_mmu_flush_tlb() in kvm_mmu_load() · 3662cb1c

由 Sheng Yang 提交于 7月 09, 2009

set_cr3() should already cover the TLB flushing.
Signed-off-by: NSheng Yang <sheng@linux.intel.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

3662cb1c

KVM: silence lapic kernel messages that can be triggered by a guest · 4088bb3c

由 Gleb Natapov 提交于 7月 08, 2009

Some Linux versions (f8) try to read EOI register that is write only.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>

4088bb3c

KVM: Reduce runnability interface with arch support code · a1b37100

由 Gleb Natapov 提交于 7月 09, 2009

Remove kvm_cpu_has_interrupt() and kvm_arch_interrupt_allowed() from
interface between general code and arch code. kvm_arch_vcpu_runnable()
checks for interrupts instead.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

a1b37100

G
KVM: Move exception handling to the same place as other events · b59bb7bd
由 Gleb Natapov 提交于 7月 09, 2009
```
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
b59bb7bd

KVM: MMU: Fix MMU_DEBUG compile breakage · a205bc19

由 Joerg Roedel 提交于 7月 09, 2009

Signed-off-by: NJoerg Roedel <joerg.roedel@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

a205bc19

KVM: add ioeventfd support · d34e6b17

由 Gregory Haskins 提交于 7月 07, 2009

ioeventfd is a mechanism to register PIO/MMIO regions to trigger an eventfd
signal when written to by a guest.  Host userspace can register any
arbitrary IO address with a corresponding eventfd and then pass the eventfd
to a specific end-point of interest for handling.

Normal IO requires a blocking round-trip since the operation may cause
side-effects in the emulated model or may return data to the caller.
Therefore, an IO in KVM traps from the guest to the host, causes a VMX/SVM
"heavy-weight" exit back to userspace, and is ultimately serviced by qemu's
device model synchronously before returning control back to the vcpu.

However, there is a subclass of IO which acts purely as a trigger for
other IO (such as to kick off an out-of-band DMA request, etc).  For these
patterns, the synchronous call is particularly expensive since we really
only want to simply get our notification transmitted asychronously and
return as quickly as possible.  All the sychronous infrastructure to ensure
proper data-dependencies are met in the normal IO case are just unecessary
overhead for signalling.  This adds additional computational load on the
system, as well as latency to the signalling path.

Therefore, we provide a mechanism for registration of an in-kernel trigger
point that allows the VCPU to only require a very brief, lightweight
exit just long enough to signal an eventfd.  This also means that any
clients compatible with the eventfd interface (which includes userspace
and kernelspace equally well) can now register to be notified. The end
result should be a more flexible and higher performance notification API
for the backend KVM hypervisor and perhipheral components.

To test this theory, we built a test-harness called "doorbell".  This
module has a function called "doorbell_ring()" which simply increments a
counter for each time the doorbell is signaled.  It supports signalling
from either an eventfd, or an ioctl().

We then wired up two paths to the doorbell: One via QEMU via a registered
io region and through the doorbell ioctl().  The other is direct via
ioeventfd.

You can download this test harness here:

ftp://ftp.novell.com/dev/ghaskins/doorbell.tar.bz2

The measured results are as follows:

qemu-mmio:       110000 iops, 9.09us rtt
ioeventfd-mmio: 200100 iops, 5.00us rtt
ioeventfd-pio:  367300 iops, 2.72us rtt

I didn't measure qemu-pio, because I have to figure out how to register a
PIO region with qemu's device model, and I got lazy.  However, for now we
can extrapolate based on the data from the NULLIO runs of +2.56us for MMIO,
and -350ns for HC, we get:

qemu-pio:      153139 iops, 6.53us rtt
ioeventfd-hc: 412585 iops, 2.37us rtt

these are just for fun, for now, until I can gather more data.

Here is a graph for your convenience:

http://developer.novell.com/wiki/images/7/76/Iofd-chart.png

The conclusion to draw is that we save about 4us by skipping the userspace
hop.

--------------------
Signed-off-by: NGregory Haskins <ghaskins@novell.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d34e6b17

KVM: make io_bus interface more robust · 090b7aff

由 Gregory Haskins 提交于 7月 07, 2009

Today kvm_io_bus_regsiter_dev() returns void and will internally BUG_ON
if it fails.  We want to create dynamic MMIO/PIO entries driven from
userspace later in the series, so we need to enhance the code to be more
robust with the following changes:

   1) Add a return value to the registration function
   2) Fix up all the callsites to check the return code, handle any
      failures, and percolate the error up to the caller.
   3) Add an unregister function that collapses holes in the array
Signed-off-by: NGregory Haskins <ghaskins@novell.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

090b7aff

KVM: PIT support for HPET legacy mode · e9f42757

由 Beth Kon 提交于 7月 07, 2009

When kvm is in hpet_legacy_mode, the hpet is providing the timer
interrupt and the pit should not be. So in legacy mode, the pit timer
is destroyed, but the *state* of the pit is maintained. So if kvm or
the guest tries to modify the state of the pit, this modification is
accepted, *except* that the timer isn't actually started. When we exit
hpet_legacy_mode, the current state of the pit (which is up to date
since we've been accepting modifications) is used to restart the pit
timer.

The saved_mode code in kvm_pit_load_count temporarily changes mode to
0xff in order to destroy the timer, but then restores the actual
value, again maintaining "current" state of the pit for possible later
reenablement.

[avi: add some reserved storage in the ioctl; make SET_PIT2 IOW]
[marcelo: fix memory corruption due to reserved storage]
Signed-off-by: NBeth Kon <eak@us.ibm.com>
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

e9f42757

KVM: Always report x2apic as supported feature · 0d1de2d9

由 Gleb Natapov 提交于 7月 12, 2009

We emulate x2apic in software, so host support is not required.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

0d1de2d9

KVM: No need to kick cpu if not in a guest mode · c7f0f24b

由 Gleb Natapov 提交于 7月 07, 2009

This will save a couple of IPIs.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Acked-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

c7f0f24b

KVM: Add trace points in irqchip code · 1000ff8d

由 Gleb Natapov 提交于 7月 07, 2009

Add tracepoint in msi/ioapic/pic set_irq() functions,
in IPI sending and in the point where IRQ is placed into
apic's IRR.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

1000ff8d

KVM: fix MMIO_CONF_BASE MSR access · f7c6d140

由 Andre Przywara 提交于 7月 02, 2009

Some Windows versions check whether the BIOS has setup MMI/O for
config space accesses on AMD Fam10h CPUs, we say "no" by returning 0 on
reads and only allow disabling of MMI/O CfgSpace setup by igoring "0" writes.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

f7c6d140

A
KVM: Trace shadow page lifecycle · f691fe1d
由 Avi Kivity 提交于 7月 06, 2009
```
Create, sync, unsync, zap.
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
f691fe1d
A
KVM: MMU: Trace guest pagetable walker · 07420171
由 Avi Kivity 提交于 7月 06, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
07420171

Revert "KVM: x86: check for cr3 validity in ioctl_set_sregs" · dc7e795e

由 Jan Kiszka 提交于 7月 01, 2009

This reverts commit 6c20e1442bb1c62914bb85b7f4a38973d2a423ba.

To my understanding, it became obsolete with the advent of the more
robust check in mmu_alloc_roots (89da4ff17f). Moreover, it prevents
the conceptually safe pattern

 1. set sregs
 2. register mem-slots
 3. run vcpu

by setting a sticky triple fault during step 1.
Signed-off-by: NJan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

dc7e795e

KVM: handle AMD microcode MSR · 6098ca93

由 Andre Przywara 提交于 7月 03, 2009

Windows 7 tries to update the CPU's microcode on some processors,
so we ignore the MSR write here. The patchlevel register is already handled
(returning 0), because the MSR number is the same as Intel's.
Signed-off-by: NAndre Przywara <andre.przywara@amd.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

6098ca93

KVM: Fix apic_mmio_write return for unaligned write · 756975bb

由 Sheng Yang 提交于 7月 06, 2009

Some in-famous OS do unaligned writing for APIC MMIO, and the return value
has been missed in recent change, then the OS hangs.
Signed-off-by: NSheng Yang <sheng@linux.intel.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

756975bb

KVM: x2apic interface to lapic · 0105d1a5

由 Gleb Natapov 提交于 7月 05, 2009

This patch implements MSR interface to local apic as defines by x2apic
Intel specification.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

0105d1a5

KVM: Add Directed EOI support to APIC emulation · fc61b800

由 Gleb Natapov 提交于 7月 05, 2009

Directed EOI is specified by x2APIC, but is available even when lapic is
in xAPIC mode.
Signed-off-by: NGleb Natapov <gleb@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

fc61b800

A
KVM: Trace apic registers using their symbolic names · cb247721
由 Avi Kivity 提交于 7月 01, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
cb247721
A
KVM: Trace mmio · aec51dc4
由 Avi Kivity 提交于 7月 01, 2009
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
aec51dc4

openanolis / cloud-kernel 大约 1 年 前同步成功

openanolis / cloud-kernel
大约 1 年前同步成功