提交 · 597b0d21626da4e6f09f132442caf0cc2b0eb47c · openanolis / cloud-kernel

03 1月, 2009 1 次提交

Fix compiler warning in arch/x86/mm/init_32.c · e8e32326

由 Ingo Brueckl 提交于 1月 02, 2009

Signed-off-by: NIngo Brueckl <ib@wupperonline.de>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

e8e32326

01 1月, 2009 2 次提交

A
take init_fs to saner place · 18d8fda7
由 Al Viro 提交于 12月 26, 2008
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
18d8fda7

shrink struct dentry · c2452f32

由 Nick Piggin 提交于 12月 01, 2008

struct dentry is one of the most critical structures in the kernel. So it's
sad to see it going neglected.

With CONFIG_PROFILING turned on (which is probably the common case at least
for distros and kernel developers), sizeof(struct dcache) == 208 here
(64-bit). This gives 19 objects per slab.

I packed d_mounted into a hole, and took another 4 bytes off the inline
name length to take the padding out from the end of the structure. This
shinks it to 200 bytes. I could have gone the other way and increased the
length to 40, but I'm aiming for a magic number, read on...

I then got rid of the d_cookie pointer. This shrinks it to 192 bytes. Rant:
why was this ever a good idea? The cookie system should increase its hash
size or use a tree or something if lookups are a problem. Also the "fast
dcookie lookups" in oprofile should be moved into the dcookie code -- how
can oprofile possibly care about the dcookie_mutex? It gets dropped after
get_dcookie() returns so it can't be providing any sort of protection.

At 192 bytes, 21 objects fit into a 4K page, saving about 3MB on my system
with ~140 000 entries allocated. 192 is also a multiple of 64, so we get
nice cacheline alignment on 64 and 32 byte line systems -- any given dentry
will now require 3 cachelines to touch all fields wheras previously it
would require 4.

I know the inline name size was chosen quite carefully, however with the
reduction in cacheline footprint, it should actually be just about as fast
to do a name lookup for a 36 character name as it was before the patch (and
faster for other sizes). The memory footprint savings for names which are
<= 32 or > 36 bytes long should more than make up for the memory cost for
33-36 byte names.

Performance is a feature...
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

c2452f32

31 12月, 2008 37 次提交

KVM: MMU: handle large host sptes on invlpg/resync · 87917239

由 Marcelo Tosatti 提交于 12月 22, 2008

The invlpg and sync walkers lack knowledge of large host sptes,
descending to non-existant pagetable level.

Stop at directory level in such case.

Fixes SMP Windows XP with hugepages.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

87917239

KVM: Add locking to virtual i8259 interrupt controller · 3f353858

由 Avi Kivity 提交于 12月 21, 2008

While most accesses to the i8259 are with the kvm mutex taken, the call
to kvm_pic_read_irq() is not. We can't easily take the kvm mutex there
since the function is called with interrupts disabled.

Fix by adding a spinlock to the virtual interrupt controller. Since we
can't send an IPI under the spinlock (we also take the same spinlock in
an irq disabled context), we defer the IPI until the spinlock is released.
Similarly, we defer irq ack notifications until after spinlock release to
avoid lock recursion.
Signed-off-by: NAvi Kivity <avi@redhat.com>

3f353858

KVM: MMU: Don't treat a global pte as such if cr4.pge is cleared · 25e23432

由 Avi Kivity 提交于 12月 21, 2008

The pte.g bit is meaningless if global pages are disabled; deferring
mmu page synchronization on these ptes will lead to the guest using stale
shadow ptes.

Fixes Vista x86 smp bootloader failure.
Signed-off-by: NAvi Kivity <avi@redhat.com>

25e23432

KVM: ia64: Fix kvm_arch_vcpu_ioctl_[gs]et_regs() · 042b26ed

由 Jes Sorensen 提交于 12月 16, 2008

Fix kvm_arch_vcpu_ioctl_[gs]et_regs() to do something meaningful on
ia64. Old versions could never have worked since they required
pointers to be set in the ioctl payload which were never being set by
the ioctl handler for get_regs.

In addition reserve extra space for future extensions.

The change of layout of struct kvm_regs doesn't require adding a new
CAP since get/set regs never worked on ia64 until now.

This version doesn't support copying the KVM kernel stack in/out of
the kernel. This should be implemented in a seperate ioctl call if
ever needed.
Signed-off-by: NJes Sorensen <jes@sgi.com>
Acked-by : Xiantao Zhang <xiantao.zhang@intel.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

042b26ed

KVM: x86: Rework user space NMI injection as KVM_CAP_USER_NMI · 4531220b

由 Jan Kiszka 提交于 12月 11, 2008

There is no point in doing the ready_for_nmi_injection/
request_nmi_window dance with user space. First, we don't do this for
in-kernel irqchip anyway, while the code path is the same as for user
space irqchip mode. And second, there is nothing to loose if a pending
NMI is overwritten by another one (in contrast to IRQs where we have to
save the number). Actually, there is even the risk of raising spurious
NMIs this way because the reason for the held-back NMI might already be
handled while processing the first one.

Therefore this patch creates a simplified user space NMI injection
interface, exporting it under KVM_CAP_USER_NMI and dropping the old
KVM_CAP_NMI capability. And this time we also take care to provide the
interface only on archs supporting NMIs via KVM (right now only x86).
Signed-off-by: NJan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

4531220b

KVM: VMX: Fix pending NMI-vs.-IRQ race for user space irqchip · 264ff01d

由 Jan Kiszka 提交于 11月 24, 2008

As with the kernel irqchip, don't allow an NMI to stomp over an already
injected IRQ; instead wait for the IRQ injection to be completed.
Signed-off-by: NJan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

264ff01d

KVM: MMU: check for present pdptr shadow page in walk_shadow · eb64f1e8

由 Marcelo Tosatti 提交于 12月 09, 2008

walk_shadow assumes the caller verified validity of the pdptr pointer in
question, which is not the case for the invlpg handler.

Fixes oops during Solaris 10 install.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

eb64f1e8

A
KVM: Consolidate userspace memory capability reporting into common code · ca9edaee
由 Avi Kivity 提交于 12月 08, 2008
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
ca9edaee

x86: KVM guest: kvm_get_tsc_khz: return khz, not lpj · e93353c9

由 Eduardo Habkost 提交于 12月 05, 2008

kvm_get_tsc_khz() currently returns the previously-calculated preset_lpj
value, but it is in loops-per-jiffy, not kHz. The current code works
correctly only when HZ=1000.
Signed-off-by: NEduardo Habkost <ehabkost@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

e93353c9

KVM: MMU: prepopulate the shadow on invlpg · ad218f85

由 Marcelo Tosatti 提交于 12月 01, 2008

If the guest executes invlpg, peek into the pagetable and attempt to
prepopulate the shadow entry.

Also stop dirty fault updates from interfering with the fork detector.

2% improvement on RHEL3/AIM7.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

ad218f85

KVM: MMU: skip global pgtables on sync due to cr3 switch · 6cffe8ca

由 Marcelo Tosatti 提交于 12月 01, 2008

Skip syncing global pages on cr3 switch (but not on cr4/cr0). This is
important for Linux 32-bit guests with PAE, where the kmap page is
marked as global.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

6cffe8ca

KVM: MMU: collapse remote TLB flushes on root sync · b1a36821

由 Marcelo Tosatti 提交于 12月 01, 2008

Collapse remote TLB flushes on root sync.

kernbench is 2.7% faster on 4-way guest. Improvements have been seen
with other loads such as AIM7.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

b1a36821

KVM: MMU: use page array in unsync walk · 60c8aec6

由 Marcelo Tosatti 提交于 12月 01, 2008

Instead of invoking the handler directly collect pages into
an array so the caller can work with it.

Simplifies TLB flush collapsing.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

60c8aec6

KVM: x86 emulator: Fix handling of VMMCALL instruction · fbce554e

由 Amit Shah 提交于 12月 04, 2008

The VMMCALL instruction doesn't get recognised and isn't processed
by the emulator.

This is seen on an Intel host that tries to execute the VMMCALL
instruction after a guest live migrates from an AMD host.
Signed-off-by: NAmit Shah <amit.shah@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

fbce554e

KVM: x86 emulator: add the emulation of shld and shrd instructions · 9bf8ea42

由 Guillaume Thouvenin 提交于 12月 04, 2008

Add emulation of shld and shrd instructions
Signed-off-by: NGuillaume Thouvenin <guillaume.thouvenin@ext.bull.net>
Signed-off-by: NAvi Kivity <avi@redhat.com>

9bf8ea42

KVM: x86 emulator: add the assembler code for three operands · d175226a

由 Guillaume Thouvenin 提交于 12月 04, 2008

Add the assembler code for instruction with three operands and one
operand is stored in ECX register
Signed-off-by: NGuillaume Thouvenin <guillaume.thouvenin@ext.bull.net>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d175226a

KVM: x86 emulator: add a new "implied 1" Src decode type · bfcadf83

由 Guillaume Thouvenin 提交于 12月 04, 2008

Add SrcOne operand type when we need to decode an implied '1' like with
regular shift instruction
Signed-off-by: NGuillaume Thouvenin <guillaume.thouvenin@ext.bull.net>
Signed-off-by: NAvi Kivity <avi@redhat.com>

bfcadf83

KVM: x86 emulator: add Src2 decode set · 0dc8d10f

由 Guillaume Thouvenin 提交于 12月 04, 2008

Instruction like shld has three operands, so we need to add a Src2
decode set. We start with Src2None, Src2CL, and Src2ImmByte, Src2One to
support shld/shrd and we will expand it later.
Signed-off-by: NGuillaume Thouvenin <guillaume.thouvenin@ext.bull.net>
Signed-off-by: NAvi Kivity <avi@redhat.com>

0dc8d10f

KVM: x86 emulator: Extend the opcode descriptor · 45ed60b3

由 Guillaume Thouvenin 提交于 12月 04, 2008

Extend the opcode descriptor to 32 bits. This is needed by the
introduction of a new Src2 operand type.
Signed-off-by: NGuillaume Thouvenin <guillaume.thouvenin@ext.bull.net>
Signed-off-by: NAvi Kivity <avi@redhat.com>

45ed60b3

KVM: ppc: mostly cosmetic updates to the exit timing accounting code · 7b701591

由 Hollis Blanchard 提交于 12月 02, 2008

The only significant changes were to kvmppc_exit_timing_write() and
kvmppc_exit_timing_show(), both of which were dramatically simplified.
Signed-off-by: NHollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

7b701591

KVM: ppc: Implement in-kernel exit timing statistics · 73e75b41

由 Hollis Blanchard 提交于 12月 02, 2008

Existing KVM statistics are either just counters (kvm_stat) reported for
KVM generally or trace based aproaches like kvm_trace.
For KVM on powerpc we had the need to track the timings of the different exit
types. While this could be achieved parsing data created with a kvm_trace
extension this adds too much overhead (at least on embedded PowerPC) slowing
down the workloads we wanted to measure.

Therefore this patch adds a in-kernel exit timing statistic to the powerpc kvm
code. These statistic is available per vm&vcpu under the kvm debugfs directory.
As this statistic is low, but still some overhead it can be enabled via a
.config entry and should be off by default.

Since this patch touched all powerpc kvm_stat code anyway this code is now
merged and simplified together with the exit timing statistic code (still
working with exit timing disabled in .config).
Signed-off-by: NChristian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Signed-off-by: NHollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

73e75b41

KVM: ppc: save and restore guest mappings on context switch · c5fbdffb

由 Hollis Blanchard 提交于 12月 02, 2008

Store shadow TLB entries in memory, but only use it on host context switch
(instead of every guest entry). This improves performance for most workloads on
440 by reducing the guest TLB miss rate.
Signed-off-by: NHollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

c5fbdffb

KVM: ppc: directly insert shadow mappings into the hardware TLB · 7924bd41

由 Hollis Blanchard 提交于 12月 02, 2008

Formerly, we used to maintain a per-vcpu shadow TLB and on every entry to the
guest would load this array into the hardware TLB. This consumed 1280 bytes of
memory (64 entries of 16 bytes plus a struct page pointer each), and also
required some assembly to loop over the array on every entry.

Instead of saving a copy in memory, we can just store shadow mappings directly
into the hardware TLB, accepting that the host kernel will clobber these as
part of the normal 440 TLB round robin. When we do that we need less than half
the memory, and we have decreased the exit handling time for all guest exits,
at the cost of increased number of TLB misses because the host overwrites some
guest entries.

These savings will be increased on processors with larger TLBs or which
implement intelligent flush instructions like tlbivax (which will avoid the
need to walk arrays in software).

In addition to that and to the code simplification, we have a greater chance of
leaving other host userspace mappings in the TLB, instead of forcing all
subsequent tasks to re-fault all their mappings.
Signed-off-by: NHollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

7924bd41

powerpc/44x: declare tlb_44x_index for use in C code · c0ca609c

由 Hollis Blanchard 提交于 12月 02, 2008

KVM currently ignores the host's round robin TLB eviction selection, instead
maintaining its own TLB state and its own round robin index. However, by
participating in the normal 44x TLB selection, we can drop the alternate TLB
processing in KVM. This results in a significant performance improvement,
since that processing currently must be done on *every* guest exit.

Accordingly, KVM needs to be able to access and increment tlb_44x_index.
(KVM on 440 cannot be a module, so there is no need to export this symbol.)
Signed-off-by: NHollis Blanchard <hollisb@us.ibm.com>
Acked-by: NJosh Boyer <jwboyer@linux.vnet.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

c0ca609c

KVM: ppc: support large host pages · 89168618

由 Hollis Blanchard 提交于 12月 02, 2008

KVM on 440 has always been able to handle large guest mappings with 4K host
pages -- we must, since the guest kernel uses 256MB mappings.

This patch makes KVM work when the host has large pages too (tested with 64K).
Signed-off-by: NHollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

89168618

KVM: VMX: fix sparse warning · efff9e53

由 Hannes Eder 提交于 11月 28, 2008

Impact: make global function static

  arch/x86/kvm/vmx.c:134:3: warning: symbol 'vmx_capability' was not declared. Should it be static?
Signed-off-by: NHannes Eder <hannes@hanneseder.net>
Signed-off-by: NAvi Kivity <avi@redhat.com>

efff9e53

A
KVM: Remove extraneous semicolon after do/while · f3fd92fb
由 Avi Kivity 提交于 11月 29, 2008
```
Notices by Guillaume Thouvenin.
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
f3fd92fb

KVM: x86 emulator: fix popf emulation · 2b48cc75

由 Avi Kivity 提交于 11月 29, 2008

Set operand type and size to get correct writeback behavior.
Signed-off-by: NAvi Kivity <avi@redhat.com>

2b48cc75

KVM: x86 emulator: fix ret emulation · cf5de4f8

由 Avi Kivity 提交于 11月 28, 2008

'ret' did not set the operand type or size for the destination, so
writeback ignored it.
Signed-off-by: NAvi Kivity <avi@redhat.com>

cf5de4f8

A
KVM: x86 emulator: switch 'pop reg' instruction to emulate_pop() · 8a09b687
由 Avi Kivity 提交于 11月 27, 2008
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
8a09b687
A
KVM: x86 emulator: allow pop from mmio · 781d0edc
由 Avi Kivity 提交于 11月 27, 2008
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
781d0edc
A
KVM: x86 emulator: Extract 'pop' sequence into a function · faa5a3ae
由 Avi Kivity 提交于 11月 27, 2008
```
Switch 'pop r/m' instruction to use the new function.
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
faa5a3ae

KVM: s390: Fix memory leak of vcpu->run · 6692cef3

由 Christian Borntraeger 提交于 11月 26, 2008

The s390 backend of kvm never calls kvm_vcpu_uninit. This causes
a memory leak of vcpu->run pages.
Lets call kvm_vcpu_uninit in kvm_arch_vcpu_destroy to free
the vcpu->run.
Signed-off-by: NChristian Borntraeger <borntraeger@de.ibm.com>
Acked-by: NCarsten Otte <cotte@de.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

6692cef3

KVM: s390: Fix refcounting and allow module unload · d329c035

由 Christian Borntraeger 提交于 11月 26, 2008

Currently it is impossible to unload the kvm module on s390.
This patch fixes kvm_arch_destroy_vm to release all cpus.
This make it possible to unload the module.

In addition we stop messing with the module refcount in arch code.
Signed-off-by: NChristian Borntraeger <borntraeger@de.ibm.com>
Acked-by: NCarsten Otte <cotte@de.ibm.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

d329c035

A
KVM: x86 emulator: consolidate emulation of two operand instructions · 6b7ad61f
由 Avi Kivity 提交于 11月 26, 2008
```
No need to repeat the same assembly block over and over.
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
6b7ad61f
A
KVM: x86 emulator: reduce duplication in one operand emulation thunks · dda96d8f
由 Avi Kivity 提交于 11月 26, 2008
```
Signed-off-by: NAvi Kivity <avi@redhat.com>
```
dda96d8f

KVM: MMU: optimize set_spte for page sync · ecc5589f

由 Marcelo Tosatti 提交于 11月 25, 2008

The write protect verification in set_spte is unnecessary for page sync.

Its guaranteed that, if the unsync spte was writable, the target page
does not have a write protected shadow (if it had, the spte would have
been write protected under mmu_lock by rmap_write_protect before).

Same reasoning applies to mark_page_dirty: the gfn has been marked as
dirty via the pagefault path.

The cost of hash table and memslot lookups are quite significant if the
workload is pagetable write intensive resulting in increased mmu_lock
contention.
Signed-off-by: NMarcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: NAvi Kivity <avi@redhat.com>

ecc5589f

openanolis / cloud-kernel 接近 2 年 前同步成功

openanolis / cloud-kernel
接近 2 年前同步成功