提交 · b931bfbf042983f311b3b09894d8030b2755a638 · openeuler / qemu

24 9月, 2015 34 次提交

vhost-user: add multiple queue support · b931bfbf

由 Changchun Ouyang 提交于 9月 23, 2015

This patch is initially based a patch from Nikolay Nikolaev.

This patch adds vhost-user multiple queue support, by creating a nc
and vhost_net pair for each queue.

Qemu exits if find that the backend can't support the number of requested
queues (by providing queues=# option). The max number is queried by a
new message, VHOST_USER_GET_QUEUE_NUM, and is sent only when protocol
feature VHOST_USER_PROTOCOL_F_MQ is present first.

The max queue check is done at vhost-user initiation stage. We initiate
one queue first, which, in the meantime, also gets the max_queues the
backend supports.

In older version, it was reported that some messages are sent more times
than necessary. Here we came an agreement with Michael that we could
categorize vhost user messages to 2 types: non-vring specific messages,
which should be sent only once, and vring specific messages, which should
be sent per queue.

Here I introduced a helper function vhost_user_one_time_request(), which
lists following messages as non-vring specific messages:

        VHOST_USER_SET_OWNER
        VHOST_USER_RESET_DEVICE
        VHOST_USER_SET_MEM_TABLE
        VHOST_USER_GET_QUEUE_NUM

For above messages, we simply ignore them when they are not sent the first
time.
Signed-off-by: NNikolay Nikolaev <n.nikolaev@virtualopensystems.com>
Signed-off-by: NChangchun Ouyang <changchun.ouyang@intel.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NJason Wang <jasowang@redhat.com>
Tested-by: NMarcel Apfelbaum <marcel@redhat.com>

b931bfbf

vhost: introduce vhost_backend_get_vq_index method · fc57fd99

由 Yuanhan Liu 提交于 9月 23, 2015

Minusing the idx with the base(dev->vq_index) for vhost-kernel, and
then adding it back for vhost-user doesn't seem right. Here introduces
a new method vhost_backend_get_vq_index() for getting the right vq
index for following vhost messages calls.
Suggested-by: NJason Wang <jasowang@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NJason Wang <jasowang@redhat.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Tested-by: NMarcel Apfelbaum <marcel@redhat.com>

fc57fd99

vhost-user: add VHOST_USER_GET_QUEUE_NUM message · e2051e9e

由 Yuanhan Liu 提交于 9月 23, 2015

This is for querying how many queues the backend supports if it has mq
support(when VHOST_USER_PROTOCOL_F_MQ flag is set from the quried
protocol features).

vhost_net_get_max_queues() is the interface to export that value, and
to tell if the backend supports # of queues user requested, which is
done in the following patch.
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Tested-by: NMarcel Apfelbaum <marcel@redhat.com>

e2051e9e

vhost: rename VHOST_RESET_OWNER to VHOST_RESET_DEVICE · d1f8b30e

由 Yuanhan Liu 提交于 9月 23, 2015

Quote from Michael:

    We really should rename VHOST_RESET_OWNER to VHOST_RESET_DEVICE.
Suggested-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMarcel Apfelbaum <marcel@redhat.com>
Tested-by: NMarcel Apfelbaum <marcel@redhat.com>

d1f8b30e

vhost-user: add protocol feature negotiation · dcb10c00

由 Michael S. Tsirkin 提交于 9月 23, 2015

Support a separate bitmask for vhost-user protocol features,
and messages to get/set protocol features.

Invoke them at init.

No features are defined yet.

[ leverage vhost_user_call for request handling -- Yuanhan Liu ]
Signed-off-by: NMichael S. Tsirkin <address@hidden>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMarcel Apfelbaum <marcel@redhat.com>
Tested-by: NMarcel Apfelbaum <marcel@redhat.com>

dcb10c00

vhost-user: use VHOST_USER_XXX macro for switch statement · 7305483a

由 Yuanhan Liu 提交于 9月 23, 2015

So that we could let vhost_user_call to handle extented requests,
such as VHOST_USER_GET/SET_PROTOCOL_FEATURES, instead of invoking
vhost_user_read/write and constructing the msg again by ourself.
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NYuanhan Liu <yuanhan.liu@linux.intel.com>
Reviewed-by: NMarcel Apfelbaum <marcel@redhat.com>
Tested-by: NMarcel Apfelbaum <marcel@redhat.com>

7305483a

virtio-ccw: enable virtio-1 · 542571d5

由 Cornelia Huck 提交于 9月 11, 2015

Let's enable revision 1 for virtio-ccw devices. We can always offer
VERSION_1 as drivers in legacy mode won't be able to see it anyway.

We have to introduce a way to set a lower maximum revision for a device
to accommodate the following cases:
- compat machines (to enforce legacy only)
- virtio-blk with scsi support (version 1 + scsi is fenced by common
  code, with a user-configured max revision of 0 we can allow scsi
  via not offering VERSION_1)
Signed-off-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

542571d5

virtio-ccw: feature bits > 31 handling · b4f8f9df

由 Cornelia Huck 提交于 9月 11, 2015

We currently switch off the VERSION_1 feature bit if the guest has
not negotiated at least revision 1. As no feature bits beyond 31 are
valid however unless VERSION_1 has been negotiated, make sure that
legacy guests never see a feature bit beyond 31.
Reviewed-by: NDavid Hildenbrand <dahi@linux.vnet.ibm.com>
Signed-off-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

b4f8f9df

virtio-ccw: support ring size changes · 79cd0c80

由 Cornelia Huck 提交于 9月 11, 2015

Wire up changing the ring size for virtio-1 devices.
Signed-off-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

79cd0c80

virtio: ring sizes vs. reset · 46c5d082

由 Cornelia Huck 提交于 9月 11, 2015

We allow guests to change the size of the virtqueue rings by supplying
a number of buffers that is different from the number of buffers the
device was initialized with. Current code has some problems, however,
since reset does not reset the ringsizes to the default values (as this
is not saved anywhere).

Let's extend the core code to keep track of the default ringsizes and
migrate them once the guest changed them for any of the virtqueues
for a device.
Reviewed-by: NJason Wang <jasowang@redhat.com>
Signed-off-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

46c5d082

pc: Introduce pc-*-2.5 machine classes · 87e896ab

由 Eduardo Habkost 提交于 9月 11, 2015

Signed-off-by: NEduardo Habkost <ehabkost@redhat.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

87e896ab

q35: Move options common to all classes to pc_i440fx_machine_options() · 254bdb1c

由 Eduardo Habkost 提交于 9月 11, 2015

The existing default_machine_opts and default_display settings will
still apply to future machine classes. So it makes sense to move them to
pc_i440fx_machine_options() instead of keeping them in a
version-specific machine_options function.
Signed-off-by: NEduardo Habkost <ehabkost@redhat.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

254bdb1c

q35: Move options common to all classes to pc_q35_machine_options() · 0b7783a7

由 Eduardo Habkost 提交于 9月 11, 2015

The existing default_machine_opts, default_display, no_floppy, and
no_tco settings will still apply to future machine classes. So it makes
sense to move them to pc_q35_machine_options() instead of keeping them
in a version-specific machine_options function.
Signed-off-by: NEduardo Habkost <ehabkost@redhat.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

0b7783a7

virtio-net: unbreak self announcement and guest offloads after migration · 1f8828ef

由 Jason Wang 提交于 9月 11, 2015

After commit 019a3edb ("virtio: make
features 64bit wide"). Device's guest_features was actually set after
vdc->load(). This breaks the assumption that device specific load()
function can check guest_features. For virtio-net, self announcement
and guest offloads won't work after migration.

Fixing this by defer them to virtio_net_load() where guest_features
were guaranteed to be set. Other virtio devices looks fine.

Fixes: 019a3edb
       ("virtio: make features 64bit wide")
Cc: qemu-stable@nongnu.org
Cc: Gerd Hoffmann <kraxel@redhat.com>
Signed-off-by: NJason Wang <jasowang@redhat.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Reviewed-by: NCornelia Huck <cornelia.huck@de.ibm.com>

1f8828ef

virtio: right size for virtio_queue_get_avail_size · 50764fc8

由 Pierre Morel 提交于 9月 10, 2015

Being working on dataplane I notice something strange:

virtio_queue_get_avail_size() used a 64bit size index
for the calculation of the available ring size.

It is quite strange but it did work with the old calculation
of the avail ring, at most with performance penalty,
and I wonder where I missed something.

This patch let use a 16bit size as defined in virtio_ring.h
Signed-off-by: NPierre Morel <pmorel@linux.vnet.ibm.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

50764fc8

vfio/pci: Add emulated PCI IDs · 89dcccc5

由 Alex Williamson 提交于 9月 23, 2015

Specifying an emulated PCI vendor/device ID can be useful for testing
various quirk paths, even though the behavior and functionality of
the device with bogus IDs is fully unsupportable. We need to use a
uint32_t for the vendor/device IDs, even though the registers
themselves are only 16-bit in order to be able to determine whether
the value is valid and user set.

The same support is added for subsystem vendor/device ID, though these
have the possibility of being useful and supported for more than a
testing tool. An emulated platform might want to impose their own
subsystem IDs or at least hide the physical subsystem ID. Windows
guests will often reinstall drivers due to a change in subsystem IDs,
something that VM users may want to avoid. Of course careful
attention would be required to ensure that guest drivers do not rely
on the subsystem ID as a basis for device driver quirks.

All of these options are added using the standard experimental option
prefix and should not be considered stable.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

89dcccc5

vfio/pci: Cache vendor and device ID · ff635e37

由 Alex Williamson 提交于 9月 23, 2015

Simplify access to commonly referenced PCI vendor and device ID by
caching it on the VFIOPCIDevice struct.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

ff635e37

vfio/pci: Move AMD device specific reset to quirks · c9c50009

由 Alex Williamson 提交于 9月 23, 2015

This is just another quirk, for reset rather than affecting memory
regions.  Move it to our new quirks file.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

c9c50009

A
vfio/pci: Remove old config window and mirror quirks · 958d5534
由 Alex Williamson 提交于 9月 23, 2015
```
These are now unused.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>
```
958d5534

vfio/pci: Config mirror quirk · 0d38fb1c

由 Alex Williamson 提交于 9月 23, 2015

Re-implement our mirror quirk using the new infrastructure.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

0d38fb1c

vfio/pci: Config window quirks · 0e54f24a

由 Alex Williamson 提交于 9月 23, 2015

Config windows make use of an address register and a data register.
In VGA cards, these are often used to provide real mode code in the
BIOS an easy way to access MMIO registers since the window often
resides in an I/O port register.  When the MMIO register has a mirror
of PCI config space, we need to trap those accesses and redirect them
to emulated config space.

The previous version of this functionality made use of a single
MemoryRegion and single match address.  This version uses separate
MemoryRegions for each of the address and data registers and allows
for multiple match addresses.  This is useful for Nvidia cards which
have two ranges which index into PCI config space.

The previous implementation is left for the follow-on patch for a more
reviewable diff.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

0e54f24a

vfio/pci: Rework RTL8168 quirk · 954258a5

由 Alex Williamson 提交于 9月 23, 2015

Another rework of this quirk, this time to update to the new quirk
structure.  We can handle the address and data registers with
separate MemoryRegions and a quirk specific data structure, making the
code much more understandable.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

954258a5

vfio/pci: Cleanup Nvidia 0x3d0 quirk · 6029a424

由 Alex Williamson 提交于 9月 23, 2015

The Nvidia 0x3d0 quirk makes use of a two separate registers and gives
us our first chance to make use of separate memory regions for each to
simplify the code a bit.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

6029a424

vfio/pci: Cleanup ATI 0x3c3 quirk · b946d286

由 Alex Williamson 提交于 9月 23, 2015

This is an easy quirk that really doesn't need a data structure if
its own.  We can pass vdev as the opaque data and access to the
MemoryRegion isn't required.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

b946d286

vfio/pci: Foundation for new quirk structure · 8c4f2348

由 Alex Williamson 提交于 9月 23, 2015

VFIOQuirk hosts a single memory region and a fixed set of data fields
that try to handle all the quirk cases, but end up making those that
don't exactly match really confusing. This patch introduces a struct
intended to provide more flexibility and simpler code. VFIOQuirk is
stripped to its basics, an opaque data pointer for quirk specific
data and a pointer to an array of MemoryRegions with a counter. This
still allows us to have common teardown routines, but adds much
greater flexibility to support multiple memory regions and quirk
specific data structures that are easier to maintain. The existing
VFIOQuirk is transformed into VFIOLegacyQuirk, which further patches
will eliminate entirely.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

8c4f2348

vfio/pci: Cleanup ROM blacklist quirk · 056dfcb6

由 Alex Williamson 提交于 9月 23, 2015

Create a vendor:device ID helper that we'll also use as we rework the
rest of the quirks. Re-reading the config entries, even if we get
more blacklist entries, is trivial overhead and only incurred during
device setup. There's no need to typedef the blacklist structure,
it's a static private data type used once. The elements get bumped
up to uint32_t to avoid future maintenance issues if PCI_ANY_ID gets
used for a blacklist entry (avoiding an actual hardware match). Our
test loop is also crying out to be simplified as a for loop.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

056dfcb6

A
vfio/pci: Split quirks to a separate file · c00d61d8
由 Alex Williamson 提交于 9月 23, 2015
```
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>
```
c00d61d8
A
vfio/pci: Extract PCI structures to a separate header · 78f33d2b
由 Alex Williamson 提交于 9月 23, 2015
```
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>
```
78f33d2b

vfio: Change polarity of our no-mmap option · 5e15d79b

由 Alex Williamson 提交于 9月 23, 2015

The default should be to allow mmap and new drivers shouldn't need to
expose an option or set it to other than the allocation default in
their initfn.  Take advantage of the experimental flag to change this
option to the correct polarity.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

5e15d79b

vfio/pci: Make interrupt bypass runtime configurable · 46746dba

由 Alex Williamson 提交于 9月 23, 2015

Tracing is more effective when we can completely disable all KVM
bypass paths. Make these runtime rather than build-time configurable.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

46746dba

vfio/pci: Rename MSI/X functions for easier tracing · 0de70dc7

由 Alex Williamson 提交于 9月 23, 2015

This allows vfio_msi* tracing.  The MSI/X interrupt tracing is also
pulled out of #ifdef DEBUG_VFIO to avoid a recompile for tracing this
path.  A few cycles to read the message is hardly anything if we're
already in QEMU.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

0de70dc7

vfio/pci: Rename INTx functions for easier tracing · 870cb6f1

由 Alex Williamson 提交于 9月 23, 2015

Rename functions and tracing callbacks so that we can trace vfio_intx*
to see all the INTx related activities.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

870cb6f1

vfio/pci: Cleanup vfio_early_setup_msix() error path · b5bd049f

由 Alex Williamson 提交于 9月 23, 2015

With the addition of the Chelsio quirk we have an error path out of
vfio_early_setup_msix() that doesn't free the allocated VFIOMSIXInfo
struct.  This doesn't introduce a leak as it still gets freed in the
vfio_put_device() path, but it's complicated and sloppy to rely on
that.  Restructure to free the allocated data on error and only link
it into the vdev on success.
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>
Reported-by: NLaszlo Ersek <lersek@redhat.com>
Reviewed-by: NLaszlo Ersek <lersek@redhat.com>

b5bd049f

vfio/pci: Cleanup RTL8168 quirk and tracing · d451008e

由 Alex Williamson 提交于 9月 23, 2015

There's quite a bit of cleanup that can be done to the RTL8168 quirk,
as well as the tracing to prevent a spew of uninteresting accesses
for anything else the driver might choose to use the window registers
for besides the MSI-X table.  There should be no functional change,
but it's now possible to get compact and useful traces by enabling
vfio_rtl8168_quirk*, ex:

vfio_rtl8168_quirk_write 0000:04:00.0 [address]: 0x1f000
vfio_rtl8168_quirk_read 0000:04:00.0 [address]: 0x8001f000
vfio_rtl8168_quirk_read 0000:04:00.0 [data]: 0xfee0100c
vfio_rtl8168_quirk_write 0000:04:00.0 [address]: 0x1f004
vfio_rtl8168_quirk_read 0000:04:00.0 [address]: 0x8001f004
vfio_rtl8168_quirk_read 0000:04:00.0 [data]: 0x0
vfio_rtl8168_quirk_write 0000:04:00.0 [address]: 0x1f008
vfio_rtl8168_quirk_read 0000:04:00.0 [address]: 0x8001f008
vfio_rtl8168_quirk_read 0000:04:00.0 [data]: 0x49b1
vfio_rtl8168_quirk_write 0000:04:00.0 [address]: 0x1f00c
vfio_rtl8168_quirk_read 0000:04:00.0 [address]: 0x8001f00c
vfio_rtl8168_quirk_read 0000:04:00.0 [data]: 0x0
Signed-off-by: NAlex Williamson <alex.williamson@redhat.com>

d451008e

23 9月, 2015 6 次提交

sPAPR: Enable EEH on VFIO PCI device only · d76548a9

由 Gavin Shan 提交于 9月 18, 2015

This checks if the PCI device retrieved from the PCI device address
is VFIO PCI device when enabling EEH functionality. If it's not
VFIO PCI device, the EEH functonality isn't enabled.
Signed-off-by: NGavin Shan <gwshan@linux.vnet.ibm.com>
Signed-off-by: NDavid Gibson <david@gibson.dropbear.id.au>

d76548a9

sPAPR: Revert don't enable EEH on emulated PCI devices · 47445c80

由 Gavin Shan 提交于 9月 18, 2015

This reverts commit 7cb18007 ("sPAPR: Don't enable EEH on emulated
PCI devices") as rtas_ibm_set_eeh_option() isn't the right place
to check if there has the corresponding PCI device for the input
address, which can be PE address, not PCI device address.
Signed-off-by: NGavin Shan <gwshan@linux.vnet.ibm.com>
Signed-off-by: NDavid Gibson <david@gibson.dropbear.id.au>

47445c80

ppc/spapr: Implement H_RANDOM hypercall in QEMU · 4d9392be

由 Thomas Huth 提交于 9月 17, 2015

The PAPR interface defines a hypercall to pass high-quality
hardware generated random numbers to guests. Recent kernels can
already provide this hypercall to the guest if the right hardware
random number generator is available. But in case the user wants
to use another source like EGD, or QEMU is running with an older
kernel, we should also have this call in QEMU, so that guests that
do not support virtio-rng yet can get good random numbers, too.

This patch now adds a new pseudo-device to QEMU that either
directly provides this hypercall to the guest or is able to
enable the in-kernel hypercall if available. The in-kernel
hypercall can be enabled with the use-kvm property, e.g.:

 qemu-system-ppc64 -device spapr-rng,use-kvm=true

For handling the hypercall in QEMU instead, a "RngBackend" is
required since the hypercall should provide "good" random data
instead of pseudo-random (like from a "simple" library function
like rand() or g_random_int()). Since there are multiple RngBackends
available, the user must select an appropriate back-end via the
"rng" property of the device, e.g.:

 qemu-system-ppc64 -object rng-random,filename=/dev/hwrng,id=gid0 \
                   -device spapr-rng,rng=gid0 ...

See http://wiki.qemu-project.org/Features-Done/VirtIORNG for
other example of specifying RngBackends.
Signed-off-by: NThomas Huth <thuth@redhat.com>
Signed-off-by: NDavid Gibson <david@gibson.dropbear.id.au>

4d9392be

ppc/spapr: Fix buffer overflow in spapr_populate_drconf_memory() · ef001f06

由 Thomas Huth 提交于 9月 15, 2015

The buffer that is allocated in spapr_populate_drconf_memory()
is used for setting both, the "ibm,dynamic-memory" and the
"ibm,associativity-lookup-arrays" property. However, only the
size of the first one is taken into account when allocating the
memory. So if the length of the second property is larger than
the length of the first one, we run into a buffer overflow here!
Fix it by taking the length of the second property into account,
too.

Fixes: "spapr: Support ibm,dynamic-reconfiguration-memory" patch
Signed-off-by: NThomas Huth <thuth@redhat.com>
Reviewed-by: NDavid Gibson <david@gibson.dropbear.id.au>
Signed-off-by: NDavid Gibson <david@gibson.dropbear.id.au>

ef001f06

spapr: Fix default NUMA node allocation for threads · 20bb648d

由 David Gibson 提交于 9月 08, 2015

At present, if guest numa nodes are requested, but the cpus in each node
are not specified, spapr just uses the default behaviour or assigning each
vcpu round-robin to nodes.

If smp_threads != 1, that will assign adjacent threads in a core to
different NUMA nodes. As well as being just weird, that's a configuration
that can't be represented in the device tree we give to the guest, which
means the guest and qemu end up with different ideas of the NUMA topology.

This patch implements mc->cpu_index_to_socket_id in the spapr code to
make sure vcpus get assigned to nodes only at the socket granularity.
Signed-off-by: NDavid Gibson <david@gibson.dropbear.id.au>
Reviewed-by: NAlexey Kardashevskiy <aik@ozlabs.ru>

20bb648d

spapr: Move memory hotplug to RTAS_LOG_V6_HP_ID_DRC_COUNT type · 0a417869

由 Bharata B Rao 提交于 8月 03, 2015

Till now memory hotplug used RTAS_LOG_V6_HP_ID_DRC_INDEX hotplug type
which meant that we generated one hotplug type of EPOW event for every
256MB (SPAPR_MEMORY_BLOCK_SIZE). This quickly overruns the kernel
rtas log buffer thus resulting in loss of memory hotplug events. Switch
to RTAS_LOG_V6_HP_ID_DRC_COUNT hotplug type for memory so that we
generate only one event per hotplug request.
Signed-off-by: NBharata B Rao <bharata@linux.vnet.ibm.com>
Reviewed-by: NMichael Roth <mdroth@linux.vnet.ibm.com>
Reviewed-by: NDavid Gibson <david@gibson.dropbear.id.au>
Signed-off-by: NDavid Gibson <david@gibson.dropbear.id.au>

0a417869