提交 · 6b1e6cc7855b09a0a9bfa1d9f30172ba366f161c · openanolis / cloud-kernel

02 8月, 2016 3 次提交

由 Jason Wang 提交于 6月 23, 2016

This patch tries to implement an device IOTLB for vhost. This could be
used with userspace(qemu) implementation of DMA remapping
to emulate an IOMMU for the guest.

The idea is simple, cache the translation in a software device IOTLB
(which is implemented as an interval tree) in vhost and use vhost_net
file descriptor for reporting IOTLB miss and IOTLB
update/invalidation. When vhost meets an IOTLB miss, the fault
address, size and access can be read from the file. After userspace
finishes the translation, it writes the translated address to the
vhost_net file to update the device IOTLB.

When device IOTLB is enabled by setting VIRTIO_F_IOMMU_PLATFORM all vq
addresses set by ioctl are treated as iova instead of virtual address and
the accessing can only be done through IOTLB instead of direct userspace
memory access. Before each round or vq processing, all vq metadata is
prefetched in device IOTLB to make sure no translation fault happens
during vq processing.

In most cases, virtqueues are contiguous even in virtual address space.
The IOTLB translation for virtqueue itself may make it a little
slower. We might add fast path cache on top of this patch.
Signed-off-by: NJason Wang <jasowang@redhat.com>
[mst: use virtio feature bit: VHOST_F_DEVICE_IOTLB -> VIRTIO_F_IOMMU_PLATFORM ]
[mst: fix build warnings ]
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
[ weiyj.lk: missing unlock on error ]
Signed-off-by: NWei Yongjun <weiyj.lk@gmail.com>

6b1e6cc7

vhost: convert pre sorted vhost memory array to interval tree · a9709d68

由 Jason Wang 提交于 6月 23, 2016

Current pre-sorted memory region array has some limitations for future
device IOTLB conversion:

1) need extra work for adding and removing a single region, and it's
   expected to be slow because of sorting or memory re-allocation.
2) need extra work of removing a large range which may intersect
   several regions with different size.
3) need trick for a replacement policy like LRU

To overcome the above shortcomings, this patch convert it to interval
tree which can easily address the above issue with almost no extra
work.

The patch could be used for:

- Extend the current API and only let the userspace to send diffs of
  memory table.
- Simplify Device IOTLB implementation.
Signed-off-by: NJason Wang <jasowang@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

a9709d68

vhost: lockless enqueuing · 04b96e55

由 Jason Wang 提交于 4月 25, 2016

We use spinlock to synchronize the work list now which may cause
unnecessary contentions. So this patch switch to use llist to remove
this contention. Pktgen tests shows about 5% improvement:

Before:
~1300000 pps
After:
~1370000 pps
Signed-off-by: NJason Wang <jasowang@redhat.com>
Reviewed-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

04b96e55

11 3月, 2016 3 次提交

vhost_net: basic polling support · 03088137

由 Jason Wang 提交于 3月 04, 2016

This patch tries to poll for new added tx buffer or socket receive
queue for a while at the end of tx/rx processing. The maximum time
spent on polling were specified through a new kind of vring ioctl.
Signed-off-by: NJason Wang <jasowang@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

03088137

vhost: introduce vhost_vq_avail_empty() · d4a60603

由 Jason Wang 提交于 3月 04, 2016

This patch introduces a helper which will return true if we're sure
that the available ring is empty for a specific vq. When we're not
sure, e.g vq access failure, return false instead. This could be used
for busy polling code to exit the busy loop.
Signed-off-by: NJason Wang <jasowang@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

d4a60603

vhost: introduce vhost_has_work() · 526d3e7f

由 Jason Wang 提交于 3月 04, 2016

This path introduces a helper which can give a hint for whether or not
there's a work queued in the work list. This could be used for busy
polling code to exit the busy loop.
Signed-off-by: NJason Wang <jasowang@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

526d3e7f

02 3月, 2016 1 次提交

vhost: rename vhost_init_used() · 80f7d030

由 Greg Kurz 提交于 2月 16, 2016

Looking at how callers use this, maybe we should just rename init_used
to vhost_vq_init_access. The _used suffix was a hint that we
access the vq used ring. But maybe what callers care about is
that it must be called after access_ok.

Also, this function manipulates the vq->is_le field which isn't related
to the vq used ring.

This patch simply renames vhost_init_used() to vhost_vq_init_access() as
suggested by Michael.

No behaviour change.
Signed-off-by: NGreg Kurz <gkurz@linux.vnet.ibm.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

80f7d030

28 10月, 2015 1 次提交

vhost: fix performance on LE hosts · e407f39a

由 Michael S. Tsirkin 提交于 10月 27, 2015

commit 2751c988 ("vhost: cross-endian
support for legacy devices") introduced a minor regression: even with
cross-endian disabled, and even on LE host, vhost_is_little_endian is
checking is_le flag so there's always a branch.

To fix, simply check virtio_legacy_is_little_endian first.

Cc: Greg Kurz <gkurz@linux.vnet.ibm.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Reviewed-by: NGreg Kurz <gkurz@linux.vnet.ibm.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

e407f39a

16 9月, 2015 1 次提交

vhost: move features to core · 4e9fa50c

由 Michael S. Tsirkin 提交于 9月 09, 2015

virtio 1 and any layout are core features, move them
there. This fixes vhost test.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

4e9fa50c

01 6月, 2015 3 次提交

vhost: cross-endian support for legacy devices · 2751c988

由 Greg Kurz 提交于 4月 24, 2015

This patch brings cross-endian support to vhost when used to implement
legacy virtio devices. Since it is a relatively rare situation, the
feature availability is controlled by a kernel config option (not set
by default).

The vq->is_le boolean field is added to cache the endianness to be
used for ring accesses. It defaults to native endian, as expected
by legacy virtio devices. When the ring gets active, we force little
endian if the device is modern. When the ring is deactivated, we
revert to the native endian default.

If cross-endian was compiled in, a vq->user_be boolean field is added
so that userspace may request a specific endianness. This field is
used to override the default when activating the ring of a legacy
device. It has no effect on modern devices.
Signed-off-by: NGreg Kurz <gkurz@linux.vnet.ibm.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Reviewed-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NDavid Gibson <david@gibson.dropbear.id.au>

2751c988

virtio: add explicit big-endian support to memory accessors · 7d824109

由 Greg Kurz 提交于 4月 24, 2015

The current memory accessors logic is:
- little endian if little_endian
- native endian (i.e. no byteswap) if !little_endian

If we want to fully support cross-endian vhost, we also need to be
able to convert to big endian.

Instead of changing the little_endian argument to some 3-value enum, this
patch changes the logic to:
- little endian if little_endian
- big endian if !little_endian

The native endian case is handled by all users with a trivial helper. This
patch doesn't change any functionality, nor it does add overhead.
Signed-off-by: NGreg Kurz <gkurz@linux.vnet.ibm.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Reviewed-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NDavid Gibson <david@gibson.dropbear.id.au>

7d824109

vhost: introduce vhost_is_little_endian() helper · ab27c07f

由 Greg Kurz 提交于 4月 24, 2015

Signed-off-by: NGreg Kurz <gkurz@linux.vnet.ibm.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Acked-by: NCornelia Huck <cornelia.huck@de.ibm.com>
Reviewed-by: NDavid Gibson <david@gibson.dropbear.id.au>

ab27c07f

09 12月, 2014 3 次提交

vhost: remove unnecessary forward declarations in vhost.h · 2e73c716

由 Jason Wang 提交于 11月 27, 2014

Signed-off-by: NJason Wang <jasowang@redhat.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

2e73c716

vhost: add memory access wrappers · e05fd12b

由 Michael S. Tsirkin 提交于 10月 24, 2014

Add guest memory access wrappers to handle virtio endianness
conversions.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Reviewed-by: NJason Wang <jasowang@redhat.com>
Reviewed-by: NCornelia Huck <cornelia.huck@de.ibm.com>

e05fd12b

vhost: make features 64 bit · bd82752a

由 Michael S. Tsirkin 提交于 10月 24, 2014

We need to use bit 32 for virtio 1.0.
Make vhost_has_feature bool to avoid discarding high bits.

Cc: Sergei Shtylyov <sergei.shtylyov@cogentembedded.com>
Cc: Ben Hutchings <ben@decadent.org.uk>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Reviewed-by: NJason Wang <jasowang@redhat.com>

bd82752a

09 6月, 2014 2 次提交

vhost: move memory pointer to VQs · 47283bef

由 Michael S. Tsirkin 提交于 6月 05, 2014

commit 2ae76693b8bcabf370b981cd00c36cd41d33fabc
    vhost: replace rcu with mutex
replaced rcu sync for memory accesses with VQ mutex locl/unlock.
This is correct since all accesses are under VQ mutex, but incomplete:
we still do useless rcu lock/unlock operations, someone might copy this
code into some other context where this won't be right.
This use of RCU is also non standard and hard to understand.
Let's copy the pointer to each VQ structure, this way
the access rules become straight-forward, and there's
no need for RCU anymore.
Reported-by: NEric Dumazet <eric.dumazet@gmail.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

47283bef

vhost: move acked_features to VQs · ea16c514

由 Michael S. Tsirkin 提交于 6月 05, 2014

Refactor code to make sure features are only accessed
under VQ mutex. This makes everything simpler, no need
for RCU here anymore.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

ea16c514

07 12月, 2013 1 次提交

vhost: remove the dead branch · 59566b6e

由 Zhi Yong Wu 提交于 12月 07, 2013

Since vhost_dev_init() forever return 0, some branches are never run,
therefore need to be removed.
Signed-off-by: NZhi Yong Wu <wuzhy@linux.vnet.ibm.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

59566b6e

11 7月, 2013 1 次提交

vhost: Remove custom vhost rcu usage · 22fa90c7

由 Asias He 提交于 5月 07, 2013

Now, vq->private_data is always accessed under vq mutex. No need to play
the vhost rcu trick.
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

22fa90c7

07 7月, 2013 1 次提交

vhost: Make vhost a separate module · 6ac1afbf

由 Asias He 提交于 5月 06, 2013

Currently, vhost-net and vhost-scsi are sharing the vhost core code.
However, vhost-scsi shares the code by including the vhost.c file
directly.

Making vhost a separate module makes it is easier to share code with
other vhost devices.
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

6ac1afbf

11 6月, 2013 1 次提交

vhost: check owner before we overwrite ubuf_info · 05c05351

由 Michael S. Tsirkin 提交于 6月 06, 2013

If device has an owner, we shouldn't touch ubuf_info
since it might be in use.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

05c05351

06 5月, 2013 4 次提交

vhost: Remove vhost_enable_zcopy in vhost.h · e40ab748

由 Asias He 提交于 5月 06, 2013

It is net.c specific.
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

e40ab748

vhost: Remove comments for hdr in vhost.h · ab00c42a

由 Asias He 提交于 5月 06, 2013

It is supposed to be removed when hdr is moved into vhost_net_virtqueue.
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

ab00c42a

vhost: Move VHOST_NET_FEATURES to net.c · 8570a6e7

由 Asias He 提交于 5月 06, 2013

vhost.h should not depend on device specific marcos like
VHOST_NET_F_VIRTIO_NET_HDR and VIRTIO_NET_F_MRG_RXBUF.
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

8570a6e7

vhost: Export vhost_dev_set_owner · 54db63c2

由 Asias He 提交于 5月 06, 2013

Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

54db63c2

01 5月, 2013 4 次提交

vhost: fix error handling in RESET_OWNER ioctl · 150b9e51

由 Michael S. Tsirkin 提交于 4月 28, 2013

RESET_OWNER ioctl would leave the fd in a bad state if
memory allocation failed: device is stopped
but owner is not reset. Make state changes
after allocating memory, such that a failed
ioctl has no effect.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

150b9e51

vhost: move per-vq net specific fields out to net · 81f95a55

由 Michael S. Tsirkin 提交于 4月 28, 2013

This will remove the need for vhost scsi to pull
in virtio-net.h.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

81f95a55

vhost: move vhost-net zerocopy fields to net.c · 2839400f

由 Asias He 提交于 4月 27, 2013

On top of 'vhost: Allow device specific fields per vq', we can move device
specific fields to device virt queue from vhost virt queue.
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

2839400f

vhost: Allow device specific fields per vq · 3ab2e420

由 Asias He 提交于 4月 27, 2013

This is useful for any device who wants device specific fields per vq.
For example, tcm_vhost wants a per vq field to track requests which are
in flight on the vq. Also, on top of this we can add patches to move
things like ubufs from vhost.h out to net.c.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NAsias He <asias@redhat.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

3ab2e420

30 1月, 2013 1 次提交

vhost_net: handle polling errors when setting backend · 2b8b328b

由 Jason Wang 提交于 1月 28, 2013

Currently, the polling errors were ignored, which can lead following issues:

- vhost remove itself unconditionally from waitqueue when stopping the poll,
  this may crash the kernel since the previous attempt of starting may fail to
  add itself to the waitqueue
- userspace may think the backend were successfully set even when the polling
  failed.

Solve this by:

- check poll->wqh before trying to remove from waitqueue
- report polling errors in vhost_poll_start(), tx_poll_start(), the return value
  will be checked and returned when userspace want to set the backend

After this fix, there still could be a polling failure after backend is set, it
will addressed by the next patch.
Signed-off-by: NJason Wang <jasowang@redhat.com>
Acked-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

2b8b328b

06 12月, 2012 1 次提交

vhost: avoid backend flush on vring ops · 935cdee7

由 Michael S. Tsirkin 提交于 12月 06, 2012

vring changes already do a flush internally where appropriate, so we do
not need a second flush.

It's currently not very expensive but a follow-up patch makes flush more
heavy-weight, so remove the extra flush here to avoid regressing
performance if call or kick fds are changed on data path.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

935cdee7

03 11月, 2012 4 次提交

vhost: move -net specific code out · b211616d

由 Michael S. Tsirkin 提交于 11月 01, 2012

Zerocopy handling code is vhost-net specific.
Move it from vhost.c/vhost.h out to net.c
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

b211616d

vhost: track zero copy failures using DMA length · c4fcb586

由 Michael S. Tsirkin 提交于 11月 01, 2012

This will be used to disable zerocopy when error rate
is high.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

c4fcb586

vhost-net: cleanup macros for DMA status tracking · 70e4cb9a

由 Michael S. Tsirkin 提交于 11月 01, 2012

Better document macros for DMA tracking. Add an
explicit one for DMA in progress instead of
relying on user supplying len != 1.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

70e4cb9a

skb: report completion status for zero copy skbs · e19d6763

由 Michael S. Tsirkin 提交于 11月 01, 2012

Even if skb is marked for zero copy, net core might still decide
to copy it later which is somewhat slower than a copy in user context:
besides copying the data we need to pin/unpin the pages.

Add a parameter reporting such cases through zero copy callback:
if this happens a lot, device can take this into account
and switch to copying in user context.

This patch updates all users but ignores the passed value for now:
it will be used by follow-up patches.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

e19d6763

22 7月, 2012 2 次提交

vhost: make vhost work queue visible · 163049ae

由 Stefan Hajnoczi 提交于 7月 21, 2012

The vhost work queue allows processing to be done in vhost worker thread
context, which uses the owner process mm.  Access to the vring and guest
memory is typically only possible from vhost worker context so it is
useful to allow work to be queued directly by users.

Currently vhost_net only uses the poll wrappers which do not expose the
work queue functions.  However, for tcm_vhost (vhost_scsi) it will be
necessary to queue custom work.
Signed-off-by: NStefan Hajnoczi <stefanha@linux.vnet.ibm.com>
Cc: Zhi Yong Wu <wuzhy@cn.ibm.com>
Cc: Michael S. Tsirkin <mst@redhat.com>
Cc: Paolo Bonzini <pbonzini@redhat.com>
Signed-off-by: NNicholas Bellinger <nab@linux-iscsi.org>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

163049ae

vhost: Separate vhost-net features from vhost features · 0dd05a3b

由 Stefan Hajnoczi 提交于 7月 21, 2012

In order for other vhost devices to use the VHOST_FEATURES bits the
vhost-net specific bits need to be moved to their own VHOST_NET_FEATURES
constant.

(Asias: Update drivers/vhost/test.c to use VHOST_NET_FEATURES)
Signed-off-by: NStefan Hajnoczi <stefanha@linux.vnet.ibm.com>
Cc: Zhi Yong Wu <wuzhy@cn.ibm.com>
Cc: Michael S. Tsirkin <mst@redhat.com>
Cc: Paolo Bonzini <pbonzini@redhat.com>
Cc: Asias He <asias@redhat.com>
Signed-off-by: NNicholas A. Bellinger <nab@risingtidesystems.com>
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>

0dd05a3b

14 4月, 2012 1 次提交

skbuff: struct ubuf_info callback type safety · ca8f4fb2

由 Michael S. Tsirkin 提交于 4月 09, 2012

The skb struct ubuf_info callback gets passed struct ubuf_info
itself, not the arg value as the field name and the function signature
seem to imply. Rename the arg field to ctx to match usage,
add documentation and change the callback argument type
to make usage clear and to have compiler check correctness.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Signed-off-by: NDavid S. Miller <davem@davemloft.net>

ca8f4fb2

28 2月, 2012 1 次提交

vhost: fix release path lockdep checks · ea5d4046

由 Michael S. Tsirkin 提交于 11月 27, 2011

We shouldn't hold any locks on release path. Pass a flag to
vhost_dev_cleanup to use the lockdep info correctly.
Signed-off-by: NMichael S. Tsirkin <mst@redhat.com>
Tested-by: NSasha Levin <levinsasha928@gmail.com>

ea5d4046

27 7月, 2011 1 次提交

atomic: use <linux/atomic.h> · 60063497

由 Arun Sharma 提交于 7月 26, 2011

This allows us to move duplicated code in <asm/atomic.h>
(atomic_inc_not_zero() for now) to <linux/atomic.h>
Signed-off-by: NArun Sharma <asharma@fb.com>
Reviewed-by: NEric Dumazet <eric.dumazet@gmail.com>
Cc: Ingo Molnar <mingo@elte.hu>
Cc: David Miller <davem@davemloft.net>
Cc: Eric Dumazet <eric.dumazet@gmail.com>
Acked-by: NMike Frysinger <vapier@gentoo.org>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

60063497

openanolis / cloud-kernel 大约 1 年 前同步成功

openanolis / cloud-kernel
大约 1 年前同步成功