提交 · 8418f22a53795f4478a302aaec3d056795f56089 · openeuler / Kernel

12 4月, 2021 40 次提交

由 Pavel Begunkov 提交于 3月 22, 2021

Extract a helper for io_work_get_acct() and io_wqe_get_acct() to avoid
duplication.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

8418f22a

io_uring: remove tctx->sqpoll · 05356d86

由 Pavel Begunkov 提交于 3月 22, 2021

struct io_uring_task::sqpoll is not used anymore, kill it
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

05356d86

io_uring: don't do extra EXITING cancellations · 68207680

由 Pavel Begunkov 提交于 3月 22, 2021

io_match_task() matches all requests with PF_EXITING task, even though
those may be valid requests. It was necessary for SQPOLL cancellation,
but now it kills all requests before exiting via
io_uring_cancel_sqpoll(), so it's not needed.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

68207680

io_uring: don't clear REQ_F_LINK_TIMEOUT · d4729fbd

由 Pavel Begunkov 提交于 3月 22, 2021

REQ_F_LINK_TIMEOUT is a hint that to look for linked timeouts to cancel,
we're leaving it even when it's already fired. Hence don't care to clear
it in io_kill_linked_timeout(), it's safe and is called only once.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

d4729fbd

io_uring: optimise io_req_task_work_add() · c15b79de

由 Pavel Begunkov 提交于 3月 19, 2021

Inline io_task_work_add() into io_req_task_work_add(). They both work
with a request, so keeping them separate doesn't make things much more
clear, but merging allows optimise it. Apart from small wins like not
reading req->ctx or not calculating @notify in the hot path, i.e. with
tctx->task_state set, it avoids doing wake_up_process() for every single
add, but only after actually done task_work_add().
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

c15b79de

io_uring: abolish old io_put_file() · e1d767f0

由 Pavel Begunkov 提交于 3月 19, 2021

io_put_file() doesn't do a good job at generating a good code. Inline
it, so we can check REQ_F_FIXED_FILE first, prioritising FIXED_FILE case
over requests without files, and saving a memory load in that case.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

e1d767f0

io_uring: optimise io_dismantle_req() fast path · 094bae49

由 Pavel Begunkov 提交于 3月 19, 2021

Reshuffle io_dismantle_req() checks to put most of slow path stuff under
a single if.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

094bae49

io_uring: inline io_clean_op()'s fast path · 68fb8979

由 Pavel Begunkov 提交于 3月 19, 2021

Inline io_clean_op(), leaving __io_clean_op() but renaming it. This will
be used in following patches.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

68fb8979

io_uring: remove __io_req_task_cancel() · 2593553a

由 Pavel Begunkov 提交于 3月 19, 2021

Both io_req_complete_failed() and __io_req_task_cancel() do the same
thing: set failure flag, put both req refs and emit an CQE. The former
one is a bit more advance as it puts req back into a req cache, so make
it to take over __io_req_task_cancel() and remove the last one.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

2593553a

io_uring: add helper flushing locked_free_list · dac7a098

由 Pavel Begunkov 提交于 3月 19, 2021

Add a new helper io_flush_cached_locked_reqs() that splices
locked_free_list to free_list, and does it right doing all sync and
invariant reinit.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

dac7a098

io_uring: refactor io_free_req_deferred() · a05432fb

由 Pavel Begunkov 提交于 3月 19, 2021

We don't care about ret value in io_free_req_deferred(), make the code a
bit more concise.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

a05432fb

io_uring: inline io_put_req and friends · 0d85035a

由 Pavel Begunkov 提交于 3月 19, 2021

One big omission is that io_put_req() haven't been marked inline, and at
least gcc 9 doesn't inline it, not to mention that it's really hot and
extra function call is intolerable, especially when it doesn't put a
final ref.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

0d85035a

io_uring: refactor rsrc refnode allocation · 8dd03afe

由 Pavel Begunkov 提交于 3月 19, 2021

There are two problems:
1) we always allocate refnodes in advance and free them if those
haven't been used. It's expensive, takes two allocations, where one of
them is percpu. And it may be pretty common not actually using them.

2) Current API with allocating a refnode and setting some of the fields
is error prone, we don't ever want to have a file node runninng fixed
buffer callback...

Solve both with pre-init/get API. Pre-init just leaves the node for
later if not used, and for get (i.e. io_rsrc_refnode_get()), you need to
explicitly pass all arguments setting callbacks/etc., so it's more
resilient.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

8dd03afe

io_uring: refactor io_flush_cached_reqs() · dd78f492

由 Pavel Begunkov 提交于 3月 19, 2021

Emphasize that return value of io_flush_cached_reqs() depends on number
of requests in the cache. It looks nicer and might help tools from
false-negative analyses.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

dd78f492

io_uring: optimise success case of __io_queue_sqe · 1840038e

由 Pavel Begunkov 提交于 3月 19, 2021

Move the case of successfully issued request by doing that check first.
It's not much of a difference, just generates slightly better code for
me.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

1840038e

io_uring: inline __io_queue_linked_timeout() · de968c18

由 Pavel Begunkov 提交于 3月 19, 2021

Inline __io_queue_linked_timeout(), we don't need it
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

de968c18

io_uring: keep io_req_free_batch() call locality · 96670657

由 Pavel Begunkov 提交于 3月 19, 2021

Don't do a function call (io_dismantle_req()) in the middle and place it
to near other function calls, otherwise may lead to excessive register
spilling.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

96670657

io_uring: optimise tctx node checks/alloc · cf27f3b1

由 Pavel Begunkov 提交于 3月 19, 2021

First of all, w need to set tctx->sqpoll only when we add a new entry
into ->xa, so move it from the hot path. Also extract a hot path for
io_uring_add_task_file() as an inline helper.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

cf27f3b1

io_uring: optimise io_uring_enter() · 33f993da

由 Pavel Begunkov 提交于 3月 19, 2021

Add unlikely annotations, because my compiler pretty much mispredicts
every first check, and apart jumping around in the fast path, it also
generates extra instructions, like in advance setting ret value.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

33f993da

io_uring: don't take ctx refs in task_work handler · 493f3b15

由 Pavel Begunkov 提交于 3月 19, 2021

__tctx_task_work() guarantees that ctx won't be killed while running
task_works, so we can remove now unnecessary ctx pinning for internally
armed polling.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

493f3b15

io_uring: transform ret == 0 for poll cancelation completions · 45ab03b1

由 Jens Axboe 提交于 2月 23, 2021

We can set canceled == true and complete out-of-line, ensure that we catch
that and correctly return -ECANCELED if the poll operation got canceled.
Signed-off-by: NJens Axboe <axboe@kernel.dk>

45ab03b1

io_uring: correct comment on poll vs iopoll · b9b0e0d3

由 Jens Axboe 提交于 2月 23, 2021

The correct function is io_iopoll_complete(), which deals with completions
of IOPOLL requests, not io_poll_complete().
Signed-off-by: NJens Axboe <axboe@kernel.dk>

b9b0e0d3

io_uring: cache async and regular file state for fixed files · 7b29f92d

由 Jens Axboe 提交于 3月 12, 2021

We have to dig quite deep to check for particularly whether or not a
file supports a fast-path nonblock attempt. For fixed files, we can do
this lookup once and cache the state instead.

This adds two new bits to track whether we support async read/write
attempt, and lines up the REQ_F_ISREG bit with those two. The file slot
re-uses the last 3 (or 2, for 32-bit) of the file pointer to cache that
state, and then we mask it in when we go and use a fixed file.
Signed-off-by: NJens Axboe <axboe@kernel.dk>

7b29f92d

io_uring: don't check for io_uring_fops for fixed files · d44f554e

由 Jens Axboe 提交于 3月 12, 2021

We don't allow them at registration time, so limit the check for needing
inflight tracking in io_file_get() to the non-fixed path.
Signed-off-by: NJens Axboe <axboe@kernel.dk>

d44f554e

io_uring: simplify io_sqd_update_thread_idle() · c9dca27d

由 Pavel Begunkov 提交于 3月 10, 2021

Use a more comprehensible() max instead of hand coding it with ifs in
io_sqd_update_thread_idle().
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

c9dca27d

io_uring: switch to atomic_t for io_kiocb reference count · abc54d63

由 Jens Axboe 提交于 2月 24, 2021

io_uring manipulates references twice for each request, and hence is very
sensitive to performance of the reference count. This commit borrows a
trick from:

commit f958d7b5
Author: Linus Torvalds <torvalds@linux-foundation.org>
Date:   Thu Apr 11 10:06:20 2019 -0700

    mm: make page ref count overflow check tighter and more explicit

and switches to atomic_t for references, while still retaining overflow
and underflow checks.

This is good for a 2-3% increase in peak IOPS on a single core. Before:

IOPS=2970879, IOS/call=31/31, inflight=128 (128)
IOPS=2952597, IOS/call=31/31, inflight=128 (128)
IOPS=2943904, IOS/call=31/31, inflight=128 (128)
IOPS=2930006, IOS/call=31/31, inflight=96 (96)

and after:

IOPS=3054354f, IOS/call=31/31, inflight=128 (128)
IOPS=3059038, IOS/call=31/31, inflight=128 (128)
IOPS=3060320, IOS/call=31/31, inflight=128 (128)
IOPS=3068256, IOS/call=31/31, inflight=96 (96)
Signed-off-by: NJens Axboe <axboe@kernel.dk>

abc54d63

io_uring: wrap io_kiocb reference count manipulation in helpers · de9b4cca

由 Jens Axboe 提交于 2月 24, 2021

No functional changes in this patch, just in preparation for handling the
references a bit more efficiently.
Signed-off-by: NJens Axboe <axboe@kernel.dk>

de9b4cca

io_uring: simplify io_resubmit_prep() · 179ae0d1

由 Pavel Begunkov 提交于 2月 28, 2021

If not for async_data NULL check, io_resubmit_prep() is already an rw
specific version of io_req_prep_async(), but slower because 1) it always
goes through io_import_iovec() even if following io_setup_async_rw() the
result 2) instead of initialising iovec/iter in-place it does it
on-stack and then copies with io_setup_async_rw().
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

179ae0d1

io_uring: merge defer_prep() and prep_async() · b7e298d2

由 Pavel Begunkov 提交于 2月 28, 2021

Merge two function and do renaming in favour of the second one, it
relays the meaning better.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

b7e298d2

io_uring: rethink def->needs_async_data · 26f0505a

由 Pavel Begunkov 提交于 2月 28, 2021

needs_async_data controls allocation of async_data, and used in two
cases. 1) when async setup requires it (by io_req_prep_async() or
handler themselves), and 2) when op always needs additional space to
operate, like timeouts do.

Opcode preps already don't bother about the second case and do
allocation unconditionally, restrict needs_async_data to the first case
only and rename it into needs_async_setup.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
[axboe: update for IOPOLL fix]
Signed-off-by: NJens Axboe <axboe@kernel.dk>

26f0505a

io_uring: untie alloc_async_data and needs_async_data · 6cb78689

由 Pavel Begunkov 提交于 2月 28, 2021

All opcode handlers pretty well know whether they need async data or
not, and can skip testing for needs_async_data. The exception is rw
the generic path, but those test the flag by hand anyway. So, check the
flag and make io_alloc_async_data() allocating unconditionally.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

6cb78689

io_uring: refactor out send/recv async setup · 2e052d44

由 Pavel Begunkov 提交于 2月 28, 2021

IORING_OP_[SEND,RECV] don't need async setup neither will get into
io_req_prep_async(). Remove them from io_req_prep_async() and remove
needs_async_data checks from the related setup functions.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

2e052d44

io_uring: use better types for cflags · 8c3f9cd1

由 Pavel Begunkov 提交于 2月 28, 2021

__io_cqring_fill_event() takes cflags as long to squeeze it into u32 in
an CQE, awhile all users pass int or unsigned. Replace it with unsigned
int and store it as u32 in struct io_completion to match CQE.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

8c3f9cd1

io_uring: refactor provide/remove buffer locking · 9fb8cb49

由 Pavel Begunkov 提交于 2月 28, 2021

Always complete request holding the mutex instead of doing that strange
dancing with conditional ordering.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

9fb8cb49

io_uring: add a helper failing not issued requests · f41db273

由 Pavel Begunkov 提交于 2月 28, 2021

Add a simple helper doing CQE posting, marking request for link-failure,
and putting both submission and completion references.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

f41db273

io_uring: further deduplicate file slot selection · dafecf19

由 Pavel Begunkov 提交于 2月 28, 2021

io_fixed_file_slot() and io_file_from_index() behave pretty similarly,
DRY and call one from another.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

dafecf19

io_uring: reuse io_req_task_queue_fail() · 2c4b8eb6

由 Pavel Begunkov 提交于 2月 28, 2021

Use io_req_task_queue_fail() on the fail path of io_req_task_queue().
It's unlikely to happen, so don't care about additional overhead, but
allows to keep all the req->result invariant in a single function.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

2c4b8eb6

io_uring: avoid taking ctx refs for task-cancel · e83acd7d

由 Pavel Begunkov 提交于 2月 28, 2021

Don't bother to take a ctx->refs for io_req_task_cancel() because it
take uring_lock before putting a request, and the context is promised to
stay alive until unlock happens.
Signed-off-by: NPavel Begunkov <asml.silence@gmail.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

e83acd7d

L

Linux 5.12-rc7 · d434405a
由 Linus Torvalds 提交于 4月 11, 2021

d434405a

Merge tag 'for-5.12-rc6-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux · 7d900724

由 Linus Torvalds 提交于 4月 11, 2021

Pull btrfs fix from David Sterba:
 "One more patch that we'd like to get to 5.12 before release.

  It's changing where and how the superblock is stored in the zoned
  mode. It is an on-disk format change but so far there are no
  implications for users as the proper mkfs support hasn't been merged
  and is waiting for the kernel side to settle.

  Until now, the superblocks were derived from the zone index, but zone
  size can differ per device. This is changed to be based on fixed
  offset values, to make it independent of the device zone size.

  The work on that got a bit delayed, we discussed the exact locations
  to support potential device sizes and usecases. (Partially delayed
  also due to my vacation.) Having that in the same release where the
  zoned mode is declared usable is highly desired, there are userspace
  projects that need to be updated to recognize the feature. Pushing
  that to the next release would make things harder to test"

* tag 'for-5.12-rc6-tag' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux:
  btrfs: zoned: move superblock logging zone location

7d900724

openeuler / Kernel 1 年多 前同步成功

openeuler / Kernel
1 年多前同步成功