提交 · b474c7ae4395ba684e85fde8f55c8cf44a39afaf · openeuler / Kernel

18 6月, 2014 3 次提交

blk-mq: bitmap tag: fix races in bt_get() function · 86fb5c56

This update fixes few issues in bt_get() function:

- list_empty(&wait.task_list) check is not protected;

- was_empty check is always true which results in *every* thread
  entering the loop resets bt_wait_state::wait_cnt counter rather
  than every bt->wake_cnt'th thread;

- 'bt_wait_state::wait_cnt' counter update is redundant, since
  it also gets reset in bt_clear_tag() function;

Cc: Christoph Hellwig <hch@infradead.org>
Cc: Ming Lei <tom.leiming@gmail.com>
Cc: Jens Axboe <axboe@kernel.dk>
Signed-off-by: NAlexander Gordeev <agordeev@redhat.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

86fb5c56

blk-mq: bitmap tag: fix race on blk_mq_bitmap_tags::wake_cnt · 2971c35f

由 Alexander Gordeev 提交于 10年前

This piece of code in bt_clear_tag() function is racy:

	bs = bt_wake_ptr(bt);
	if (bs && atomic_dec_and_test(&bs->wait_cnt)) {
		atomic_set(&bs->wait_cnt, bt->wake_cnt);
 		wake_up(&bs->wait);
	}

Since nothing prevents bt_wake_ptr() from returning the very
same 'bs' address on multiple CPUs, the following scenario is
possible:

    CPU1                                CPU2
    ----                                ----

0.  bs = bt_wake_ptr(bt);               bs = bt_wake_ptr(bt);
1.  atomic_dec_and_test(&bs->wait_cnt)
2.                                      atomic_dec_and_test(&bs->wait_cnt)
3.  atomic_set(&bs->wait_cnt, bt->wake_cnt);

If the decrement in [1] yields zero then for some amount of time
the decrement in [2] results in a negative/overflow value, which
is not expected. The follow-up assignment in [3] overwrites the
invalid value with the batch value (and likely prevents the issue
from being severe) which is still incorrect and should be a lesser.

Cc: Ming Lei <tom.leiming@gmail.com>
Cc: Jens Axboe <axboe@kernel.dk>
Signed-off-by: NAlexander Gordeev <agordeev@redhat.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

2971c35f

blk-mq: bitmap tag: fix races on shared ::wake_index fields · 8537b120

由 Alexander Gordeev 提交于 10年前

Fix racy updates of shared blk_mq_bitmap_tags::wake_index
and blk_mq_hw_ctx::wake_index fields.

Cc: Ming Lei <tom.leiming@gmail.com>
Signed-off-by: NAlexander Gordeev <agordeev@redhat.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

8537b120

14 6月, 2014 2 次提交

C
blk-mq: merge blk_mq_drain_queue and __blk_mq_drain_queue · 95ed0681
由 Christoph Hellwig 提交于 10年前
```
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>
```
95ed0681

blk-mq: properly drain stopped queues · 8f5280f4

由 Christoph Hellwig 提交于 10年前

If we need to drain a queue we need to run all queues, even if they
are marked stopped to make sure the driver has a chance to error out
on all queued requests.

This fixes surprise removal with scsi-mq.
Reported-by: NBart Van Assche <bvanassche@acm.org>
Tested-by: NBart Van Assche <bvanassche@acm.org>
Signed-off-by: NJens Axboe <axboe@fb.com>

8f5280f4

12 6月, 2014 2 次提交

block: remove WQ_POWER_EFFICIENT from kblockd · 28747fcd

由 Matias Bjørling 提交于 10年前

blk-mq issues async requests through kblockd. To issue a work request on
a specific CPU, kblockd_schedule_delayed_work_on is used. However, the
specific CPU choice may not be honored, if the power_efficient option
for workqueues is set. blk-mq requires that we have strict per-cpu
scheduling, so it wont work properly if kblockd is marked
POWER_EFFICIENT and power_efficient is set.

Remove the kblockd WQ_POWER_EFFICIENT flag to prevent this behavior.
This essentially reverts part of commit 695588f9, which added
the WQ_POWER_EFFICIENT marker to kblockd.
Signed-off-by: NMatias Bjørling <m@bjorling.me>
Signed-off-by: NJens Axboe <axboe@fb.com>

28747fcd

block: remove elv_abort_queue and blk_abort_flushes · 2940474a

由 Christoph Hellwig 提交于 10年前

elv_abort_queue has no callers, and blk_abort_flushes is only called by
elv_abort_queue.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

2940474a

11 6月, 2014 3 次提交

block: add __init to blkcg_policy_register · a2d445d4

由 Fabian Frederick 提交于 10年前

blkcg_policy_register is only called by
__init functions:

__init cfq_init
__init throtl_init

Cc: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: NFabian Frederick <fabf@skynet.be>
Signed-off-by: NJens Axboe <axboe@fb.com>

a2d445d4

block: add __init to elv_register · b5097e95

由 Fabian Frederick 提交于 10年前

elv_register is only called by elevator init functions:

__init cfq_init
__init deadline_init
__init noop_init

Cc: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: NFabian Frederick <fabf@skynet.be>
Signed-off-by: NJens Axboe <axboe@fb.com>

b5097e95

block: ensure that bio_add_page() always accepts a page for an empty bio · 58a4915a

由 Jens Axboe 提交于 10年前

With commit 762380ad added support for chunk sizes and no merging
across them, it broke the rule of always allowing adding of a single
page to an empty bio. So relax the restriction a bit to allow for that,
similarly to what we have always done.

This fixes a crash with mkfs.xfs and 512b sector sizes on NVMe.
Reported-by: NKeith Busch <keith.busch@intel.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

58a4915a

10 6月, 2014 1 次提交

blk-mq: add timer in blk_mq_start_request · 2b8393b4

由 Ming Lei 提交于 10年前

This way will become consistent with non-mq case, also
avoid to update rq->deadline twice for mq.

The comment said: "We do this early, to ensure we are on
the right CPU.", but no percpu stuff is used in blk_add_timer(),
so it isn't necessary. Even when inserting from plug list, there
is no such guarantee at all.
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

2b8393b4

09 6月, 2014 2 次提交

blk-mq: always initialize request->start_time · 3ee32372

由 Jens Axboe 提交于 10年前

The blk-mq core only initializes this if io stats are enabled, since
blk-mq only reads the field in that case. But drivers could
potentially use it internally, so ensure that we always set it to
the current time when the request is allocated.
Reported-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

3ee32372

block: blk-exec.c: Cleaning up local variable address returnd · de83953f

由 Rickard Strandqvist 提交于 10年前

Address of local variable assigned to a function parameter

This was partly found using a static code analysis program called cppcheck.
Signed-off-by: NRickard Strandqvist <rickard_strandqvist@spectrumdigital.se>
Signed-off-by: NJens Axboe <axboe@fb.com>

de83953f

07 6月, 2014 3 次提交

mm: convert some level-less printks to pr_* · b1de0d13

由 Mitchel Humpherys 提交于 10年前

printk is meant to be used with an associated log level.  There are some
instances of printk scattered around the mm code where the log level is
missing.  Add a log level and adhere to suggestions by
scripts/checkpatch.pl by moving to the pr_* macros.

Also add the typical pr_fmt definition so that print statements can be
easily traced back to the modules where they occur, correlated one with
another, etc.  This will require the removal of some (now redundant)
prefixes on a few print statements.
Signed-off-by: NMitchel Humpherys <mitchelh@codeaurora.org>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

b1de0d13

blk-mq: ->timeout should be cleared in blk_mq_rq_ctx_init() · f6be4fb4

由 Jens Axboe 提交于 10年前

It'll be used in blk_mq_start_request() to set a potential timeout
for the request, so clear it to zero at alloc time to ensure that
we know if someone has set it or not.

Fixes random early timeouts on NVMe testing.
Signed-off-by: NJens Axboe <axboe@fb.com>

f6be4fb4

blk-mq: don't allow queue entering for a dying queue · 3b632cf0

由 Keith Busch 提交于 10年前

If the queue is going away, don't let new allocs or queueing
happen on it. Go through the normal wait process, and exit with
ENODEV in that case.
Signed-off-by: NKeith Busch <keith.busch@intel.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

3b632cf0

06 6月, 2014 3 次提交

blk-mq: bump max tag depth to 10K tags · a4391c64

由 Jens Axboe 提交于 10年前

For some scsi-mq cases, the tag map can be huge. So increase the
max number of tags we support.

Additionally, don't fail with EINVAL if a user requests too many
tags. Warn that the tag depth has been adjusted down, and store
the new value inside the tag_set passed in.
Signed-off-by: NJens Axboe <axboe@fb.com>

a4391c64

block: add blk_rq_set_block_pc() · f27b087b

由 Jens Axboe 提交于 10年前

With the optimizations around not clearing the full request at alloc
time, we are leaving some of the needed init for REQ_TYPE_BLOCK_PC
up to the user allocating the request.

Add a blk_rq_set_block_pc() that sets the command type to
REQ_TYPE_BLOCK_PC, and properly initializes the members associated
with this type of request. Update callers to use this function instead
of manipulating rq->cmd_type directly.

Includes fixes from Christoph Hellwig <hch@lst.de> for my half-assed
attempt.
Signed-off-by: NJens Axboe <axboe@fb.com>

f27b087b

block: add notion of a chunk size for request merging · 762380ad

由 Jens Axboe 提交于 10年前

Some drivers have different limits on what size a request should
optimally be, depending on the offset of the request. Similar to
dividing a device into chunks. Add a setting that allows the driver
to inform the block layer of such a chunk size. The block layer will
then prevent merging across the chunks.

This is needed to optimally support NVMe with a non-zero stripe size.
Signed-off-by: NJens Axboe <axboe@fb.com>

762380ad

05 6月, 2014 2 次提交

block: mq flush: clear flush_rq's tag in flush_end_io() · 14b83e17

由 Ming Lei 提交于 10年前

blk_mq_tag_to_rq() needs to be able to tell if it should return
the original request, or the flush request if we are doing a flush
sequence. Clear the flush tag when IO completes for a flush, since
that is what we are comparing against.
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

14b83e17

blk-mq: let blk_mq_tag_to_rq() take blk_mq_tags as the main parameter · 0e62f51f

由 Jens Axboe 提交于 10年前

We currently pass in the hardware queue, and get the tags from there.
But from scsi-mq, with a shared tag space, it's a lot more convenient
to pass in the blk_mq_tags instead as the hardware queue isn't always
directly available. So instead of having to re-map to a given
hardware queue from rq->mq_ctx, just pass in the tags structure.
Signed-off-by: NJens Axboe <axboe@fb.com>

0e62f51f

04 6月, 2014 5 次提交

blk-mq: fix regression from commit · f899fed4

由 Jens Axboe 提交于 10年前

When the code was collapsed to avoid duplication, the recent patch
for ensuring that a queue is idled before free was dropped, which was
added by commit 19c5d84f.

Add back the blk_mq_tag_idle(), to ensure we don't leak a reference
to an active queue when it is freed.
Signed-off-by: NJens Axboe <axboe@fb.com>

f899fed4

blk-mq: handle NULL req return from blk_map_request in single queue mode · ff87bcec

由 Jens Axboe 提交于 10年前

blk_mq_map_request() can return NULL if we fail entering the queue
(dying, or removed), in which case it has already ended IO on the
bio. So nothing more to do, except just return.
Signed-off-by: NJens Axboe <axboe@fb.com>

ff87bcec

blk-mq: fix sparse warning on missed __percpu annotation · e6cdb092

由 Ming Lei 提交于 10年前

'struct blk_mq_ctx' is  __percpu, so add the annotation
and fix the sparse warning reported from Fengguang:

	[block:for-linus 2/3] block/blk-mq.h:75:16: sparse: incorrect
	type in initializer (different address spaces)
Reported-by: Nkbuild test robot <fengguang.wu@intel.com>
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

e6cdb092

blk-mq: fix schedule from atomic context · cb96a42c

由 Ming Lei 提交于 10年前

blk_mq_put_ctx() has to be called before io_schedule() in
bt_get().

This patch fixes the problem by taking similar approach from
percpu_ida allocation for the situation.
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

cb96a42c

blk-mq: move blk_mq_get_ctx/blk_mq_put_ctx to mq private header · 1aecfe48

由 Ming Lei 提交于 10年前

The blk-mq tag code need these helpers.
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

1aecfe48

31 5月, 2014 4 次提交

blk-mq: push IPI or local end_io decision to __blk_mq_complete_request() · ed851860

由 Jens Axboe 提交于 10年前

We have callers outside of the blk-mq proper (like timeouts) that
want to call __blk_mq_complete_request(), so rename the function
and put the decision code for whether to use ->softirq_done_fn
or blk_mq_endio() into __blk_mq_complete_request().

This also makes the interface more logical again.
blk_mq_complete_request() attempts to atomically mark the request
completed, and calls __blk_mq_complete_request() if successful.
__blk_mq_complete_request() then just ends the request.
Signed-off-by: NJens Axboe <axboe@fb.com>

ed851860

blk-mq: remember to start timeout handler for direct queue · feff6894

由 Jens Axboe 提交于 10年前

Commit 07068d5b added a direct-to-hw-queue mode, but this mode
needs to remember to add the request timeout handler as well.
Without it, we don't track timeouts for these requests.
Signed-off-by: NJens Axboe <axboe@fb.com>

feff6894

block: ensure that the timer is always added · c7bca418

由 Jens Axboe 提交于 10年前

Commit f793aa53 relaxed the timer addition a little too much.
If the timer isn't pending, we always need to add it.
Signed-off-by: NJens Axboe <axboe@fb.com>

c7bca418

blk-mq: blk_mq_unregister_hctx() can be static · ee3c5db0

由 Fengguang Wu 提交于 10年前

CC: Jens Axboe <axboe@kernel.dk>
Signed-off-by: NFengguang Wu <fengguang.wu@intel.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

ee3c5db0

30 5月, 2014 4 次提交

blk-mq: make the sysfs mq/ layout reflect current mappings · 67aec14c

由 Jens Axboe 提交于 10年前

Currently blk-mq registers all the hardware queues in sysfs,
regardless of whether it uses them (e.g. they have CPU mappings)
or not. The unused hardware queues lack the cpux/ directories,
and the other sysfs entries (like active, pending, etc) are all
zeroes.

Change this so that sysfs correctly reflects the current mappings
of the hardware queues.
Signed-off-by: NJens Axboe <axboe@fb.com>

67aec14c

blk-mq: blk_mq_tag_to_rq should handle flush request · 22302375

由 Shaohua Li 提交于 10年前

flush request is special, which borrows the tag from the parent
request. Hence blk_mq_tag_to_rq needs special handling to return
the flush request from the tag.
Signed-off-by: NShaohua Li <shli@fusionio.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

22302375

block: remove dead code in scsi_ioctl:blk_verify_command · da52f22f

由 Dave Jones 提交于 10年前

filter gets assigned the address of blk_default_cmd_filter on
entry to this function, so the !filter condition can never be true.
Signed-off-by: NDave Jones <davej@redhat.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

da52f22f

blk-mq: request initialization optimizations · 4b570521

由 Jens Axboe 提交于 10年前

We currently clear a lot more than we need to, so make that a bit
more clever. Make some of the init dependent on features, like
only setting start_time if we are going to use it.
Signed-off-by: NJens Axboe <axboe@fb.com>

4b570521

29 5月, 2014 4 次提交

block: add queue flag for disabling SG merging · 05f1dd53

由 Jens Axboe 提交于 10年前

If devices are not SG starved, we waste a lot of time potentially
collapsing SG segments. Enough that 1.5% of the CPU time goes
to this, at only 400K IOPS. Add a queue flag, QUEUE_FLAG_NO_SG_MERGE,
which just returns the number of vectors in a bio instead of looping
over all segments and checking for collapsible ones.

Add a BLK_MQ_F_SG_MERGE flag so that drivers can opt-in on the sg
merging, if they so desire.
Signed-off-by: NJens Axboe <axboe@fb.com>

05f1dd53

block: remove 'magic' from struct blk_plug · 4d92a9be

由 Jens Axboe 提交于 10年前

I don't think we've ever caught any bugs with this, and there's the
list poisoning for the plug lists to catch uninitialized cases.
So remove the magic member and save 8 bytes in the struct.
Signed-off-by: NJens Axboe <axboe@fb.com>

4d92a9be

blk-mq: remove alloc_hctx and free_hctx methods · cdef54dd

由 Christoph Hellwig 提交于 10年前

There is no need for drivers to control hardware context allocation
now that we do the context to node mapping in common code.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

cdef54dd

blk-mq: add file comments and update copyright notices · 75bb4625

由 Jens Axboe 提交于 10年前

None of the blk-mq files have an explanatory comment at the top
for what that particular file does. Add that and add appropriate
copyright notices as well.
Signed-off-by: NJens Axboe <axboe@fb.com>

75bb4625

28 5月, 2014 2 次提交

blk-mq: remove blk_mq_alloc_request_pinned · d852564f

由 Christoph Hellwig 提交于 10年前

We now only have one caller left and can open code it there in a cleaner
way.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

d852564f

blk-mq: do not use blk_mq_alloc_request_pinned in blk_mq_map_request · 793597a6

由 Christoph Hellwig 提交于 10年前

We already do a non-blocking allocation in blk_mq_map_request, no need
to repeat it.  Just call __blk_mq_alloc_request to wait directly.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

793597a6

openeuler / Kernel 1 年多 前同步成功

openeuler / Kernel
1 年多前同步成功