提交 · 793597a6a95675f4f85671cf747c1d92e7dbc295 · openeuler / raspberrypi-kernel

28 5月, 2014 7 次提交

blk-mq: do not use blk_mq_alloc_request_pinned in blk_mq_map_request · 793597a6

由 Christoph Hellwig 提交于 5月 27, 2014

We already do a non-blocking allocation in blk_mq_map_request, no need
to repeat it.  Just call __blk_mq_alloc_request to wait directly.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

793597a6

blk-mq: remove blk_mq_wait_for_tags · a3bd7756

由 Christoph Hellwig 提交于 5月 27, 2014

The current logic for blocking tag allocation is rather confusing, as we
first allocated and then free again a tag in blk_mq_wait_for_tags, just
to attempt a non-blocking allocation and then repeat if someone else
managed to grab the tag before us.

Instead change blk_mq_alloc_request_pinned to simply do a blocking tag
allocation itself and use the request we get back from it.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

a3bd7756

blk-mq: initialize request in __blk_mq_alloc_request · 5dee8577

由 Christoph Hellwig 提交于 5月 27, 2014

Both callers if __blk_mq_alloc_request want to initialize the request, so
lift it into the common path.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

5dee8577

blk-mq: merge blk_mq_alloc_reserved_request into blk_mq_alloc_request · 4ce01dd1

由 Christoph Hellwig 提交于 5月 27, 2014

Instead of having two almost identical copies of the same code just let
the callers pass in the reserved flag directly.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

4ce01dd1

blk-mq: add helper to insert requests from irq context · 6fca6a61

由 Christoph Hellwig 提交于 5月 28, 2014

Both the cache flush state machine and the SCSI midlayer want to submit
requests from irq context, and the current per-request requeue_work
unfortunately causes corruption due to sharing with the csd field for
flushes.  Replace them with a per-request_queue list of requests to
be requeued.

Based on an earlier test by Ming Lei.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Reported-by: NMing Lei <tom.leiming@gmail.com>
Tested-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

6fca6a61

blk-mq: allow non-softirq completions · 95f09684

由 Jens Axboe 提交于 5月 27, 2014

Right now we export two ways of completing a request:

1) blk_mq_complete_request(). This uses an IPI (if needed) and
   completes through q->softirq_done_fn(). It also works with
   timeouts.

2) blk_mq_end_io(). This completes inline, and ignores any timeout
   state of the request.

Let blk_mq_complete_request() handle non-softirq_done_fn completions
as well, by just completing inline. If a driver has enough completion
ports to place completions correctly, it need not define a
mq_ops->complete() and we can avoid an indirect function call by
doing the completion inline.
Signed-off-by: NJens Axboe <axboe@fb.com>

95f09684

blk-mq: pass in suggested NUMA node to ->alloc_hctx() · f14bbe77

由 Jens Axboe 提交于 5月 27, 2014

Drivers currently have to figure this out on their own, and they
are missing information to do it properly. The ones that did
attempt to do it, do it wrong.

So just pass in the suggested node directly to the alloc
function.
Signed-off-by: NJens Axboe <axboe@fb.com>

f14bbe77

27 5月, 2014 4 次提交

block: only allocate/free mq_usage_counter in blk-mq · 3d2936f4

由 Ming Lei 提交于 5月 27, 2014

The percpu counter is only used for blk-mq, so move
its allocation and free inside blk-mq, and don't
allocate it for legacy queue device.
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

3d2936f4

blk-mq: avoid code duplication · 624dbe47

由 Ming Lei 提交于 5月 27, 2014

blk_mq_exit_hw_queues() and blk_mq_free_hw_queues()
are introduced to avoid code duplication.
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

624dbe47

blk-mq: fix leak of hctx->ctx_map · 1f9f07e9

由 Ming Lei 提交于 5月 27, 2014

hctx->ctx_map should have been freed inside blk_mq_free_queue().
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

1f9f07e9

blk-mq: idle all hardware contexts before freeing a queue · 19c5d84f

由 Christoph Hellwig 提交于 5月 26, 2014

Without this we can leak the active_queues reference if a command is
freed while it is considered active.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

19c5d84f

24 5月, 2014 1 次提交

blk-mq: allow setting of per-request timeouts · c22d9d8a

由 Jens Axboe 提交于 5月 23, 2014

Currently blk-mq uses the queue timeout for all requests. But
for some commands, drivers may want to set a specific timeout
for special requests. Allow this to be passed in through
request->timeout, and use it if set.
Signed-off-by: NJens Axboe <axboe@fb.com>

c22d9d8a

23 5月, 2014 1 次提交

blk-mq: split make request handler for multi and single queue · 07068d5b

由 Jens Axboe 提交于 5月 22, 2014

We want slightly different behavior from them:

- On single queue devices, we currently use the per-process plug
  for deferred IO and for merging.

- On multi queue devices, we don't use the per-process plug, but
  we want to go straight to hardware for SYNC IO.

Split blk_mq_make_request() into a blk_sq_make_request() for single
queue devices, and retain blk_mq_make_request() for multi queue
devices. Then we don't need multiple checks for q->nr_hw_queues
in the request mapping.
Signed-off-by: NJens Axboe <axboe@fb.com>

07068d5b

22 5月, 2014 2 次提交

blk-mq: save memory by freeing requests on unused hardware queues · 484b4061

由 Jens Axboe 提交于 5月 21, 2014

Depending on the topology of the machine and the number of queues
exposed by a device, we can end up in a situation where some of
the hardware queues are unused (as in, they don't map to any
software queues). For this case, free up the memory used by the
request map, as we will not use it. This can be a substantial
amount of memory, depending on the number of queues vs CPUs and
the queue depth of the device.
Signed-off-by: NJens Axboe <axboe@fb.com>

484b4061

blk-mq: allow the hctx cpu hotplug notifier to return errors · e814e71b

由 Jens Axboe 提交于 5月 21, 2014

Prepare this for the next patch which adds more smarts in the
plugging logic, so that we can save some memory.
Signed-off-by: NJens Axboe <axboe@fb.com>

e814e71b

21 5月, 2014 3 次提交

blk-mq: Micro-optimize blk_queue_nomerges() check · da41a589

由 Robert Elliott 提交于 5月 20, 2014

In blk_mq_make_request(), do the blk_queue_nomerges() check
outside the call to blk_attempt_plug_merge() to eliminate
function call overhead when nomerges=2 (disabled)
Signed-off-by: NRobert Elliott <elliott@hp.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

da41a589

blk-mq: initialize q->nr_requests after calling blk_queue_make_request() · eba71768

由 Jens Axboe 提交于 5月 20, 2014

blk_queue_make_requests() overwrites our set value for q->nr_requests,
turning it into the default of 128. Set this appropriately after
initializing queue values in blk_queue_make_request().
Signed-off-by: NJens Axboe <axboe@fb.com>

eba71768

blk-mq: allow changing of queue depth through sysfs · e3a2b3f9

由 Jens Axboe 提交于 5月 20, 2014

For request_fn based devices, the block layer exports a 'nr_requests'
file through sysfs to allow adjusting of queue depth on the fly.
Currently this returns -EINVAL for blk-mq, since it's not wired up.
Wire this up for blk-mq, so that it now also always dynamic
adjustments of the allowed queue depth for any given block device
managed by blk-mq.
Signed-off-by: NJens Axboe <axboe@fb.com>

e3a2b3f9

20 5月, 2014 1 次提交

blk-mq: switch ctx pending map to the sparser blk_align_bitmap · 1429d7c9

由 Jens Axboe 提交于 5月 19, 2014

Each hardware queue has a bitmap of software queues with pending
requests. When new IO is queued on a software queue, the bit is
set, and when IO is pruned on a hardware queue run, the bit is
cleared. This causes a lot of traffic. Switch this from the regular
BITS_PER_LONG bitmap to a sparser layout, similarly to what was
done for blk-mq tagging.

20% performance increase was observed for single threaded IO, and
about 15% performanc increase on multiple threads driving the
same device.
Signed-off-by: NJens Axboe <axboe@fb.com>

1429d7c9

14 5月, 2014 1 次提交

blk-mq: improve support for shared tags maps · 0d2602ca

由 Jens Axboe 提交于 5月 13, 2014

This adds support for active queue tracking, meaning that the
blk-mq tagging maintains a count of active users of a tag set.
This allows us to maintain a notion of fairness between users,
so that we can distribute the tag depth evenly without starving
some users while allowing others to try unfair deep queues.

If sharing of a tag set is detected, each hardware queue will
track the depth of its own queue. And if this exceeds the total
depth divided by the number of active queues, the user is actively
throttled down.

The active queue count is done lazily to avoid bouncing that data
between submitter and completer. Each hardware queue gets marked
active when it allocates its first tag, and gets marked inactive
when 1) the last tag is cleared, and 2) the queue timeout grace
period has passed.
Signed-off-by: NJens Axboe <axboe@fb.com>

0d2602ca

10 5月, 2014 1 次提交

blk-mq: fix race in IO start accounting · cf4b50af

由 Jens Axboe 提交于 5月 09, 2014

Commit c6d600c6 opened up a small race where we could attempt to
account IO completion on a request, racing with IO start accounting.
Fix this up by ensuring that we've accounted for IO start before
inserting the request.
Signed-off-by: NJens Axboe <axboe@fb.com>

cf4b50af

09 5月, 2014 3 次提交

blk-mq: implement new and more efficient tagging scheme · 4bb659b1

由 Jens Axboe 提交于 5月 09, 2014

blk-mq currently uses percpu_ida for tag allocation. But that only
works well if the ratio between tag space and number of CPUs is
sufficiently high. For most devices and systems, that is not the
case. The end result if that we either only utilize the tag space
partially, or we end up attempting to fully exhaust it and run
into lots of lock contention with stealing between CPUs. This is
not optimal.

This new tagging scheme is a hybrid bitmap allocator. It uses
two tricks to both be SMP friendly and allow full exhaustion
of the space:

1) We cache the last allocated (or freed) tag on a per blk-mq
   software context basis. This allows us to limit the space
   we have to search. The key element here is not caching it
   in the shared tag structure, otherwise we end up dirtying
   more shared cache lines on each allocate/free operation.

2) The tag space is split into cache line sized groups, and
   each context will start off randomly in that space. Even up
   to full utilization of the space, this divides the tag users
   efficiently into cache line groups, avoiding dirtying the same
   one both between allocators and between allocator and freeer.

This scheme shows drastically better behaviour, both on small
tag spaces but on large ones as well. It has been tested extensively
to show better performance for all the cases blk-mq cares about.
Signed-off-by: NJens Axboe <axboe@fb.com>

4bb659b1

blk-mq: initialize struct request fields individually · af76e555

由 Christoph Hellwig 提交于 5月 06, 2014

This allows us to avoid a non-atomic memset over ->atomic_flags as well
as killing lots of duplicate initializations.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

af76e555

J
blk-mq: update a hotplug comment for grammar · 9fccfed8
由 Jens Axboe 提交于 5月 08, 2014
```
Signed-off-by: NJens Axboe <axboe@fb.com>
```
9fccfed8

08 5月, 2014 1 次提交

blk-mq: add basic round-robin of what CPU to queue workqueue work on · 506e931f

由 Jens Axboe 提交于 5月 07, 2014

Right now we just pick the first CPU in the mask, but that can
easily overload that one. Add some basic batching and round-robin
all the entries in the mask instead.
Signed-off-by: NJens Axboe <axboe@fb.com>

506e931f

03 5月, 2014 1 次提交

blk-mq: remove extra requeue trace · 74814b1c

由 Jens Axboe 提交于 5月 02, 2014

We already issue a blktrace requeue event in
__blk_mq_requeue_request(), don't do it from the original caller
as well.
Signed-off-by: NJens Axboe <axboe@fb.com>

74814b1c

01 5月, 2014 2 次提交

blk-mq: refactor request insertion/merging · c6d600c6

由 Jens Axboe 提交于 4月 30, 2014

Refactor the logic around adding a new bio to a software queue,
so we nest the ctx->lock where we really need it (merge and
insertion) and don't hold it when we don't (init and IO start
accounting).
Signed-off-by: NJens Axboe <axboe@fb.com>

c6d600c6

J
blk-mq remove debug BUG_ON() when draining software queues · 98bc1f27
由 Jens Axboe 提交于 4月 30, 2014
```
It's never been of any use, lets get rid of it.
Signed-off-by: NJens Axboe <axboe@fb.com>
```
98bc1f27

30 4月, 2014 1 次提交

blk-mq: fix waiting for reserved tags · 5810d903

由 Jens Axboe 提交于 4月 29, 2014

blk_mq_wait_for_tags() is only able to wait for "normal" tags,
not reserved tags. Pass in which one we should attempt to get
a tag for, so that waiting for reserved tags will work.

Reserved tags are used for internal commands, which are usually
serialized. Hence no waiting generally takes place, but we should
ensure that it actually works if users need that functionality.
Signed-off-by: NJens Axboe <axboe@fb.com>

5810d903

25 4月, 2014 1 次提交

blk-mq: respect rq_affinity · 38535201

由 Christoph Hellwig 提交于 4月 25, 2014

The blk-mq code is using it's own version of the I/O completion affinity
tunables, which causes a few issues:

 - the rq_affinity sysfs file doesn't work for blk-mq devices, even if it
   still is present, thus breaking existing tuning setups.
 - the rq_affinity = 1 mode, which is the defauly for legacy request based
   drivers isn't implemented at all.
 - blk-mq drivers don't implement any completion affinity with the default
   flag settings.

This patches removes the blk-mq ipi_redirect flag and sysfs file, as well
as the internal BLK_MQ_F_SHOULD_IPI flag and replaces it with code that
respects the queue-wide rq_affinity flags and also implements the
rq_affinity = 1 mode.

This means I/O completion affinity can now only be tuned block-queue wide
instead of per context, which seems more sensible to me anyway.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

38535201

24 4月, 2014 3 次提交

blk-mq: fix race with timeouts and requeue events · 87ee7b11

由 Jens Axboe 提交于 4月 24, 2014

If a requeue event races with a timeout, we can get into the
situation where we attempt to complete a request from the
timeout handler when it's not start anymore. This causes a crash.
So have the timeout handler check that REQ_ATOM_STARTED is still
set on the request - if not, we ignore the event. If this happens,
the request has now been marked as complete. As a consequence, we
need to ensure to clear REQ_ATOM_COMPLETE in blk_mq_start_request(),
as to maintain proper request state.
Signed-off-by: NJens Axboe <axboe@fb.com>

87ee7b11

Revert "blk-mq: initialize req->q in allocation" · 70ab0b2d

由 Jens Axboe 提交于 4月 24, 2014

This reverts commit 6a3c8a3a.

We need selective clearing of the request to make the init-at-free
time completely safe. Otherwise we end up stomping on
rq->atomic_flags, which we don't want to do.

70ab0b2d

blk-mq: fix leak of set->tags · 981bd189

由 Ming Lei 提交于 4月 24, 2014

set->tags should be freed in blk_mq_free_tag_set().
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

981bd189

22 4月, 2014 4 次提交

blk-mq: initialize req->q in allocation · 6a3c8a3a

由 Ming Lei 提交于 4月 19, 2014

The patch basically reverts the patch of(blk-mq:
initialize request on allocation) in Jens's tree(already
in -next), and only initialize req->q in allocation
for two reasons:

	- presumed cache hotness on completion
	- blk_rq_tagged(rq) depends on reset of req->mq_ctx
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

6a3c8a3a

blk-mq: user (1 << order) to implement order_to_size() · 4ca08500

由 Ming Lei 提交于 4月 19, 2014

Cc: Jörg-Volker Peetz <jvpeetz@web.de>
Cc: Max Filippov <jcmvbkbc@gmail.com>
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

4ca08500

blk-mq: fix allocation of set->tags · 48479005

由 Ming Lei 提交于 4月 19, 2014

type of set->tags is struct blk_mq_tags **.
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

48479005

blk-mq: free hctx->ctx_map when init failed · 11471e0d

由 Ming Lei 提交于 4月 19, 2014

Avoid memory leak in the failure path.
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NMing Lei <tom.leiming@gmail.com>
Signed-off-by: NJens Axboe <axboe@fb.com>

11471e0d

17 4月, 2014 3 次提交

blk-mq: add blk_mq_requeue_request · ed0791b2

由 Christoph Hellwig 提交于 4月 16, 2014

This allows to requeue a request that has been accepted by ->queue_rq
earlier.  This is needed by the SCSI layer in various error conditions.

The existing internal blk_mq_requeue_request is renamed to
__blk_mq_requeue_request as it is a lower level building block for this
funtionality.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

ed0791b2

blk-mq: add blk_mq_start_hw_queues · 2f268556

由 Christoph Hellwig 提交于 4月 16, 2014

Add a helper to unconditionally kick contexts of a queue.  This will
be needed by the SCSI layer to provide fair queueing between multiple
devices on a single host.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <axboe@fb.com>

2f268556

blk-mq: add blk_mq_delay_queue · 70f4db63

由 Christoph Hellwig 提交于 4月 16, 2014

Add a blk-mq equivalent to blk_delay_queue so that the scsi layer can ask
to be kicked again after a delay.
Signed-off-by: NChristoph Hellwig <hch@lst.de>

Modified by me to kill the unnecessary preempt disable/enable
in the delayed workqueue handler.
Signed-off-by: NJens Axboe <axboe@fb.com>

70f4db63