提交 · fe0714377ee2ca161bf2afb7773e22f15f1786d4 · openeuler / raspberrypi-kernel

01 10月, 2010 6 次提交

blkio: Recalculate the throttled bio dispatch time upon throttle limit change · fe071437

由 Vivek Goyal 提交于 10月 01, 2010

o Currently any cgroup throttle limit changes are processed asynchronousy and
the change does not take affect till a new bio is dispatched from same group.

o It might happen that a user sets a redicuously low limit on throttling.
Say 1 bytes per second on reads. In such cases simple operations like mount
a disk can wait for a very long time.

o Once bio is throttled, there is no easy way to come out of that wait even if
user increases the read limit later.

o This patch fixes it. Now if a user changes the cgroup limits, we recalculate
the bio dispatch time according to new limits.

o Can't take queueu lock under blkcg_lock, hence after the change I wake
up the dispatch thread again which recalculates the time. So there are some
variables being synchronized across two threads without lock and I had to
make use of barriers. Hoping I have used barriers correctly. Any review of
memory barrier code especially will help.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

fe071437

blkio: Add root group to td->tg_list · 02977e4a

由 Vivek Goyal 提交于 10月 01, 2010

o Currently all the dynamically allocated groups, except root grp is added
  to td->tg_list. This was not a problem so far but in next patch I will
  travel through td->tg_list to process any updates of limits on the group.
  If root group is not in tg_list, then root group's updates are not
  processed.

o It is better to root group also to tg_list instead of doing special
  processing for it during limit updates.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

02977e4a

blkio: deletion of a cgroup was causes oops · 61014e96

由 Vivek Goyal 提交于 10月 01, 2010

o Now a cgroup list of blkg elements can contain blkg from multiple policies.
Before sending an unlink event, make sure blkg belongs to they policy. If
policy does not own the blkg, do not send update for this blkg.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

61014e96

blkio: Do not export throttle files if CONFIG_BLK_DEV_THROTTLING=n · 13f98250

由 Vivek Goyal 提交于 10月 01, 2010

Currently throttling related files were visible even if user had disabled
throttling using config options. It was switching off background throttling
of bio but not the cgroup files. This patch fixes it.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

13f98250

block: set the bounce_pfn to the actual DMA limit rather than to max memory · efb012b3

由 Malahal Naineni 提交于 10月 01, 2010

The bounce_pfn of the request queue in 64 bit systems is set to the
current max_low_pfn. Adding more memory later makes this incorrect.
Memory allocated beyond this boot time max_low_pfn appear to require
bounce buffers (bounce buffers are actually not allocated but used in
calculating segments that may result in "over max segments limit"
errors).
Signed-off-by: NMalahal Naineni <malahal@us.ibm.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

efb012b3

block: revert bad fix for memory hotplug causing bounces · 260a67a9

由 Jens Axboe 提交于 10月 01, 2010

Revert "block: set the bounce_pfn to the actual DMA limit rather than to max memory"

This reverts commit c49825fa.
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

260a67a9

25 9月, 2010 1 次提交

block: set the bounce_pfn to the actual DMA limit rather than to max memory · c49825fa

由 Malahal Naineni 提交于 9月 24, 2010

The bounce_pfn of the request queue in 64 bit systems is set to the
current max_low_pfn. Adding more memory later makes this incorrect.
Memory allocated beyond this boot time max_low_pfn appear to require
bounce buffers (bounce buffers are actually not allocated but used in
calculating segments that may result in "over max segments limit"
errors).
Signed-off-by: NMalahal Naineni <malahal@us.ibm.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

c49825fa

24 9月, 2010 1 次提交

block: Prevent hang_check firing during long I/O · 4b197769

由 Mark Lord 提交于 9月 24, 2010

During long I/O operations, the hang_check timer may fire,
trigger stack dumps that unnecessarily alarm the user.

Eg.  hdparm --security-erase NULL /dev/sdb  ## can take *hours* to complete

So, if hang_check is armed, we should wake up periodically
to prevent it from triggering.  This patch uses a wake-up interval
equal to half the hang_check timer period, which keeps overhead low enough.
Signed-off-by: NMark Lord <mlord@pobox.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

4b197769

20 9月, 2010 1 次提交

cfq: improve fsync performance for small files · 749ef9f8

由 Corrado Zoccolo 提交于 9月 20, 2010

Fsync performance for small files achieved by cfq on high-end disks is
lower than what deadline can achieve, due to idling introduced between
the sync write happening in process context and the journal commit.

Moreover, when competing with a sequential reader, a process writing
small files and fsync-ing them is starved.

This patch fixes the two problems by:
- marking journal commits as WRITE_SYNC, so that they get the REQ_NOIDLE
  flag set,
- force all queues that have REQ_NOIDLE requests to be put in the noidle
  tree.

Having the queue associated to the fsync-ing process and the one associated
 to journal commits in the noidle tree allows:
- switching between them without idling,
- fairness vs. competing idling queues, since they will be serviced only
  after the noidle tree expires its slice.
Acked-by: NVivek Goyal <vgoyal@redhat.com>
Reviewed-by: NJeff Moyer <jmoyer@redhat.com>
Tested-by: NJeff Moyer <jmoyer@redhat.com>
Signed-off-by: NCorrado Zoccolo <czoccolo@gmail.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

749ef9f8

17 9月, 2010 1 次提交

block: Fix race during disk initialization · 01ea5063

由 Signed-off-by: Jan Kara 提交于 9月 16, 2010

When a new disk is being discovered, add_disk() first ties the bdev to gendisk
(via register_disk()->blkdev_get()) and only after that calls
bdi_register_bdev(). Because register_disk() also creates disk's kobject, it
can happen that userspace manages to open and modify the device's data (or
inode) before its BDI is properly initialized leading to a warning in
__mark_inode_dirty().

Fix the problem by registering BDI early enough.

This patch addresses https://bugzilla.kernel.org/show_bug.cgi?id=16312

Cc: stable@kernel.org
Reported-by: NLarry Finger <Larry.Finger@lwfinger.net>
Signed-off-by: NJan Kara <jack@suse.cz>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

01ea5063

16 9月, 2010 6 次提交

blkio: Implementation of IOPS limit logic · 8e89d13f

由 Vivek Goyal 提交于 9月 15, 2010

o core logic of implementing IOPS throttling.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

8e89d13f

blk-cgroup: cgroup changes for IOPS limit support · 7702e8f4

由 Vivek Goyal 提交于 9月 15, 2010

o cgroup changes for IOPS throttling rules.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

7702e8f4

blkio: Core implementation of throttle policy · e43473b7

由 Vivek Goyal 提交于 9月 15, 2010

o Actual implementation of throttling policy in block layer. Currently it
  implements READ and WRITE bytes per second throttling logic. IOPS throttling
  comes in later patches.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

e43473b7

blk-cgroup: Introduce cgroup changes for throttling policy · 4c9eefa1

由 Vivek Goyal 提交于 9月 15, 2010

o cgroup chagnes for throttle policy.

o Introduces READ and WRITE bytes per second throttling rules.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

4c9eefa1

blk-cgroup: Prepare the base for supporting more than one IO control policies · 062a644d

由 Vivek Goyal 提交于 9月 15, 2010

o This patch prepares the base for introducing new IO control policies.
  Currently all the code is written knowing there is only one policy
  and that is proportional bandwidth. Creating infrastructure for newer
  policies to come in.

o Also there were many functions which were generated using macro. It was
  very confusing. Got rid of those.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

062a644d

blk-cgroup: Kill the header printed at the start of blkio.weight_device file · af41d7bd

由 Vivek Goyal 提交于 9月 15, 2010

o Kill extra "dev weight" header which is printed when somebody reads
  blkio.weight_device file. This really seems to be out of convention. No other
  blkio files are printing any header at the start of file. I think it is ok
  to just print values and how to interpret values should be part of
  documentation.
Signed-off-by: NVivek Goyal <vgoyal@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

af41d7bd

15 9月, 2010 3 次提交

init: add support for root devices specified by partition UUID · b5af921e

由 Will Drewry 提交于 8月 31, 2010

This is the third patch in a series which adds support for
storing partition metadata, optionally, off of the hd_struct.

One major use for that data is being able to resolve partition
by other identities than just the index on a block device.  Device
enumeration varies by platform and there's a benefit to being able
to use something like EFI GPT's GUIDs to determine the correct
block device and partition to mount as the root.

This change adds that support to root= by adding support for
the following syntax:

  root=PARTUUID=hex-uuid
Signed-off-by: NWill Drewry <wad@chromium.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

b5af921e

block, partition: add partition_meta_info to hd_struct · 6d1d8050

由 Will Drewry 提交于 8月 31, 2010

I'm reposting this patch series as v4 since there have been no additional
comments, and I cleaned up one extra bit of unneeded code (in 3/3). The patches
are against Linus's tree: 2bfc96a1
(2.6.36-rc3).

Would this patchset be suitable for inclusion in an mm branch?

This changes adds a partition_meta_info struct which itself contains a
union of structures that provide partition table specific metadata.

This change leaves the union empty. The subsequent patch includes an
implementation for CONFIG_EFI_PARTITION-based metadata.
Signed-off-by: NWill Drewry <wad@chromium.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

6d1d8050

block: fix an address space warning in blk-map.c · 14417799

由 Namhyung Kim 提交于 9月 15, 2010

Change type of 2nd parameter of blk_rq_aligned() into unsigned long
and remove unnecessary casting. Now we can call it with 'uaddr'
instead of 'ubuf' in __blk_rq_map_user() so that it can remove
following warnings from sparse:

 block/blk-map.c:57:31: warning: incorrect type in argument 2 (different address spaces)
 block/blk-map.c:57:31:    expected void *addr
 block/blk-map.c:57:31:    got void [noderef] <asn:1>*ubuf

However blk_rq_map_kern() needs one more local variable to handle it.
Signed-off-by: NNamhyung Kim <namhyung@gmail.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

14417799

14 9月, 2010 1 次提交

block: block_dump: Add number of sectors to debug output · 8dcbdc74

由 San Mehat 提交于 9月 14, 2010

Signed-off-by: NSan Mehat <san@android.com>
Signed-off-by: NLinus Walleij <linus.walleij@stericsson.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

8dcbdc74

11 9月, 2010 2 次提交

block/scsi: Provide a limit on the number of integrity segments · 13f05c8d

由 Martin K. Petersen 提交于 9月 10, 2010

Some controllers have a hardware limit on the number of protection
information scatter-gather list segments they can handle.

Introduce a max_integrity_segments limit in the block layer and provide
a new scsi_host_template setting that allows HBA drivers to provide a
value suitable for the hardware.

Add support for honoring the integrity segment limit when merging both
bios and requests.
Signed-off-by: NMartin K. Petersen <martin.petersen@oracle.com>
Signed-off-by: NJens Axboe <axboe@carl.home.kernel.dk>

13f05c8d

Consolidate min_not_zero · c8bf1336

由 Martin K. Petersen 提交于 9月 10, 2010

We have several users of min_not_zero, each of them using their own
definition.  Move the define to kernel.h.
Signed-off-by: NMartin K. Petersen <martin.petersen@oracle.com>
Signed-off-by: NJens Axboe <axboe@carl.home.kernel.dk>

c8bf1336

12 8月, 2010 1 次提交

block: add secure discard · 8d57a98c

由 Adrian Hunter 提交于 8月 11, 2010

Secure discard is the same as discard except that all copies of the
discarded sectors (perhaps created by garbage collection) must also be
erased.
Signed-off-by: NAdrian Hunter <adrian.hunter@nokia.com>
Acked-by: NJens Axboe <axboe@kernel.dk>
Cc: Kyungmin Park <kmpark@infradead.org>
Cc: Madhusudhan Chikkature <madhu.cr@ti.com>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Ben Gardiner <bengardiner@nanometrics.ca>
Cc: <linux-mmc@vger.kernel.org>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

8d57a98c

09 8月, 2010 2 次提交

blkdev: fix blkdev_issue_zeroout return value · 18edc8ea

由 Dmitry Monakhov 提交于 8月 06, 2010

- If function called without barrier option retvalue is incorrect
Signed-off-by: NDmitry Monakhov <dmonakhov@openvz.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

18edc8ea

block: update request stacking methods to support discards · 3383977f

由 ike Snitzer 提交于 8月 08, 2010

Propagate REQ_DISCARD in cmd_flags when cloning a discard request.
Skip blk_rq_check_limits's existing checks for discard requests because
discard limits will have already been checked in blkdev_issue_discard.
Signed-off-by: NMike Snitzer <snitzer@redhat.com>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

3383977f

08 8月, 2010 15 次提交

block: set up rq->rq_disk properly for flush requests · 16f2319f

由 FUJITA Tomonori 提交于 7月 09, 2010

q->bar_rq.rq_disk is NULL. Use the rq_disk of the original request
instead.
Signed-off-by: NFUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

16f2319f

block: set REQ_TYPE_FS on flush requests · 28e18d01

由 FUJITA Tomonori 提交于 7月 09, 2010

the block layer doesn't set rq->cmd_type on flush requests. By
definition, it should be REQ_TYPE_FS (the lower layers build a command
and interpret the result of it, that is, the block layer doesn't know
the details).
Signed-off-by: NFUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

28e18d01

block: fix problem with sending down discard that isn't of correct granularity · 10d1f9e2

由 Jens Axboe 提交于 7月 15, 2010

If the queue doesn't have a limit set, or it just set UINT_MAX like
we default to, we coud be sending down a discard request that isn't
of the correct granularity if the block size is > 512b.

Fix this by adjusting max_discard_sectors down to the proper
alignment.
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

10d1f9e2

blkdev: check for valid request queue before issuing flush · f10d9f61

由 Dave Chinner 提交于 7月 13, 2010

Issuing a blkdev_issue_flush() on an unconfigured loop device causes a panic as
q->make_request_fn is not configured. This can occur when trying to mount the
unconfigured loop device as an XFS filesystem. There are no guards that catch
the bio before the request function is called because we don't add a payload to
the bio. Instead, manually check this case as soon as we have a pointer to the
queue to flush.
Signed-off-by: NDave Chinner <dchinner@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

f10d9f61

block: remove BKL from partition ioctls · 15392efb

由 Arnd Bergmann 提交于 7月 07, 2010

The blkpg_ioctl and blkdev_reread_part access fields of
the bdev and gendisk structures, yet they always do so
under the protection of bdev->bd_mutex, which seems
sufficient.
Signed-off-by: NArnd Bergmann <arnd@arndb.de>
cked-by: NChristoph Hellwig <hch@infradead.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

15392efb

block: remove BKL from BLKROSET and BLKFLSBUF · 6de43703

由 Arnd Bergmann 提交于 7月 07, 2010

We only call the functions set_device_ro(),
invalidate_bdev(), sync_filesystem() and sync_blockdev()
while holding the BKL in these commands. All
of these are also done in other code paths without
the BKL, which leads me to the conclusion that
the BKL is not needed here either.

The reason we hold it here is that it was originally
pushed down into the ioctl function from vfs_ioctl.
Signed-off-by: NArnd Bergmann <arnd@arndb.de>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

6de43703

block: push BKL into blktrace ioctls · 62c2a7d9

由 Arnd Bergmann 提交于 7月 07, 2010

The blktrace driver currently needs the BKL, but
we should not need to take that in the block layer,
so just push it down into the driver itself.

It is quite likely that the BKL is not actually
required in blktrace code and could be removed
in a follow-on patch.
Signed-off-by: NArnd Bergmann <arnd@arndb.de>
Acked-by: NChristoph Hellwig <hch@infradead.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

62c2a7d9

block: push down BKL into .locked_ioctl · 8a6cfeb6

由 Arnd Bergmann 提交于 7月 08, 2010

As a preparation for the removal of the big kernel
lock in the block layer, this removes the BKL
from the common ioctl handling code, moving it
into every single driver still using it.
Signed-off-by: NArnd Bergmann <arnd@arndb.de>
Acked-by: NChristoph Hellwig <hch@infradead.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

8a6cfeb6

block: remove q->prepare_flush_fn completely · 00fff265

由 FUJITA Tomonori 提交于 7月 03, 2010

This removes q->prepare_flush_fn completely (changes the
blk_queue_ordered API).
Signed-off-by: NFUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

00fff265

block: permit PREFLUSH and POSTFLUSH without prepare_flush_fn · b6a90315

由 FUJITA Tomonori 提交于 7月 03, 2010

This is preparation for removing q->prepare_flush_fn.

Temporarily, blk_queue_ordered() permits QUEUE_ORDERED_DO_PREFLUSH and
QUEUE_ORDERED_DO_POSTFLUSH without prepare_flush_fn.
Signed-off-by: NFUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

b6a90315

block: introduce REQ_FLUSH flag · 8749534f

由 FUJITA Tomonori 提交于 7月 03, 2010

SCSI-ml needs a way to mark a request as flush request in
q->prepare_flush_fn because it needs to identify them later (e.g. in
q->request_fn or prep_rq_fn).

queue_flush sets REQ_HARDBARRIER in rq->cmd_flags however the block
layer also sends normal REQ_TYPE_FS requests with REQ_HARDBARRIER. So
SCSI-ml can't use REQ_HARDBARRIER to identify flush requests.

We could change the block layer to clear REQ_HARDBARRIER bit before
sending non flush requests to the lower layers. However, intorudcing
the new flag looks cleaner (surely easier).
Signed-off-by: NFUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>
Cc: James Bottomley <James.Bottomley@suse.de>
Cc: David S. Miller <davem@davemloft.net>
Cc: Rusty Russell <rusty@rustcorp.com.au>
Cc: Alasdair G Kergon <agk@redhat.com>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

8749534f

J
block: implement an unprep function corresponding directly to prep · 28018c24
由 James Bottomley 提交于 7月 01, 2010
```
Reviewed-by: NFUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>
```
28018c24

block: fixup missing conversion from BIO_RW_DISCARD to REQ_DISCARD · 3ffb52e7

由 Jens Axboe 提交于 6月 29, 2010

Didn't cause a merge conflict, so fixed this one up manually
post merge.
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

3ffb52e7

gcc-4.6: block: fix unused but set variables in blk-merge · 2c8919de

由 Andi Kleen 提交于 6月 21, 2010

Just some dead code.
Signed-off-by: NAndi Kleen <ak@linux.intel.com>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

2c8919de

block: don't allocate a payload for discard request · 66ac0280

由 Christoph Hellwig 提交于 6月 18, 2010

Allocating a fixed payload for discard requests always was a horrible hack,
and it's not coming to byte us when adding support for discard in DM/MD.

So change the code to leave the allocation of a payload to the lowlevel
driver. Unfortunately that means we'll need another hack, which allows
us to update the various block layer length fields indicating that we
have a payload. Instead of hiding this in sd.c, which we already partially
do for UNMAP support add a documented helper in the core block layer for it.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Acked-by: NMike Snitzer <snitzer@redhat.com>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

66ac0280