提交 · fab363b5ff502d1b39ddcfec04271f5858d9f26e · openanolis / cloud-kernel

03 7月, 2012 1 次提交

md: make 'name' arg to md_register_thread non-optional. · 0232605d

由 NeilBrown 提交于 7月 03, 2012

Having the 'name' arg optional and defaulting to the current
personality name is no necessary and leads to errors, as when
changing the level of an array we can end up using the
name of the old level instead of the new one.

So make it non-optional and always explicitly pass the name
of the level that the array will be.
Reported-by: Nmajianpeng <majianpeng@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

0232605d

31 5月, 2012 1 次提交

md: raid1/raid10: fix problem with merge_bvec_fn · aba336bd

由 NeilBrown 提交于 5月 31, 2012

The new merge_bvec_fn which calls the corresponding function
in subsidiary devices requires that mddev->merge_check_needed
be set if any child has a merge_bvec_fn.

However were were only setting that when a device was hot-added,
not when a device was present from the start.

This bug was introduced in 3.4 so patch is suitable for 3.4.y
kernels.  However that are conflicts in raid10.c so a separate
patch will be needed for 3.4.y.

Cc: stable@vger.kernel.org
Reported-by: NSebastian Riemer <sebastian.riemer@profitbricks.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

aba336bd

22 5月, 2012 3 次提交

MD RAID1: Further conditionalize 'fullsync' · 4f0a5e01

由 Jonathan Brassow 提交于 5月 22, 2012

A RAID1 device does not necessarily need a fullsync if the bitmap can be used instead.

Similar to commit d6b212f4 in raid5.c, if a raid1
device can be brought back (i.e. from a transient failure) it shouldn't need a
complete resync.  Provided the bitmap is not to old, it will have recorded the areas
of the disk that need recovery.
Signed-off-by: NJonathan Brassow <jbrassow@redhat.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

4f0a5e01

md: allow array to be resized while bitmap is present. · a4a6125a

由 NeilBrown 提交于 5月 22, 2012

Now that bitmaps can be resized, we can allow an array to be resized
while the bitmap is present.

This only covers resizing that involves changing the effective size
of member devices, not resizing that changes the number of devices.
Signed-off-by: NNeilBrown <neilb@suse.de>

a4a6125a

md/raid1: allow fix_read_error to read from recovering device. · da8840a7

由 majianpeng 提交于 5月 22, 2012

When attempting to fix a read error, it is acceptable to read from a
device that is recovering, provided the recovery has got past the
place we are reading from.  This makes the test for "can we read from
here" the same as the test in read_balance.
Signed-off-by: Nmajianpeng <majianpeng@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

da8840a7

21 5月, 2012 1 次提交

md: add possibility to change data-offset for devices. · c6563a8c

由 NeilBrown 提交于 5月 21, 2012

When reshaping we can avoid costly intermediate backup by
changing the 'start' address of the array on the device
(if there is enough room).

So as a first step, allow such a change to be requested
through sysfs, and recorded in v1.x metadata.

(As we didn't previous check that all 'pad' fields were zero,
 we need a new FEATURE flag for this.
 A (belatedly) check that all remaining 'pad' fields are
 zero to avoid a repeat of this)

The new data offset must be requested separately for each device.
This allows each to have a different change in the data offset.
This is not likely to be used often but as data_offset can be
set per-device, new_data_offset should be too.

This patch also removes the 'acknowledged' arg to rdev_set_badblocks as
it is never used and never will be.  At the same time we add a new
arg ('in_new') which is currently always zero but will be used more
soon.

When a reshape finishes we will need to update the data_offset
and rdev->sectors.  So provide an exported function to do that.
Signed-off-by: NNeilBrown <neilb@suse.de>

c6563a8c

12 4月, 2012 1 次提交

md/raid1,raid10: Fix calculation of 'vcnt' when processing error recovery. · f4380a91

由 majianpeng 提交于 4月 12, 2012

If r1bio->sectors % 8 != 0,then the memcmp and a later
memcpy will omit the last bio_vec.

This is suitable for any stable kernel since 3.1 when bad-block
management was introduced.

Cc: stable@vger.kernel.org
Signed-off-by: Nmajianpeng <majianpeng@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

f4380a91

03 4月, 2012 2 次提交

md/raid1,raid10: don't compare excess byte during consistency check. · 5020ad7d

由 NeilBrown 提交于 4月 02, 2012

When comparing two pages read from different legs of a mirror, only
compare the bytes that were read, not the whole page.

In most cases we read a whole page, but in some cases with
bad blocks or odd sizes devices we might read fewer than that.

This bug has been present "forever" but at worst it might cause
a report of two many mismatches and generate a little bit
extra resync IO, so there is no need to back-port to -stable
kernels.
Reported-by: Nmajianpeng <majianpeng@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

5020ad7d

md/raid1:Remove unnecessary rcu_dereference(conf->mirrors[i].rdev). · a42f9d83

由 majianpeng 提交于 4月 02, 2012

Because rde->nr_pending > 0,so can not remove this disk.
And in any case, we aren't holding rcu_read_lock()
Signed-off-by: Nmajianpeng <majianpeng@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

a42f9d83

02 4月, 2012 1 次提交
- M
  md/raid1: If md_integrity_register() failed,run() must free the mem · 5220ea1e
  由 majianpeng 提交于 4月 02, 2012
```
Signed-off-by: Nmajianpeng <majianpeng@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
```
  5220ea1e
19 3月, 2012 3 次提交

md/raid1: handle merge_bvec_fn in member devices. · 6b740b8d

由 NeilBrown 提交于 3月 19, 2012

Currently we don't honour merge_bvec_fn in member devices so if there
is one, we force all requests to be single-page at most.
This is not ideal.

So create a raid1 merge_bvec_fn to check that function in children
as well.

This introduces a small problem.  There is no locking around calls
the ->merge_bvec_fn and subsequent calls to ->make_request.  So a
device added between these could end up getting a request which
violates its merge_bvec_fn.

Currently the best we can do is synchronize_sched().  This will work
providing no preemption happens.  If there is is preemption, we just
have to hope that new devices are largely consistent with old devices.
Signed-off-by: NNeilBrown <neilb@suse.de>

6b740b8d

md: tidy up rdev_for_each usage. · dafb20fa

由 NeilBrown 提交于 3月 19, 2012

md.h has an 'rdev_for_each()' macro for iterating the rdevs in an
mddev.  However it uses the 'safe' version of list_for_each_entry,
and so requires the extra variable, but doesn't include 'safe' in the
name, which is useful documentation.

Consequently some places use this safe version without needing it, and
many use an explicity list_for_each entry.

So:
 - rename rdev_for_each to rdev_for_each_safe
 - create a new rdev_for_each which uses the plain
   list_for_each_entry,
 - use the 'safe' version only where needed, and convert all other
   list_for_each_entry calls to use rdev_for_each.
Signed-off-by: NNeilBrown <neilb@suse.de>

dafb20fa

md/raid1,raid10: avoid deadlock during resync/recovery. · d6b42dcb

由 NeilBrown 提交于 3月 19, 2012

If RAID1 or RAID10 is used under LVM or some other stacking
block device, it is possible to enter a deadlock during
resync or recovery.
This can happen if the upper level block device creates
two requests to the RAID1 or RAID10.  The first request gets
processed, blocks recovery and queue requests for underlying
requests in current->bio_list.  A resync request then starts
which will wait for those requests and block new IO.

But then the second request to the RAID1/10 will be attempted
and it cannot progress until the resync request completes,
which cannot progress until the underlying device requests complete,
which are on a queue behind that second request.

So allow that second request to proceed even though there is
a resync request about to start.

This is suitable for any -stable kernel.

Cc: stable@vger.kernel.org
Reported-by: NRay Morris <support@bettercgi.com>
Tested-by: NRay Morris <support@bettercgi.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

d6b42dcb

13 2月, 2012 1 次提交

md/raid1: fix buglet in md_raid1_contested. · f53e29fc

由 NeilBrown 提交于 2月 13, 2012

Since we added 'replacement' capability, RAID1 can have twice
as many devices as ->raid_disks indicates.
So md_raid1_congested needs to check that many possible devices,
not just ->raid_disks many.
Signed-off-by: NNeilBrown <neilb@suse.de>

f53e29fc

11 1月, 2012 1 次提交

md/raid1: perform bad-block tests for WriteMostly devices too. · 307729c8

由 NeilBrown 提交于 1月 09, 2012

We normally try to avoid reading from write-mostly devices, but when
we do we really have to check for bad blocks and be sure not to
try reading them.

With the current code, best_good_sectors might not get set and that
causes zero-length read requests to be send down which is very
confusing.

This bug was introduced in commit d2eb35ac and so the patch
is suitable for 3.1.x and 3.2.x
Reported-and-tested-by: NMichał Mirosław <mirq-linux@rere.qmqm.pl>
Reported-and-tested-by: NArt -kwaak- van Breemen <ard@telegraafnet.nl>
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@vger.kernel.org

307729c8

23 12月, 2011 8 次提交

md/raid1: Mark device want_replacement when we see a write error. · 19d67169

由 NeilBrown 提交于 12月 23, 2011

Now that WantReplacement drives are replaced cleanly, mark a drive
as want_replacement when we see a write error.  It might get failed soon so
the WantReplacement flag is irrelevant, but if the write error is recorded
in the bad block log, we still want to activate any spare that might
be available.
Signed-off-by: NNeilBrown <neilb@suse.de>

19d67169

md/raid1: If there is a spare and a want_replacement device, start replacement. · 7ef449d1

由 NeilBrown 提交于 12月 23, 2011

When attempting to add a spare to a RAID1 array, also consider
adding it as a replacement for a want_replacement device.
Signed-off-by: NNeilBrown <neilb@suse.de>

7ef449d1

md/raid1: recognise replacements when assembling arrays. · c19d5798

由 NeilBrown 提交于 12月 23, 2011

If a Replacement is seen, file it as such.

If we see two replacements (or two normal devices) for the one slot,
abort.
Signed-off-by: NNeilBrown <neilb@suse.de>

c19d5798

md/raid1: handle activation of replacement device when recovery completes. · 8c7a2c2b

由 NeilBrown 提交于 12月 23, 2011

When recovery completes ->spare_active is called.
This checks if the replacement is ready and if so it fails
the original.
Signed-off-by: NNeilBrown <neilb@suse.de>

8c7a2c2b

md/raid1: Allow a failed replacement device to be removed. · b014f14c

由 NeilBrown 提交于 12月 23, 2011

Replacement devices are stored at a different offset, so look
there too.
Signed-off-by: NNeilBrown <neilb@suse.de>

b014f14c

md/raid1: Allocate spare to store replacement devices and their bios. · 8f19ccb2

由 NeilBrown 提交于 12月 23, 2011

In RAID1, a replacement is much like a normal device, so we just
double the size of the relevant arrays and look at all possible
devices for reads and writes.

This means that the array looks like it is now double the size in some
way - we need to be careful about that.
In particular, we checking if the array is still degraded while
creating a recovery request we need to only consider the first 'half'
- i.e. the real (non-replacement) devices.
Signed-off-by: NNeilBrown <neilb@suse.de>

8f19ccb2

md/raid1: Replace use of mddev->raid_disks with conf->raid_disks. · 30194636

由 NeilBrown 提交于 12月 23, 2011

In general mddev->raid_disks can change unexpectedly while
conf->raid_disks will only change in a very controlled way.  So change
some uses of one to the other.

The use of mddev->raid_disks will not cause actually problems but
this way is more consistent and safer in the long term.
Signed-off-by: NNeilBrown <neilb@suse.de>

30194636

md: change hot_remove_disk to take an rdev rather than a number. · b8321b68

由 NeilBrown 提交于 12月 23, 2011

Soon an array will be able to have multiple devices with the
same raid_disk number (an original and a replacement).  So removing
a device based on the number won't work.  So pass the actual device
handle instead.
Reviewed-by: NDan Williams <dan.j.williams@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

b8321b68

01 11月, 2011 1 次提交

md: Add module.h to all files using it implicitly · 056075c7

由 Paul Gortmaker 提交于 7月 03, 2011

A pending cleanup will mean that module.h won't be implicitly
everywhere anymore. Make sure the modular drivers in md dir
are actually calling out for <module.h> explicitly in advance.
Signed-off-by: NPaul Gortmaker <paul.gortmaker@windriver.com>

056075c7

26 10月, 2011 1 次提交

md: Fix some bugs in recovery_disabled handling. · d890fa2b

由 NeilBrown 提交于 10月 26, 2011

In 3.0 we changed the way recovery_disabled was handle so that instead
of testing against zero, we test an mddev-> value against a conf->
value.
Two problems:
  1/ one place in raid1 was missed and still sets to '1'.
  2/ We didn't explicitly set the conf-> value at array creation
     time.
     It defaulted to '0' just like the mddev value does so they
     could appear equal and thus disable recovery.
     This did not affect normal 'md' as it calls bind_rdev_to_array
     which changes the mddev value.  However the dmraid interface
     doesn't call this and so doesn't change ->recovery_disabled; so at
     array start all recovery is incorrectly disabled.

So initialise the 'conf' value to one less that the mddev value, so
the will only be the same when explicitly set that way.
Reported-by: NJonathan Brassow <jbrassow@redhat.com>
Signed-off-by: NNeilBrown  <neilb@suse.de>

d890fa2b

24 10月, 2011 1 次提交

block: Remove the control of complete cpu from bio. · 9562ad9a

由 Tao Ma 提交于 10月 24, 2011

bio originally has the functionality to set the complete cpu, but
it is broken.

Chirstoph said that "This code is unused, and from the all the
discussions lately pretty obviously broken.  The only thing keeping
it serves is creating more confusion and possibly more bugs."

And Jens replied with "We can kill bio_set_completion_cpu(). I'm fine
with leaving cpu control to the request based drivers, they are the
only ones that can toggle the setting anyway".

So this patch tries to remove all the work of controling complete cpu
from a bio.

Cc: Shaohua Li <shaohua.li@intel.com>
Cc: Christoph Hellwig <hch@infradead.org>
Signed-off-by: NTao Ma <boyu.mt@taobao.com>
Signed-off-by: NJens Axboe <axboe@kernel.dk>

9562ad9a

11 10月, 2011 7 次提交

md: add proper write-congestion reporting to RAID1 and RAID10. · 34db0cd6

由 NeilBrown 提交于 10月 11, 2011

RAID1 and RAID10 handle write requests by queuing them for handling by
a separate thread. This is because when a write-intent-bitmap is
active we might need to update the bitmap first, so it is good to
queue a lot of writes, then do one big bitmap update for them all.

However writeback request devices to appear to be congested after a
while so it can make some guesstimate of throughput. The infinite
queue defeats that (note that RAID5 has already has a finite queue so
it doesn't suffer from this problem).

So impose a limit on the number of pending write requests. By default
it is 1024 which seems to be generally suitable. Make it configurable
via module option just in case someone finds a regression.
Signed-off-by: NNeilBrown <neilb@suse.de>

34db0cd6

N
md: rename "mdk_personality" to "md_personality" · 84fc4b56
由 NeilBrown 提交于 10月 11, 2011
```
"mdk" doesn't mean anything any more.
Signed-off-by: NNeilBrown <neilb@suse.de>
```
84fc4b56
N
md/raid1: typedef removal: conf_t -> struct r1conf · e8096360
由 NeilBrown 提交于 10月 11, 2011
```
Signed-off-by: NNeilBrown <neilb@suse.de>
```
e8096360
N
md: remove typedefs: mirror_info_t -> struct mirror_info · 0f6d02d5
由 NeilBrown 提交于 10月 11, 2011
```
Signed-off-by: NNeilBrown <neilb@suse.de>
```
0f6d02d5
N
md: remove typedefs: r10bio_t -> struct r10bio and r1bio_t -> struct r1bio · 9f2c9d12
由 NeilBrown 提交于 10月 11, 2011
```
Signed-off-by: NNeilBrown <neilb@suse.de>
```
9f2c9d12

md: remove typedefs: mddev_t -> struct mddev · fd01b88c

由 NeilBrown 提交于 10月 11, 2011

Having mddev_t and 'struct mddev_s' is ugly and not preferred
Signed-off-by: NNeilBrown <neilb@suse.de>

fd01b88c

md: removing typedefs: mdk_rdev_t -> struct md_rdev · 3cb03002

由 NeilBrown 提交于 10月 11, 2011

The typedefs are just annoying. 'mdk' probably refers to 'md_k.h'
which used to be an include file that defined this thing.
Signed-off-by: NNeilBrown <neilb@suse.de>

3cb03002

07 10月, 2011 3 次提交

N
md: remove PRINTK and dprintk debugging and use pr_debug · 36a4e1fe
由 NeilBrown 提交于 10月 07, 2011
```
Being able to dynamically enable these make them much more useful.
Signed-off-by: NNeilBrown <neilb@suse.de>
```
36a4e1fe

md/raid1/ avoid bio search in end_sync_read() · 0fc280f6

由 NeilBrown 提交于 10月 07, 2011

We know which device we just read from so we don't need to
search the bios to find out.  Just use ->read_disk.
Signed-off-by: NNeilBrown <neilb@suse.de>

0fc280f6

md/raid1: factor out common bio handling code · ba3ae3be

由 Namhyung Kim 提交于 10月 07, 2011

When normal-write and sync-read/write bio completes, we should
find out the disk number the bio belongs to. Factor those common
code out to a separate function.
Signed-off-by: NNamhyung Kim <namhyung@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

ba3ae3be

21 9月, 2011 1 次提交

md: Avoid waking up a thread after it has been freed. · 01f96c0a

由 NeilBrown 提交于 9月 21, 2011

Two related problems:

1/ some error paths call "md_unregister_thread(mddev->thread)"
   without subsequently clearing ->thread.  A subsequent call
   to mddev_unlock will try to wake the thread, and crash.

2/ Most calls to md_wakeup_thread are protected against the thread
   disappeared either by:
      - holding the ->mutex
      - having an active request, so something else must be keeping
        the array active.
   However mddev_unlock calls md_wakeup_thread after dropping the
   mutex and without any certainty of an active request, so the
   ->thread could theoretically disappear.
   So we need a spinlock to provide some protections.

So change md_unregister_thread to take a pointer to the thread
pointer, and ensure that it always does the required locking, and
clears the pointer properly.
Reported-by: N"Moshe Melnikov" <moshe@zadarastorage.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
cc: stable@kernel.org

01f96c0a

12 9月, 2011 1 次提交

block: remove support for bio remapping from ->make_request · 5a7bbad2

由 Christoph Hellwig 提交于 9月 12, 2011

There is very little benefit in allowing to let a ->make_request
instance update the bios device and sector and loop around it in
__generic_make_request when we can archive the same through calling
generic_make_request from the driver and letting the loop in
generic_make_request handle it.

Note that various drivers got the return value from ->make_request and
returned non-zero values for errors.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Acked-by: NNeilBrown <neilb@suse.de>
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

5a7bbad2

10 9月, 2011 1 次提交

md/raid1,10: Remove use-after-free bug in make_request. · 079fa166

由 NeilBrown 提交于 9月 10, 2011

A single request to RAID1 or RAID10 might result in multiple
requests if there are known bad blocks that need to be avoided.

To detect if we need to submit another write request we test:
 	if (sectors_handled < (bio->bi_size >> 9)) {

However this is after we call **_write_done() so the 'bio' no longer
belongs to us - the writes could have completed and the bio freed.

So move the **_write_done call until after the test against
bio->bi_size.

This addresses https://bugzilla.kernel.org/show_bug.cgi?id=41862Reported-by: NBruno Wolff III <bruno@wolff.to>
Tested-by: NBruno Wolff III <bruno@wolff.to>
Signed-off-by: NNeilBrown <neilb@suse.de>

079fa166

28 7月, 2011 1 次提交

md/raid1: factor several functions out or raid1d() · 62096bce

由 NeilBrown 提交于 7月 28, 2011

raid1d is too big with several deep branches.
So separate them out into their own functions.
Signed-off-by: NNeilBrown <neilb@suse.de>
Reviewed-by: NNamhyung Kim <namhyung@gmail.com>

62096bce

openanolis / cloud-kernel 1 年多 前同步成功

openanolis / cloud-kernel
1 年多前同步成功