提交 · 75a73a29e520a6ce982b0da6dd8b7560ae3faa90 · openanolis / cloud-kernel

18 5月, 2010 20 次提交

md: restore ability of spare drives to spin down. · 75a73a29

由 NeilBrown 提交于 5月 07, 2010

Some time ago we stopped the clean/active metadata updates
from being written to a 'spare' device in most cases so that
it could spin down and say spun down.  Device failure/removal
etc are still recorded on spares.

However commit 51d5668c broke this 50% of the time,
depending on whether the event count is even or odd.
The change log entry said:

   This means that the alignment between 'odd/even' and
    'clean/dirty' might take a little longer to attain,

how ever the code makes no attempt to create that alignment, so it
could take arbitrarily long.

So when we find that clean/dirty is not aligned with odd/even,
force a second metadata-update immediately.  There are already cases
where a second metadata-update is needed immediately (e.g. when a
device fails during the metadata update).  We just piggy-back on that.
Reported-by: NJoe Bryant <tenminjoe@yahoo.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@kernel.org

75a73a29

md: allow integers to be passed to md/level · f2859af6

由 Dan Williams 提交于 5月 02, 2010

e.g. allow md to interpret 'echo 4 > md/level' as a request for raid4.
Signed-off-by: NDan Williams <dan.j.williams@intel.com>

f2859af6

md: notify mdstat waiters of level change · bb7f8d22

由 Dan Williams 提交于 5月 01, 2010

Level modifications change the output of mdstat. The mdmon manager
thread is interested in these events for external metadata management.
Signed-off-by: NDan Williams <dan.j.williams@intel.com>

bb7f8d22

md: don't unregister the thread in mddev_suspend · 9e35b99c

由 NeilBrown 提交于 4月 06, 2010

This is
 - unnecessary because mddev_suspend is always followed by a call to
   ->stop, and each ->stop unregisters the thread, and
 - a problem as it makes it awkwards to suspend and then resume a
   device as we will want later.
Signed-off-by: NNeilBrown <neilb@suse.de>

9e35b99c

md: factor out init code for an mddev · fafd7fb0

由 NeilBrown 提交于 4月 01, 2010

This is a simple factorisation that makes mddev_find easier to read.
Signed-off-by: NNeilBrown <neilb@suse.de>

fafd7fb0

md: pass mddev to make_request functions rather than request_queue · 21a52c6d

由 NeilBrown 提交于 4月 01, 2010

We used to pass the personality make_request function direct
to the block layer so the first argument had to be a queue.
But now we have the intermediary md_make_request so it makes
at lot more sense to pass a struct mddev_s.
It makes it possible to have an mddev without its own queue too.
Signed-off-by: NNeilBrown <neilb@suse.de>

21a52c6d

md: call md_stop_writes from md_stop · cca9cf90

由 NeilBrown 提交于 4月 01, 2010

This moves the call to the other side of set_readonly, but that should
not be an issue.
This encapsulates in 'md_stop' all of the functionality for internally
stopping the array, leaving all the interactions with externalities
(sysfs, request_queue, gendisk) in do_md_stop.
Signed-off-by: NNeilBrown <neilb@suse.de>

cca9cf90

md: split md_set_readonly out of do_md_stop · a4bd82d0

由 NeilBrown 提交于 3月 29, 2010

Using do_md_stop to set an array to read-only is a little confusing.
Now most of the common code has been factored out, split
md_set_readonly off in to a separate function.
Signed-off-by: NNeilBrown <neilb@suse.de>

a4bd82d0

md: factor md_stop_writes out of do_md_stop. · a047e125

由 NeilBrown 提交于 3月 29, 2010

Further refactoring of do_md_stop.
This one requires some explanation as it takes code from different
places in do_md_stop, so some re-ordering happens.

We only get into this part of do_md_stop if there are no active opens
of the device, so no writes can be happening and the device must have
been flushed.  In md_stop_writes we want to stop any internal sources
of writes - i.e. resync - and flush out the metadata.

The only code that was previously before some of this code is
code to clean up the queue, the mddev, the gendisk, or sysfs, all
of which is probably better after code that makes active changes (i.e.
triggers writes).
Signed-off-by: NNeilBrown <neilb@suse.de>

a047e125

md: start to refactor do_md_stop · 6177b472

由 NeilBrown 提交于 3月 29, 2010

do_md_stop is large and clunky, so hard to understand.

This is a first step of refactoring, pulling two simple
sub-functions out.
Signed-off-by: NNeilBrown <neilb@suse.de>

6177b472

md: factor do_md_run to separate accesses to ->gendisk · fe60b014

由 NeilBrown 提交于 3月 29, 2010

As part of relaxing the binding between an mddev and gendisk,
we separate do_md_run into two functions.
  md_run does all the work internal to md
  do_md_run calls md_run and makes and changes to gendisk
     that are required.
Signed-off-by: NNeilBrown <neilb@suse.de>

fe60b014

md: remove ->changed and related code. · b821eaa5

由 NeilBrown 提交于 3月 29, 2010

We set ->changed to 1 and call check_disk_change at the end
of md_open so that bd_invalidated would be set and thus
partition rescan would happen appropriately.

Now that we call revalidate_disk directly, which sets bd_invalidates,
that indirection is no longer needed and can be removed.
Signed-off-by: NNeilBrown <neilb@suse.de>

b821eaa5

md: don't reference gendisk in getgeo · 49ce6cea

由 NeilBrown 提交于 3月 29, 2010

Using ->array_sectors rather than get_capacity() is more
direct and is a step towards relaxing the tight connection
between mddev and gendisk.
Signed-off-by: NNeilBrown <neilb@suse.de>

49ce6cea

md: move io accounting out of personalities into md_make_request · 49077326

由 NeilBrown 提交于 3月 25, 2010

While I generally prefer letting personalities do as much as possible,
given that we have a central md_make_request anyway we may as well use
it to simplify code.
Also this centralises knowledge of ->gendisk which will help later.
Signed-off-by: NNeilBrown <neilb@suse.de>

49077326

md: notify level changes through sysfs. · 5cac7861

由 Maciej Trela 提交于 4月 14, 2010

Level changes can be very significant, so make sure
to notify them via sysfs.
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

5cac7861

md: Relax checks on ->max_disks when external metadata handling is used. · 233fca36

由 NeilBrown 提交于 4月 14, 2010

When metadata is being managed by user-space, md doesn't know
what the maximum number of devices allowed in an array is
so ->max_disks is 0.  In this case we should allow any (+ve)
number of disks.
Signed-off-by: NNeilBrown <neilb@suse.de>

233fca36

md: Correctly handle device removal via sysfs · b7103107

由 Maciej Trela 提交于 4月 14, 2010

Writing "none" to "../md/dev-xx/slot" removes that device
from being an active part of the array, but it didn't
set ->raid_disk to -1 to record this fact.
Signed-off-by: NMaciej Trela <Maciej.Trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

b7103107

T
md: Add support for Raid5->Raid0 and Raid10->Raid0 takeover · 9af204cf
由 Trela, Maciej 提交于 3月 08, 2010
```
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
```
9af204cf

md:Add support for Raid0->Raid5 takeover · 54071b38

由 Trela Maciej 提交于 3月 08, 2010

Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

54071b38

md: discard StateChanged device flag. · c0cc75f8

由 NeilBrown 提交于 3月 22, 2010

This was needed when sysfs files could only be 'notified'
from process context.  Now that we have sys_notify_direct,
we can call it directly from an interrupt.
Signed-off-by: NNeilBrown <neilb@suse.de>

c0cc75f8

17 5月, 2010 2 次提交

md: manage redundancy group in sysfs when changing level. · a64c876f

由 NeilBrown 提交于 4月 14, 2010

Some levels expect the 'redundancy group' to be present,
others don't.
So when we change level of an array we might need to
add or remove this group.

This requires fixing up the current practice of overloading ->private
to indicate (when ->pers == NULL) that something needs to be removed.
So create a new ->to_remove to fill that role.

When changing levels, we may need to add or remove attributes.  When
changing RAID5 -> RAID6, we both add and remove the same thing.  It is
important to catch this and optimise it out as the removal is delayed
until a lock is released, so trying to add immediately would cause
problems.


Cc: stable@kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

a64c876f

md: remove unneeded sysfs files more promptly · b6eb127d

由 NeilBrown 提交于 4月 15, 2010

When an array is stopped we need to remove some
sysfs files which are dependent on the type of array.

We need to delay that deletion as deleting them while holding
reconfig_mutex can lead to deadlocks.

We currently delay them until the array is completely destroyed.
However it is possible to deactivate and then reactivate the array.
It is also possible to need to remove sysfs files when changing level,
which can potentially happen several times before an array is
destroyed.

So we need to delete these files more promptly: as soon as
reconfig_mutex is dropped.

We need to ensure this happens before do_md_run can restart the array,
so we use open_mutex for some extra locking.  This is not deadlock
prone.

Cc: stable@kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

b6eb127d

12 5月, 2010 1 次提交

md: set mddev readonly flag on blkdev BLKROSET ioctl · e2218350

由 Dan Williams 提交于 5月 12, 2010

When the user sets the block device to readwrite then the mddev should
follow suit.  Otherwise, the BUG_ON in md_write_start() will be set to
trigger.

The reverse direction, setting mddev->ro to match a set readonly
request, can be ignored because the blkdev level readonly flag precludes
the need to have mddev->ro set correctly.  Nevermind the fact that
setting mddev->ro to 1 may fail if the array is in use.

Cc: <stable@kernel.org>
Signed-off-by: NDan Williams <dan.j.williams@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

e2218350

10 2月, 2010 1 次提交

md: fix some lockdep issues between md and sysfs. · ef286f6f

由 NeilBrown 提交于 2月 09, 2010

======
This fix is related to
    http://bugzilla.kernel.org/show_bug.cgi?id=15142
but does not address that exact issue.
======

sysfs does like attributes being removed while they are being accessed
(i.e. read or written) and waits for the access to complete.

As accessing some md attributes takes the same lock that is held while
removing those attributes a deadlock can occur.

This patch addresses 3 issues in md that could lead to this deadlock.

Two relate to calling flush_scheduled_work while the lock is held.
This is probably a bad idea in general and as we use schedule_work to
delete various sysfs objects it is particularly bad.

In one case flush_scheduled_work is called from md_alloc (called by
md_probe) called from do_md_run which holds the lock.  This call is
only present to ensure that ->gendisk is set.  However we can be sure
that gendisk is always set (though possibly we couldn't when that code
was originally written.  This is because do_md_run is called in three
different contexts:
  1/ from md_ioctl.  This requires that md_open has succeeded, and it
     fails if ->gendisk is not set.
  2/ from writing a sysfs attribute.  This can only happen if the
     mddev has been registered in sysfs which happens in md_alloc
     after ->gendisk has been set.
  3/ from autorun_array which is only called by autorun_devices, which
     checks for ->gendisk to be set before calling autorun_array.
So the call to md_probe in do_md_run can be removed, and the check on
->gendisk can also go.


In the other case flush_scheduled_work is being called in do_md_stop,
purportedly to wait for all md_delayed_delete calls (which delete the
component rdevs) to complete.  However there really isn't any need to
wait for them - they have already been disconnected in all important
ways.

The third issue is that raid5->stop() removes some attribute names
while the lock is held.  There is already some infrastructure in place
to delay attribute removal until after the lock is released (using
schedule_work).  So extend that infrastructure to remove the
raid5_attrs_group.

This does not address all lockdep issues related to the sysfs
"s_active" lock.  The rest can be address by splitting that lockdep
context between symlinks and non-symlinks which hopefully will happen.
Signed-off-by: NNeilBrown <neilb@suse.de>

ef286f6f

30 12月, 2009 5 次提交

md: allow a resync that is waiting for other resync to complete, to be aborted. · 404e4b43

由 NeilBrown 提交于 12月 30, 2009

If two arrays share a device, then they will not both resync at the
same time.  One will wait for the other to complete.
While waiting, the MD_RECOVERY_INTR flag is not checked so a device
failure, which would make the resync pointless, does not cause the
resync to abort, so the failed device cannot be removed (as it cannot
be remove while a resync is happening).

So add a test for MD_RECOVERY_INTR.
Reported-by: NBrett Russ <bruss@netezza.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

404e4b43

md: remove unnecessary code from do_md_run · 7fb9dadc

由 NeilBrown 提交于 12月 30, 2009

Since commit dfc70645,
->hot_remove_disks has not removed non-failed devices from
an array until recovery is no longer possible.
So the code in do_md_run to get around the fact that
md_check_recovery (which calls ->hot_remove_disks) would
remove partially-in-sync devices is no longer needed.

So remove it.
Signed-off-by: NNeilBrown <neilb@suse.de>

7fb9dadc

md: make recovery started by do_md_run() visible via sync_action · a2d79c32

由 Dan Williams 提交于 12月 21, 2009

By default md_do_sync() will perform recovery if no other actions are
specified.  However, action_show() relies on MD_RECOVERY_RECOVER to be
set otherwise it returns 'idle'.  So, add a missing set
MD_RECOVERY_RECOVER when starting recovery.
Signed-off-by: NDan Williams <dan.j.williams@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

a2d79c32

md: fix small irregularity with start_ro module parameter · 0f9552b5

由 NeilBrown 提交于 12月 30, 2009

The start_ro modules parameter can be used to force arrays to be
started in 'auto-readonly' in which they are read-only until the first
write.  This ensures that no resync/recovery happens until something
else writes to the device.  This is important for resume-from-disk
off an md array.

However if an array is started 'readonly' (by writing 'readonly' to
the 'array_state' sysfs attribute) we want it to be really 'readonly',
not 'auto-readonly'.

So strengthen the condition to only set auto-readonly if the
array is not already read-only.
Signed-off-by: NNeilBrown <neilb@suse.de>

0f9552b5

md: Fix unfortunate interaction with evms · cbd19983

由 NeilBrown 提交于 12月 30, 2009

evms configures md arrays by:
  open device
  send ioctl
  close device

for each different ioctl needed.
Since 2.6.29, the device can disappear after the 'close'
unless a significant configuration has happened to the device.
The change made by "SET_ARRAY_INFO" can too minor to stop the device
from disappearing, but important enough that losing the change is bad.

So: make sure SET_ARRAY_INFO sets mddev->ctime, and keep the device
active as long as ctime is non-zero (it gets zeroed with lots of other
things when the array is stopped).

This is suitable for -stable kernels since 2.6.29.
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@kernel.org

cbd19983

16 12月, 2009 2 次提交

drivers/md/md.c: use %pU to print UUIDs · 7b75c2f8

由 Joe Perches 提交于 12月 14, 2009

Signed-off-by: NJoe Perches <joe@perches.com>
Cc: Neil Brown <neilb@suse.de>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

7b75c2f8

tree-wide: convert open calls to remove spaces to skip_spaces() lib function · e7d2860b

由 André Goddard Rosa 提交于 12月 14, 2009

Makes use of skip_spaces() defined in lib/string.c for removing leading
spaces from strings all over the tree.

It decreases lib.a code size by 47 bytes and reuses the function tree-wide:
   text    data     bss     dec     hex filename
  64688     584     592   65864   10148 (TOTALS-BEFORE)
  64641     584     592   65817   10119 (TOTALS-AFTER)

Also, while at it, if we see (*str && isspace(*str)), we can be sure to
remove the first condition (*str) as the second one (isspace(*str)) also
evaluates to 0 whenever *str == 0, making it redundant. In other words,
"a char equals zero is never a space".

Julia Lawall tried the semantic patch (http://coccinelle.lip6.fr) below,
and found occurrences of this pattern on 3 more files:
    drivers/leds/led-class.c
    drivers/leds/ledtrig-timer.c
    drivers/video/output.c

@@
expression str;
@@

( // ignore skip_spaces cases
while (*str &&  isspace(*str)) { \(str++;\|++str;\) }
|
- *str &&
isspace(*str)
)
Signed-off-by: NAndré Goddard Rosa <andre.goddard@gmail.com>
Cc: Julia Lawall <julia@diku.dk>
Cc: Martin Schwidefsky <schwidefsky@de.ibm.com>
Cc: Jeff Dike <jdike@addtoit.com>
Cc: Ingo Molnar <mingo@elte.hu>
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: "H. Peter Anvin" <hpa@zytor.com>
Cc: Richard Purdie <rpurdie@rpsys.net>
Cc: Neil Brown <neilb@suse.de>
Cc: Kyle McMartin <kyle@mcmartin.ca>
Cc: Henrique de Moraes Holschuh <hmh@hmh.eng.br>
Cc: David Howells <dhowells@redhat.com>
Cc: <linux-ext4@vger.kernel.org>
Cc: Samuel Ortiz <samuel@sortiz.org>
Cc: Patrick McHardy <kaber@trash.net>
Cc: Takashi Iwai <tiwai@suse.de>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

e7d2860b

14 12月, 2009 9 次提交

md: add 'recovery_start' per-device sysfs attribute · 06e3c817

由 Dan Williams 提交于 12月 12, 2009

Enable external metadata arrays to manage rebuild checkpointing via a
md/dev-XXX/recovery_start attribute which reflects rdev->recovery_offset

Also update resync_start_store to allow 'none' to be written, for
consistency.
Signed-off-by: NDan Williams <dan.j.williams@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

06e3c817

md: rcu_read_lock() walk of mddev->disks in md_do_sync() · 4e59ca7d

由 Dan Williams 提交于 12月 12, 2009

Other walks of this list are either under rcu_read_lock() or the list
mutation lock (mddev_lock()). This protects against the improbable case of a
disk being removed from the array at the start of md_do_sync().
Signed-off-by: NDan Williams <dan.j.williams@intel.com>

4e59ca7d

md: integrate spares into array at earliest opportunity. · 93be75ff

由 NeilBrown 提交于 12月 14, 2009

As v1.x metadata can record that a member of the array is
not completely recovered, it make sense to record that a
spare has become a regular member of the array at the earliest
opportunity.
So remove the tests on "recovery_offset > 0" in super_1_sync
as they really aren't needed, and schedule a metadata update
immediately after adding spares to a degraded array.

This means that if a crash happens immediately after a recovery
starts, the new device will be included in the array and recovery will
continue from wherever it was up to.  Previously this didn't happen
unless recovery was at least 1/16 of the way through.
Signed-off-by: NNeilBrown <neilb@suse.de>

93be75ff

md: move compat_ioctl handling into md.c · aa98aa31

由 Arnd Bergmann 提交于 12月 14, 2009

The RAID ioctls are only implemented in md.c, so the
handling for them should also be moved there from
fs/compat_ioctl.c.
Signed-off-by: NArnd Bergmann <arnd@arndb.de>
Cc: Neil Brown <neilb@suse.de>
Cc: Andre Noll <maan@systemlinux.org>
Cc: linux-raid@vger.kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

aa98aa31

N
md: add MODULE_DESCRIPTION for all md related modules. · 0efb9e61
由 NeilBrown 提交于 12月 14, 2009
```
Suggested by  Oren Held <orenhe@il.ibm.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
```
0efb9e61

raid: improve MD/raid10 handling of correctable read errors. · 1e50915f

由 Robert Becker 提交于 12月 14, 2009

We've noticed severe lasting performance degradation of our raid
arrays when we have drives that yield large amounts of media errors.
The raid10 module will queue each failed read for retry, and also
will attempt call fix_read_error() to perform the read recovery.
Read recovery is performed while the array is frozen, so repeated
recovery attempts can degrade the performance of the array for
extended periods of time.

With this patch I propose adding a per md device max number of
corrected read attempts.  Each rdev will maintain a count of
read correction attempts in the rdev->read_errors field (not
used currently for raid10). When we enter fix_read_error()
we'll check to see when the last read error occurred, and
divide the read error count by 2 for every hour since the
last read error. If at that point our read error count
exceeds the read error threshold, we'll fail the raid device.

In addition in this patch I add sysfs nodes (get/set) for
the per md max_read_errors attribute, the rdev->read_errors
attribute, and added some printk's to indicate when
fix_read_error fails to repair an rdev.

For testing I used debugfs->fail_make_request to inject
IO errors to the rdev while doing IO to the raid array.
Signed-off-by: NRobert Becker <Rob.Becker@riverbed.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

1e50915f

md: support updating bitmap parameters via sysfs. · 43a70507

由 NeilBrown 提交于 12月 14, 2009

A new attribute directory 'bitmap' in 'md' is created which
contains files for configuring the bitmap.
'location' identifies where the bitmap is, either 'none',
or 'file' or 'sector offset from metadata'.
Writing 'location' can create or remove a bitmap.
Adding a 'file' bitmap this way is not yet supported.
'chunksize' and 'time_base' must be set before 'location'
can be set.

'chunksize' can be set before creating a bitmap, but is
currently always over-ridden by the bitmap superblock.

'time_base' and 'backlog' can be updated at any time.
Signed-off-by: NNeilBrown <neilb@suse.de>
Reviewed-by: NAndre Noll <maan@systemlinux.org>

43a70507

md: factor out parsing of fixed-point numbers · 72e02075

由 NeilBrown 提交于 12月 14, 2009

safe_delay_store can parse fixed point numbers (for fractions
of a second).  We will want to do that for another sysfs
file soon, so factor out the code.
Signed-off-by: NNeilBrown <neilb@suse.de>

72e02075

md: move offset, daemon_sleep and chunksize out of bitmap structure · 42a04b50

由 NeilBrown 提交于 12月 14, 2009

... and into bitmap_info.  These are all configuration parameters
that need to be set before the bitmap is created.
Signed-off-by: NNeilBrown <neilb@suse.de>

42a04b50

openanolis / cloud-kernel 1 年多 前同步成功

openanolis / cloud-kernel
1 年多前同步成功