提交 · 4fe6a816707aace9e8e297b708411c5930537793 · openeuler / raspberrypi-kernel

19 3月, 2014 12 次提交

bcache: Add a real GC_MARK_RECLAIMABLE · 4fe6a816

由 Kent Overstreet 提交于 3月 13, 2014

This means the garbage collection code can better check for data and metadata
pointers to the same buckets.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

4fe6a816

bcache: Add bch_keylist_init_single() · c13f3af9

由 Kent Overstreet 提交于 1月 08, 2014

This will potentially save us an allocation when we've got inode/dirent bkeys
that don't fit in the keylist's inline keys.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

c13f3af9

bcache: Improve priority_stats · 15754020

由 Kent Overstreet 提交于 2月 25, 2014

Break down data into clean data/dirty data/metadata.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

15754020

bcache: Better alloc tracepoints · 7159b1ad

由 Kent Overstreet 提交于 2月 12, 2014

Change the invalidate tracepoint to indicate how much data we're invalidating,
and change the alloc tracepoints to indicate what offset they're for.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

7159b1ad

bcache: Kill dead cgroup code · 3f5e0a34

由 Kent Overstreet 提交于 1月 23, 2014

This hasn't been used or even enabled in ages.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

3f5e0a34

N
bcache: stop moving_gc marking buckets that can't be moved. · 3f6ef381
由 Nicholas Swenson 提交于 1月 23, 2014
```
Signed-off-by: NNicholas Swenson <nks@daterainc.com>
```
3f6ef381

bcache: Fix moving_pred() · 10d9dcf6

由 Kent Overstreet 提交于 2月 17, 2014

Avoid a potential null pointer deref (e.g. from check keys for cache misses)
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

10d9dcf6

bcache: Fix moving_gc deadlocking with a foreground write · da415a09

由 Nicholas Swenson 提交于 1月 09, 2014

Deadlock happened because a foreground write slept, waiting for a bucket
to be allocated. Normally the gc would mark buckets available for invalidation.
But the moving_gc was stuck waiting for outstanding writes to complete.
These writes used the bcache_wq, the same queue foreground writes used.

This fix gives moving_gc its own work queue, so it was still finish moving
even if foreground writes are stuck waiting for allocation. It also makes
work queue a parameter to the data_insert path, so moving_gc can use its
workqueue for writes.
Signed-off-by: NNicholas Swenson <nks@daterainc.com>
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

da415a09

bcache: Fix discard granularity · 90db6919

由 Kent Overstreet 提交于 2月 10, 2014

blk_stack_limits() doesn't like a discard granularity of 0.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

90db6919

bcache: Fix another bug recovering from unclean shutdown · 487dded8

由 Kent Overstreet 提交于 3月 17, 2014

The on disk bucket gens are allowed to be out of date, when we reuse buckets
that didn't have any live data in them. To deal with this, the initial gc has to
update the bucket gen when we find a pointer gen newer than the bucket's gen.

Unfortunately we weren't doing this for pointers in the journal that we're about
to replay.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

487dded8

bcache: Fix a bug recovering from unclean shutdown · 0bd143fd

由 Kent Overstreet 提交于 3月 04, 2014

The code to fixup incorrect bucket prios incorrectly did not skip btree node
freeing keys
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

0bd143fd

bcache: Fix a journalling reclaim after recovery bug · 27201cfd

由 Kent Overstreet 提交于 3月 13, 2014

On recovery we weren't correctly keeping track of what journal buckets had open
journal entries, thus it was possible for them to be overwritten until we'd
written all new journal entries.
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

27201cfd

18 3月, 2014 2 次提交
- K
  bcache: Fix a null ptr deref in journal replay · 65ddf45a
  由 Kent Overstreet 提交于 2月 24, 2014
```
Signed-off-by: NKent Overstreet <kmo@daterainc.com>
```
  65ddf45a
- K
  bcache: Fix a lockdep splat in an error path · 4fa03402
  由 Kent Overstreet 提交于 3月 17, 2014
```
Signed-off-by: NKent Overstreet <kmo@daterainc.com>
```
  4fa03402
26 2月, 2014 2 次提交

bcache: Fix a shutdown bug · dabb4433

由 Kent Overstreet 提交于 2月 19, 2014

Shutdown wasn't cancelling/waiting on journal_write_work()
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

dabb4433

bcache: Fix flash_dev_cache_miss() for real this time · 1b4eaf3d

由 Kent Overstreet 提交于 1月 16, 2014

The code was using sectors to count the number of sectors it was zeroing... but
then it passed it to bio_advance()... after it had been set to 0. Amusing...
Signed-off-by: NKent Overstreet <kmo@daterainc.com>

1b4eaf3d

19 2月, 2014 1 次提交

bcache: Fix another compiler warning on m68k · 85cbe1f8

由 Kent Overstreet 提交于 2月 17, 2014

Use a bigger hammer this time
Signed-off-by: NKent Overstreet <kmo@daterainc.com>
Cc: linux-stable <stable@vger.kernel.org>

85cbe1f8

13 2月, 2014 1 次提交

md/raid5: Fix CPU hotplug callback registration · 789b5e03

由 Oleg Nesterov 提交于 2月 06, 2014

Subsystems that want to register CPU hotplug callbacks, as well as perform
initialization for the CPUs that are already online, often do it as shown
below:

	get_online_cpus();

	for_each_online_cpu(cpu)
		init_cpu(cpu);

	register_cpu_notifier(&foobar_cpu_notifier);

	put_online_cpus();

This is wrong, since it is prone to ABBA deadlocks involving the
cpu_add_remove_lock and the cpu_hotplug.lock (when running concurrently
with CPU hotplug operations).

Interestingly, the raid5 code can actually prevent double initialization and
hence can use the following simplified form of callback registration:

	register_cpu_notifier(&foobar_cpu_notifier);

	get_online_cpus();

	for_each_online_cpu(cpu)
		init_cpu(cpu);

	put_online_cpus();

A hotplug operation that occurs between registering the notifier and calling
get_online_cpus(), won't disrupt anything, because the code takes care to
perform the memory allocations only once.

So reorganize the code in raid5 this way to fix the deadlock with callback
registration.

Cc: linux-raid@vger.kernel.org
Cc: stable@vger.kernel.org (v2.6.32+)
Fixes: 36d1c647Signed-off-by: NOleg Nesterov <oleg@redhat.com>
[Srivatsa: Fixed the unregister_cpu_notifier() deadlock, added the
free_scratch_buffer() helper to condense code further and wrote the changelog.]
Signed-off-by: NSrivatsa S. Bhat <srivatsa.bhat@linux.vnet.ibm.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

789b5e03

11 2月, 2014 1 次提交

drivers/md/bcache/extents.c: use %zi to format size_t · bd180b4e

由 Geert Uytterhoeven 提交于 2月 10, 2014

drivers/md/bcache/extents.c: In function `btree_ptr_bad_expensive':
drivers/md/bcache/extents.c:196: warning: format `%li' expects type `long int', but argument 4 has type `size_t'
Signed-off-by: NGeert Uytterhoeven <geert@linux-m68k.org>
Cc: Kent Overstreet <kmo@daterainc.com>
Cc: Neil Brown <neilb@suse.de>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

bd180b4e

05 2月, 2014 1 次提交

md/raid1: restore ability for check and repair to fix read errors. · 1877db75

由 NeilBrown 提交于 2月 05, 2014

commit 30bc9b53
    md/raid1: fix bio handling problems in process_checks()

Move the bio_reset() to a point before where BIO_UPTODATE is checked,
so that check now always report that the bio is uptodate, even if it is not.

This causes process_check() to sometimes treat read-errors as
successful matches so the good data isn't written out.

This patch preserves the flag until it is needed.

Bug was introduced in 3.11, but backported to 3.10-stable (as it fixed
an even worse bug).  So suitable for any -stable since 3.10.
Reported-and-tested-by: NMichael Tokarev <mjt@tls.msk.ru>
Cc: stable@vger.kernel.org (3.10+)
Fixed: 30bc9b53Signed-off-by: NNeilBrown <neilb@suse.de>

1877db75

30 1月, 2014 3 次提交

N
bcache: bugfix - gc thread now gets woken when cache is full · e3b4825b
由 Nicholas Swenson 提交于 12月 12, 2013
```
Signed-off-by: NNicholas Swenson <nks@daterainc.com>
```
e3b4825b
K
bcache: Minor fixes from kbuild robot · 3572324a
由 Kent Overstreet 提交于 1月 10, 2014
```
Signed-off-by: NKent Overstreet <kmo@daterainc.com>
```
3572324a

bcache: fix BUG_ON due to integer overflow with GC_SECTORS_USED · 94717447

由 Darrick J. Wong 提交于 1月 28, 2014

The BUG_ON at the end of __bch_btree_mark_key can be triggered due to
an integer overflow error:

BITMASK(GC_SECTORS_USED, struct bucket, gc_mark, 2, 13);
...
SET_GC_SECTORS_USED(g, min_t(unsigned,
	     GC_SECTORS_USED(g) + KEY_SIZE(k),
	     (1 << 14) - 1));
BUG_ON(!GC_SECTORS_USED(g));

In bcache.h, the SECTORS_USED bitfield is defined to be 13 bits wide.
While the SET_ code tries to ensure that the field doesn't overflow by
clamping it to (1<<14)-1 == 16383, this is incorrect because 16383
requires 14 bits.  Therefore, if GC_SECTORS_USED() + KEY_SIZE() =
8192, the SET_ statement tries to store 8192 into a 13-bit field.  In
a 13-bit field, 8192 becomes zero, thus triggering the BUG_ON.

Therefore, create a field width constant and a max value constant, and
use those to create the bitfield and check the inputs to
SET_GC_SECTORS_USED.  Arguably the BITMASK() template ought to have
BUG_ON checks for too-large values, but that's a separate patch.
Signed-off-by: NDarrick J. Wong <darrick.wong@oracle.com>

94717447

22 1月, 2014 3 次提交

dm log userspace: allow mark requests to piggyback on flush requests · 5066a4df

由 Dongmao Zhang 提交于 1月 15, 2014

In the cluster evironment, cluster write has poor performance because
userspace_flush() has to contact a userspace program (cmirrord) for
clear/mark/flush requests.  But both mark and flush requests require
cmirrord to communicate the message to all the cluster nodes for each
flush call.  This behaviour is really slow.

To address this we now merge mark and flush requests together to reduce
the kernel-userspace-kernel time.  We allow a new directive,
"integrated_flush" that can be used to instruct the kernel log code to
combine flush and mark requests when directed by userspace.  If not
directed by userspace (due to an older version of the userspace code
perhaps), the kernel will function as it did previously - preserving
backwards compatibility.  Additionally, flush requests are performed
lazily when only clear requests exist.
Signed-off-by: NDongmao Zhang <dmzhang@suse.com>
Signed-off-by: NJonathan Brassow <jbrassow@redhat.com>
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

5066a4df

md/raid5: close recently introduced race in stripe_head management. · 7da9d450

由 NeilBrown 提交于 1月 22, 2014

As release_stripe and __release_stripe decrement ->count and then
manipulate ->lru both under ->device_lock, it is important that
get_active_stripe() increments ->count and clears ->lru also under
->device_lock.

However we currently list_del_init ->lru under the lock, but increment
the ->count outside the lock.  This can lead to races and list
corruption.

So move the atomic_inc(&sh->count) up inside the ->device_lock
protected region.

Note that we still increment ->count without device lock in the case
where get_free_stripe() was called, and in fact don't take
->device_lock at all in that path.
This is safe because if the stripe_head can be found by
get_free_stripe, then the hash lock assures us the no-one else could
possibly be calling release_stripe() at the same time.

Fixes: 566c09c5
Cc: stable@vger.kernel.org (3.13)
Reported-and-tested-by: NIan Kumlien <ian.kumlien@gmail.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

7da9d450

dm space map metadata: fix bug in resizing of thin metadata · fca02843

由 Joe Thornber 提交于 1月 21, 2014

This bug was introduced in commit 7e664b3d ("dm space map metadata:
fix extending the space map").

When extending a dm-thin metadata volume we:

- Switch the space map into a simple bootstrap mode, which allocates
  all space linearly from the newly added space.
- Add new bitmap entries for the new space
- Increment the reference counts for those newly allocated bitmap
  entries
- Commit changes to disk
- Switch back out of bootstrap mode.

But, the disk commit may allocate space itself, if so this fact will be
lost when switching out of bootstrap mode.

The bug exhibited itself as an error when the bitmap_root, with an
erroneous ref count of 0, was subsequently decremented as part of a
later disk commit.  This would cause the disk commit to fail, and thinp
to enter read_only mode.  The metadata was not damaged (thin_check
passed).

The fix is to put the increments + commit into a loop, running until
the commit has not allocated extra space.  In practise this loop only
runs twice.

With this fix the following device mapper testsuite test passes:
 dmtest run --suite thin-provisioning -n thin_remove_works_after_resize
Signed-off-by: NJoe Thornber <ejt@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>
Cc: stable@vger.kernel.org # depends on commit 7e664b3d

fca02843

17 1月, 2014 1 次提交

dm cache: add policy name to status output · 2e68c4e6

由 Mike Snitzer 提交于 1月 15, 2014

The cache's policy may have been established using the "default" alias,
which is currently the "mq" policy but the default policy may change in
the future. It is useful to know exactly which policy is being used.

Add a 'real' member to the dm_cache_policy_type structure and have the
"default" dm_cache_policy_type point to the real "mq"
dm_cache_policy_type. Update dm_cache_policy_get_name() to check if
real is set, if so report the name of the real policy (not the alias).
Requested-by: NJonathan Brassow <jbrassow@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

2e68c4e6

16 1月, 2014 3 次提交

dm thin: fix pool feature parsing · 74aa45c3

由 Mike Snitzer 提交于 1月 15, 2014

Commit 787a996c ("dm thin: add error_if_no_space feature")
mistakenly forgot to increase the number of feature args supported.
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

74aa45c3

md/raid5: fix long-standing problem with bitmap handling on write failure. · 9f97e4b1

由 NeilBrown 提交于 1月 16, 2014

Before a write starts we set a bit in the write-intent bitmap.
When the write completes we clear that bit if the write was successful
to all devices.  However if the write wasn't fully successful we
should not clear the bit.  If the faulty drive is subsequently
re-added, the fact that the bit is still set ensure that we will
re-write the data that is missing.

This logic is mediated by the STRIPE_DEGRADED flag - we only clear the
bitmap bit when this flag is not set.
Currently we correctly set the flag if a write starts when some
devices are failed or missing.  But we do *not* set the flag if some
device failed during the write attempt.
This is wrong and can result in clearing the bit inappropriately.

So: set the flag when a write fails.

This bug has been present since bitmaps were introduces, so the fix is
suitable for any -stable kernel.
Reported-by: NEthan Wilson <ethan.wilson@shiftmail.org>
Cc: stable@vger.kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

9f97e4b1

md: check command validity early in md_ioctl(). · cb335f88

由 Nicolas Schichan 提交于 1月 15, 2014

Verify that the cmd parameter passed to md_ioctl() is valid before
doing anything.

This fixes mddev->hold_active being set to 0 when an invalid ioctl
command is passed to md_ioctl() before the array has been configured.

Clearing mddev->hold_active in that case can lead to a livelock
situation when an invalid ioctl number is given to md_ioctl() by a
process when the mddev is currently being opened by another process:

Process 1				Process 2
---------				---------

md_alloc()
  mddev_find()
  -> returns a new mddev with
     hold_active == UNTIL_IOCTL
  add_disk()
  -> sends KOBJ_ADD uevent

					(sees KOBJ_ADD uevent for device)
                    			md_open()
                    			md_ioctl(INVALID_IOCTL)
                    			-> returns ENODEV and clears
                       			   mddev->hold_active
                    			md_release()
                      			md_put()
                      			-> deletes the mddev as
                         		   hold_active is 0

md_open()
  mddev_find()
  -> returns a newly
    allocated mddev with
    mddev->gendisk == NULL
-> returns with ERESTARTSYS
   (kernel restarts the open syscall)
Signed-off-by: NNicolas Schichan <nschichan@freebox.fr>
Signed-off-by: NNeilBrown <neilb@suse.de>

cb335f88

15 1月, 2014 5 次提交

dm sysfs: fix a module unload race · 2995fa78

由 Mikulas Patocka 提交于 1月 13, 2014

This reverts commit be35f486 ("dm: wait until embedded kobject is
released before destroying a device") and provides an improved fix.

The kobject release code that calls the completion must be placed in a
non-module file, otherwise there is a module unload race (if the process
calling dm_kobject_release is preempted and the DM module unloaded after
the completion is triggered, but before dm_kobject_release returns).

To fix this race, this patch moves the completion code to dm-builtin.c
which is always compiled directly into the kernel if BLK_DEV_DM is
selected.

The patch introduces a new dm_kobject_holder structure, its purpose is
to keep the completion and kobject in one place, so that it can be
accessed from non-module code without the need to export the layout of
struct mapped_device to that code.
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>
Cc: stable@vger.kernel.org

2995fa78

dm snapshot: use dm-bufio prefetch · 55b082e6

由 Mikulas Patocka 提交于 1月 13, 2014

This patch modifies dm-snapshot so that it prefetches the buffers when
loading the exceptions.

The number of buffers read ahead is specified in the DM_PREFETCH_CHUNKS
macro.  The current value for DM_PREFETCH_CHUNKS (12) was found to
provide the best performance on a single 15k SCSI spindle.  In the
future we may modify this default or make it configurable.

Also, introduce the function dm_bufio_set_minimum_buffers to setup
bufio's number of internal buffers before freeing happens.  dm-bufio may
hold more buffers if enough memory is available.  There is no guarantee
that the specified number of buffers will be available - if you need a
guarantee, use the argument reserved_buffers for
dm_bufio_client_create.
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

55b082e6

dm snapshot: use dm-bufio · 55494bf2

由 Mikulas Patocka 提交于 1月 13, 2014

Use dm-bufio for initial loading of the exceptions.
Introduce a new function dm_bufio_forget that frees the given buffer.
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

55494bf2

dm snapshot: prepare for switch to using dm-bufio · 2cadabd5

由 Mikulas Patocka 提交于 1月 13, 2014

Change the functions get_exception, read_exception and insert_exceptions
so that ps->area is passed as an argument.

This patch doesn't change any functionality, but it refactors the code
to allow for a cleaner switch over to using dm-bufio.
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

2cadabd5

dm snapshot: use GFP_KERNEL when initializing exceptions · 119bc547

由 Mikulas Patocka 提交于 1月 13, 2014

The list of initial exceptions is loaded in the target constructor.  We
are allowed to allocate memory with GFP_KERNEL at this point.  So,
change alloc_completed_exception to use GFP_KERNEL when being called
from the constructor.
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>

119bc547

14 1月, 2014 5 次提交

md: ensure metadata is writen after raid level change. · 830778a1

由 NeilBrown 提交于 1月 14, 2014

level_store() currently does not make sure the metadata is
updates to reflect the new raid level.  It simply sets MD_CHANGE_DEVS.

Any level with a ->thread will quickly notice this and update the
metadata.  However RAID0 and Linear do not have a thread so no
metadata update happens until the array is stopped.  At that point the
metadata is written.

This is later that we would like.  While the delay doesn't risk any
data it can cause confusion.  So if there is no md thread, immediately
update the metadata after a level change.
Reported-by: NRichard Michael <rmichael@edgeofthenet.org>
Signed-off-by: NNeilBrown <neilb@suse.de>

830778a1

md/raid10: avoid fullsync when not necessary. · 0b59bb64

由 NeilBrown 提交于 1月 14, 2014

This is the raid10 equivalent of

commit 4f0a5e01
    MD RAID1: Further conditionalize 'fullsync'

If a device in a newly assembled array is not fully recovered we
currently do a fully resync by don't need to.
Signed-off-by: NNeilBrown <neilb@suse.de>

0b59bb64

md: allow a partially recovered device to be hot-added to an array. · 7eb41885

由 NeilBrown 提交于 1月 14, 2014

When adding a new device into an array it is normally important to
clear any stale data from ->recovery_offset else the new device may
not be recovered properly.

However when re-adding a device which is known to be nearly in-sync,
this is not needed and can be detrimental.  The (bitmap-based)
resync will still happen, and further recovery is only needed from
where-ever it was already up to.

So if save_raid_disk is set, signifying a re-add, don't clear
->recovery_offset.
Signed-off-by: NNeilBrown <neilb@suse.de>

7eb41885

md: Change handling of save_raid_disk and metadata update during recovery. · f466722c

由 NeilBrown 提交于 12月 09, 2013

Since commit d70ed2e4
   MD: Allow restarting an interrupted incremental recovery.

we don't write out the metadata to devices while they are recovering.
This had a good reason, but has unfortunate consequences.  This patch
changes things to make them work better.

At issue is what happens if the array is shut down while a recovery is
happening, particularly a bitmap-guided recovery.
Ideally the recovery should pick up where it left off.
However the metadata cannot represent the state "A recovery is in
process which is guided by the bitmap".

Before the above mentioned commit, we wrote metadata to the device
which said "this is being recovered and it is up to <here>".  So after
a restart, a full recovery (not bitmap-guided) would happen from
where-ever it was up to.

After the commit the metadata wasn't updated so it still said "This
device is fully in sync with <this> event count".  That leads to a
bitmap-based recovery following the whole bitmap, which should be a
lot less work than a full recovery from some starting point.  So this
was an improvement.

However updates some metadata but not all leads to other problems.
In particular, the metadata written to the fully-up-to-date device
record that the array has all devices present (even though some are
recovering).  So on restart, mdadm wants to find all devices and
expects them to have current event counts.
Obviously it doesn't (some have old event counts) so (when assembling
with --incremental) it waits indefinitely for the rest of the expected
devices.

It really is wrong to not update all the metadata together.  Do that
is bound to cause confusion.
Instead, we should make it possible to record the truth in the
metadata.  i.e. we need to be able to record that a device is being
recovered based on the bitmap.
We already have a Feature flag to say that recovery is happening.  We
now add another one to say that it is a bitmap-based recovery.

With this we can remove the code that disables the write-out of
metadata on some devices.

So this patch:
 - moves the setting of 'saved_raid_disk' from add_new_disk to
   the validate_super methods.  This makes sure it is always set
   properly, both when adding a new device to an array, and when
   assembling an array from a collection of devices.
 - Adds a metadata flag MD_FEATURE_RECOVERY_BITMAP which is only
   used if MD_FEATURE_RECOVERY_OFFSET is set, and record that a
   bitmap-based recovery is allowed.
   This is only present in v1.x metadata. v0.90 doesn't support
   devices which are in the middle of recovery at all.
 - Only skips writing metadata to Faulty devices.

 - Also allows rdev state to be set to "-insync" via sysfs.
   This can be used for external-metadata arrays.  When the
   'role' is set the device is assumed to be in-sync.  If, after
   setting the role, we set the state to "-insync", the role is
   moved to saved_raid_disk which effectively says the device is
   partly in-sync with that slot and needs a bitmap recovery.

Cc: Andrei Warkentin <andreiw@vmware.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

f466722c

md: fix problem when adding device to read-only array with bitmap. · 8313b8e5

由 NeilBrown 提交于 12月 12, 2013

If an array is started degraded, and then the missing device
is found it can be re-added and a minimal bitmap-based recovery
will bring it fully up-to-date.

If the array is read-only a recovery would not be allowed.
But also if the array is read-only and the missing device was
present very recently, then there could be no need for any
recovery at all, so we simply include the device in the read-only
array without any recovery.

However... if the missing device was removed a little longer ago
it could be missing some updates, but if a bitmap is present it will
be conditionally accepted pending a bitmap-based update.  We don't
currently detect this case properly and will include that old
device into the read-only array with no recovery even though it really
needs a recovery.

This patch keeps track of whether a bitmap-based-recovery is really
needed or not in the new Bitmap_sync rdev flag.  If that is set,
then the device will not be added to a read-only array.

Cc: Andrei Warkentin <andreiw@vmware.com>
Fixes: d70ed2e4
Cc: stable@vger.kernel.org (3.2+)
Signed-off-by: NNeilBrown <neilb@suse.de>

8313b8e5