提交 · b821eaa572fd737faaf6928ba046e571526c36c6 · openeuler / Kernel

18 5月, 2010 17 次提交

md: remove ->changed and related code. · b821eaa5

由 NeilBrown 提交于 3月 29, 2010

We set ->changed to 1 and call check_disk_change at the end
of md_open so that bd_invalidated would be set and thus
partition rescan would happen appropriately.

Now that we call revalidate_disk directly, which sets bd_invalidates,
that indirection is no longer needed and can be removed.
Signed-off-by: NNeilBrown <neilb@suse.de>

b821eaa5

md: don't reference gendisk in getgeo · 49ce6cea

由 NeilBrown 提交于 3月 29, 2010

Using ->array_sectors rather than get_capacity() is more
direct and is a step towards relaxing the tight connection
between mddev and gendisk.
Signed-off-by: NNeilBrown <neilb@suse.de>

49ce6cea

md: move io accounting out of personalities into md_make_request · 49077326

由 NeilBrown 提交于 3月 25, 2010

While I generally prefer letting personalities do as much as possible,
given that we have a central md_make_request anyway we may as well use
it to simplify code.
Also this centralises knowledge of ->gendisk which will help later.
Signed-off-by: NNeilBrown <neilb@suse.de>

49077326

md/raid5: small tidyup in raid5_align_endio · 2b7f2228

由 NeilBrown 提交于 3月 25, 2010

Diving through ->queue to find mddev is unnecessarily complex - there
is an easier path to finding mddev, so use that.
Signed-off-by: NNeilBrown <neilb@suse.de>

2b7f2228

md: add support for raid5 to raid4 conversion · a78d38a1

由 NeilBrown 提交于 3月 22, 2010

This is unlikely to be wanted, but we may as well provide it
for completeness.
Signed-off-by: NNeilBrown <neilb@suse.de>

a78d38a1

md: notify level changes through sysfs. · 5cac7861

由 Maciej Trela 提交于 4月 14, 2010

Level changes can be very significant, so make sure
to notify them via sysfs.
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

5cac7861

md: Relax checks on ->max_disks when external metadata handling is used. · 233fca36

由 NeilBrown 提交于 4月 14, 2010

When metadata is being managed by user-space, md doesn't know
what the maximum number of devices allowed in an array is
so ->max_disks is 0.  In this case we should allow any (+ve)
number of disks.
Signed-off-by: NNeilBrown <neilb@suse.de>

233fca36

md: Correctly handle device removal via sysfs · b7103107

由 Maciej Trela 提交于 4月 14, 2010

Writing "none" to "../md/dev-xx/slot" removes that device
from being an active part of the array, but it didn't
set ->raid_disk to -1 to record this fact.
Signed-off-by: NMaciej Trela <Maciej.Trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

b7103107

md: Add support for Raid0->Raid10 takeover · dab8b292

由 Trela, Maciej 提交于 3月 08, 2010

Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

dab8b292

T
md: Add support for Raid5->Raid0 and Raid10->Raid0 takeover · 9af204cf
由 Trela, Maciej 提交于 3月 08, 2010
```
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
```
9af204cf

md:Add support for Raid0->Raid5 takeover · 54071b38

由 Trela Maciej 提交于 3月 08, 2010

Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

54071b38

md: don't use mddev->raid_disks in raid0 or raid10 while array is active. · 84707f38

由 NeilBrown 提交于 3月 16, 2010

In a subsequent patch we will make it possible to change
mddev->raid_disks while a RAID0 or RAID10 array is active.  This is
part of the process of reshaping such an array.

This means that we cannot use this value while processes requests
(it is OK to use it during initialisation as we are locked against
changes then).
Both RAID0 and RAID10 have the same value stored in the private data
structure, so use that value instead.
Signed-off-by: NNeilBrown <neilb@suse.de>

84707f38

md: discard StateChanged device flag. · c0cc75f8

由 NeilBrown 提交于 3月 22, 2010

This was needed when sysfs files could only be 'notified'
from process context.  Now that we have sys_notify_direct,
we can call it directly from an interrupt.
Signed-off-by: NNeilBrown <neilb@suse.de>

c0cc75f8

drivers/md: Remove unnecessary casts of void * · 7b92813c

由 H Hartley Sweeten 提交于 3月 08, 2010

void pointers do not need to be cast to other pointer types.
Signed-off-by: NH Hartley Sweeten <hsweeten@visionengravers.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

7b92813c

md: expose max value of behind writes counter · 696fcd53

由 Paul Clements 提交于 3月 08, 2010

Keep track of the maximum number of concurrent write-behind requests
for an md array and exposed this number in sysfs at
   md/bitmap/max_backlog_used

Writing any value to this file will clear it.

This allows userspace to be involved in tuning bitmap/backlog.
Signed-off-by: NPaul Clements <paul.clements@steeleye.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

696fcd53

md: remove some dead fields from mddev_s · ee8b81b0

由 NeilBrown 提交于 3月 08, 2010

These fields have never been used.
commit 4b6d287f
added them, but also added identical files to bitmap_super_s,
and only used the latter.

So remove these unused fields.
Signed-off-by: NNeilBrown <neilb@suse.de>

ee8b81b0

md/raid1: fix counting of write targets. · 964147d5

由 NeilBrown 提交于 5月 18, 2010

There is a very small race window when writing to a
RAID1 such that if a device is marked faulty at exactly the wrong
time, the write-in-progress will not be sent to the device,
but the bitmap (if present) will be updated to say that
the write was sent.

Then if the device turned out to still be usable as was re-added
to the array, the bitmap-based-resync would skip resyncing that
block, possibly leading to corruption.  This would only be a problem
if no further writes were issued to that area of the device (i.e.
that bitmap chunk).

Suitable for any pending -stable kernel.

Cc: stable@kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

964147d5

17 5月, 2010 3 次提交

md: manage redundancy group in sysfs when changing level. · a64c876f

由 NeilBrown 提交于 4月 14, 2010

Some levels expect the 'redundancy group' to be present,
others don't.
So when we change level of an array we might need to
add or remove this group.

This requires fixing up the current practice of overloading ->private
to indicate (when ->pers == NULL) that something needs to be removed.
So create a new ->to_remove to fill that role.

When changing levels, we may need to add or remove attributes.  When
changing RAID5 -> RAID6, we both add and remove the same thing.  It is
important to catch this and optimise it out as the removal is delayed
until a lock is released, so trying to add immediately would cause
problems.


Cc: stable@kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

a64c876f

md: remove unneeded sysfs files more promptly · b6eb127d

由 NeilBrown 提交于 4月 15, 2010

When an array is stopped we need to remove some
sysfs files which are dependent on the type of array.

We need to delay that deletion as deleting them while holding
reconfig_mutex can lead to deadlocks.

We currently delay them until the array is completely destroyed.
However it is possible to deactivate and then reactivate the array.
It is also possible to need to remove sysfs files when changing level,
which can potentially happen several times before an array is
destroyed.

So we need to delete these files more promptly: as soon as
reconfig_mutex is dropped.

We need to ensure this happens before do_md_run can restart the array,
so we use open_mutex for some extra locking.  This is not deadlock
prone.

Cc: stable@kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

b6eb127d

md/linear: avoid possible oops and array stop · ef2f80ff

由 NeilBrown 提交于 5月 17, 2010

Since commit ef286f6f
it has been important that each personality clears
->private in the ->stop() function, or sets it to a
attribute group to be removed.
linear.c doesn't.  This can sometimes lead to an oops,
though it doesn't always.

Suitable for 2.6.33-stable and 2.6.34.
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@kernel.org

ef2f80ff

12 5月, 2010 1 次提交

md: set mddev readonly flag on blkdev BLKROSET ioctl · e2218350

由 Dan Williams 提交于 5月 12, 2010

When the user sets the block device to readwrite then the mddev should
follow suit.  Otherwise, the BUG_ON in md_write_start() will be set to
trigger.

The reverse direction, setting mddev->ro to match a set readonly
request, can be ignored because the blkdev level readonly flag precludes
the need to have mddev->ro set correctly.  Nevermind the fact that
setting mddev->ro to 1 may fail if the array is in use.

Cc: <stable@kernel.org>
Signed-off-by: NDan Williams <dan.j.williams@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

e2218350

16 3月, 2010 1 次提交

md: deal with merge_bvec_fn in component devices better. · 627a2d3c

由 NeilBrown 提交于 3月 08, 2010

If a component device has a merge_bvec_fn then as we never call it
we must ensure we never need to.  Currently this is done by setting
max_sector to 1 PAGE, however this does not stop a bio being created
with several sub-page iovecs that would violate the merge_bvec_fn.

So instead set max_segments to 1 and set the segment boundary to the
same as a page boundary to ensure there is only ever one single-page
segment of IO requested at a time.

This can particularly be an issue when 'xen' is used as it is
known to submit multiple small buffers in a single bio.
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@kernel.org

627a2d3c

06 3月, 2010 14 次提交

dm raid1: fix deadlock when suspending failed device · f0703040

由 Takahiro Yasui 提交于 3月 06, 2010

To prevent deadlock, bios in the hold list should be flushed before
dm_rh_stop_recovery() is called in mirror_suspend().

The recovery can't start because there are pending bios and therefore
dm_rh_stop_recovery deadlocks.

When there are pending bios in the hold list, the recovery waits for
the completion of the bios after recovery_count is acquired.
The recovery_count is released when the recovery finished, however,
the bios in the hold list are processed after dm_rh_stop_recovery() in
mirror_presuspend(). dm_rh_stop_recovery() also acquires recovery_count,
then deadlock occurs.
Signed-off-by: NTakahiro Yasui <tyasui@redhat.com>
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>
Reviewed-by: NMikulas Patocka <mpatocka@redhat.com>

f0703040

dm: eliminate some holes data structures · 924e600d

由 Mike Snitzer 提交于 3月 06, 2010

Eliminate a 4-byte hole in 'struct dm_io_memory' by moving 'offset' above the
'ptr' to which it applies (size reduced from 24 to 16 bytes).  And by
association, 1-4 byte hole is eliminated in 'struct dm_io_request' (size
reduced from 56 to 48 bytes).

Eliminate all 6 4-byte holes and 1 cache-line in 'struct dm_snapshot' (size
reduced from 392 to 368 bytes).
Signed-off-by: NMike Snitzer <snitzer@redhat.com>
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>

924e600d

dm ioctl: introduce flag indicating uevent was generated · 3abf85b5

由 Peter Rajnoha 提交于 3月 06, 2010

Set a new DM_UEVENT_GENERATED_FLAG when returning from ioctls to
indicate that a uevent was actually generated.  This tells the userspace
caller that it may need to wait for the event to be processed.
Signed-off-by: NPeter Rajnoha <prajnoha@redhat.com>
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>

3abf85b5

dm: free dm_io before bio_endio not after · a97f925a

由 Mikulas Patocka 提交于 3月 06, 2010

Free the dm_io structure before calling bio_endio() instead of after it,
to ensure that the io_pool containing it is not referenced after it is
freed.

This partially fixes a problem described here
  https://www.redhat.com/archives/dm-devel/2010-February/msg00109.html

thread 1:
bio_endio(bio, io_error);
/* scheduling happens */
					thread 2:
					close the device
					remove the device
thread 1:
free_io(md, io);

Thread 2, when removing the device, sees non-empty md->io_pool (because the
io hasn't been freed by thread 1 yet) and may crash with BUG in mempool_free.
Thread 1 may also crash, when freeing into a nonexisting mempool.

To fix this we must make sure that bio_endio() is the last call and
the md structure is not accessed afterwards.

There is another bio_endio in process_barrier, but it is called from the thread
and the thread is destroyed prior to freeing the mempools, so this call is
not affected by the bug.

A similar bug exists with module unloads - the module may be unloaded
immediately after bio_endio - but that is more difficult to fix.
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Cc: stable@kernel.org
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>

a97f925a

dm table: remove unused dm_get_device range parameters · 8215d6ec

由 Nikanth Karthikesan 提交于 3月 06, 2010

Remove unused parameters(start and len) of dm_get_device()
and fix the callers.
Signed-off-by: NNikanth Karthikesan <knikanth@suse.de>
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>

8215d6ec

dm ioctl: only issue uevent on resume if state changed · 0f3649a9

由 Mike Snitzer 提交于 3月 06, 2010

Only issue a uevent on a resume if the state of the device changed,
i.e. if it was suspended and/or its table was replaced.
Signed-off-by: NDave Wysochanski <dwysocha@redhat.com>
Signed-off-by: NMike Snitzer <snitzer@redhat.com>
Cc: stable@kernel.org
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>

0f3649a9

dm raid1: always return error if all legs fail · ede5ea0b

由 Mikulas Patocka 提交于 3月 06, 2010

If all mirror legs fail, always return an error instead of holding the
bio, even if the handle_errors option was set.  At present it is the
responsibility of the driver underneath us to deal with retries,
multipath etc.

The patch adds the bio to the failures list instead of holding it
directly.  do_failures tests first if all legs failed and, if so,
returns the bio with -EIO.  If any leg is still alive and handle_errors
is set, do_failures calls hold_bio.
Reviewed-by: NTakahiro Yasui <tyasui@redhat.com>
Signed-off-by: NMikulas Patocka <mpatocka@redhat.com>
Signed-off-by: NAlasdair G Kergon <agk@redhat.com>

ede5ea0b

dm mpath: refactor pg_init · fb612642