提交 · b63d7c2e29bf9cc94989806f2df0cfca4976b830 · openeuler / raspberrypi-kernel

26 7月, 2010 13 次提交

md/bitmap: clean up plugging calls. · b63d7c2e

由 NeilBrown 提交于 6月 01, 2010

1/ use md_unplug in bitmap.c as we will soon be using bitmaps under
  arrays with no queue attached.

2/ Don't bother plugging the queue when we set a bit in the bitmap.
   The reason for this was to encourage as many bits as possible to
   get set before we unplug and write stuff out.
   However every personality already plugs the queue after
   bitmap_startwrite either directly (raid1/raid10) or be setting
   STRIPE_BIT_DELAY which causes the queue to be plugged later
   (raid5).
Signed-off-by: NNeilBrown <neilb@suse.de>

b63d7c2e

md/bitmap: reduce dependence on sysfs. · 5ff5afff

由 NeilBrown 提交于 6月 01, 2010

For dm-raid45 we will want to use bitmaps in dm-targets which don't
have entries in sysfs, so cope with the mddev not living in sysfs.
Signed-off-by: NNeilBrown <neilb@suse.de>

5ff5afff

md/bitmap: white space clean up and similar. · ac2f40be

由 NeilBrown 提交于 6月 01, 2010

Fixes some whitespace problems
Fixed some checkpatch.pl complaints.
Replaced kmalloc ... memset(0), with kzalloc
Fixed an unlikely memory leak on an error path.
Reformatted a number of 'if/else' sets, sometimes
replacing goto with an else clause.
Removed some old comments and commented-out code.
Signed-off-by: NNeilBrown <neilb@suse.de>

ac2f40be

md/raid5: export raid5 unplugging interface. · 9f7c2220

由 NeilBrown 提交于 7月 26, 2010

Also remove remaining accesses to ->queue and ->gendisk when ->queue
is NULL (As it is in a DM target).
Signed-off-by: NNeilBrown <neilb@suse.de>

9f7c2220

md/plug: optionally use plugger to unplug an array during resync/recovery. · 252ac522

由 NeilBrown 提交于 6月 01, 2010

If an array doesn't have a 'queue' then md_do_sync cannot
unplug it.
In that case it will have a 'plugger', so make that available
to the mddev, and use it to unplug the array if needed.
Signed-off-by: NNeilBrown <neilb@suse.de>

252ac522

md/raid5: add simple plugging infrastructure. · 2ac87401

由 NeilBrown 提交于 6月 01, 2010

md/raid5 uses the plugging infrastructure provided by the block layer
and 'struct request_queue'.  However when we plug raid5 under dm there
is no request queue so we cannot use that.

So create a similar infrastructure that is much lighter weight and use
it for raid5.
Signed-off-by: NNeilBrown <neilb@suse.de>

2ac87401

md/raid5: export is_congested test · 11d8a6e3

由 NeilBrown 提交于 7月 26, 2010

the dm module will need this for dm-raid45.

Also only access ->queue->backing_dev_info->congested_fn
if ->queue actually exists.  It won't in a dm target.
Signed-off-by: NNeilBrown <neilb@suse.de>

11d8a6e3

raid5: Don't set read-ahead when there is no queue · 4a5add49

由 NeilBrown 提交于 6月 01, 2010

dm-raid456 does not provide a 'queue' for raid5 to use,
so we must make raid5 stop depending on the queue.

First: read_ahead
dm handles read-ahead adjustment fully in userspace, so
simply don't do any readahead adjustments if there is
no queue.

Also re-arrange code slightly so all the accesses to ->queue are
together.

Finally, move the blk_queue_merge_bvec function into the 'if' as
the ->split_io setting in dm-raid456 has the same effect.
Signed-off-by: NNeilBrown <neilb@suse.de>

4a5add49

md: add support for raising dm events. · 768a418d

由 NeilBrown 提交于 7月 26, 2010

dm uses scheduled work to raise events to user-space.
So allow md device to have work_structs and schedule them on an error.
Signed-off-by: NNeilBrown <neilb@suse.de>

768a418d

md: export various start/stop interfaces · 390ee602

由 NeilBrown 提交于 6月 01, 2010

export entry points for starting and stopping md arrays.
This will be used by a module to make md/raid5 work under
dm.
Also stop calling md_stop_writes from md_stop, as that won't
work well with dm - it will want to call the two separately.
Signed-off-by: NNeilBrown <neilb@suse.de>

390ee602

md: split out md_rdev_init · e8bb9a83

由 NeilBrown 提交于 6月 01, 2010

This functionality will be needed separately in a subsequent patch, so
split it into it's own exported function.
Signed-off-by: NNeilBrown <neilb@suse.de>

e8bb9a83

md: be more careful setting MD_CHANGE_CLEAN · 676e42d8

由 NeilBrown 提交于 6月 01, 2010

When MD_CHANGE_CLEAN is set we might block in md_write_start.
So we should only set it when fairly sure that something will clear
it.

There are two places where it is set so as to encourage a metadata
update to record the progress of resync/recovery.  This should only
be done if the internal metadata update mechanisms are in use, which
can be tested by by inspecting '->persistent'.
Signed-off-by: NNeilBrown <neilb@suse.de>

676e42d8

md/raid5: ensure we create a unique name for kmem_cache when mddev has no gendisk · f4be6b43

由 NeilBrown 提交于 6月 01, 2010

We will shortly allow md devices with no gendisk (they are attached to
a dm-target instead).  That will cause mdname() to return 'mdX'.
There is one place where mdname really needs to be unique: when
creating the name for a slab cache.
So in that case, if there is no gendisk, you the address of the mddev
formatted in HEX to provide a unique name.
Signed-off-by: NNeilBrown <neilb@suse.de>

f4be6b43

21 7月, 2010 2 次提交

md/raid5: factor out code for changing size of stripe cache. · c41d4ac4

由 NeilBrown 提交于 6月 01, 2010

Separate the actual 'change' code from the sysfs interface
so that it can eventually be called internally.
Signed-off-by: NNeilBrown <neilb@suse.de>

c41d4ac4

md: reduce dependence on sysfs. · 00bcb4ac

由 NeilBrown 提交于 6月 01, 2010

We will want md devices to live as dm targets where sysfs is not
visible.  So allow md to not connect to sysfs.
Signed-off-by: NNeilBrown <neilb@suse.de>

00bcb4ac

24 6月, 2010 12 次提交

md/raid5: don't include 'spare' drives when reshaping to fewer devices. · 3424bf6a

由 NeilBrown 提交于 6月 17, 2010

There are few situations where it would make any sense to add a spare
when reducing the number of devices in an array, but it is
conceivable:  A 6 drive RAID6 with two missing devices could be
reshaped to a 5 drive RAID6, and a spare could become available
just in time for the reshape, but not early enough to have been
recovered first.  'freezing' recovery can make this easy to
do without any races.

However doing such a thing is a bad idea.  md will not record the
partially-recovered state of the 'spare' and when the reshape
finished it will think that the spare is still spare.
Easiest way to avoid this confusion is to simply disallow it.
Signed-off-by: NNeilBrown <neilb@suse.de>

3424bf6a

md/raid5: add a missing 'continue' in a loop. · 2f115882

由 NeilBrown 提交于 6月 17, 2010

As the comment says, the tail of this loop only applies to devices
that are not fully in sync, so if In_sync was set, we should avoid
the rest of the loop.

This bug will hardly ever cause an actual problem.  The worst it
can do is allow an array to be assembled that is dirty and degraded,
which is not generally a good idea (without warning the sysadmin
first).

This will only happen if the array is RAID4 or a RAID5/6 in an
intermediate state during a reshape and so has one drive that is
all 'parity' - no data - while some other device has failed.

This is certainly possible, but not at all common.
Signed-off-by: NNeilBrown <neilb@suse.de>

2f115882

md/raid5: Allow recovered part of partially recovered devices to be in-sync · 415e72d0

由 NeilBrown 提交于 6月 17, 2010

During a recovery of reshape the early part of some devices might be
in-sync while the later parts are not.
We we know we are looking at an early part it is good to treat that
part as in-sync for stripe calculations.

This is particularly important for a reshape which suffers device
failure.  Treating the data as in-sync can mean the difference between
data-safety and data-loss.
Signed-off-by: NNeilBrown <neilb@suse.de>

415e72d0

md/raid5: More careful check for "has array failed". · 674806d6

由 NeilBrown 提交于 6月 16, 2010

When we are reshaping an array, the device failure combinations
that cause us to decide that the array as failed are more subtle.

In particular, any 'spare' will be fully in-sync in the section
of the array that has already been reshaped, thus failures that
affect only that section are less critical.

So encode this subtlety in a new function and call it as appropriate.

The case that showed this problem was a 4 drive RAID5 to 8 drive RAID6
conversion where the last two devices failed.
This resulted in:

  good good good good incomplete good good failed failed

while converting a 5-drive RAID6 to 8 drive RAID5
The incomplete device causes the whole array to look bad,
bad as it was actually good for the section that had been
converted to 8-drives, all the data was actually safe.
Reported-by: NTerry Morris <tbmorris@tbmorris.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

674806d6

md: Don't update ->recovery_offset when reshaping an array to fewer devices. · 70fffd0b

由 NeilBrown 提交于 6月 16, 2010

When an array is reshaped to have fewer devices, the reshape proceeds
from the end of the devices to the beginning.

If a device happens to be non-In_sync (which is possible but rare)
we would normally update the ->recovery_offset as the reshape
progresses. However that would be wrong as the recover_offset records
that the early part of the device is in_sync, while in fact it would
only be the later part that is in_sync, and in any case the offset
number would be measured from the wrong end of the device.

Relatedly, if after a reshape a spare is discovered to not be
recoverred all the way to the end, not allow spare_active
to incorporate it in the array.

This becomes relevant in the following sample scenario:

A 4 drive RAID5 is converted to a 6 drive RAID6 in a combined
operation.
The RAID5->RAID6 conversion will cause a 5 drive to be included as a
spare, then the 5drive -> 6drive reshape will effectively rebuild that
spare as it progresses.  The 6th drive is treated as in_sync the whole
time as there is never any case that we might consider reading from
it, but must not because there is no valid data.

If we interrupt this reshape part-way through and reverse it to return
to a 5-drive RAID6 (or event a 4-drive RAID5), we don't want to update
the recovery_offset - as that would be wrong - and we don't want to
include that spare as active in the 5-drive RAID6 when the reversed
reshape completed and it will be mostly out-of-sync still.
Signed-off-by: NNeilBrown <neilb@suse.de>

70fffd0b

md/raid5: avoid oops when number of devices is reduced then increased. · e4e11e38

由 NeilBrown 提交于 6月 16, 2010

The entries in the stripe_cache maintained by raid5 are enlarged
when we increased the number of devices in the array, but not
shrunk when we reduce the number of devices.
So if entries are added after reducing the number of devices, we
much ensure to initialise the whole entry, not just the part that
is currently relevant.  Otherwise if we enlarge the array again,
we will reference uninitialised values.

As grow_buffers/shrink_buffer now want to use a count that is stored
explicity in the raid_conf, they should get it from there rather than
being passed it as a parameter.
Signed-off-by: NNeilBrown <neilb@suse.de>

e4e11e38

md: enable raid4->raid0 takeover · 049d6c1e

由 Maciej Trela 提交于 6月 16, 2010

Only level 5 with layout=PARITY_N can be taken over to raid0 now.
Lets allow level 4 either.
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

049d6c1e

md: clear layout after ->raid0 takeover · 001048a3

由 Maciej Trela 提交于 6月 16, 2010

After takeover from raid5/10 -> raid0 mddev->layout is not cleared.
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

001048a3

md: fix raid10 takeover: use new_layout for setup_conf · f73ea873

由 Maciej Trela 提交于 6月 16, 2010

Use mddev->new_layout in setup_conf.
Also use new_chunk, and don't set ->degraded in takeover().  That
gets set in run()
Signed-off-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

f73ea873

md: fix handling of array level takeover that re-arranges devices. · e93f68a1

由 NeilBrown 提交于 6月 15, 2010

Most array level changes leave the list of devices largely unchanged,
possibly causing one at the end to become redundant.
However conversions between RAID0 and RAID10 need to renumber
all devices (except 0).

This renumbering is currently being done in the ->run method when the
new personality takes over.  However this is too late as the common
code in md.c might already have invalidated some of the devices if
they had a ->raid_disk number that appeared to high.

Moving it into the ->takeover method is too early as the array is
still active at that time and wrong ->raid_disk numbers could cause
confusion.

So add a ->new_raid_disk field to mdk_rdev_s and use it to communicate
the new raid_disk number.
Now the common code knows exactly which devices need to be renumbered,
and which can be invalidated, and can do it all at a convenient time
when the array is suspend.
It can also update some symlinks in sysfs which previously were not be
updated correctly.
Reported-by: NMaciej Trela <maciej.trela@intel.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

e93f68a1

md: raid10: Fix null pointer dereference in fix_read_error() · 0544a21d

由 Prasanna S. Panchamukhi 提交于 6月 24, 2010

Such NULL pointer dereference can occur when the driver was fixing the
read errors/bad blocks and the disk was physically removed
causing a system crash. This patch check if the
rcu_dereference() returns valid rdev before accessing it in fix_read_error().

Cc: stable@kernel.org
Signed-off-by: NPrasanna S. Panchamukhi <prasanna.panchamukhi@riverbed.com>
Signed-off-by: NRob Becker <rbecker@riverbed.com>
Signed-off-by: NNeilBrown <neilb@suse.de>

0544a21d

Restore partition detection of newly created md arrays. · f3b99be1

由 NeilBrown 提交于 6月 24, 2010

Commit  b821eaa5 broke partition
detection for md arrays.

The logic was almost right.  However if revalidate_disk is called
when the device is not yet open, bdev->bd_disk won't be set, so the
flush_disk() Call will not set bd_invalidated.

So when md_open is called we still need to ensure that
->bd_invalidated gets set.  This is easily done with a call to
check_disk_size_change in the place where the offending commit removed
check_disk_change.  At the important times, the size will have changed
from 0 to non-zero, so check_disk_size_change will set bd_invalidated.
Tested-by: NDuncan <1i5t5.duncan@cox.net>
Reported-by: NDuncan <1i5t5.duncan@cox.net>
Signed-off-by: NNeilBrown <neilb@suse.de>

f3b99be1

28 5月, 2010 1 次提交

md: convert cpu notifier to return encapsulate errno value · 55af6bb5

由 Akinobu Mita 提交于 5月 26, 2010

By the previous modification, the cpu notifier can return encapsulate
errno value.  This converts the cpu notifiers for raid5.
Signed-off-by: NAkinobu Mita <akinobu.mita@gmail.com>
Cc: Neil Brown <neilb@suse.de>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

55af6bb5

22 5月, 2010 2 次提交

sanitize vfs_fsync calling conventions · 8018ab05

由 Christoph Hellwig 提交于 3月 22, 2010

Now that the last user passing a NULL file pointer is gone we can remove
the redundant dentry argument and associated hacks inside vfs_fsynmc_range.

The next step will be removig the dentry argument from ->fsync, but given
the luck with the last round of method prototype changes I'd rather
defer this until after the main merge window.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

8018ab05

sysfs: Implement sysfs tagged directory support. · 3ff195b0

由 Eric W. Biederman 提交于 3月 30, 2010

The problem.  When implementing a network namespace I need to be able
to have multiple network devices with the same name.  Currently this
is a problem for /sys/class/net/*, /sys/devices/virtual/net/*, and
potentially a few other directories of the form /sys/ ... /net/*.

What this patch does is to add an additional tag field to the
sysfs dirent structure.  For directories that should show different
contents depending on the context such as /sys/class/net/, and
/sys/devices/virtual/net/ this tag field is used to specify the
context in which those directories should be visible.  Effectively
this is the same as creating multiple distinct directories with
the same name but internally to sysfs the result is nicer.

I am calling the concept of a single directory that looks like multiple
directories all at the same path in the filesystem tagged directories.

For the networking namespace the set of directories whose contents I need
to filter with tags can depend on the presence or absence of hotplug
hardware or which modules are currently loaded.  Which means I need
a simple race free way to setup those directories as tagged.

To achieve a reace free design all tagged directories are created
and managed by sysfs itself.

Users of this interface:
- define a type in the sysfs_tag_type enumeration.
- call sysfs_register_ns_types with the type and it's operations
- sysfs_exit_ns when an individual tag is no longer valid

- Implement mount_ns() which returns the ns of the calling process
  so we can attach it to a sysfs superblock.
- Implement ktype.namespace() which returns the ns of a syfs kobject.

Everything else is left up to sysfs and the driver layer.

For the network namespace mount_ns and namespace() are essentially
one line functions, and look to remain that.

Tags are currently represented a const void * pointers as that is
both generic, prevides enough information for equality comparisons,
and is trivial to create for current users, as it is just the
existing namespace pointer.

The work needed in sysfs is more extensive.  At each directory
or symlink creating I need to check if the directory it is being
created in is a tagged directory and if so generate the appropriate
tag to place on the sysfs_dirent.  Likewise at each symlink or
directory removal I need to check if the sysfs directory it is
being removed from is a tagged directory and if so figure out
which tag goes along with the name I am deleting.

Currently only directories which hold kobjects, and
symlinks are supported.  There is not enough information
in the current file attribute interfaces to give us anything
to discriminate on which makes it useless, and there are
no potential users which makes it an uninteresting problem
to solve.
Signed-off-by: NEric W. Biederman <ebiederm@xmission.com>
Signed-off-by: NBenjamin Thery <benjamin.thery@bull.net>
Signed-off-by: NGreg Kroah-Hartman <gregkh@suse.de>

3ff195b0

18 5月, 2010 10 次提交

md: don't insist on valid event count for spare devices. · be6800a7

由 NeilBrown 提交于 5月 18, 2010

Devices which know that they are spares do not really need to have
an event count that matches the rest of the array, so there are no
data-in-sync issues. It is enough that the uuid matches.
So remove the requirement that the event count is up-to-date.

We currently still write out and event count on spares, but this
allows us in a year or 3 to stop doing that completely.
Signed-off-by: NNeilBrown <neilb@suse.de>

be6800a7

md: simplify updating of event count to sometimes avoid updating spares. · a8707c08

由 NeilBrown 提交于 5月 18, 2010

When updating the event count for a simple clean <-> dirty transition,
we try to avoid updating the spares so they can safely spin-down.
As the event_counts across an array must be +/- 1, this means
decrementing the event_count on a dirty->clean transition.
This is not always safe and we have to avoid the unsafe time.
We current do this with a misguided idea about it being safe or
not depending on whether the event_count is odd or even.  This
approach only works reliably in a few common instances, but easily
falls down.

So instead, simply keep internal state concerning whether it is safe
or not, and always assume it is not safe when an array is first
assembled.
Signed-off-by: NNeilBrown <neilb@suse.de>

a8707c08

md/raid6: Fix raid-6 read-error correction in degraded state · 7b0bb536

由 Gabriele A. Trombetti 提交于 4月 28, 2010

Fix: Raid-6 was not trying to correct a read-error when in
singly-degraded state and was instead dropping one more device, going to
doubly-degraded state. This patch fixes this behaviour.
Tested-by: NJanos Haar <janos.haar@netcenter.hu>
Signed-off-by: NGabriele A. Trombetti <g.trombetti.lkrnl1213@logicschema.com>
Reported-by: NJanos Haar <janos.haar@netcenter.hu>
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@kernel.org

7b0bb536

md: restore ability of spare drives to spin down. · 75a73a29

由 NeilBrown 提交于 5月 07, 2010

Some time ago we stopped the clean/active metadata updates
from being written to a 'spare' device in most cases so that
it could spin down and say spun down.  Device failure/removal
etc are still recorded on spares.

However commit 51d5668c broke this 50% of the time,
depending on whether the event count is even or odd.
The change log entry said:

   This means that the alignment between 'odd/even' and
    'clean/dirty' might take a little longer to attain,

how ever the code makes no attempt to create that alignment, so it
could take arbitrarily long.

So when we find that clean/dirty is not aligned with odd/even,
force a second metadata-update immediately.  There are already cases
where a second metadata-update is needed immediately (e.g. when a
device fails during the metadata update).  We just piggy-back on that.
Reported-by: NJoe Bryant <tenminjoe@yahoo.com>
Signed-off-by: NNeilBrown <neilb@suse.de>
Cc: stable@kernel.org

75a73a29

md: Fix read balancing in RAID1 and RAID10 on drives > 2TB · af3a2cd6

由 NeilBrown 提交于 5月 08, 2010

read_balance uses a "unsigned long" for a sector number which
will get truncated beyond 2TB.
This will cause read-balancing to be non-optimal, and can cause
data to be read from the 'wrong' branch during a resync.  This has a
very small chance of returning wrong data.
Reported-by: NJordan Russell <jr-list-2010@quo.to>
Cc: stable@kernel.org
Signed-off-by: NNeilBrown <neilb@suse.de>

af3a2cd6

N
md/linear: standardise all printk messages · 2dc40f80
由 NeilBrown 提交于 5月 03, 2010
```
  md/linear:mdname:
Signed-off-by: NNeilBrown <neilb@suse.de>
```
2dc40f80

md/raid0: tidy up printk messages. · b5a20961

由 NeilBrown 提交于 5月 03, 2010

All messages now start
   md/raid0:md-device-name:
Signed-off-by: NNeilBrown <neilb@suse.de>

b5a20961

md/raid10: tidy up printk messages. · 128595ed

由 NeilBrown 提交于 5月 03, 2010

All raid10 printk messages now start
   md/raid10:md-device-name:
Signed-off-by: NNeilBrown <neilb@suse.de>

128595ed

md/raid1: improve printk messages · 9dd1e2fa

由 NeilBrown 提交于 5月 03, 2010

Make sure the array name is included in a uniform way in all printk
messages.
Signed-off-by: NNeilBrown <neilb@suse.de>

9dd1e2fa

md/raid5: improve consistency of error messages. · 0c55e022

由 NeilBrown 提交于 5月 03, 2010

Many 'printk' messages from the raid456 module mention 'raid5' even
though it may be a 'raid6' or even 'raid4' array.  This can cause
confusion.
Also the actual array name is not always reported and when it is
it is not reported consistently.

So change all the messages to start:
    md/raid:%s:
where '%s' becomes e.g. md3 to identify the particular array.
Signed-off-by: NNeilBrown <neilb@suse.de>

0c55e022