提交 · b4d7c3c9456a311a45bc1ef8944b5ba5b176244f · openeuler / raspberrypi-kernel

24 7月, 2012 31 次提交

Btrfs: kill free_space pointer from inode structure · b4d7c3c9

由 Li Zefan 提交于 7月 09, 2012

Inodes always allocate free space with BTRFS_BLOCK_GROUP_DATA type,
which means every inode has the same BTRFS_I(inode)->free_space pointer.

This shrinks struct btrfs_inode by 4 bytes (or 8 bytes on 64 bits).
Signed-off-by: NLi Zefan <lizefan@huawei.com>

b4d7c3c9

A
btrfs read error corrected message floods the console during recovery · d5b025d5
由 Anand Jain 提交于 7月 02, 2012
```
Changing printk_in_rcu to printk_ratelimited_in_rcu will suffice
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>
```
d5b025d5

Btrfs: fix buffer leak in btrfs_next_old_leaf · e6466e35

由 Jan Schmidt 提交于 7月 04, 2012

When calling btrfs_next_old_leaf, we were leaking an extent buffer in the
rare case of using the deadlock avoidance code needed for the tree mod log.
Signed-off-by: NJan Schmidt <list.btrfs@jan-o-sch.net>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

e6466e35

Btrfs: do not count in readonly bytes · f6175efa

由 Liu Bo 提交于 7月 06, 2012

If a block group is ro, do not count its entries in when we dump space info.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

f6175efa

Btrfs: add ro notification to dump_space_info · 799ffc3c

由 Liu Bo 提交于 7月 06, 2012

Block group has ro attributes, make dump_space_info show it.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

799ffc3c

Btrfs: fix a bug of writting free space cache during balance · cf7c1ef6

由 Liu Bo 提交于 7月 06, 2012

Here is the whole story:
1)
A free space cache consists of two parts:
o  free space cache inode, which is special becase it's stored in root tree.
o  free space info, which is stored as the above inode's file data.

But we only build up another new inode and does not flush its free space info
onto disk when we _clear and setup_ free space cache, and this ends up with
that the block group cache's cache_state remains DC_SETUP instead of DC_WRITTEN.

And holding DC_SETUP means that we will not truncate this free space cache inode,
which means the disk offset of its file extent will remain _unchanged_ at least
until next transaction finishes committing itself.

2)
We can set a block group readonly when we relocate the block group.

However,
if the readonly block group covers the disk offset where our free space cache
inode is going to write, it will force the free space cache inode into
cow_file_range() and it'll end up hitting a BUG_ON.

3)
Due to the above analysis, we fix this bug by adding the missing dirty flag.

4)
However, it's not over, there is still another case, nospace_cache.

With nospace_cache, we do not want to set dirty flag, instead we just truncate
free space cache inode and bail out with setting cache state DC_WRITTEN.

We can benifit from it since it saves us another 'pre-allocation' part which
usually costs a lot.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NMiao Xie <miaox@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

cf7c1ef6

Btrfs: do not abort transaction in prealloc case · 06789384

由 Liu Bo 提交于 7月 06, 2012

During disk balance, we prealloc new file extent for file data relocation,
but we may fail in 'no available space' case, and it leads to flipping btrfs
into readonly.

It is not necessary to bail out and abort transaction since we do have several
ways to rescue ourselves from ENOSPC case.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

06789384

Btrfs: kill root from btrfs_is_free_space_inode · 83eea1f1

由 Liu Bo 提交于 7月 10, 2012

Since root can be fetched via BTRFS_I macro directly, we can save an args
for btrfs_is_free_space_inode().
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

83eea1f1

Btrfs: fix btrfs_is_free_space_inode to recognize btree inode · 51a8cf9d

由 Liu Bo 提交于 7月 10, 2012

For btree inode, its root is also 'tree root', so btree inode can be
misunderstood as a free space inode.

We should add one more check for btree inode.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

51a8cf9d

Btrfs: avoid I/O repair BUG() from btree_read_extent_buffer_pages() · c0901581

由 Stefan Behrens 提交于 7月 10, 2012

From btree_read_extent_buffer_pages(), currently repair_io_failure()
can be called with mirror_num being zero when submit_one_bio() returned
an error before. This used to cause a BUG_ON(!mirror_num) in
repair_io_failure() and indeed this is not a case that needs the I/O
repair code to rewrite disk blocks.
This commit prevents calling repair_io_failure() in this case and thus
avoids the BUG_ON() and malfunction.
Signed-off-by: NStefan Behrens <sbehrens@giantdisaster.de>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

c0901581

Btrfs: rework shrink_delalloc · f4c738c2

由 Josef Bacik 提交于 7月 02, 2012

So shrink_delalloc has grown all sorts of cruft over the years thanks to
many reworkings of how we track enospc. What happens now as we fill up the
disk is we will loop for freaking ever hoping to reclaim a arbitrary amount
of space of metadata, this was from when everybody flushed at the same time.
Now we only have people flushing one at a time. So instead of trying to
reclaim a huge amount of space, just try to flush a decent chunk of space,
and stop looping as soon as we have enough free space to satisfy our
reservation. This makes xfstests 224 go much faster. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

f4c738c2

Btrfs: do not set subvolume flags in readonly mode · b9ca0664

由 Liu Bo 提交于 6月 29, 2012

$ mkfs.btrfs /dev/sdb7
$ btrfstune -S1 /dev/sdb7
$ mount /dev/sdb7 /mnt/btrfs
mount: block device /dev/sdb7 is write-protected, mounting read-only
$ btrfs dev add /dev/sdb8 /mnt/btrfs/

Now we get a btrfs in which mnt flags has readonly but sb flags does
not.  So for those ioctls that only check sb flags with MS_RDONLY, it
is going to be a problem.
Setting subvolume flags is such an ioctl, we should use mnt_want_write_file()
to check RO flags.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>

b9ca0664

Btrfs: use mnt_want_write_file instead of mnt_want_write · e54bfa31

由 Liu Bo 提交于 6月 29, 2012

mnt_want_write_file is faster when file has been opened for write.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>

e54bfa31

Btrfs: remove redundant r/o check for superblock · 768e9dfe

由 Liu Bo 提交于 6月 29, 2012

mnt_want_write() and mnt_want_write_file() will check sb->s_flags with
MS_RDONLY, and we don't need to do it ourselves.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>

768e9dfe

Btrfs: check write access to mount earlier while creating snapshots · a874a63e

由 Liu Bo 提交于 6月 29, 2012

Move check of write access to mount into upper functions so that we can
use mnt_want_write_file instead, which is faster than mnt_want_write.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>

a874a63e

Btrfs: fix typo in cow_file_range_async and async_cow_submit · 287082b0

由 Liu Bo 提交于 6月 28, 2012

It should be 10 * 1024 * 1024.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJiri Kosina <jkosina@suse.cz>

287082b0

Btrfs: change how we indicate we're adding csums · 0e721106

由 Josef Bacik 提交于 6月 26, 2012

There is weird logic I had to put in place to make sure that when we were
adding csums that we'd used the delalloc block rsv instead of the global
block rsv. Part of this meant that we had to free up our transaction
reservation before we ran the delayed refs since csum deletion happens
during the delayed ref work. The problem with this is that when we release
a reservation we will add it to the global reserve if it is not full in
order to keep us going along longer before we have to force a transaction
commit. By releasing our reservation before we run delayed refs we don't
get the opportunity to drain down the global reserve for the work we did, so
we won't refill it as often. This isn't a problem per-se, it just results
in us possibly committing transactions more and more often, and in rare
cases could cause those WARN_ON()'s to pop in use_block_rsv because we ran
out of space in our block rsv.

This also helps us by holding onto space while the delayed refs run so we
don't end up with as many people trying to do things at the same time, which
again will help us not force commits or hit the use_block_rsv warnings.
Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

0e721106

Btrfs: return error of btrfs_update_inode() to caller · b9959295

由 Tsutomu Itoh 提交于 6月 25, 2012

We didn't check error of btrfs_update_inode(), but that error looks
easy to bubble back up.
Reviewed-by: NDavid Sterba <dsterba@suse.cz>
Signed-off-by: NTsutomu Itoh <t-itoh@jp.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

b9959295

Btrfs: fix error handling in __add_reloc_root() · 23291a04

由 Dan Carpenter 提交于 6月 25, 2012

We dereferenced "node" in the error message after freeing it.  Also
btrfs_panic() can return so we should return an error code instead of
continuing.
Signed-off-by: NDan Carpenter <dan.carpenter@oracle.com>

23291a04

Btrfs: do not ignore errors from btrfs_cleanup_fs_roots() when mounting · 44c44af2

由 Ilya Dryomov 提交于 6月 22, 2012

There used to be a BUG_ON(ret) there before EH patch (79787eaa) went in.
Bail out with EINVAL.

Cc: David Sterba <dsterba@suse.cz>
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

44c44af2

I
Btrfs: do not return EINVAL instead of ENOMEM from open_ctree() · fed425c7
由 Ilya Dryomov 提交于 6月 22, 2012
```
When bailing from open_ctree() err is returned, not ret.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
```
fed425c7

Btrfs: add DEVICE_READY ioctl · 02db0844

由 Josef Bacik 提交于 6月 21, 2012

This will be used in conjunction with btrfs device ready <dev>.  This is
needed for initrd's to have a nice and lightweight way to tell if all of the
devices needed for a file system are in the cache currently.  This keeps
them from having to do mount+sleep loops waiting for devices to show up.
Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

02db0844

Btrfs: flush delayed inodes if we're short on space · 96c3f433

由 Josef Bacik 提交于 6月 21, 2012

Those crazy gentoo guys have been complaining about ENOSPC errors on their
portage volumes. This is because doing things like untar tends to create
lots of new files which will soak up all the reservation space in the
delayed inodes. Usually this gets papered over by the fact that we will try
and commit the transaction, however if this happens in the wrong spot or we
choose not to commit the transaction you will be screwed. So add the
ability to expclitly flush delayed inodes to free up space. Please test
this out guys to make sure it works since as usual I cannot reproduce.
Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

96c3f433

btrfs: join DEV_STATS ioctls to one · b27f7c0c

由 David Sterba 提交于 6月 22, 2012

Commit c11d2c23 (Btrfs: add ioctl to get and reset the device
stats) introduced two ioctls doing almost the same thing distinguished
by just the ioctl number which encodes "do reset after read". I have
suggested

http://www.mail-archive.com/linux-btrfs@vger.kernel.org/msg16604.html

to implement it via the ioctl args. This hasn't happen, and I think we
should use a more clean way to pass flags and should not waste ioctl
numbers.

CC: Stefan Behrens <sbehrens@giantdisaster.de>
Signed-off-by: NDavid Sterba <dsterba@suse.cz>

b27f7c0c

btrfs: ignore unfragmented file checks in defrag when compression enabled - rebased · a43a2111

由 Andrew Mahone 提交于 6月 19, 2012

Rebased on btrfs-next and retested.

Inform should_defrag_range if BTRFS_DEFRAG_RANGE_COMPRESS is set. If so, skip
checks for adjacent extents and extent size when deciding whether to defrag,
as these can prevent an uncompressed and unfragmented file from being
compressed as requested.
Signed-off-by: NAndrew Mahone <andrew.mahone@gmail.com>

a43a2111

Btrfs: small naming cleanup in join_transaction() · e4b50e14

由 Dan Carpenter 提交于 6月 19, 2012

"root->fs_info" and "fs_info" are the same, but "fs_info" is prefered
because it is shorter and that's what is used in the rest of the
function.
Signed-off-by: NDan Carpenter <dan.carpenter@oracle.com>

e4b50e14

Btrfs: don't update atime on RO subvolumes · 2bc55652

由 Alexander Block 提交于 6月 15, 2012

Before the update_time inode operation was indroduced, it was
not possible to prevent updates of atime on RO subvolumes. VFS
was only able to check for RO on the mount, but did not know
anything about btrfs subvolumes.

btrfs_update_time does now check if the root is RO and skip
updating of times.
Signed-off-by: NAlexander Block <ablock84@googlemail.com>

2bc55652

Btrfs: allow mount -o remount,compress=no · 063849ea

由 Arnd Hannemann 提交于 4月 16, 2012

Btrfs allows to turn on compression on a mounted and used filesystem
by issuing mount -o remount,compress=lzo.
This patch allows to turn compression off again
while the filesystem is mounted. As suggested by David Sterba
if the compress-force option was set, it is implicitly cleared
if compression is turned off.
Tested-by: NDavid Sterba <dsterba@suse.cz>
Signed-off-by: NArnd Hannemann <arnd@arndnet.de>

063849ea

Btrfs: remove ->dirty_inode · c5c3c5f3

由 Josef Bacik 提交于 4月 05, 2012

We do all of our inode updating when we change it, and now that we do
->update_time we don't need ->dirty_inode for atime updates anymore, so just
remove it.  Thanks,
Signed-off-by: NJosef Bacik <josef@redhat.com>

c5c3c5f3

Btrfs: reduce calls to wake_up on uncontended locks · cbea5ac1

由 Chris Mason 提交于 7月 23, 2012

The btrfs locks were unconditionally calling wake_up as the
locks were released.  This lead to extra thrashing on the waitqueue,
especially for locks that were dominated by readers.
Signed-off-by: NChris Mason <chris.mason@fusionio.com>

cbea5ac1

Btrfs: don't wait around for new log writers on an SSD · e39e64ac

由 Chris Mason 提交于 7月 23, 2012

Waiting on spindles improves performance, but ssds want all the
IO as quickly as we can push it down.
Signed-off-by: NChris Mason <chris.mason@fusionio.com>

e39e64ac

03 7月, 2012 9 次提交

Btrfs: run delayed directory updates during log replay · b6305567

由 Chris Mason 提交于 7月 02, 2012

While we are resolving directory modifications in the
tree log, we are triggering delayed metadata updates to
the filesystem btrees.

This commit forces the delayed updates to run so the
replay code can find any modifications done.  It stops
us from crashing because the directory deleltion replay
expects items to be removed immediately from the tree.
Signed-off-by: NChris Mason <chris.mason@fusionio.com>
cc: stable@kernel.org

b6305567

Btrfs: hold a ref on the inode during writepages · 7fd1a3f7

由 Josef Bacik 提交于 6月 27, 2012

We can race with unlink and not actually be able to do our igrab in
btrfs_add_ordered_extent. This will result in all sorts of problems.
Instead of doing the complicated work to try and handle returning an error
properly from btrfs_add_ordered_extent, just hold a ref to the inode during
writepages. If we cannot grab a ref we know we're freeing this inode anyway
and can just drop the dirty pages on the floor, because screw them we're
going to invalidate them anyway. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

7fd1a3f7

Btrfs: fix tree log remove space corner case · bdb7d303

由 Josef Bacik 提交于 6月 27, 2012

The tree log stuff can have allocated space that we end up having split
across a bitmap and a real extent. The free space code does not deal with
this, it assumes that if it finds an extent or bitmap entry that the entire
range must fall within the entry it finds. This isn't necessarily the case,
so rework the remove function so it can handle this case properly. This
fixed two panics the user hit, first in the case where the space was
initially in a bitmap and then in an extent entry, and then the reverse
case. Thanks,
Reported-and-tested-by: NShaun Reich <sreich@kde.org>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

bdb7d303

Btrfs: fix wrong check during log recovery · 6bf02314

由 Liu Bo 提交于 6月 25, 2012

When we're evicting an inode during log recovery, we need to ensure that the inode
is not in orphan state any more, which means inode's run_time flags has _no_
BTRFS_INODE_HAS_ORPHAN_ITEM.  Thus, the BUG_ON was triggered because of a wrong
check for the flags.
Reviewed-by: NDavid Sterba <dsterba@suse.cz>
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

6bf02314

Btrfs: use _IOR for BTRFS_IOC_SUBVOL_GETFLAGS · d3a94048

由 Alexander Block 提交于 6月 25, 2012

We used the wrong ioctl macro for the getflags ioctl before.
As we don't have the set/getflags ioctls in the user space ioctl.h
at the moment, it's safe to fix it now.
Reviewed-by: NDavid Sterba <dsterba@suse.cz>
Signed-off-by: NAlexander Block <ablock84@googlemail.com>

d3a94048

Btrfs: resume balance on rw (re)mounts properly · 2b6ba629

由 Ilya Dryomov 提交于 6月 22, 2012

This introduces btrfs_resume_balance_async(), which, given that
restriper state was recovered earlier by btrfs_recover_balance(),
resumes balance in btrfs-balance kthread.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

2b6ba629

Btrfs: restore restriper state on all mounts · 68310a5e

由 Ilya Dryomov 提交于 6月 22, 2012

Fix a bug that triggered asserts in btrfs_balance() in both normal and
resume modes -- restriper state was not properly restored on read-only
mounts. This factors out resuming code from btrfs_restore_balance(),
which is now also called earlier in the mount sequence to avoid the
problem of some early writes getting the old profile.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

68310a5e

Btrfs: fix dio write vs buffered read race · c3473e83

由 Josef Bacik 提交于 6月 19, 2012

Miao pointed out there's a problem with mixing dio writes and buffered
reads. If the read happens between us invalidating the page range and
actually locking the extent we can bring in pages into page cache. Then
once the write finishes if somebody tries to read again it will just find
uptodate pages and we'll read stale data. So we need to lock the extent and
check for uptodate bits in the range. If there are uptodate bits we need to
unlock and invalidate again. This will keep this race from happening since
we will hold the extent locked until we create the ordered extent, and then
teh read side always waits for ordered extents. There was also a race in
how we updated i_size, previously we were relying on the generic DIO stuff
to adjust the i_size after the DIO had completed, but this happens outside
of the extent lock which means reads could come in and not see the updated
i_size. So instead move this work into where we create the extents, and
then this way the update ordered i_size stuff works properly in the endio
handlers. Thanks,
Signed-off-by: NJosef Bacik <josef@redhat.com>

c3473e83

Btrfs: don't count I/O statistic read errors for missing devices · 597a60fa

由 Stefan Behrens 提交于 6月 14, 2012

It is normal behaviour of the low level btrfs function btrfs_map_bio()
to complete a bio with -EIO if the device is missing, instead of just
preventing the bio creation in an earlier step.
This used to cause I/O statistic read error increments and annoying
printk_ratelimited messages. This commit fixes the issue.
Signed-off-by: NStefan Behrens <sbehrens@giantdisaster.de>
Reported-by: NCarey Underwood <cwillu@cwillu.com>

597a60fa