提交 · c97c2916e25c56e878e3e94efd449e2d688fcb31 · gsplhtlxg / clone-Linux

17 8月, 2011 6 次提交

Btrfs: use plain page_address() in header fields setget functions · c97c2916

由 Li Zefan 提交于 8月 03, 2011

We've stopped using highmem for extent buffers.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

c97c2916

Btrfs: forced readonly when btrfs_drop_snapshot() fails · cb1b69f4

由 Tsutomu Itoh 提交于 8月 09, 2011

The filesystem turns readonly instead of returning the error to the
caller when detected error in btrfs_drop_snapshot().
and, because the caller doesn't check the error, the function type is
changed to 'void'.
Signed-off-by: NTsutomu Itoh <t-itoh@jp.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

cb1b69f4

Btrfs: check if there is enough space for balancing smarter · cdcb725c

由 liubo 提交于 8月 03, 2011

When checking if there is enough space for balancing a block group,
since we do not take raid types into consideration, we do not account
corrent amounts of space that we needed.  This makes us do some extra
work before we get ENOSPC.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

cdcb725c

Btrfs: fix a bug of balance on full multi-disk partitions · 38c01b96

由 liubo 提交于 8月 02, 2011

When balancing, we'll first try to shrink devices for some space,
but if it is working on a full multi-disk partition with raid protection,
we may encounter a bug, that is, while shrinking, total_bytes may be less
than bytes_used, and btrfs may allocate a dev extent that accesses out of
device's bounds.

Then we will not be able to write or read the data which stores at the end
of the device, and get the followings:

device fsid 0939f071-7ea3-46c8-95df-f176d773bfb6 devid 1 transid 10 /dev/sdb5
Btrfs detected SSD devices, enabling SSD mode
btrfs: relocating block group 476315648 flags 9
btrfs: found 4 extents
attempt to access beyond end of device
sdb5: rw=145, want=546176, limit=546147
attempt to access beyond end of device
sdb5: rw=145, want=546304, limit=546147
attempt to access beyond end of device
sdb5: rw=145, want=546432, limit=546147
attempt to access beyond end of device
sdb5: rw=145, want=546560, limit=546147
attempt to access beyond end of device
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

38c01b96

Btrfs: fix an oops of log replay · 34f3e4f2

由 liubo 提交于 8月 06, 2011

When btrfs recovers from a crash, it may hit the oops below:

------------[ cut here ]------------
kernel BUG at fs/btrfs/inode.c:4580!
[...]
RIP: 0010:[<ffffffffa03df251>]  [<ffffffffa03df251>] btrfs_add_link+0x161/0x1c0 [btrfs]
[...]
Call Trace:
 [<ffffffffa03e7b31>] ? btrfs_inode_ref_index+0x31/0x80 [btrfs]
 [<ffffffffa04054e9>] add_inode_ref+0x319/0x3f0 [btrfs]
 [<ffffffffa0407087>] replay_one_buffer+0x2c7/0x390 [btrfs]
 [<ffffffffa040444a>] walk_down_log_tree+0x32a/0x480 [btrfs]
 [<ffffffffa0404695>] walk_log_tree+0xf5/0x240 [btrfs]
 [<ffffffffa0406cc0>] btrfs_recover_log_trees+0x250/0x350 [btrfs]
 [<ffffffffa0406dc0>] ? btrfs_recover_log_trees+0x350/0x350 [btrfs]
 [<ffffffffa03d18b2>] open_ctree+0x1442/0x17d0 [btrfs]
[...]

This comes from that while replaying an inode ref item, we forget to
check those old conflicting DIR_ITEM and DIR_INDEX items in fs/file tree,
then we will come to conflict corners which lead to BUG_ON().
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Tested-by: NAndy Lutomirski <luto@mit.edu>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

34f3e4f2

Btrfs: detect wether a device supports discard · d5e2003c

由 Josef Bacik 提交于 8月 04, 2011

We have a problem where if a user specifies discard but doesn't actually support
it we will return EOPNOTSUPP from btrfs_discard_extent. This is a problem
because this gets called (in a fashion) from the tree log recovery code, which
has a nice little BUG_ON(ret) after it, which causes us to fail the tree log
replay. So instead detect wether our devices support discard when we're adding
them and then don't issue discards if we know that the device doesn't support
it. And just for good measure set ret = 0 in btrfs_issue_discard just in case
we still get EOPNOTSUPP so we don't screw anybody up like this again. Thanks,
Signed-off-by: NJosef Bacik <josef@redhat.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

d5e2003c

06 8月, 2011 1 次提交

Btrfs: force unplugs when switching from high to regular priority bios · 2ab1ba68

由 Chris Mason 提交于 8月 04, 2011

Btrfs does bio submissions from a worker thread, and each device
has a list of high priority bios and regular priority bios.

Synchronous writes go to the high priority thread while async writes
go to regular list.  This commit brings back an explicit unplug
any time we switch from high to regular priority, which makes it
easier for the block layer to give us low latencies.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

2ab1ba68

02 8月, 2011 25 次提交

Btrfs: don't call writepages from within write_full_page · 0d10ee2e

由 Josef Bacik 提交于 8月 01, 2011

When doing a writepage we call writepages to try and write out any other dirty
pages in the area. This could cause problems where we commit a transaction and
then have somebody else dirtying metadata in the area as we could end up writing
out a lot more than we care about, which could cause latency on anybody who is
waiting for the transaction to completely finish committing. Thanks,
Signed-off-by: NJosef Bacik <josef@redhat.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

0d10ee2e

Btrfs: Remove unused variable 'last_index' in file.c · 341d14f1

由 Mitch Harder 提交于 7月 12, 2011

The variable 'last_index' is calculated in the __btrfs_buffered_write
function and passed as a parameter to the prepare_pages function,
but is not used anywhere in the prepare_pages function.

Remove instances of 'last_index' in these functions.
Signed-off-by: NMitch Harder <mitch.harder@sabayonlinux.org>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

341d14f1

Btrfs: clean up for find_first_extent_bit() · 69261c4b

由 Xiao Guangrong 提交于 7月 14, 2011

find_first_extent_bit() and find_first_extent_bit_state() share
most of the code, and we can just make the former call the latter.
Signed-off-by: NXiao Guangrong <xiaoguangrong@cn.fujitsu.com>
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

69261c4b

Btrfs: clean up for wait_extent_bit() · ded91f08

由 Xiao Guangrong 提交于 7月 14, 2011

We can just use cond_resched_lock().
Signed-off-by: NXiao Guangrong <xiaoguangrong@cn.fujitsu.com>
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

ded91f08

Btrfs: clean up for insert_state() · 3150b699

由 Xiao Guangrong 提交于 7月 14, 2011

Don't duplicate set_state_bits().
Signed-off-by: NXiao Guangrong <xiaoguangrong@cn.fujitsu.com>
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

3150b699

Btrfs: remove unused members from struct extent_state · 3a6d457e

由 Xiao Guangrong 提交于 7月 14, 2011

These members are not used at all.
Signed-off-by: NXiao Guangrong <xiaoguangrong@cn.fujitsu.com>
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

3a6d457e

Btrfs: clean up code for merging extent maps · 4d2c8f62

由 Li Zefan 提交于 7月 14, 2011

unpin_extent_cache() and add_extent_mapping() shares the same code
that merges extent maps.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

4d2c8f62

Btrfs: clean up code for extent_map lookup · ed64f066

由 Li Zefan 提交于 7月 14, 2011

lookup_extent_map() and search_extent_map() can share most of code.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

ed64f066

Btrfs: clean up search_extent_mapping() · 7e016a03

由 Li Zefan 提交于 7月 14, 2011

rb_node returned by __tree_search() can be a valid pointer or NULL,
but won't be some errno.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

7e016a03

Btrfs: remove redundant code for dir item lookup · 85d85a74

由 Li Zefan 提交于 7月 14, 2011

When we search a dir item with a specific hash code, we can
just return NULL without further checking if btrfs_search_slot()
returns 1.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

85d85a74

Btrfs: make acl functions really no-op if acl is not enabled · 9b89d95a

由 Li Zefan 提交于 7月 14, 2011

So there's no overhead for something we don't use.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

9b89d95a

Btrfs: remove remaining ref-cache code · 15de900d

由 Li Zefan 提交于 7月 14, 2011

Since commit f2a97a9d
("btrfs: remove all unused functions"), there's no extern functions
at all in ref-cache.c, so just remove the remaining dead code.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

15de900d

Btrfs: remove a BUG_ON() in btrfs_commit_transaction() · b9c8300c

由 Li Zefan 提交于 7月 14, 2011

wait_for_commit() always returns 0.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

b9c8300c

Btrfs: use wait_event() · 72d63ed6

由 Li Zefan 提交于 7月 14, 2011

Use wait_event() when possible to avoid code duplication.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

72d63ed6

Btrfs: check the nodatasum flag when writing compressed files · e55179b3

由 Li Zefan 提交于 7月 14, 2011

If mounting with nodatasum option, we won't csum file data for
general write or direct-io write, and this rule should also be
applied when writing compressed files.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

e55179b3

Btrfs: copy string correctly in INO_LOOKUP ioctl · 77906a50

由 Li Zefan 提交于 7月 14, 2011

Memory areas [ptr, ptr+total_len] and [name, name+total_len]
may overlap, so it's wrong to use memcpy().
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

77906a50

Btrfs: don't print the leaf if we had an error · b783e62d

由 Josef Bacik 提交于 7月 13, 2011

In __btrfs_free_extent we will print the leaf if we fail to find the extent we
wanted, but the problem is if we get an error we won't have a leaf so often this
leads to a NULL pointer dereference and we lose the error that actually
occurred. So only print the leaf if ret > 0, which means we didn't find the
item we were looking for but we didn't error either. This way the error is
preserved.
Signed-off-by: NJosef Bacik <josef@redhat.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

b783e62d

btrfs: make btrfs_set_root_node void · bf5f32ec

由 Mark Fasheh 提交于 7月 14, 2011

This is fairly trivial - btrfs_set_root_node() - always returns zero so we
can just make it void.  All callers ignore the return code now anyway.  I
also made sure to check that none of the functions that
btrfs_set_root_node() calls returns an error that we might have needed to
catch and pass back.
Signed-off-by: NMark Fasheh <mfasheh@suse.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

bf5f32ec

Btrfs: fix oops while writing data to SSD partitions · ff1f2b44

由 liubo 提交于 7月 27, 2011

Here I have a two SSD-partitions btrfs, and they are defaultly set to
"data=raid0, metadata=raid1", then I try to fill my btrfs partition
till "No space left on device", via "dd if=/dev/zero of=/mnt/btrfs/tmp".

I get an oops panic from kernel BUG at fs/btrfs/extent-tree.c:5199!, which
refers to find_free_extent's
BUG_ON(index != get_block_group_index(block_group));

In SSD mode, in order to find enough space to alloc, we may check the
block_group cache which has been checked sometime before, but the index is not
updated, where it hits the BUG_ON.
Signed-off-by: NLiu Bo <liubo2009@cn.fujitsu.com>
Acked-by: NJosef Bacik <josef@redhat.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

ff1f2b44

Btrfs: Protect the readonly flag of block group · 61cfea9b

由 WuBo 提交于 7月 26, 2011

The access for ro in btrfs_block_group_cache should be protected
because of the racy lock in relocation.
Signed-off-by: NWu Bo <wu.bo@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

61cfea9b

btrfs: Make extent-io callbacks that never fail return void · 1bf85046

由 Jeff Mahoney 提交于 7月 21, 2011

The set/clear bit and the extent split/merge hooks only ever return 0.

 Changing them to return void simplifies the error handling cases later.

 This patch changes the hook prototypes, the single implementation of each,
 and the functions that call them to return void instead.

 Since all four of these hooks execute under a spinlock, they're necessarily
 simple.
Signed-off-by: NJeff Mahoney <jeffm@suse.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

1bf85046

Btrfs: fix readahead in file defrag · b6973aa6

由 Li Zefan 提交于 7月 20, 2011

We passed the wrong value to btrfs_force_ra(). Fix this by changing
the argument of btrfs_force_ra() from last_index to nr_page.
Signed-off-by: NLi Zefan <lizf@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

b6973aa6

Btrfs: return error to caller when btrfs_unlink() failes · b532402e

由 Tsutomu Itoh 提交于 7月 19, 2011

When btrfs_unlink_inode() and btrfs_orphan_add() in btrfs_unlink()
are error, the error code is returned to the caller instead of
BUG_ON().
Signed-off-by: NTsutomu Itoh <t-itoh@jp.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

b532402e

Btrfs:don't check the return value of __btrfs_add_inode_defrag · a0f98dde

由 Wanlong Gao 提交于 7月 18, 2011

Don't need to check the return value of __btrfs_add_inode_defrag(),
since it will always return 0.
Signed-off-by: NWanlong Gao <gaowanlong@cn.fujitsu.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

a0f98dde

Merge branch 'alloc_path' of... · b43b31bd

由 Chris Mason 提交于 8月 01, 2011

Merge branch 'alloc_path' of git://git.kernel.org/pub/scm/linux/kernel/git/mfasheh/btrfs-error-handling into for-linus

b43b31bd

28 7月, 2011 8 次提交

C

Merge branch 'integration' into for-linus · ff95acb6
由 Chris Mason 提交于 7月 27, 2011

ff95acb6

Btrfs: make sure reserve_metadata_bytes doesn't leak out strange errors · 75c195a2

由 Chris Mason 提交于 7月 27, 2011

The btrfs transaction code will return any errors that come from
reserve_metadata_bytes.  We need to make sure we don't return funny
things like 1 or EAGAIN.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

75c195a2

Btrfs: use the commit_root for reading free_space_inode crcs · 2cf8572d

由 Chris Mason 提交于 7月 26, 2011

Now that we are using regular file crcs for the free space cache,
we can deadlock if we try to read the free_space_inode while we are
updating the crc tree.

This commit fixes things by using the commit_root to read the crcs.  This is
safe because we the free space cache file would already be loaded if
that block group had been changed in the current transaction.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

2cf8572d

Btrfs: reduce extent_state lock contention for metadata · 19b6caf4

由 Chris Mason 提交于 7月 25, 2011

For metadata buffers that don't straddle pages (all of them), btrfs
can safely use the page uptodate bits and extent_buffer uptodate bit
instead of needing to use the extent_state tree.

This greatly reduces contention on the state tree lock.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

19b6caf4

Btrfs: remove lockdep magic from btrfs_next_leaf · 31533fb2

由 Chris Mason 提交于 7月 26, 2011

Before the reader/writer locks, btrfs_next_leaf needed to keep
the path blocking to avoid making lockdep upset.

Now that btrfs_next_leaf only takes read locks, this isn't required.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

31533fb2

Btrfs: make a lockdep class for each root · 85d4e461

由 Chris Mason 提交于 7月 26, 2011

This patch was originally from Tejun Heo. lockdep complains about the btrfs
locking because we sometimes take btree locks from two different trees at the
same time. The current classes are based only on level in the btree, which
isn't enough information for lockdep to figure out if the lock is safe.

This patch makes a class for each type of tree, and lumps all the FS trees that
actually have files and directories into the same class.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

85d4e461

Btrfs: switch the btrfs tree locks to reader/writer · bd681513

由 Chris Mason 提交于 7月 16, 2011

The btrfs metadata btree is the source of significant
lock contention, especially in the root node.   This
commit changes our locking to use a reader/writer
lock.

The lock is built on top of rw spinlocks, and it
extends the lock tracking to remember if we have a
read lock or a write lock when we go to blocking.  Atomics
count the number of blocking readers or writers at any
given time.

It removes all of the adaptive spinning from the old code
and uses only the spinning/blocking hints inside of btrfs
to decide when it should continue spinning.

In read heavy workloads this is dramatically faster.  In write
heavy workloads we're still faster because of less contention
on the root node lock.

We suffer slightly in dbench because we schedule more often
during write locks, but all other benchmarks so far are improved.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

bd681513

Btrfs: fix deadlock when throttling transactions · 81317fde

由 Josef Bacik 提交于 7月 24, 2011

Hit this nice little deadlock.  What happens is this

__btrfs_end_transaction with throttle set, --use_count so it equals 0
  btrfs_commit_transaction
    <somebody else actually manages to start the commit>
    btrfs_end_transaction --use_count so now its -1 <== BAD
      we just return and wait on the transaction

This is bad because we just return after our use_count is -1 and don't let go
of our num_writer count on the transaction, so the guy committing the
transaction just sits there forever.  Fix this by inc'ing our use_count if we're
going to call commit_transaction so that if we call btrfs_end_transaction it's
valid.  Thanks,
Signed-off-by: NJosef Bacik <josef@redhat.com>
Signed-off-by: NChris Mason <chris.mason@oracle.com>

81317fde