提交 · 180591bcfed1a2cec048abb21d3dab840625caab · openanolis / cloud-kernel

06 1月, 2009 13 次提交

Btrfs: Use btrfs_join_transaction to avoid deadlocks during snapshot creation · 180591bc

由 Yan Zheng 提交于 1月 06, 2009

Snapshot creation happens at a specific time during transaction commit.  We
need to make sure the code called by snapshot creation doesn't wait
for the running transaction to commit.

This changes btrfs_delete_inode and finish_pending_snaps to use
btrfs_join_transaction instead of btrfs_start_transaction to avoid deadlocks.

It would be better if btrfs_delete_inode didn't use the join, but the
call path that triggers it is:

btrfs_commit_transaction->create_pending_snapshots->
create_pending_snapshot->btrfs_lookup_dentry->
fixup_tree_root_location->btrfs_read_fs_root->
btrfs_read_fs_root_no_name->btrfs_orphan_cleanup->iput

This will be fixed in a later patch by moving the orphan cleanup to the
cleaner thread.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

180591bc

C
Btrfs: drop remaining LINUX_KERNEL_VERSION checks and compat code · 9ca03b99
由 Chris Mason 提交于 1月 06, 2009
```
Signed-off-by: NChris Mason <chris.mason@oracle.com>
```
9ca03b99

Btrfs: drop EXPORT symbols from extent_io.c · 43b774ba

由 Chris Mason 提交于 1月 05, 2009

They should stay out until this is turned into generic code.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

43b774ba

Btrfs: Fix checkpatch.pl warnings · d397712b

由 Chris Mason 提交于 1月 05, 2009

There were many, most are fixed now.  struct-funcs.c generates some warnings
but these are bogus.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

d397712b

Btrfs: Fix free block discard calls down to the block layer · 1f3c79a2

由 Liu Hui 提交于 1月 05, 2009

This is a patch to fix discard semantic to make Btrfs work with FTL and SSD.
We can improve FTL's performance by telling it which sectors are freed by file
system. But if we don't tell FTL the information of free sectors in proper
time, the transaction mechanism of Btrfs will be destroyed and Btrfs could not
roll back the previous transaction under the power loss condition.

There are some problems in the old implementation:
1, In __free_extent(), the pinned down extents should not be discarded.
2, In free_extents(), the free extents are all pinned, so they need to
be discarded in transaction committing time instead of free_extents().
3, The reserved extent used by log tree should be discard too.

This patch change discard behavior as follows:
1, For the extents which need to be free at once,
   we discard them in update_block_group().
2, Delay discarding the pinned extent in btrfs_finish_extent_commit()
   when committing transaction.
3, Remove discarding from free_extents() and __free_extent()
4, Add discard interface into btrfs_free_reserved_extent()
5, Discard sectors before updating the free space cache, otherwise,
   FTL will destroy file system data.

1f3c79a2

Btrfs: avoid orphan inode caused by log replay · ec051c0f

由 Yan Zheng 提交于 1月 05, 2009

drop_one_dir_item does not properly update inode's link count. It can be
reproduced by executing following commands:

#touch test
#sync
#rm -f test
#dd if=/dev/zero bs=4k count=1 of=test conv=fsync
#echo b > /proc/sysrq-trigger

This fixes it by adding an BTRFS_ORPHAN_ITEM_KEY for the inode
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

ec051c0f

Btrfs: avoid potential super block corruption · 2d69a0f8

由 Yan Zheng 提交于 1月 05, 2009

The data in fs_info->super_for_commit are zeros before the
first transaction commit. If tree log sync and system crash
both occur before the first transaction commit, super block
will get corrupted.

This fixes it by properly filling in the super_for_commit field at
open time.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

2d69a0f8

S
Btrfs: do not call kfree if kmalloc failed in btrfs_sysfs_add_super · dd3fd8bd
由 Shen Feng 提交于 1月 05, 2009
```
Signed-off-by: NShen Feng <shen@cn.fujitsu.com>
```
dd3fd8bd

Btrfs: fix a memory leak in btrfs_get_sb · 1f483660

由 Shen Feng 提交于 1月 05, 2009

subvol_name should be freed if error occurs.
Signed-off-by: NShen Feng <shen@cn.fujitsu.com>

1f483660

Btrfs: Fix typo in clear_state_cb · c584482b

由 Liu Hui 提交于 1月 05, 2009

In clear_state_cb, we should check 'tree->ops->clear_bit_hook' instead
of 'tree->ops->set_bit_hook'.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

c584482b

Y
Btrfs: Fix memset length in btrfs_file_write · 9aead435
由 yanhai zhu 提交于 1月 05, 2009
```
Signed-off-by: NChris Mason <chris.mason@oracle.com>
```
9aead435

Btrfs: update directory's size when creating subvol/snapshot · 52c26179

由 Yan Zheng 提交于 1月 05, 2009

Make sure directory's size properly updated when creating
subvol/snapshot.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

52c26179

Btrfs: add permission checks to the ioctls · e441d54d

由 Chris Mason 提交于 1月 05, 2009

Only root can add/remove devices
Only root can defrag subtrees
Only files open for writing can be defragged
Only files open for writing can be the destination for a clone
Signed-off-by: NChris Mason <chris.mason@oracle.com>

e441d54d

20 12月, 2008 4 次提交

fs/9p: change simple_strtol to simple_strtoul · f1d9e458

由 Julia Lawall 提交于 12月 19, 2008

Since v9ses->uid is unsigned, it would seem better to use simple_strtoul that
simple_strtol.

A simplified version of the semantic patch that makes this change is as
follows: (http://www.emn.fr/x-info/coccinelle/)

// <smpl>
@r2@
long e;
position p;
@@

e = simple_strtol@p(...)

@@
position p != r2.p;
type T;
T e;
@@

e =
- simple_strtol@p
+ simple_strtoul
  (...)
// </smpl>
Signed-off-by: NJulia Lawall <julia@diku.dk>
Acked-by: NEric Van Hensbergen <ericvh@gmail.com>

f1d9e458

9p: convert d_iname references to d_name.name · 7dd0cdc5

由 Wu Fengguang 提交于 12月 19, 2008

d_iname is rubbish for long file names.
Use d_name.name in printks instead.
Signed-off-by: NWu Fengguang <wfg@linux.intel.com>
Acked-by: NEric Van Hensbergen <ericvh@gmail.com>

7dd0cdc5

D
9p: Remove potentially bad parameter from function entry debug print. · 6ff23207
由 Duane Griffin 提交于 12月 19, 2008
```
Signed-off-by: NDuane Griffin <duaneg@dghda.com>
Signed-off-by: NEric Van Hensbergen <ericvh@gmail.com>
```
6ff23207
C
Btrfs: Fix compile warning around num_online_cpus() in a min statement · b34b086c
由 Chris Mason 提交于 12月 19, 2008
```
Signed-off-by: NChris Mason <chris.mason@oracle.com>
```
b34b086c

19 12月, 2008 3 次提交

Btrfs: set EXTENT_BOUNDARY bit before marking extent delalloc. · 1f80e4db

由 Yan Zheng 提交于 12月 19, 2008

There is a race in relocate_inode_pages, it happens when
find_delalloc_range finds the delalloc extent before the
boundary bit is set. Thank you,
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

1f80e4db

Btrfs: properly update block accounting for metadata · 34bf63c4

由 Yan Zheng 提交于 12月 19, 2008

This adds the missing block accounting code to finish_current_insert and makes
block accounting for root item properly protected by the delalloc spin lock.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

34bf63c4

Btrfs: Add missing mnt_drop_write in ioctl.c · ab67b7c1

由 Yan Zheng 提交于 12月 19, 2008

This patch adds the missing mnt_drop_write to match
mnt_want_write in btrfs_ioctl_defrag and btrfs_ioctl_clone
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

ab67b7c1

18 12月, 2008 1 次提交

cifs: fix buffer overrun in parse_DFS_referrals · 331c3135

由 Jeff Layton 提交于 12月 17, 2008

While testing a kernel with memory poisoning enabled, I saw some warnings
about the redzone getting clobbered when chasing DFS referrals. The
buffer allocation for the unicode converted version of the searchName is
too small and needs to take null termination into account.
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Acked-by: NSteve French <sfrench@us.ibm.com>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

331c3135

17 12月, 2008 1 次提交
- Y
  Btrfs: fix return value from btrfs_listxattr when buffer size is too small · b16281c3
  由 Yehuda Sadeh Weinraub 提交于 12月 17, 2008
```
The return value was being overwritten.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
```
  b16281c3
18 12月, 2008 1 次提交

Btrfs: shift all end_io work to thread pools · cad321ad

由 Chris Mason 提交于 12月 17, 2008

bio_end_io for reads without checksumming on and btree writes were
happening without using async thread pools.  This means the extent_io.c
code had to use spin_lock_irq and friends on the rb tree locks for
extent state.

There were some irq safe vs unsafe lock inversions between the delallock
lock and the extent state locks.  This patch gets rid of them by moving
all end_io code into the thread pools.

To avoid contention and deadlocks between the data end_io processing and the
metadata end_io processing yet another thread pool is added to finish
off metadata writes.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

cad321ad

17 12月, 2008 4 次提交

Btrfs: properly check free space for tree balancing · 87b29b20

由 Yan Zheng 提交于 12月 17, 2008

btrfs_insert_empty_items takes the space needed by the btrfs_item
structure into account when calculating the required free space.

So the tree balancing code shouldn't add sizeof(struct btrfs_item)
to the size when checking the free space. This patch removes these
superfluous additions.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

87b29b20

ocfs2: Add JBD2 compat feature bit. · a9772189

由 Joel Becker 提交于 12月 16, 2008

Define the OCFS2_FEATURE_COMPAT_JBD2 bit in the filesystem header.
Signed-off-by: NJoel Becker <joel.becker@oracle.com>
Signed-off-by: NMark Fasheh <mfasheh@suse.com>

a9772189

ocfs2: Always update xattr search when creating bucket. · 83099bc6

由 Tao Ma 提交于 12月 05, 2008

When we create xattr bucket during the process of xattr set, we always
need to update the ocfs2_xattr_search since even if the bucket size is
the same as block size, the offset will change because of the removal
of the ocfs2_xattr_block header.
Signed-off-by: NTao Ma <tao.ma@oracle.com>
Signed-off-by: NMark Fasheh <mfasheh@suse.com>

83099bc6

Btrfs: delete checksum items before marking blocks free · dcbdd4dc

由 Chris Mason 提交于 12月 16, 2008

Btrfs maintains a cache of blocks available for allocation in ram.  The
code that frees extents was marking the extents free and then deleting
the checksum items.

This meant it was possible the extent would be reallocated before the
checksum item was actually deleted, leading to races and other
problems as the checksums were updated for the newly allocated extent.

The fix is to delete the checksum before marking the extent free.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

dcbdd4dc

16 12月, 2008 2 次提交

Btrfs: Don't use spin*lock_irq for the delalloc lock · 75eff68e

由 Chris Mason 提交于 12月 15, 2008

The delalloc lock doesn't need to have irqs disabled, nobody that
changes the number of delalloc bytes in the FS is running with irqs off.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

75eff68e

Btrfs: Fix compressed writes on truncated pages · 42dc7bab

由 Chris Mason 提交于 12月 15, 2008

The compression code was using isize to limit the amount of data it
sent through zlib.  But, it wasn't properly limiting the looping to
just the pages inside i_size.  The end result was trying to compress
too many pages, including those that had not been setup and properly locked
down.  This made the compression code oops while trying find_get_page on a
page that didn't exist.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

42dc7bab

12 12月, 2008 4 次提交

Btrfs: fix nodatasum handling in balancing code · 17d217fe

由 Yan Zheng 提交于 12月 12, 2008

Checksums on data can be disabled by mount option, so it's
possible some data extents don't have checksums or have
invalid checksums. This causes trouble for data relocation.
This patch contains following things to make data relocation
work.

1) make nodatasum/nodatacow mount option only affects new
files. Checksums and COW on data are only controlled by the
inode flags.

2) check the existence of checksum in the nodatacow checker.
If checksums exist, force COW the data extent. This ensure that
checksum for a given block is either valid or does not exist.

3) update data relocation code to properly handle the case
of checksum missing.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

17d217fe

Btrfs: shared seed device · e4404d6e

由 Yan Zheng 提交于 12月 12, 2008

This patch makes seed device possible to be shared by
multiple mounted file systems. The sharing is achieved
by cloning seed device's btrfs_fs_devices structure.
Thanks you,
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

e4404d6e

Btrfs: fix leaking block group on balance · d2fb3437

由 Yan Zheng 提交于 12月 11, 2008

The block group structs are referenced in many different
places, and it's not safe to free while balancing.  So, those block
group structs were simply leaked instead.

This patch replaces the block group pointer in the inode with the starting byte
offset of the block group and adds reference counting to the block group
struct.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

d2fb3437

Btrfs: mnt_drop_write in ioctl_trans_end · cfc8ea87

由 Sage Weil 提交于 12月 11, 2008

Add missing mnt_drop_write to match the mnt_want_write in
btrfs_ioctl_trans_start.
Signed-off-by: NSage Weil <sage@newdream.net>

cfc8ea87

11 12月, 2008 6 次提交

Btrfs: Add checking of csum tree in balancing code · 0403e47e

由 Yan Zheng 提交于 12月 10, 2008

This updates the space balancing code for the
new checksum format.
Signed-off-by: NYan Zheng <zheng.yan@oracle.com>

0403e47e

KSYM_SYMBOL_LEN fixes · 9c246247

由 Hugh Dickins 提交于 12月 09, 2008

Miles Lane tailing /sys files hit a BUG which Pekka Enberg has tracked
to my 966c8c12 sprint_symbol(): use
less stack exposing a bug in slub's list_locations() -
kallsyms_lookup() writes a 0 to namebuf[KSYM_NAME_LEN-1], but that was
beyond the end of page provided.

The 100 slop which list_locations() allows at end of page looks roughly
enough for all the other stuff it might print after the symbol before
it checks again: break out KSYM_SYMBOL_LEN earlier than before.

Latencytop and ftrace and are using KSYM_NAME_LEN buffers where they
need KSYM_SYMBOL_LEN buffers, and vmallocinfo a 2*KSYM_NAME_LEN buffer
where it wants a KSYM_SYMBOL_LEN buffer: fix those before anyone copies
them.

[akpm@linux-foundation.org: ftrace.h needs module.h]
Signed-off-by: NHugh Dickins <hugh@veritas.com>
Cc: Christoph Lameter <cl@linux-foundation.org>
Cc Miles Lane <miles.lane@gmail.com>
Acked-by: NPekka Enberg <penberg@cs.helsinki.fi>
Acked-by: NSteven Rostedt <srostedt@redhat.com>
Acked-by: NFrederic Weisbecker <fweisbec@gmail.com>
Cc: Rusty Russell <rusty@rustcorp.com.au>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

9c246247

inotify: fix IN_ONESHOT unmount event watcher · 6ee5a399

由 Dmitri Monakhov 提交于 12月 09, 2008

On umount two event will be dispatched to watcher:

1: inotify_dev_queue_event(.., IN_UNMOUNT,..)
2: remove_watch(watch, dev)
    ->inotify_dev_queue_event(.., IN_IGNORED, ..)

But if watcher has IN_ONESHOT bit set then the watcher will be released
inside first event.  Which result in accessing invalid object later.  IMHO
it is not pure regression.  This bug wasn't triggered while initial
inotify interface testing phase because of another bug in IN_ONESHOT
handling logic :)

  commit ac74c00e
  Author: Ulisses Furquim <ulissesf@gmail.com>
  Date:   Fri Feb 8 04:18:16 2008 -0800
    inotify: fix check for one-shot watches before destroying them
    As the IN_ONESHOT bit is never set when an event is sent we must check it
    in the watch's mask and not in the event's mask.

TESTCASE:
mkdir mnt
mount -ttmpfs none mnt
mkdir mnt/d
./inotify mnt/d&
umount mnt ## << lockup or crash here

TESTSOURCE:
/* gcc -oinotify inotify.c */
#include <stdio.h>
#include <stdlib.h>
#include <sys/inotify.h>

int main(int argc, char **argv)
{
        char buf[1024];
        struct inotify_event *ie;
        char *p;
        int i;
        ssize_t l;

        p = argv[1];
        i = inotify_init();
        inotify_add_watch(i, p, ~0);

        l = read(i, buf, sizeof(buf));
        printf("read %d bytes\n", l);
        ie = (struct inotify_event *) buf;
        printf("event mask: %d\n", ie->mask);
	return 0;
}
Signed-off-by: NDmitri Monakhov <dmonakhov@openvz.org>
Cc: John McCutchan <ttb@tentacle.dhs.org>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Robert Love <rlove@google.com>
Cc: Ulisses Furquim <ulissesf@gmail.com>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

6ee5a399

pagemap: fix 32-bit pagemap regression · 49c50342

由 Matt Mackall 提交于 12月 09, 2008

The large pages fix from bcf8039e broke 32-bit pagemap by pulling the
pagemap entry code out into a function with the wrong return type.
Pagemap entries are 64 bits on all systems and unsigned long is only 32
bits on 32-bit systems.
Signed-off-by: NMatt Mackall <mpm@selenic.com>
Reported-by: NDoug Graham <dgraham@nortel.com>
Cc: Alexey Dobriyan <adobriyan@gmail.com>
Cc: Dave Hansen <dave@linux.vnet.ibm.com>
Cc: <stable@kernel.org>		[2.6.26.x, 2.6.27.x]
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

49c50342

revert "percpu_counter: new function percpu_counter_sum_and_set" · 02d21168

由 Andrew Morton 提交于 12月 09, 2008

Revert

    commit e8ced39d
    Author: Mingming Cao <cmm@us.ibm.com>
    Date:   Fri Jul 11 19:27:31 2008 -0400

        percpu_counter: new function percpu_counter_sum_and_set

As described in

	revert "percpu counter: clean up percpu_counter_sum_and_set()"

the new percpu_counter_sum_and_set() is racy against updates to the
cpu-local accumulators on other CPUs.  Revert that change.

This means that ext4 will be slow again.  But correct.
Reported-by: NEric Dumazet <dada1@cosmosbay.com>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>
Cc: Mingming Cao <cmm@us.ibm.com>
Cc: <linux-ext4@vger.kernel.org>
Cc: <stable@kernel.org>		[2.6.27.x]
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

02d21168

revert "percpu counter: clean up percpu_counter_sum_and_set()" · 71c5576f

由 Andrew Morton 提交于 12月 09, 2008

Revert

    commit 1f7c14c6
    Author: Mingming Cao <cmm@us.ibm.com>
    Date:   Thu Oct 9 12:50:59 2008 -0400

        percpu counter: clean up percpu_counter_sum_and_set()

Before this patch we had the following:

percpu_counter_sum(): return the percpu_counter's value

percpu_counter_sum_and_set(): return the percpu_counter's value, copying
that value into the central value and zeroing the per-cpu counters before
returning.

After this patch, percpu_counter_sum_and_set() has gone, and
percpu_counter_sum() gets the old percpu_counter_sum_and_set()
functionality.

Problem is, as Eric points out, the old percpu_counter_sum_and_set()
functionality was racy and wrong.  It zeroes out counters on "other" cpus,
without holding any locks which will prevent races agaist updates from
those other CPUS.

This patch reverts 1f7c14c6.  This means
that percpu_counter_sum_and_set() still has the race, but
percpu_counter_sum() does not.

Note that this is not a simple revert - ext4 has since started using
percpu_counter_sum() for its dirty_blocks counter as well.

Note that this revert patch changes percpu_counter_sum() semantics.

Before the patch, a call to percpu_counter_sum() will bring the counter's
central counter mostly up-to-date, so a following percpu_counter_read()
will return a close value.

After this patch, a call to percpu_counter_sum() will leave the counter's
central accumulator unaltered, so a subsequent call to
percpu_counter_read() can now return a significantly inaccurate result.

If there is any code in the tree which was introduced after
e8ced39d was merged, and which depends
upon the new percpu_counter_sum() semantics, that code will break.
Reported-by: NEric Dumazet <dada1@cosmosbay.com>
Cc: "David S. Miller" <davem@davemloft.net>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>
Cc: Mingming Cao <cmm@us.ibm.com>
Cc: <linux-ext4@vger.kernel.org>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

71c5576f

10 12月, 2008 1 次提交

Btrfs: Delete csum items when freeing extents · 459931ec

由 Chris Mason 提交于 12月 10, 2008

This finishes off the new checksumming code by removing csum items
for extents that are no longer in use.

The trick is doing it without racing because a single csum item may
hold csums for more than one extent.  Extra checks are added to
btrfs_csum_file_blocks to make sure that we are using the correct
csum item after dropping locks.

A new btrfs_split_item is added to split a single csum item so it
can be split without dropping the leaf lock.  This is used to
remove csum bytes from the middle of an item.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

459931ec

openanolis / cloud-kernel 大约 1 年 前同步成功

openanolis / cloud-kernel
大约 1 年前同步成功