提交 · 15e3004a0eb2c4879bc666d0e6a3acd973fb226c · openeuler / Kernel

09 10月, 2012 28 次提交

Btrfs: cleanup pages properly when ENOMEM in compression · 15e3004a

由 Josef Bacik 提交于 10月 05, 2012

We were freeing non-existent pages which was causing a panic for a user who
was suffering from ENOMEM. This patch fixes the problem. Thanks,
Reported-by: NJérôme Poulin <jeromepoulin@gmail.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

15e3004a

Btrfs: make filesystem read-only when submitting barrier fails · 5af3e8cc

由 Stefan Behrens 提交于 8月 01, 2012

So far the return code of barrier_all_devices() is ignored, which
means that errors are ignored. The result can be a corrupt
filesystem which is not consistent.
This commit adds code to evaluate the return code of
barrier_all_devices(). The normal btrfs_error() mechanism is used to
switch the filesystem into read-only mode when errors are detected.

In order to decide whether barrier_all_devices() should return
error or success, the number of disks that are allowed to fail the
barrier submission is calculated. This calculation accounts for the
worst RAID level of metadata, system and data. If single, dup or
RAID0 is in use, a single disk error is already considered to be
fatal. Otherwise a single disk error is tolerated.

The calculation of the number of disks that are tolerated to fail
the barrier operation is performed when the filesystem gets mounted,
when a balance operation is started and finished, and when devices
are added or removed.
Signed-off-by: NStefan Behrens <sbehrens@giantdisaster.de>

5af3e8cc

Btrfs: detect corrupted filesystem after write I/O errors · 62856a9b

由 Stefan Behrens 提交于 7月 31, 2012

In check-integrity, detect when a superblock is written that points
to blocks that have not been written to disk due to I/O write errors.
Signed-off-by: NStefan Behrens <sbehrens@giantdisaster.de>

62856a9b

Btrfs: make compress and nodatacow mount options mutually exclusive · bedb2cca

由 Andrei Popa 提交于 9月 20, 2012

If a filesystem is mounted with compression and then remounted by adding nodatacow,
the compression is disabled but the compress flag is still visible.
Also, if a filesystem is mounted with nodatacow and then remounted with compression,
nodatacow flag is still present but it's not active.
This patch:
- removes compress flags and notifies that the compression has been disabled if the
  filesystem is mounted with nodatacow
- removes nodatacow and nodatasum flags if mounted with compress.
Signed-off-by: NAndrei Popa <andrei.popa@i-neo.ro>

bedb2cca

btrfs: fix message printing · 48940662

由 Daniel J Blueman 提交于 5月 07, 2012

Fix various messages to include newline and module prefix.
Signed-off-by: NDaniel J Blueman <daniel@quora.org>

48940662

Btrfs: don't bother committing delayed inode updates when fsyncing · 94edf4ae

由 Josef Bacik 提交于 9月 25, 2012

We can just copy the in memory inode into the tree log directly, no sense in
updating the fs tree so we can copy it into the tree log tree. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

94edf4ae

btrfs: move inline function code to header file · 479ed9ab

由 Robin Dong 提交于 9月 29, 2012

When building btrfs from kernel code, it will report:

fs/btrfs/extent_io.h:281: warning: 'extent_buffer_page' declared inline after being called
fs/btrfs/extent_io.h:281: warning: previous declaration of 'extent_buffer_page' was here
fs/btrfs/extent_io.h:280: warning: 'num_extent_pages' declared inline after being called
fs/btrfs/extent_io.h:280: warning: previous declaration of 'num_extent_pages' was here

because of the wrong declaration of inline functions.
Signed-off-by: NRobin Dong <sanbai@taobao.com>

479ed9ab

Btrfs: remove unnecessary IS_ERR in bio_readpage_error() · 7a2d6a64

由 Tsutomu Itoh 提交于 10月 01, 2012

Because the value of extent_map is only a correct value or NULL,
so IS_ERR is unnecessary.
Signed-off-by: NTsutomu Itoh <t-itoh@jp.fujitsu.com>

7a2d6a64

btrfs: remove unused function btrfs_insert_some_items() · 8d1a1317

由 Robin Dong 提交于 9月 29, 2012

The function btrfs_insert_some_items() would not be called by any other functions,
so remove it.
Signed-off-by: NRobin Dong <sanbai@taobao.com>

8d1a1317

Btrfs: don't commit instead of overcommitting · 44734ed1

由 Josef Bacik 提交于 9月 28, 2012

I don't think we have the same problem that this was supposed to fix
originally since we can allocate chunks in the enospc path now. This code
is causing us to constantly commit the transaction as we get close to using
all of our available space in our currently allocated chunks, instead of
allocating another chunk and carrying on with life, which is not nice for
performance. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

44734ed1

Btrfs: confirmation of value is added before trace_btrfs_get_extent() is called · f0bd95ea

由 Tsutomu Itoh 提交于 10月 01, 2012

We should confirm the value of extent_map before calling
trace_btrfs_get_extent() because the value of extent_map has the
possibility of NULL.
Signed-off-by: NTsutomu Itoh <t-itoh@jp.fujitsu.com>

f0bd95ea

Btrfs: be smarter about dropping things from the tree log · 18ec90d6

由 Josef Bacik 提交于 9月 28, 2012

When we truncate existing items in the tree log we've been searching for
each individual item and removing them. This is unnecessary churn and
searching, just keep track of the slot we are on and how many items we need
to delete and delete them all at once. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

18ec90d6

Btrfs: don't lookup csums for prealloc extents · 6f1fed77

由 Josef Bacik 提交于 9月 26, 2012

The tree logging stuff was looking up csums to copy over for prealloc
extents which is just work we don't need to be doing.  Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

6f1fed77

Btrfs: cache extent state when writing out dirty metadata pages · e6138876

由 Josef Bacik 提交于 9月 27, 2012

Everytime we write out dirty pages we search for an offset in the tree,
convert the bits in the state, and then when we wait we search for the
offset again and clear the bits. So for every dirty range in the io tree we
are doing 4 rb searches, which is suboptimal. With this patch we are only
doing 2 searches for every cycle (modulo weird things happening). Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

e6138876

Btrfs: do not hold the file extent leaf locked when adding extent item · ce195332

由 Josef Bacik 提交于 9月 25, 2012

For some reason we unlock everything except the leaf we are on, set the path
blocking and then add the extent item for the extent we just finished
writing. I can't for the life of me figure out why we would want to do
this, and the history doesn't really indicate that there was a real reason
for it, so just remove it. This will reduce our tree lock contention on
heavy writes. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

ce195332

Btrfs: do not async metadata csumming in certain situations · de0022b9

由 Josef Bacik 提交于 9月 25, 2012

There are a coule scenarios where farming metadata csumming off to an async
thread doesn't help. The first is if our processor supports crc32c, in
which case the csumming will be fast and so the overhead of the async model
is not worth the cost. The other case is for our tree log. We will be
making that stuff dirty and writing it out and waiting for it immediately.
Even with software crc32c this gives me a ~15% increase in speed with O_SYNC
workloads. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

de0022b9

btrfs: fix min csum item size warnings in 32bit · 221b8318

由 Zach Brown 提交于 9月 20, 2012

commit 7ca4be45 limited csum items to
PAGE_CACHE_SIZE.  It used min() with incompatible types in 32bit which
generates warnings:

fs/btrfs/file-item.c: In function ‘btrfs_csum_file_blocks’:
fs/btrfs/file-item.c:717: warning: comparison of distinct pointer types lacks a cast

This uses min_t(u32,) to fix the warnings.  u32 seemed reasonable
because btrfs_root->leafsize is u32 and PAGE_CACHE_SIZE is unsigned
long.
Signed-off-by: NZach Brown <zab@zabbo.net>

221b8318

Btrfs: run delayed refs first when out of space · 67b0fd63

由 Josef Bacik 提交于 9月 24, 2012

Running delayed refs is faster than running delalloc, so lets do that first
to try and reclaim space.  This makes my fs_mark test about 20% faster.
Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

67b0fd63

Btrfs: fix orphan transaction on the freezed filesystem · 354aa0fb

由 Miao Xie 提交于 9月 20, 2012

With the following debug patch:

 static int btrfs_freeze(struct super_block *sb)
 {
+ 	struct btrfs_fs_info *fs_info = btrfs_sb(sb);
+	struct btrfs_transaction *trans;
+
+	spin_lock(&fs_info->trans_lock);
+	trans = fs_info->running_transaction;
+	if (trans) {
+		printk("Transid %llu, use_count %d, num_writer %d\n",
+			trans->transid, atomic_read(&trans->use_count),
+			atomic_read(&trans->num_writers));
+	}
+	spin_unlock(&fs_info->trans_lock);
 	return 0;
 }

I found there was a orphan transaction after the freeze operation was done.

It is because the transaction may not be committed when the transaction handle
end even though it is the last handle of the current transaction. This design
avoid committing the transaction frequently, but also introduce the above
problem.

So I add btrfs_attach_transaction() which can catch the current transaction
and commit it. If there is no transaction, it will return ENOENT, and do not
anything.

This function also can be used to instead of btrfs_join_transaction_freeze()
because it don't increase the writer counter and don't start a new transaction,
so it also can fix the deadlock between sync and freeze.

Besides that, it is used to instead of btrfs_join_transaction() in
transaction_kthread(), because if there is no transaction, the transaction
kthread needn't anything.
Signed-off-by: NMiao Xie <miaox@cn.fujitsu.com>

354aa0fb

Btrfs: add a type field for the transaction handle · a698d075

由 Miao Xie 提交于 9月 20, 2012

This patch add a type field into the transaction handle structure,
in this way, we needn't implement various end-transaction functions
and can make the code more simple and readable.
Signed-off-by: NMiao Xie <miaox@cn.fujitsu.com>

a698d075

Btrfs: fix memory leak in start_transaction() · e8830e60

由 Miao Xie 提交于 9月 19, 2012

This patch fixes memory leak of the transaction handle which happened
when starting transaction failed on a freezed fs.
Signed-off-by: NMiao Xie <miaox@cn.fujitsu.com>

e8830e60

btrfs: extended inode ref iteration · d24bec3a

由 Mark Fasheh 提交于 8月 08, 2012

The iterate_irefs in backref.c is used to build path components from inode
refs. This patch adds code to iterate extended refs as well.

I had modify the callback function signature to abstract out some of the
differences between ref structures. iref_to_path() also needed similar
changes.
Signed-off-by: NMark Fasheh <mfasheh@suse.de>

d24bec3a

btrfs: extended inode refs · f186373f

由 Mark Fasheh 提交于 8月 08, 2012

This patch adds basic support for extended inode refs. This includes support
for link and unlink of the refs, which basically gets us support for rename
as well.

Inode creation does not need changing - extended refs are only added after
the ref array is full.
Signed-off-by: NMark Fasheh <mfasheh@suse.de>

f186373f

btrfs: improved readablity for add_inode_ref · 5a1d7843

由 Jan Schmidt 提交于 8月 17, 2012

Moved part of the code into a sub function and replaced most of the gotos
by ifs, hoping that it will be easier to read now.
Signed-off-by: NJan Schmidt <list.btrfs@jan-o-sch.net>
Signed-off-by: NMark Fasheh <mfasheh@suse.de>

5a1d7843

Btrfs: handle not finding the extent exactly when logging changed extents · 0aa4a17d

由 Josef Bacik 提交于 9月 19, 2012

I started hitting warnings when running xfstest 68 in a loop because there
were EM's that were not lined up properly with the physical extents. This
is ok, if we do something like punch a hole or write to a preallocated space
or something like that we can have an EM that doesn't cover the entire
physical extent. So fix the tree logging stuff to cope with this case so we
don't just commit the transaction. With this patch I no longer see the
warnings from the tree logging code. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

0aa4a17d

btrfs: move transaction aborts to the point of failure · 005d6427

由 David Sterba 提交于 9月 18, 2012

Call btrfs_abort_transaction as early as possible when an error
condition is detected, that way the line number reported is useful
and we're not clueless anymore which error path led to the abort.
Signed-off-by: NDavid Sterba <dsterba@suse.cz>

005d6427

Btrfs: fix the missing error information in create_pending_snapshot() · 8732d44f

由 Miao Xie 提交于 9月 17, 2012

The macro btrfs_abort_transaction() can get the line number of the code
where the problem happens, so we should invoke it in the place that the
error occurs, or we will lose the line number.
Reported-by: NDavid Sterba <dave@jikos.cz>
Signed-off-by: NMiao Xie <miaox@cn.fujitsu.com>

8732d44f

Btrfs: fix off-by-one in file clone · aa42ffd9

由 Liu Bo 提交于 9月 18, 2012

Btrfs uses inclusive range end for lock_extent(), unlock_extent() and
related functions, so we made off-by-one errors in file clone.

This fixes it and also fixes some style problems.
Signed-off-by: NLiu Bo <bo.li.liu@oracle.com>

aa42ffd9

04 10月, 2012 12 次提交

btrfs: allow setting NOCOW for a zero sized file via ioctl · 7e97b8da

由 David Sterba 提交于 9月 07, 2012

Hi,

the patch si simple, but it has user visible impact and I'm not quite sure how
to resolve it.

In short, $subj says it, chattr -C supports it and we want to use it.

The conditions that acutally allow to change the NOCOW flag are clear. What if
I try to set the flag on a file that is not empty? Options:

1) whole ioctl will fail, EINVAL
2.1) ioctl will succeed, the NOCOW flag will be silently removed, but the file
     will stay COW-ed and checksummed
2.2) ioctl will succeed, flag will not be removed and a syslog message will
     warn that the COW flag has not been changed
2.2.1) dtto, no syslog message

Man page of chattr states that

 "If it is set on a file which already has data blocks, it is undefined when
 the blocks assigned to the file will be fully stable."

Yes, it's undefined and with current implementation it'll never happen. So from
this end, the user cannot expect anything. I'm trying to find a reasonable
behaviour, so that a command like 'chattr -R -aijS +C' to tweak a broad set of
flags in a deep directory does not fail unnecessarily and does not pollute the
log.

My personal preference is 2.2.1, but my dev's oppinion is skewed, not counting
the fact that I know the code and otherwise would look there before consulting
the documentation.

The patch implements 2.2.1.

david

-------------8<-------------------
From: David Sterba <dsterba@suse.cz>

It's safe to turn off checksums for a zero sized file.

http://thread.gmane.org/gmane.comp.file-systems.btrfs/18030

"We cannot switch on NODATASUM for a file that already has extents that
are checksummed. The invariant here is that either all the extents or
none are checksummed.

Theoretically it's possible to add/remove all checksums from a given
file, but it's a potentially longtime operation, the file has to be in
some intermediate state where the checksums partially exist but have to
be ignored (for the csum->nocsum) until the file is fully converted,
this brings more special cases to extent handling, it has to survive
power failure and remain consistent, and probably needs to be restarted
after next mount."
Signed-off-by: NDavid Sterba <dsterba@suse.cz>

7e97b8da

Btrfs: fix punch hole when no extent exists · c3308f84

由 Josef Bacik 提交于 9月 14, 2012

I saw the warning in btrfs_drop_extent_cache where our end is less than our
start while running xfstests 68 in a loop.  This is because we
unconditionally do drop_end = min(end, extent_end) in
__btrfs_drop_extents(), even though we may not have found an extent in the
range we were looking to drop.  So keep track of wether or not we found
something, and if we didn't just use our end.  Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

c3308f84

Btrfs: don't do anything in our ->freeze_fs and ->unfreeze_fs · 926ced12

由 Josef Bacik 提交于 9月 14, 2012

We do not need to do anything special to freeze or unfreeze, it's all taken
care of by the generic work, and what we currently have is wrong anyway
since we shouldn't be returnning to userspace with mutexes held anyway.
Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

926ced12

Btrfs: remove unused write cache pages hook · 892951a9

由 Josef Bacik 提交于 9月 14, 2012

The btree inode has it's own write cache pages so we can remove this write
cache pages hook as it's not used.  Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

892951a9

Btrfs: fix race when getting the eb out of page->private · b5bae261

由 Josef Bacik 提交于 9月 14, 2012

We can race when checking wether PagePrivate is set on a page and we
actually have an eb saved in the pages private pointer. We could have
easily written out this page and released it in the time that we did the
pagevec lookup and actually got around to looking at this page. So use
mapping->private_lock to ensure we get a consistent view of the
page->private pointer. This is inline with the alloc and releasepage paths
which use private_lock when manipulating page->private. Thanks,
Reported-by: NDavid Sterba <dave@jikos.cz>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

b5bae261

Btrfs: do not hold the write_lock on the extent tree while logging · ff44c6e3

由 Josef Bacik 提交于 9月 14, 2012

Dave Sterba pointed out a sleeping while atomic bug while doing fsync. This
is because I'm an idiot and didn't realize that rwlock's were spin locks, so
we've been holding this thing while doing allocations and such which is not
good. This patch fixes this by dropping the write lock before we do
anything heavy and re-acquire it when it is done. We also need to take a
ref on the em's in case their corresponding pages are evicted and mark them
as being logged so that releasepage does not remove them and doesn't remove
them from our local list. Thanks,
Reported-by: NDave Sterba <dave@jikos.cz>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

ff44c6e3

Btrfs: fix race with freeze and free space inodes · 98114659

由 Josef Bacik 提交于 9月 14, 2012

So we start our freeze, somebody comes in and does an fsync() on a file
where we have to commit a transaction for whatever reason, and we will
deadlock because the freeze is waiting on FS_FREEZE people to stop writing
to the file system, but the transaction is waiting for its free space inodes
to be written out, which are in turn waiting on sb_start_intwrite while
trying to write the file extents. To fix this we'll just skip the
sb_start_intwrite() if we TRANS_JOIN_NOLOCK since we're being waited on by a
transaction commit so we're safe wrt to freeze and this will keep us from
deadlocking. Thanks,
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

98114659

L
Btrfs: kill obsolete arguments in btrfs_wait_ordered_extents · 6bbe3a9c
由 Liu Bo 提交于 9月 14, 2012
```
nocow_only is now an obsolete argument.
Signed-off-by: NLiu Bo <bo.li.liu@oracle.com>
```
6bbe3a9c

Btrfs: cleanup fs_info->hashers · 2e90cf85

由 Liu Bo 提交于 9月 14, 2012

fs_info->hashers is now an obsolete one.
Signed-off-by: NLiu Bo <bo.li.liu@oracle.com>

2e90cf85

Btrfs: cleanup for duplicated code in find_free_extent · ab26e9d6

由 Liu Bo 提交于 9月 14, 2012

There is already an 'add free space' phrase in front of this one, we
needn't to redo it.
Signed-off-by: NLiu Bo <bo.li.liu@oracle.com>

ab26e9d6

Btrfs: fix race in sync and freeze again · 60376ce4

由 Josef Bacik 提交于 9月 14, 2012

I screwed this up, there is a race between checking if there is a running
transaction and actually starting a transaction in sync where we could race
with a freezer and get ourselves into trouble. To fix this we need to make
a new join type to only do the try lock on the freeze stuff. If it fails
we'll return EPERM and just return from sync. This fixes a hang Liu Bo
reported when running xfstest 68 in a loop. Thanks,
Reported-by: NLiu Bo <bo.li.liu@oracle.com>
Signed-off-by: NJosef Bacik <jbacik@fusionio.com>

60376ce4

btrfs: return EPERM upon rmdir on a subvolume · b3ae244e

由 David Sterba 提交于 9月 13, 2012

A subvolume cannot be deleted via rmdir, but the error code ENOTEMPTY
is confusing. Return EPERM instead, as this is not permitted.
Signed-off-by: NDavid Sterba <dsterba@suse.cz>

b3ae244e

openeuler / Kernel 1 年多 前同步成功

openeuler / Kernel
1 年多前同步成功