提交 · 65bfa6580791f8c01fbc9cd8bd73d92aea53723f · openeuler / raspberrypi-kernel

02 2月, 2016 12 次提交

Btrfs: btrfs_ioctl_clone: Truncate complete page after performing clone operation · 65bfa658

由 Chandan Rajendra 提交于 1月 21, 2016

In subpagesize-blocksize scenario, the "destination offset" argument passed to
the btrfs_ioctl_clone() can be aligned to sectorsize but may not be
necessarily aligned to the machine's page size. In such cases,
truncate_inode_pages_range() ends up zeroing out the partial page and future
read operations will return incorrect data. Hence this commit explicitly
rounds down the "destination offset" to the machine's page size.
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

65bfa658

Btrfs: Clean pte corresponding to page straddling i_size · 27772b68

由 Chandan Rajendra 提交于 1月 21, 2016

When extending a file by either "truncate up" or by writing beyond i_size, the
page which had i_size needs to be marked "read only" so that future writes to
the page via mmap interface causes btrfs_page_mkwrite() to be invoked. If not,
a write performed after extending the file via the mmap interface will find
the page to be writaeable and continue writing to the page without invoking
btrfs_page_mkwrite() i.e. we end up writing to a file without reserving disk
space.
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

27772b68

Btrfs: Fix block size returned to user space · 5a2834f8

由 Chandan Rajendra 提交于 1月 21, 2016

btrfs_getattr() returns PAGE_CACHE_SIZE as the block size. Since
generic_fillattr() already does the right thing (by obtaining block size
from inode->i_blkbits), just remove the statement from btrfs_getattr.
Reviewed-by: NJosef Bacik <jbacik@fb.com>
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

5a2834f8

Btrfs: Limit inline extents to root->sectorsize · 0c29ba99

由 Chandan Rajendra 提交于 1月 21, 2016

cow_file_range_inline() limits the size of an inline extent to
PAGE_CACHE_SIZE. This breaks in subpagesize-blocksize scenarios. Fix this by
comparing against root->sectorsize.
Reviewed-by: NJosef Bacik <jbacik@fb.com>
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

0c29ba99

Btrfs: btrfs_submit_direct_hook: Handle map_length < bio vector length · 5f4dc8fc

由 Chandan Rajendra 提交于 1月 21, 2016

In subpagesize-blocksize scenario, map_length can be less than the length of a
bio vector. Such a condition may cause btrfs_submit_direct_hook() to submit a
zero length bio. Fix this by comparing map_length against block size rather
than with bv_len.
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

5f4dc8fc

Btrfs: Use (eb->start, seq) as search key for tree modification log · 298cfd36

由 Chandan Rajendra 提交于 1月 21, 2016

In subpagesize-blocksize a page can map multiple extent buffers and hence
using (page index, seq) as the search key is incorrect. For example, searching
through tree modification log tree can return an entry associated with the
first extent buffer mapped by the page (if such an entry exists), when we are
actually searching for entries associated with extent buffers that are mapped
at position 2 or more in the page.
Reviewed-by: NLiu Bo <bo.li.liu@oracle.com>
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

298cfd36

Btrfs: Search for all ordered extents that could span across a page · dbfdb6d1

由 Chandan Rajendra 提交于 1月 21, 2016

In subpagesize-blocksize scenario it is not sufficient to search using the
first byte of the page to make sure that there are no ordered extents
present across the page. Fix this.
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

dbfdb6d1

Btrfs: btrfs_page_mkwrite: Reserve space in sectorsized units · d0b7da88

由 Chandan Rajendra 提交于 1月 21, 2016

In subpagesize-blocksize scenario, if i_size occurs in a block which is not
the last block in the page, then the space to be reserved should be calculated
appropriately.
Reviewed-by: NLiu Bo <bo.li.liu@oracle.com>
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

d0b7da88

Btrfs: fallocate: Work with sectorsized blocks · 9703fefe

由 Chandan Rajendra 提交于 1月 21, 2016

While at it, this commit changes btrfs_truncate_page() to truncate sectorsized
blocks instead of pages. Hence the function has been renamed to
btrfs_truncate_block().
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

9703fefe

Btrfs: Direct I/O read: Work on sectorsized blocks · 2dabb324

由 Chandan Rajendra 提交于 1月 21, 2016

The direct I/O read's endio and corresponding repair functions work on
page sized blocks. This commit adds the ability for direct I/O read to work on
subpagesized blocks.
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

2dabb324

Btrfs: Compute and look up csums based on sectorsized blocks · c40a3d38

由 Chandan Rajendra 提交于 1月 21, 2016

Checksums are applicable to sectorsize units. The current code uses
bio->bv_len units to compute and look up checksums. This works on machines
where sectorsize == PAGE_SIZE. This patch makes the checksum computation and
look up code to work with sectorsize units.
Reviewed-by: NLiu Bo <bo.li.liu@oracle.com>
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

c40a3d38

Btrfs: __btrfs_buffered_write: Reserve/release extents aligned to block size · 2e78c927

由 Chandan Rajendra 提交于 1月 21, 2016

Currently, the code reserves/releases extents in multiples of PAGE_CACHE_SIZE
units. Fix this by doing reservation/releases in block size units.
Signed-off-by: NChandan Rajendra <chandan@linux.vnet.ibm.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>

2e78c927

30 1月, 2016 1 次提交

Revert "btrfs: synchronize incompat feature bits with sysfs files" · e410e34f

由 Chris Mason 提交于 1月 29, 2016

This reverts commit 14e46e04.

This ends up doing sysfs operations from deep in balance (where we
should be GFP_NOFS) and under heavy balance load, we're making races
against sysfs internals.

Revert it for now while we figure things out.
Signed-off-by: NChris Mason <clm@fb.com>

e410e34f

27 1月, 2016 3 次提交

C
btrfs: don't use GFP_HIGHMEM for free-space-tree bitmap kzalloc · e1c0ebad
由 Chris Mason 提交于 1月 27, 2016
```
This was copied incorrectly from the __vmalloc call.
Signed-off-by: NChris Mason <clm@fb.com>
```
e1c0ebad

Merge branch 'dev/fst-followup' of... · d32a4e34

由 Chris Mason 提交于 1月 27, 2016

Merge branch 'dev/fst-followup' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux into for-linus-4.5

d32a4e34

btrfs: sysfs: check initialization state before updating features · bf609206

由 David Sterba 提交于 1月 27, 2016

If the mount phase is not finished, we can't update the sysfs files.
Reported-by: NChris Mason <clm@fb.com>
Signed-off-by: NDavid Sterba <dsterba@suse.com>
Signed-off-by: NChris Mason <clm@fb.com>

bf609206

26 1月, 2016 4 次提交

Revert "btrfs: clear PF_NOFREEZE in cleaner_kthread()" · 80ad623e

由 David Sterba 提交于 1月 25, 2016

This reverts commit 69624913. The
cleaner thread can block freezing when there's a snapshot cleaning in
progress and the other threads get suspended first. From the logs
provided by Martin we're waiting for reading extent pages:

kernel: PM: Syncing filesystems ... done.
kernel: Freezing user space processes ... (elapsed 0.015 seconds) done.
kernel: Freezing remaining freezable tasks ...
kernel: Freezing of tasks failed after 20.003 seconds (1 tasks refusing to freeze, wq_busy=0):
kernel: btrfs-cleaner   D ffff88033dd13bc0     0   152      2 0x00000000
kernel: ffff88032ebc2e00 ffff88032e750000 ffff88032e74fa50 7fffffffffffffff
kernel: ffffffff814a58df 0000000000000002 ffffea000934d580 ffffffff814a5451
kernel: 7fffffffffffffff ffffffff814a6e8f 0000000000000000 0000000000000020
kernel: Call Trace:
kernel: [<ffffffff814a58df>] ? bit_wait+0x2c/0x2c
kernel: [<ffffffff814a5451>] ? schedule+0x6f/0x7c
kernel: [<ffffffff814a6e8f>] ? schedule_timeout+0x2f/0xd8
kernel: [<ffffffff81076f94>] ? timekeeping_get_ns+0xa/0x2e
kernel: [<ffffffff81077603>] ? ktime_get+0x36/0x44
kernel: [<ffffffff814a4f6c>] ? io_schedule_timeout+0x94/0xf2
kernel: [<ffffffff814a4f6c>] ? io_schedule_timeout+0x94/0xf2
kernel: [<ffffffff814a590b>] ? bit_wait_io+0x2c/0x30
kernel: [<ffffffff814a5694>] ? __wait_on_bit+0x41/0x73
kernel: [<ffffffff8109eba8>] ? wait_on_page_bit+0x6d/0x72
kernel: [<ffffffff8105d718>] ? autoremove_wake_function+0x2a/0x2a
kernel: [<ffffffff811a02d7>] ? read_extent_buffer_pages+0x1bd/0x203
kernel: [<ffffffff8117d9e9>] ? free_root_pointers+0x4c/0x4c
kernel: [<ffffffff8117e831>] ? btree_read_extent_buffer_pages.constprop.57+0x5a/0xe9
kernel: [<ffffffff8117f4f3>] ? read_tree_block+0x2d/0x45
kernel: [<ffffffff8116782a>] ? read_block_for_search.isra.34+0x22a/0x26b
kernel: [<ffffffff811656c3>] ? btrfs_set_path_blocking+0x1e/0x4a
kernel: [<ffffffff8116919b>] ? btrfs_search_slot+0x648/0x736
kernel: [<ffffffff81170559>] ? btrfs_lookup_extent_info+0xb7/0x2c7
kernel: [<ffffffff81170ee5>] ? walk_down_proc+0x9c/0x1ae
kernel: [<ffffffff81171c9d>] ? walk_down_tree+0x40/0xa4
kernel: [<ffffffff8117375f>] ? btrfs_drop_snapshot+0x2da/0x664
kernel: [<ffffffff8104ff21>] ? finish_task_switch+0x126/0x167
kernel: [<ffffffff811850f8>] ? btrfs_clean_one_deleted_snapshot+0xa6/0xb0
kernel: [<ffffffff8117eaba>] ? cleaner_kthread+0x13e/0x17b
kernel: [<ffffffff8117e97c>] ? btrfs_item_end+0x33/0x33
kernel: [<ffffffff8104d256>] ? kthread+0x95/0x9d
kernel: [<ffffffff8104d1c1>] ? kthread_parkme+0x16/0x16
kernel: [<ffffffff814a7b5f>] ? ret_from_fork+0x3f/0x70
kernel: [<ffffffff8104d1c1>] ? kthread_parkme+0x16/0x16

As this affects a released kernel (4.4) we need a minimal fix for
stable kernels.

Bugzilla: https://bugzilla.kernel.org/show_bug.cgi?id=108361Reported-by: NMartin Ziegler <ziegler@uni-freiburg.de>
CC: stable@vger.kernel.org # 4.4
CC: Jiri Kosina <jkosina@suse.cz>
Signed-off-by: NDavid Sterba <dsterba@suse.com>
Signed-off-by: NChris Mason <clm@fb.com>

80ad623e

btrfs: async-thread: Fix a use-after-free error for trace · 0a95b851

由 Qu Wenruo 提交于 1月 22, 2016

Parameter of trace_btrfs_work_queued() can be freed in its workqueue.
So no one use use that pointer after queue_work().

Fix the user-after-free bug by move the trace line before queue_work().
Reported-by: NDave Jones <davej@codemonkey.org.uk>
Signed-off-by: NQu Wenruo <quwenruo@cn.fujitsu.com>
Reviewed-by: NDavid Sterba <dsterba@suse.com>
Signed-off-by: NChris Mason <clm@fb.com>

0a95b851

Btrfs: fix race between fsync and lockless direct IO writes · de0ee0ed

由 Filipe Manana 提交于 1月 21, 2016

An fsync, using the fast path, can race with a concurrent lockless direct
IO write and end up logging a file extent item that points to an extent
that wasn't written to yet. This is because the fast fsync path collects
ordered extents into a local list and then collects all the new extent
maps to log file extent items based on them, while the direct IO write
path creates the new extent map before it creates the corresponding
ordered extent (and submitting the respective bio(s)).

So fix this by making the direct IO write path create ordered extents
before the extent maps and make the fast fsync path collect any new
ordered extents after it collects the extent maps.
Note that making the fsync handler call inode_dio_wait() (after acquiring
the inode's i_mutex) would not work and lead to a deadlock when doing
AIO, as through AIO we end up in a path where the fsync handler is called
(through dio_aio_complete_work() -> dio_complete() -> vfs_fsync_range())
before the inode's dio counter is decremented (inode_dio_wait() waits
for this counter to have a value of zero).
Signed-off-by: NFilipe Manana <fdmanana@suse.com>
Signed-off-by: NChris Mason <clm@fb.com>

de0ee0ed

Merge branch 'fix/fst-sysfs' of... · 6b5aa88c

由 Chris Mason 提交于 1月 25, 2016

Merge branch 'fix/fst-sysfs' of git://git.kernel.org/pub/scm/linux/kernel/git/kdave/linux into for-linus-4.5
Signed-off-by: NChris Mason <clm@fb.com>

6b5aa88c

25 1月, 2016 2 次提交
- D
  btrfs: add free space tree to the cow-only list · 3e4c5efb
  由 David Sterba 提交于 1月 25, 2016
```
Signed-off-by: NDavid Sterba <dsterba@suse.com>
```
  3e4c5efb
- D
  btrfs: add free space tree to lockdep classes · 6b20e0ad
  由 David Sterba 提交于 1月 25, 2016
```
Signed-off-by: NDavid Sterba <dsterba@suse.com>
```
  6b20e0ad
23 1月, 2016 1 次提交

btrfs: tweak free space tree bitmap allocation · 79b134a2

由 David Sterba 提交于 1月 22, 2016

The requested bitmap size varies, observed numbers were < 4K up to 16K.
Using vmalloc unconditionally would be too heavy, we'll try contiguous
allocations first and fall back to vmalloc if there's no contig memory.
Signed-off-by: NDavid Sterba <dsterba@suse.com>

79b134a2

22 1月, 2016 4 次提交

btrfs: tests: switch to GFP_KERNEL · 8cce83ba

由 David Sterba 提交于 1月 22, 2016

There's no reason to do GFP_NOFS in tests, it's not data-heavy and
memory allocation failures would affect only developers or testers.
Signed-off-by: NDavid Sterba <dsterba@suse.com>

8cce83ba

btrfs: synchronize incompat feature bits with sysfs files · 14e46e04

由 David Sterba 提交于 1月 21, 2016

The files under /sys/fs/UUID/features get out of sync with the actual
incompat bits set for the filesystem if they change after mount (eg. the
LZO compression).

Synchronize the feature bits with the sysfs files representing them
right after we set/clear them.
Signed-off-by: NDavid Sterba <dsterba@suse.com>

14e46e04

btrfs: sysfs: introduce helper for syncing bits with sysfs files · 444e7516

由 David Sterba 提交于 1月 21, 2016

The files under /sys/fs/UUID/features get out of sync with the actual
incompat bits set for the filesystem if they change after mount. We're
going to sync them and need a helper to do that.
Signed-off-by: NDavid Sterba <dsterba@suse.com>

444e7516

btrfs: sysfs: add free-space-tree bit attribute · 3b5bb73b

由 David Sterba 提交于 1月 21, 2016

The incompat bit representing the newly added free space tree feature is
missing. Right now it will be listed only among features supported by
the module, not per-fs.
Signed-off-by: NDavid Sterba <dsterba@suse.com>

3b5bb73b

21 1月, 2016 1 次提交
- D
  btrfs: sysfs: fix typo in compat_ro attribute definition · ba2d0840
  由 David Sterba 提交于 1月 20, 2016
```
Signed-off-by: NDavid Sterba <dsterba@suse.com>
```
  ba2d0840
20 1月, 2016 12 次提交

btrfs: raid56: Use raid_write_end_io for scrub · a6111d11

由 Zhao Lei 提交于 1月 12, 2016

No need to create additional end_io function for scrub, it increased
code size and introduced some un-unified lines, as:
raid_write_parity_end_io():
        int err = bio->bi_error;
        if (bio->bi_error)
raid_write_end_io():
        int err = bio->bi_error;
        if (err)

This patch combines them.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

a6111d11

btrfs: Remove unnecessary ClearPageUptodate for raid56 · 748f4ef4

由 Zhao Lei 提交于 1月 12, 2016

PageUptodate flag already initialized to 0 for new page,
no need to set it again.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

748f4ef4

btrfs: use rbio->nr_pages to reduce calculation · 915e2290

由 Zhao Lei 提交于 3月 03, 2015

We can use rbio->stripe_npages to reduce unnecessary calculation in
many code place.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

915e2290

btrfs: Use unified stripe_page's index calculation · b7178a5f

由 Zhao Lei 提交于 3月 03, 2015

We are using different index calculation method for stripe_page in
current code:
1: (rbio->stripe_len / PAGE_CACHE_SIZE) * stripe_index + page_index
2: DIV_ROUND_UP(rbio->stripe_len, PAGE_CACHE_SIZE) * stripe_index + page_index
3: DIV_ROUND_UP(rbio->stripe_len * stripe_index, PAGE_CACHE_SIZE) + page_index
...

They can get same result when stripe_len align to PAGE_CACHE_SIZE,
this is why current code can work, intruduce and use a common function
for calculation is a better choose.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

b7178a5f

btrfs: Fix calculation of rbio->dbitmap's size calculation · bfca9a6d

由 Zhao Lei 提交于 12月 08, 2014

Current code is trying to calculate rbio->dbitmap's size to make it
align to sizeof(long), but implement haven't achived this object,
it is align to sizeof(char) instead.
This patch fixed above calculation, and use sizeof(long) instead of
fixed "8" to increate compatibility.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

bfca9a6d

btrfs: Fix no_space in write and rm loop · e1746e83

由 Zhao Lei 提交于 12月 01, 2015

I see no_space in v4.4-rc1 again in xfstests generic/102.
It happened randomly in some node only.
(one of 4 phy-node, and a kvm with non-virtio block driver)

By bisect, we can found the first-bad is:
 commit bdced438 ("block: setup bi_phys_segments after splitting")'
But above patch only triggered the bug by making bio operation
faster(or slower).

Main reason is in our space_allocating code, we need to commit
page writeback before wait it complish, this patch fixed above
bug.

BTW, there is another reason for generic/102 fail, caused by
disable default mixed-blockgroup, I'll fix it in xfstests.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

e1746e83

btrfs: merge functions for wait snapshot creation · 0bc19f90

由 Zhao Lei 提交于 1月 06, 2016

wait_for_snapshot_creation() is in same group with oher two:
 btrfs_start_write_no_snapshoting()
 btrfs_end_write_no_snapshoting()

Rename wait_for_snapshot_creation() and move it into same place
with other two.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

0bc19f90

btrfs: delete unused argument in btrfs_copy_from_user · ee22f0c4

由 Zhao Lei 提交于 1月 06, 2016

size_t write_bytes is not necessary for btrfs_copy_from_user(),
delete it.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

ee22f0c4

btrfs: Use direct way to determine raid56 write/recover mode · ad1ba2a0

由 Zhao Lei 提交于 12月 15, 2015

Old code used bbio->raid_map to determine whether in raid56
write/recover operation, because we didn't't have bbio->map_type.

Now we have direct way for this condition, rid of using
the function-relative data, and make the code more readable.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

ad1ba2a0

btrfs: Small cleanup for get index_srcdev loop · 94a97dfe

由 Zhao Lei 提交于 12月 09, 2015

1: Adjust condition in loop to make less TAB
2: Move btrfs_put_bbio()'s line for combine, and makes logic clean.
Signed-off-by: NZhao Lei <zhaolei@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

94a97dfe

btrfs: Enhance chunk validation check · f04b772b

由 Qu Wenruo 提交于 12月 15, 2015

Enhance chunk validation:
1) Num_stripes
   We already have such check but it's only in super block sys chunk
   array.
   Now check all on-disk chunks.

2) Chunk logical
   It should be aligned to sector size.
   This behavior should be *DOUBLE CHECKED* for 64K sector size like
   PPC64 or AArch64.
   Maybe we can found some hidden bugs.

3) Chunk length
   Same as chunk logical, should be aligned to sector size.

4) Stripe length
   It should be power of 2.

5) Chunk type
   Any bit out of TYPE_MAS | PROFILE_MASK is invalid.

With all these much restrict rules, several fuzzed image reported in
mail list should no longer cause kernel panic.
Reported-by: NVegard Nossum <vegard.nossum@oracle.com>
Signed-off-by: NQu Wenruo <quwenruo@cn.fujitsu.com>
Signed-off-by: NChris Mason <clm@fb.com>

f04b772b

btrfs: Enhance super validation check · 319e4d06

由 Qu Wenruo 提交于 12月 15, 2015

Enhance btrfs_check_super_valid() function by the following points:
1) Restrict sector/node size check
   Not the old max/min valid check, but also check if it's a power of 2.
   So some bogus number like 12K node size won't pass now.

2) Super flag check
   For now, there is still some inconsistency between kernel and
   btrfs-progs super flags.
   And considering btrfs-progs may add new flags for super block, this
   check will only output warning.

3) Better root alignment check
   Now root bytenr is checked against sector size.

4) Move some check into btrfs_check_super_valid().
   Like node size vs leaf size check, and PAGESIZE vs sectorsize check.
   And magic number check.
Reported-by: NVegard Nossum <vegard.nossum@oracle.com>
Signed-off-by: NQu Wenruo <quwenruo@cn.fujitsu.com>
Reviewed-by: NDavid Sterba <dsterba@suse.com>
Signed-off-by: NChris Mason <clm@fb.com>

319e4d06