提交 · 9e38d86ff2d8a8db99570e982230861046df32b5 · openeuler / Kernel

26 10月, 2010 28 次提交

fs: Implement lazy LRU updates for inodes · 9e38d86f

由 Nick Piggin 提交于 10月 23, 2010

Convert the inode LRU to use lazy updates to reduce lock and
cacheline traffic.  We avoid moving inodes around in the LRU list
during iget/iput operations so these frequent operations don't need
to access the LRUs. Instead, we defer the refcount checks to
reclaim-time and use a per-inode state flag, I_REFERENCED, to tell
reclaim that iget has touched the inode in the past. This means that
only reclaim should be touching the LRU with any frequency, hence
significantly reducing lock acquisitions and the amount contention
on LRU updates.

This also removes the inode_in_use list, which means we now only
have one list for tracking the inode LRU status. This makes it much
simpler to split out the LRU list operations under it's own lock.
Signed-off-by: NNick Piggin <npiggin@suse.de>
Signed-off-by: NDave Chinner <dchinner@redhat.com>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

9e38d86f

fs: Convert nr_inodes and nr_unused to per-cpu counters · cffbc8aa

由 Dave Chinner 提交于 10月 23, 2010

The number of inodes allocated does not need to be tied to the
addition or removal of an inode to/from a list. If we are not tied
to a list lock, we could update the counters when inodes are
initialised or destroyed, but to do that we need to convert the
counters to be per-cpu (i.e. independent of a lock). This means that
we have the freedom to change the list/locking implementation
without needing to care about the counters.

Based on a patch originally from Eric Dumazet.

[AV: cleaned up a bit, fixed build breakage on weird configs
Signed-off-by: NDave Chinner <dchinner@redhat.com>
Reviewed-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

cffbc8aa

vfs: fix infinite loop caused by clone_mnt race · be1a16a0

由 Miklos Szeredi 提交于 10月 05, 2010

If clone_mnt() happens while mnt_make_readonly() is running, the
cloned mount might have MNT_WRITE_HOLD flag set, which results in
mnt_want_write() spinning forever on this mount.

Needs CAP_SYS_ADMIN to trigger deliberately and unlikely to happen
accidentally.  But if it does happen it can hang the machine.
Signed-off-by: NMiklos Szeredi <mszeredi@suse.cz>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

be1a16a0

A
switch hfs to hlist_add_fake() · 89b0fc38
由 Al Viro 提交于 10月 23, 2010
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
89b0fc38

list.h: new helper - hlist_add_fake() · 756acc2d

由 Al Viro 提交于 10月 23, 2010

Make node look as if it was on hlist, with hlist_del()
working correctly.  Usable without any locking...

Convert a couple of places where we want to do that to
inode->i_hash.
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

756acc2d

new helper: inode_unhashed() · 1d3382cb

由 Al Viro 提交于 10月 23, 2010

note: for race-free uses you inode_lock held
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

1d3382cb

A
unexport invalidate_inodes · a8dade34
由 Al Viro 提交于 10月 24, 2010
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
a8dade34
A
smbfs never retains inodes with zero refcount in the first place · 61ebdb42
由 Al Viro 提交于 10月 24, 2010
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
61ebdb42

ntfs: don't call invalidate_inodes() · 70fd136e

由 Al Viro 提交于 10月 24, 2010

We are in fill_super(); again, no inodes with zero i_count could
be around until we set MS_ACTIVE.
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

70fd136e

gfs2: invalidate_inodes() is no-op there · 9dcefee5

由 Al Viro 提交于 10月 24, 2010

In fill_super() we hadn't MS_ACTIVE set yet, so there won't
be any inodes with zero i_count sitting around.

In put_super() we already have MS_ACTIVE removed *and* we
had called invalidate_inodes() since then.  So again there
won't be any inodes with zero i_count...
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

9dcefee5

ext2_remount: don't bother with invalidate_inodes() · 8e3b9a07

由 Al Viro 提交于 10月 24, 2010

It's pointless - we *do* have busy inodes (root directory,
for one), so that call will fail and attempt to change
XIP flag will be ignored.
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

8e3b9a07

fs/buffer.c: call __block_write_begin() if we have page · 309f77ad

由 Namhyung Kim 提交于 10月 25, 2010

If we have the appropriate page already, call __block_write_begin()
directly instead of releasing and regrabbing it inside of
block_write_begin().
Signed-off-by: NNamhyung Kim <namhyung@gmail.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

309f77ad

lockdep: fixup checking of dir inode annotation · a3314a0e

由 Namhyung Kim 提交于 10月 11, 2010

Since inode->i_mode shares its bits for S_IFMT, S_ISDIR should be
used to distinguish whether it is a dir or not.
Signed-off-by: NNamhyung Kim <namhyung@gmail.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

a3314a0e

aio: bump i_count instead of using igrab · 306fb097

由 Chris Mason 提交于 8月 23, 2010

The aio batching code is using igrab to get an extra reference on the
inode so it can safely batch.  igrab will go ahead and take the global
inode spinlock, which can be a bottleneck on large machines doing lots
of AIO.

In this case, igrab isn't required because we already have a reference
on the file handle.  It is safe to just bump the i_count directly
on the inode.

Benchmarking shows this patch brings IOP/s on tons of flash up by about
2.5X.
Signed-off-by: NChris Mason <chris.mason@oracle.com>

306fb097

fs/buffer.c: remove duplicated assignment on b_private · 8358e7d7

由 Namhyung Kim 提交于 10月 16, 2010

bh->b_private is initialized within init_buffer(), thus the
assignment should be redundant. Remove it.
Signed-off-by: NNamhyung Kim <namhyung@gmail.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

8358e7d7

fs: move exportfs since it is not a networking filesystem · bb1e5f8c

由 Randy Dunlap 提交于 10月 15, 2010

Move the EXPORTFS kconfig symbol out of the NETWORK_FILESYSTEMS block
since it provides a library function that can be (and is) used by other
(non-network) filesystems.

This also eliminates a kconfig dependency warning:

warning: (XFS_FS && BLOCK || NFSD && NETWORK_FILESYSTEMS && INET && FILE_LOCKING && BKL) selects EXPORTFS which has unmet direct dependencies (NETWORK_FILESYSTEMS)
Signed-off-by: NRandy Dunlap <randy.dunlap@oracle.com>
Cc: Dave Chinner <david@fromorbit.com>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Alex Elder <aelder@sgi.com>
Cc: xfs-masters@oss.sgi.com
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

bb1e5f8c

hfs: use sync_dirty_buffer · 3072b90c

由 Christoph Hellwig 提交于 10月 06, 2010

Use sync_dirty_buffer instead of the incorrect opencoding it.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

3072b90c

vfs: introduce FMODE_UNSIGNED_OFFSET for allowing negative f_pos · 4a3956c7

由 KAMEZAWA Hiroyuki 提交于 10月 01, 2010

Now, rw_verify_area() checsk f_pos is negative or not.  And if negative,
returns -EINVAL.

But, some special files as /dev/(k)mem and /proc/<pid>/mem etc..  has
negative offsets.  And we can't do any access via read/write to the
file(device).

So introduce FMODE_UNSIGNED_OFFSET to allow negative file offsets.
Signed-off-by: NWu Fengguang <fengguang.wu@intel.com>
Signed-off-by: NKAMEZAWA Hiroyuki <kamezawa.hiroyu@jp.fujitsu.com>
Cc: Al Viro <viro@ZenIV.linux.org.uk>
Cc: Heiko Carstens <heiko.carstens@de.ibm.com>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

4a3956c7

hostfs: fix UML crash: remove f_spare from hostfs · ba10f486

由 Richard Weinberger 提交于 10月 19, 2010

365b1818 ("add f_flags to struct statfs(64)") resized f_spare within
struct statfs which caused a UML crash.  There is no need to copy f_spare.
Signed-off-by: NRichard Weinberger <richard@nod.at>
Reported-by: NToralf Förster <toralf.foerster@gmx.de>
Tested-by: NToralf Förster <toralf.foerster@gmx.de>
Cc: Christoph Hellwig <hch@lst.de>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Jeff Dike <jdike@addtoit.com>
Signed-off-by: NAndrew Morton <akpm@linux-foundation.org>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

ba10f486

affs: testing the wrong variable · 0e45b67d

由 Dan Carpenter 提交于 8月 25, 2010

The intent was to verify that bh = affs_bread_ino(...) returned a valid
pointer.  We checked "ext_bh" earlier in the function and it's valid
here.
Signed-off-by: NDan Carpenter <error27@gmail.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

0e45b67d

fs: allow for more than 2^31 files · 7e360c38

由 Eric Dumazet 提交于 10月 05, 2010

Andrew,

Could you please review this patch, you probably are the right guy to
take it, because it crosses fs and net trees.

Note : /proc/sys/fs/file-nr is a read-only file, so this patch doesnt
depend on previous patch (sysctl: fix min/max handling in
__do_proc_doulongvec_minmax())

Thanks !

[PATCH V4] fs: allow for more than 2^31 files

Robin Holt tried to boot a 16TB system and found af_unix was overflowing
a 32bit value :

<quote>

We were seeing a failure which prevented boot.  The kernel was incapable
of creating either a named pipe or unix domain socket.  This comes down
to a common kernel function called unix_create1() which does:

        atomic_inc(&unix_nr_socks);
        if (atomic_read(&unix_nr_socks) > 2 * get_max_files())
                goto out;

The function get_max_files() is a simple return of files_stat.max_files.
files_stat.max_files is a signed integer and is computed in
fs/file_table.c's files_init().

        n = (mempages * (PAGE_SIZE / 1024)) / 10;
        files_stat.max_files = n;

In our case, mempages (total_ram_pages) is approx 3,758,096,384
(0xe0000000).  That leaves max_files at approximately 1,503,238,553.
This causes 2 * get_max_files() to integer overflow.

</quote>

Fix is to let /proc/sys/fs/file-nr & /proc/sys/fs/file-max use long
integers, and change af_unix to use an atomic_long_t instead of
atomic_t.

get_max_files() is changed to return an unsigned long.
get_nr_files() is changed to return a long.

unix_nr_socks is changed from atomic_t to atomic_long_t, while not
strictly needed to address Robin problem.

Before patch (on a 64bit kernel) :
# echo 2147483648 >/proc/sys/fs/file-max
# cat /proc/sys/fs/file-max
-18446744071562067968

After patch:
# echo 2147483648 >/proc/sys/fs/file-max
# cat /proc/sys/fs/file-max
2147483648
# cat /proc/sys/fs/file-nr
704     0       2147483648
Reported-by: NRobin Holt <holt@sgi.com>
Signed-off-by: NEric Dumazet <eric.dumazet@gmail.com>
Acked-by: NDavid Miller <davem@davemloft.net>
Reviewed-by: NRobin Holt <holt@sgi.com>
Tested-by: NRobin Holt <holt@sgi.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

7e360c38

isofs: Fix isofs_get_blocks for 8TB files · fde214d4

由 Jan Kara 提交于 10月 04, 2010

Currently isofs_get_blocks() was limited to handle only 4TB files on 32-bit
architectures because of unnecessary use of iblock variable which was signed
long. Just remove the variable. The error messages that were using this
variable should have rather used b_off anyway because that is the block we
are currently mapping.
Signed-off-by: NJan Kara <jack@suse.cz>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

fde214d4

fs: kill block_prepare_write · ebdec241

由 Christoph Hellwig 提交于 10月 06, 2010

__block_write_begin and block_prepare_write are identical except for slightly
different calling conventions.  Convert all callers to the __block_write_begin
calling conventions and drop block_prepare_write.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

ebdec241

fs: mark destroy_inode static · 56b0dacf

由 Christoph Hellwig 提交于 10月 06, 2010

Hugetlbfs used to need it, but after the destroy_inode and evict_inode
changes it's not required anymore.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

56b0dacf

fs: add sync_inode_metadata · c3765016

由 Christoph Hellwig 提交于 10月 06, 2010

Add a new helper to write out the inode using the writeback code,
that is including the correct dirty bit and list manipulation.  A few
of filesystems already opencode this, and a lot of others should be
using it instead of using write_inode_now which also writes out the
data.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

c3765016

fs: move permission check back into __lookup_hash · 81fca444

由 Christoph Hellwig 提交于 10月 06, 2010

The caller that didn't need it is gone.
Signed-off-by: NChristoph Hellwig <hch@lst.de>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

81fca444

exofs: Remove inode->i_count manipulation in exofs_new_inode · fe2fd9ed

由 Boaz Harrosh 提交于 10月 16, 2010

exofs_new_inode() was incrementing the inode->i_count and
decrementing it in create_done(), in a bad attempt to make sure
the inode will still be there when the asynchronous create_done()
finally arrives. This was very stupid because iput() was not called,
and if it was actually needed, it would leak the inode.

However all this is not needed, because at exofs_evict_inode()
we already wait for create_done() by waiting for the
object_created event. Therefore remove the superfluous ref counting
and just Thicken the comment at exofs_evict_inode() a bit.

While at it change places that open coded wait_obj_created()
to call the already available wrapper.

CC: Dave Chinner <dchinner@redhat.com>
CC: Christoph Hellwig <hch@lst.de>
CC: Nick Piggin <npiggin@kernel.dk>
Signed-off-by: NBoaz Harrosh <bharrosh@panasas.com>

fe2fd9ed

fs/exofs: typo fix of faild to failed · 571f7f46

由 Joe Perches 提交于 10月 21, 2010

Signed-off-by: NJoe Perches <joe@perches.com>
Signed-off-by: NBoaz Harrosh <bharrosh@panasas.com>

571f7f46

25 10月, 2010 4 次提交

Coda: replace BKL with mutex · da47c19e

由 Yoshihisa Abe 提交于 10月 25, 2010

Replace the BKL with a mutex to protect the venus_comm structure which
binds the mountpoint with the character device and holds the upcall
queues.
Signed-off-by: NYoshihisa Abe <yoshiabe@cs.cmu.edu>
Signed-off-by: NJan Harkes <jaharkes@cs.cmu.edu>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

da47c19e

Coda: push BKL regions into coda_upcall() · f7cc02b8

由 Yoshihisa Abe 提交于 10月 25, 2010

Now that shared inode state is locked using the cii->c_lock, the BKL is
only used to protect the upcall queues used to communicate with the
userspace cache manager. The remaining state is all local and we can
push the lock further down into coda_upcall().
Signed-off-by: NYoshihisa Abe <yoshiabe@cs.cmu.edu>
Signed-off-by: NJan Harkes <jaharkes@cs.cmu.edu>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

f7cc02b8

Coda: add spin lock to protect accesses to struct coda_inode_info. · b5ce1d83

由 Yoshihisa Abe 提交于 10月 25, 2010

We mostly need it to protect cached user permissions. The c_flags field
is advisory, reading the wrong value is harmless and in the worst case
we hit a slow path where we have to make an extra upcall to the
userspace cache manager when revalidating a dentry or inode.
Signed-off-by: NYoshihisa Abe <yoshiabe@cs.cmu.edu>
Signed-off-by: NJan Harkes <jaharkes@cs.cmu.edu>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

b5ce1d83

Revert "block: fix accounting bug on cross partition merges" · f253b86b

由 Jens Axboe 提交于 10月 24, 2010

This reverts commit 7681bfee.

Conflicts:

	include/linux/genhd.h

It has numerous issues with the cleanup path and non-elevator
devices. Revert it for now so we can come up with a clean
version without rushing things.
Signed-off-by: NJens Axboe <jaxboe@fusionio.com>

f253b86b

23 10月, 2010 8 次提交

Revert "tty: Add a new file /proc/tty/consoles" · 6c2754c2

由 Linus Torvalds 提交于 10月 23, 2010

This reverts commit f4a3e0bc.  Jiri
Sladby points out that the tty structure we're using may already be
gone, and Al Viro doesn't hold back in complaining about the random
loading of 'filp->private_data' which doesn't have to be a pointer at
all, nor does checking the magic field for TTY_MAGIC prove anything.

Belated review by Al:

 "a) global variable depending on stdin of the last opener? Affecting
     output of read(2)? Really?

  b) iterator is broken; list should be locked in ->start(), unlocked in
     ->stop() and *NOT* unlocked/relocked in ->next()

  c) ->show() ought to do nothing in case of ->device == NULL, instead
     of skipping those in ->next()/->start()

  d) regardless of the merits of the bright idea about asterisk at that
     line in output *and* regardless of (a), the implementation is not
     only atrociously ugly, it's actually very likely to be a roothole.
     Verifying that Cthulhu knows what number happens to be address of a
     tty_struct by blindly dereferencing memory at that address...
     Ouch.

  Please revert that crap."

And Christoph pipes in and NAK's the approach of walking fd tables etc
too.  So it's pretty unanimous.
Noticed-by: NJri Slaby <jslaby@suse.cz>
Requested-by: NAl Viro <viro@zeniv.linux.org.uk>
Cc: Greg Kroah-Hartman <gregkh@suse.de>
Cc: Werner Fink <werner@suse.de>
Cc: Alan Cox <alan@lxorguk.ukuu.org.uk>
Cc: Christoph Hellwig <hch@infradead.org>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

6c2754c2

ocfs2: drop the BLKDEV_IFL_WAIT flag · f8cae0f0

由 Linus Torvalds 提交于 10月 22, 2010

Commit dd3932ed ("block: remove BLKDEV_IFL_WAIT") had removed the
flag argument to blkdev_issue_flush(), but the ocfs2 merge brought in a
new one. It didn't cause a merge conflict, so the merges silently
worked out fine, but the result didn't actually compile.
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

f8cae0f0

nilfs2: eliminate sparse warning - "context imbalance" · 6b81e14e

由 Jiro SEKIBA 提交于 10月 14, 2010

insert sparse annotations to fix following sparse warning.

fs/nilfs2/segment.c:2681:3: warning: context imbalance in 'nilfs_segctor_kill_thread' - unexpected unlock

nilfs_segctor_kill_thread is only called inside sc_state_lock lock.
sparse doesn't detect the context and warn "unexpected unlock".
__acquires/__releases pretend to lock/unlock the sc_state_lock for sparse.
Signed-off-by: NJiro SEKIBA <jir@unicus.jp>
Signed-off-by: NRyusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>

6b81e14e

nilfs2: eliminate sparse warnings - "symbol not declared" · abc0b50b

由 Jiro SEKIBA 提交于 10月 08, 2010

change nilfs_dat_commit_free and nilfs_inode_cachep static
to fix following warnings

fs/nilfs2/super.c:72:19: warning: symbol 'nilfs_inode_cachep' was not declared. Should it be static?
fs/nilfs2/dat.c:106:6: warning: symbol 'nilfs_dat_commit_free' was not declared. Should it be static?
Signed-off-by: NJiro SEKIBA <jir@unicus.jp>
Signed-off-by: NRyusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>

abc0b50b

nilfs2: get rid of bdi from nilfs object · 026a7d63

由 Ryusuke Konishi 提交于 10月 07, 2010

Nilfs now can use sb->s_bdi to get backing_dev_info, so we use it
instead of ns_bdi on the nilfs object and remove ns_bdi.
Signed-off-by: NRyusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>

026a7d63

nilfs2: add bdev freeze/thaw support · 5beb6e0b

由 Ryusuke Konishi 提交于 9月 20, 2010

Nilfs hasn't supported the freeze/thaw feature because it didn't work
due to the peculiar design that multiple super block instances could
be allocated for a device. This limitation was removed by the patch
"nilfs2: do not allocate multiple super block instances for a device".

So now this adds the freeze/thaw support to nilfs.
Signed-off-by: NRyusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>

5beb6e0b

nilfs2: accept 64-bit checkpoint numbers in cp mount option · c05dbfc2

由 Ryusuke Konishi 提交于 9月 16, 2010

The current implementation doesn't mount snapshots with checkpoint
numbers larger than INT_MAX since it uses match_int() for parsing
"cp=" mount option.

This uses simple_strtoull() for the conversion to resolve the issue.
Signed-off-by: NRyusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>

c05dbfc2

nilfs2: remove own inode allocator and destructor for metadata files · 2879ed66

由 Ryusuke Konishi 提交于 9月 05, 2010

This finally removes own inode allocator and destructor functions for
metadata files.  Several routines, nilfs_mdt_new(),
nilfs_mdt_new_common(), nilfs_mdt_clear(), nilfs_mdt_destroy(), and
nilfs_alloc_inode_common() will be gone.
Signed-off-by: NRyusuke Konishi <konishi.ryusuke@lab.ntt.co.jp>

2879ed66

openeuler / Kernel 大约 1 年 前同步成功

openeuler / Kernel
大约 1 年前同步成功