提交 · 15570086b590a69d59183b08a7770e316cca20a7 · openanolis / cloud-kernel

03 9月, 2013 1 次提交

vfs: reimplement d_rcu_to_refcount() using lockref_get_or_lock() · 15570086

由 Linus Torvalds 提交于 9月 02, 2013

This moves __d_rcu_to_refcount() from <linux/dcache.h> into fs/namei.c
and re-implements it using the lockref infrastructure instead.  It also
adds a lot of comments about what is actually going on, because turning
a dentry that was looked up using RCU into a long-lived reference
counted entry is one of the more subtle parts of the rcu walk.

We also used to be _particularly_ subtle in unlazy_walk() where we
re-validate both the dentry and its parent using the same sequence
count.  We used to do it by nesting the locks and then verifying the
sequence count just once.

That was silly, because nested locking is expensive, but the sequence
count check is not.  So this just re-validates the dentry and the parent
separately, avoiding the nested locking, and making the lockref lookup
possible.
Acked-by: NWaiman Long <waiman.long@hp.com>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

15570086

29 8月, 2013 2 次提交

vfs: make the dentry cache use the lockref infrastructure · 98474236

由 Waiman Long 提交于 8月 28, 2013

This just replaces the dentry count/lock combination with the lockref
structure that contains both a count and a spinlock, and does the
mechanical conversion to use the lockref infrastructure.

There are no semantic changes here, it's purely syntactic.  The
reference lockref implementation uses the spinlock exactly the same way
that the old dcache code did, and the bulk of this patch is just
expanding the internal "d_count" use in the dcache code to use
"d_lockref.count" instead.

This is purely preparation for the real change to make the reference
count updates be lockless during the 3.12 merge window.

[ As with the previous commit, this is a rewritten version of a concept
  originally from Waiman, so credit goes to him, blame for any errors
  goes to me.

  Waiman's patch had some semantic differences for taking advantage of
  the lockless update in dget_parent(), while this patch is
  intentionally a pure search-and-replace change with no semantic
  changes.     - Linus ]
Signed-off-by: NWaiman Long <Waiman.Long@hp.com>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

98474236

Revert "fs: Allow unprivileged linkat(..., AT_EMPTY_PATH) aka flink" · f0cc6ffb

由 Linus Torvalds 提交于 8月 28, 2013

This reverts commit bb2314b4.

It wasn't necessarily wrong per se, but we're still busily discussing
the exact details of this all, so I'm going to revert it for now.

It's true that you can already do flink() through /proc and that flink()
isn't new.  But as Brad Spengler points out, some secure environments do
not mount proc, and flink adds a new interface that can avoid path
lookup of the source for those kinds of environments.

We may re-do this (and even mark it for stable backporting back in 3.11
and possibly earlier) once the whole discussion about the interface is done.

Cc: Andy Lutomirski <luto@amacapital.net>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Oleg Nesterov <oleg@redhat.com>
Cc: Brad Spengler <spender@grsecurity.net>
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

f0cc6ffb

05 8月, 2013 1 次提交

fs: Allow unprivileged linkat(..., AT_EMPTY_PATH) aka flink · bb2314b4

由 Andy Lutomirski 提交于 8月 01, 2013

Every now and then someone proposes a new flink syscall, and this spawns
a long discussion of whether it would be a security problem.  I think
that this is missing the point: flink is *already* allowed without
privilege as long as /proc is mounted -- it's called AT_SYMLINK_FOLLOW.

Now that O_TMPFILE is here, the ability to create a file with O_TMPFILE,
write it, and link it in is very convenient.  The only problem is that
it requires that /proc be mounted so that you can do:

linkat(AT_FDCWD, "/proc/self/fd/<tmpfd>", dfd, path, AT_SYMLINK_NOFOLLOW)

This sucks -- it's much nicer to do:

linkat(tmpfd, "", dfd, path, AT_EMPTY_PATH)

Let's allow it.

If this turns out to be excessively scary, it we could instead require
that the inode in question be I_LINKABLE, but this seems pointless given
the /proc situation
Signed-off-by: NAndy Lutomirski <luto@amacapital.net>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

bb2314b4

13 7月, 2013 1 次提交

Safer ABI for O_TMPFILE · bb458c64

由 Al Viro 提交于 7月 13, 2013

[suggested by Rasmus Villemoes] make O_DIRECTORY | O_RDWR part of O_TMPFILE;
that will fail on old kernels in a lot more cases than what I came up with.
And make sure O_CREAT doesn't get there...
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

bb458c64

29 6月, 2013 5 次提交

Don't pass inode to ->d_hash() and ->d_compare() · da53be12

由 Linus Torvalds 提交于 5月 21, 2013

Instances either don't look at it at all (the majority of cases) or
only want it to find the superblock (which can be had as dentry->d_sb).
A few cases that want more are actually safe with dentry->d_inode -
the only precaution needed is the check that it hadn't been replaced with
NULL by rmdir() or by overwriting rename(), which case should be simply
treated as cache miss.
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

da53be12

allow the temp files created by open() to be linked to · f4e0c30c

由 Al Viro 提交于 6月 11, 2013

O_TMPFILE | O_CREAT => linkat() with AT_SYMLINK_FOLLOW and /proc/self/fd/<n>
as oldpath (i.e. flink()) will create a link
O_TMPFILE | O_CREAT | O_EXCL => ENOENT on attempt to link those guys
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

f4e0c30c

A
[O_TMPFILE] it's still short a few helpers, but infrastructure should be OK now... · 60545d0d
由 Al Viro 提交于 6月 07, 2013
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
60545d0d
A
allow build_open_flags() to return an error · f9652e10
由 Al Viro 提交于 6月 11, 2013
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
f9652e10

do_last(): fix missing checks for LAST_BIND case · bc77daa7

由 Al Viro 提交于 6月 06, 2013

/proc/self/cwd with O_CREAT should fail with EISDIR.  /proc/self/exe, OTOH,
should fail with ENOTDIR when opened with O_DIRECTORY.
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

bc77daa7

15 6月, 2013 1 次提交

use can_lookup() instead of direct checks of ->i_op->lookup · 05252901

由 Al Viro 提交于 6月 06, 2013

a couple of places got missed back when Linus has introduced that one...
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

05252901

08 5月, 2013 1 次提交

audit: vfs: fix audit_inode call in O_CREAT case of do_last · 33e2208a

由 Jeff Layton 提交于 4月 12, 2013

Jiri reported a regression in auditing of open(..., O_CREAT) syscalls.
In older kernels, creating a file with open(..., O_CREAT) created
audit_name records that looked like this:

type=PATH msg=audit(1360255720.628:64): item=1 name="/abc/foo" inode=138810 dev=fd:00 mode=0100640 ouid=0 ogid=0 rdev=00:00 obj=unconfined_u:object_r:default_t:s0
type=PATH msg=audit(1360255720.628:64): item=0 name="/abc/" inode=138635 dev=fd:00 mode=040750 ouid=0 ogid=0 rdev=00:00 obj=unconfined_u:object_r:default_t:s0

...in recent kernels though, they look like this:

type=PATH msg=audit(1360255402.886:12574): item=2 name=(null) inode=264599 dev=fd:00 mode=0100640 ouid=0 ogid=0 rdev=00:00 obj=unconfined_u:object_r:default_t:s0
type=PATH msg=audit(1360255402.886:12574): item=1 name=(null) inode=264598 dev=fd:00 mode=040750 ouid=0 ogid=0 rdev=00:00 obj=unconfined_u:object_r:default_t:s0
type=PATH msg=audit(1360255402.886:12574): item=0 name="/abc/foo" inode=264598 dev=fd:00 mode=040750 ouid=0 ogid=0 rdev=00:00 obj=unconfined_u:object_r:default_t:s0

Richard bisected to determine that the problems started with commit
bfcec708, but the log messages have changed with some later
audit-related patches.

The problem is that this audit_inode call is passing in the parent of
the dentry being opened, but audit_inode is being called with the parent
flag false. This causes later audit_inode and audit_inode_child calls to
match the wrong entry in the audit_names list.

This patch simply sets the flag to properly indicate that this inode
represents the parent. With this, the audit_names entries are back to
looking like they did before.

Cc: <stable@vger.kernel.org> # v3.7+
Reported-by: NJiri Jaburek <jjaburek@redhat.com>
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Test By: Richard Guy Briggs <rbriggs@redhat.com>
Signed-off-by: NEric Paris <eparis@redhat.com>

33e2208a

09 3月, 2013 1 次提交

vfs: don't BUG_ON() if following a /proc fd pseudo-symlink results in a symlink · 7b54c165

由 Linus Torvalds 提交于 3月 08, 2013

It's "normal" - it can happen if the file descriptor you followed was
opened with O_NOFOLLOW.
Reported-by: NDave Jones <davej@redhat.com>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: stable@kernel.org
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

7b54c165

02 3月, 2013 1 次提交
- A
  constify path_get/path_put and fs_struct.c stuff · dcf787f3
  由 Al Viro 提交于 3月 01, 2013
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  dcf787f3
26 2月, 2013 1 次提交

vfs: kill FS_REVAL_DOT by adding a d_weak_revalidate dentry op · ecf3d1f1

由 Jeff Layton 提交于 2月 20, 2013

The following set of operations on a NFS client and server will cause

    server# mkdir a
    client# cd a
    server# mv a a.bak
    client# sleep 30  # (or whatever the dir attrcache timeout is)
    client# stat .
    stat: cannot stat `.': Stale NFS file handle

Obviously, we should not be getting an ESTALE error back there since the
inode still exists on the server. The problem is that the lookup code
will call d_revalidate on the dentry that "." refers to, because NFS has
FS_REVAL_DOT set.

nfs_lookup_revalidate will see that the parent directory has changed and
will try to reverify the dentry by redoing a LOOKUP. That of course
fails, so the lookup code returns ESTALE.

The problem here is that d_revalidate is really a bad fit for this case.
What we really want to know at this point is whether the inode is still
good or not, but we don't really care what name it goes by or whether
the dcache is still valid.

Add a new d_op->d_weak_revalidate operation and have complete_walk call
that instead of d_revalidate. The intent there is to allow for a
"weaker" d_revalidate that just checks to see whether the inode is still
good. This is also gives us an opportunity to kill off the FS_REVAL_DOT
special casing.

[AV: changed method name, added note in porting, fixed confusion re
having it possibly called from RCU mode (it won't be)]

Cc: NeilBrown <neilb@suse.de>
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

ecf3d1f1

23 2月, 2013 6 次提交
- A
  lookup_slow: get rid of name argument · cc2a5271
  由 Al Viro 提交于 1月 24, 2013
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  cc2a5271
- A
  lookup_fast: get rid of name argument · e97cdc87
  由 Al Viro 提交于 1月 24, 2013
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  e97cdc87
- A
  get rid of name and type arguments of walk_component() · 21b9b073
  由 Al Viro 提交于 1月 24, 2013
```
... always can be found in nameidata now.
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  21b9b073
- A
  link_path_walk(): move assignments to nd->last/nd->last_type up · 5f4a6a69
  由 Al Viro 提交于 1月 24, 2013
```
... and clean the main loop a bit
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  5f4a6a69
- A
  propagate error from get_empty_filp() to its callers · 1afc99be
  由 Al Viro 提交于 2月 14, 2013
```
Based on parts from Anatol's patch (the rest is the next commit).
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  1afc99be
- A
  new helper: file_inode(file) · 496ad9aa
  由 Al Viro 提交于 1月 23, 2013
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  496ad9aa
21 12月, 2012 12 次提交

vfs: fix renameat to retry on ESTALE errors · c6a94284

由 Jeff Layton 提交于 12月 11, 2012

...as always, rename is the messiest of the bunch. We have to track
whether to retry or not via a separate flag since the error handling
is already quite complex.
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

c6a94284

J
vfs: make do_unlinkat retry once on ESTALE errors · 5d18f813
由 Jeff Layton 提交于 12月 20, 2012
```
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
5d18f813

vfs: make do_rmdir retry once on ESTALE errors · c6ee9206

由 Jeff Layton 提交于 12月 20, 2012

Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

c6ee9206

vfs: add a flags argument to user_path_parent · 9e790bd6

由 Jeff Layton 提交于 12月 11, 2012

...so we can pass in LOOKUP_REVAL. For now, nothing does yet.
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

9e790bd6

vfs: fix linkat to retry once on ESTALE errors · 442e31ca

由 Jeff Layton 提交于 12月 20, 2012

Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

442e31ca

vfs: fix symlinkat to retry on ESTALE errors · f46d3567

由 Jeff Layton 提交于 12月 11, 2012

Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

f46d3567

J
vfs: fix mkdirat to retry once on an ESTALE error · b76d8b82
由 Jeff Layton 提交于 12月 20, 2012
```
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
b76d8b82

vfs: fix mknodat to retry on ESTALE errors · 972567f1

由 Jeff Layton 提交于 12月 20, 2012

Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

972567f1

vfs: turn is_dir argument to kern_path_create into a lookup_flags arg · 1ac12b4b

由 Jeff Layton 提交于 12月 11, 2012

Where we can pass in LOOKUP_DIRECTORY or LOOKUP_REVAL. Any other flags
passed in here are currently ignored.
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

1ac12b4b

vfs: remove DCACHE_NEED_LOOKUP · 39e3c955

由 Jeff Layton 提交于 11月 28, 2012

The code that relied on that flag was ripped out of btrfs quite some
time ago, and never added back. Josef indicated that he was going to
take a different approach to the problem in btrfs, and that we
could just eliminate this flag.

Cc: Josef Bacik <jbacik@fusionio.com>
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

39e3c955

A
path_init(): make -ENOTDIR failure exits consistent · 741b7c3f
由 Al Viro 提交于 12月 20, 2012
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
741b7c3f

vfs: remove unneeded permission check from path_init · 582aa64a

由 Jeff Layton 提交于 12月 11, 2012

When path_init is called with a valid dfd, that code checks permissions
on the open directory fd and returns an error if the check fails. This
permission check is redundant, however.

Both callers of path_init immediately call link_path_walk afterward. The
first thing that link_path_walk does for pathnames that do not consist
only of slashes is to check for exec permissions at the starting point of
the path walk. And this check in path_init() is on the path taken only
when *name != '/' && *name != '\0'.

In most cases, these checks are very quick, but when the dfd is for a
file on a NFS mount with the actimeo=0, each permission check goes
out onto the wire. The result is 2 identical ACCESS calls.

Given that these codepaths are fairly "hot", I think it makes sense to
eliminate the permission check in path_init and simply assume that the
caller will eventually check the permissions before proceeding.
Reported-by: NDave Wysochanski <dwysocha@redhat.com>
Signed-off-by: NJeff Layton <jlayton@redhat.com>
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>

582aa64a

30 11月, 2012 1 次提交
- A
  lookup_one_len: don't accept . and .. · 21d8a15a
  由 Al Viro 提交于 11月 29, 2012
```
Signed-off-by: NAl Viro <viro@zeniv.linux.org.uk>
```
  21d8a15a
27 10月, 2012 1 次提交

VFS: don't do protected {sym,hard}links by default · 561ec64a

由 Linus Torvalds 提交于 10月 26, 2012

In commit 800179c9 ("This adds symlink and hardlink restrictions to
the Linux VFS"), the new link protections were enabled by default, in
the hope that no actual application would care, despite it being
technically against legacy UNIX (and documented POSIX) behavior.

However, it does turn out to break some applications.  It's rare, and
it's unfortunate, but it's unacceptable to break existing systems, so
we'll have to default to legacy behavior.

In particular, it has broken the way AFD distributes files, see

  http://www.dwd.de/AFD/

along with some legacy scripts.

Distributions can end up setting this at initrd time or in system
scripts: if you have security problems due to link attacks during your
early boot sequence, you have bigger problems than some kernel sysctl
setting. Do:

	echo 1 > /proc/sys/fs/protected_symlinks
	echo 1 > /proc/sys/fs/protected_hardlinks

to re-enable the link protections.

Alternatively, we may at some point introduce a kernel config option
that sets these kinds of "more secure but not traditional" behavioural
options automatically.
Reported-by: NNick Bowler <nbowler@elliptictech.com>
Reported-by: NHolger Kiehl <Holger.Kiehl@dwd.de>
Cc: Kees Cook <keescook@chromium.org>
Cc: Ingo Molnar <mingo@elte.hu>
Cc: Andrew Morton <akpm@linux-foundation.org>
Cc: Al Viro <viro@zeniv.linux.org.uk>
Cc: Alan Cox <alan@lxorguk.ukuu.org.uk>
Cc: Theodore Ts'o <tytso@mit.edu>
Cc: stable@kernel.org # v3.6
Signed-off-by: NLinus Torvalds <torvalds@linux-foundation.org>

561ec64a

13 10月, 2012 5 次提交

vfs: embed struct filename inside of names_cache allocation if possible · 7950e385