提交 · 54ceac4515986030c2502960be620198dd8fe25b · openeuler / raspberrypi-kernel

23 9月, 2006 20 次提交

NFS: Share NFS superblocks per-protocol per-server per-FSID · 54ceac45

由 David Howells 提交于 8月 22, 2006

The attached patch makes NFS share superblocks between mounts from the same
server and FSID over the same protocol.

It does this by creating each superblock with a false root and returning the
real root dentry in the vfsmount presented by get_sb(). The root dentry set
starts off as an anonymous dentry if we don't already have the dentry for its
inode, otherwise it simply returns the dentry we already have.

We may thus end up with several trees of dentries in the superblock, and if at
some later point one of anonymous tree roots is discovered by normal filesystem
activity to be located in another tree within the superblock, the anonymous
root is named and materialises attached to the second tree at the appropriate
point.

Why do it this way? Why not pass an extra argument to the mount() syscall to
indicate the subpath and then pathwalk from the server root to the desired
directory? You can't guarantee this will work for two reasons:

 (1) The root and intervening nodes may not be accessible to the client.

     With NFS2 and NFS3, for instance, mountd is called on the server to get
     the filehandle for the tip of a path. mountd won't give us handles for
     anything we don't have permission to access, and so we can't set up NFS
     inodes for such nodes, and so can't easily set up dentries (we'd have to
     have ghost inodes or something).

     With this patch we don't actually create dentries until we get handles
     from the server that we can use to set up their inodes, and we don't
     actually bind them into the tree until we know for sure where they go.

 (2) Inaccessible symbolic links.

     If we're asked to mount two exports from the server, eg:

	mount warthog:/warthog/aaa/xxx /mmm
	mount warthog:/warthog/bbb/yyy /nnn

     We may not be able to access anything nearer the root than xxx and yyy,
     but we may find out later that /mmm/www/yyy, say, is actually the same
     directory as the one mounted on /nnn. What we might then find out, for
     example, is that /warthog/bbb was actually a symbolic link to
     /warthog/aaa/xxx/www, but we can't actually determine that by talking to
     the server until /warthog is made available by NFS.

     This would lead to having constructed an errneous dentry tree which we
     can't easily fix. We can end up with a dentry marked as a directory when
     it should actually be a symlink, or we could end up with an apparently
     hardlinked directory.

     With this patch we need not make assumptions about the type of a dentry
     for which we can't retrieve information, nor need we assume we know its
     place in the grand scheme of things until we actually see that place.

This patch reduces the possibility of aliasing in the inode and page caches for
inodes that may be accessed by more than one NFS export. It also reduces the
number of superblocks required for NFS where there are many NFS exports being
used from a server (home directory server + autofs for example).

This in turn makes it simpler to do local caching of network filesystems, as it
can then be guaranteed that there won't be links from multiple inodes in
separate superblocks to the same cache file.

Obviously, cache aliasing between different levels of NFS protocol could still
be a problem, but at least that gives us another key to use when indexing the
cache.

This patch makes the following changes:

 (1) The server record construction/destruction has been abstracted out into
     its own set of functions to make things easier to get right.  These have
     been moved into fs/nfs/client.c.

     All the code in fs/nfs/client.c has to do with the management of
     connections to servers, and doesn't touch superblocks in any way; the
     remaining code in fs/nfs/super.c has to do with VFS superblock management.

 (2) The sequence of events undertaken by NFS mount is now reordered:

     (a) A volume representation (struct nfs_server) is allocated.

     (b) A server representation (struct nfs_client) is acquired.  This may be
     	 allocated or shared, and is keyed on server address, port and NFS
     	 version.

     (c) If allocated, the client representation is initialised.  The state
     	 member variable of nfs_client is used to prevent a race during
     	 initialisation from two mounts.

     (d) For NFS4 a simple pathwalk is performed, walking from FH to FH to find
     	 the root filehandle for the mount (fs/nfs/getroot.c).  For NFS2/3 we
     	 are given the root FH in advance.

     (e) The volume FSID is probed for on the root FH.

     (f) The volume representation is initialised from the FSINFO record
     	 retrieved on the root FH.

     (g) sget() is called to acquire a superblock.  This may be allocated or
     	 shared, keyed on client pointer and FSID.

     (h) If allocated, the superblock is initialised.

     (i) If the superblock is shared, then the new nfs_server record is
     	 discarded.

     (j) The root dentry for this mount is looked up from the root FH.

     (k) The root dentry for this mount is assigned to the vfsmount.

 (3) nfs_readdir_lookup() creates dentries for each of the entries readdir()
     returns; this function now attaches disconnected trees from alternate
     roots that happen to be discovered attached to a directory being read (in
     the same way nfs_lookup() is made to do for lookup ops).

     The new d_materialise_unique() function is now used to do this, thus
     permitting the whole thing to be done under one set of locks, and thus
     avoiding any race between mount and lookup operations on the same
     directory.

 (4) The client management code uses a new debug facility: NFSDBG_CLIENT which
     is set by echoing 1024 to /proc/net/sunrpc/nfs_debug.

 (5) Clone mounts are now called xdev mounts.

 (6) Use the dentry passed to the statfs() op as the handle for retrieving fs
     statistics rather than the root dentry of the superblock (which is now a
     dummy).
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

54ceac45

NFS: Start rpciod in server common management · cf6d7b5d

由 David Howells 提交于 8月 22, 2006

Start rpciod in the server common (nfs_client struct) management code rather
than in the superblock management code. This means we only need to "start" it
once per server instead of once per superblock.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

cf6d7b5d

NFS: Eliminate client_sys in favour of cl_rpcclient · 5006a76c

由 David Howells 提交于 8月 22, 2006

Eliminate nfs_server::client_sys in favour of nfs_client::cl_rpcclient as we
only really need one per server that we're talking to since it doesn't have any
security on it.

The retransmission management variables are also moved to the common struct as
they're required to set up the cl_rpcclient connection.

The NFS2/3 client and client_acl connections are thenceforth derived by cloning
the cl_rpcclient connection and post-applying the authorisation flavour.

The code for setting up the initial common connection has been moved to
client.c as nfs_create_rpc_client(). All the NFS program definition tables are
also moved there as that's where they're now required rather than super.c.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

5006a76c

NFS: Move rpc_ops from nfs_server to nfs_client · 8fa5c000

由 David Howells 提交于 8月 22, 2006

Move the rpc_ops from the nfs_server struct to the nfs_client struct as they're
common to all server records of a particular NFS protocol version.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

8fa5c000

NFS: Make better use of inode* dereferencing macros · 1f163415

由 David Howells 提交于 8月 22, 2006

Make better use of inode* dereferencing macros to hide dereferencing chains
(including NFS_PROTO and NFS_CLIENT).
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

1f163415

NFS: Maintain a common server record for NFS2/3 as well as for NFS4 · 27951bd2

由 David Howells 提交于 8月 22, 2006

Maintain a common server record for NFS2/3 as well as for NFS4 so that common
stuff can be moved there from struct nfs_server.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

27951bd2

NFS: Add extra const qualifiers · 509de811

由 David Howells 提交于 8月 22, 2006

Add some extra const qualifiers into NFS.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

509de811

NFS: Use the dentry superblock directly in nfs_statfs() · 0c7d90cf

由 David Howells 提交于 8月 22, 2006

Use the nominated dentry's superblock directly in the NFS statfs() op to get a
file handle, rather than using s_root (which will become a dummy dentry in a
future patch).
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

0c7d90cf

NFS: Generalise the nfs_client structure · 24c8dbbb

由 David Howells 提交于 8月 22, 2006

Generalise the nfs_client structure by:

 (1) Moving nfs_client to a more general place (nfs_fs_sb.h).

 (2) Renaming its maintenance routines to be non-NFS4 specific.

 (3) Move those maintenance routines to a new non-NFS4 specific file (client.c)
     and move the declarations to internal.h.

 (4) Make nfs_find/get_client() take a full sockaddr_in to include the port
     number (will be required for NFS2/3).

 (5) Make nfs_find/get_client() take the NFS protocol version (again will be
     required to differentiate NFS2, 3 & 4 client records).

Also:

 (6) Make nfs_client construction proceed akin to inodes, marking them as under
     construction and providing a function to indicate completion.

 (7) Make nfs_get_client() wait interruptibly if it finds a client that it can
     share, but that client is currently being constructed.

 (8) Make nfs4_create_client() use (6) and (7) instead of locking cl_sem.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

24c8dbbb

NFS: Add a server capabilities NFS RPC op · e9326dca

由 David Howells 提交于 8月 22, 2006

Add a set_capabilities NFS RPC op so that the server capabilities can be set.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

e9326dca

NFS: Add a lookupfh NFS RPC op · 2b3de441

由 David Howells 提交于 8月 22, 2006

Add a lookup filehandle NFS RPC op so that a file handle can be looked up
without requiring dentries and inodes and other VFS stuff when doing an NFS4
pathwalk during mounting.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

2b3de441

NFS: Return an error when starting the idmapping pipe · b7162792

由 David Howells 提交于 8月 22, 2006

Return an error when starting the idmapping pipe so that we can detect it
failing.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

b7162792

NFS: Rename nfs_server::nfs4_state · 7539bbab

由 David Howells 提交于 8月 22, 2006

Rename nfs_server::nfs4_state to nfs_client as it will be used to represent the
client state for NFS2 and NFS3 also.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

7539bbab

NFS: Rename struct nfs4_client to struct nfs_client · adfa6f98

由 David Howells 提交于 8月 22, 2006

Rename struct nfs4_client to struct nfs_client so that it can become the basis
for a general client record for NFS2 and NFS3 in addition to NFS4.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

adfa6f98

NFS: Fix NFS4 callback up/down prototypes · 5ae1fbce

由 David Howells 提交于 8月 22, 2006

Make the nfs_callback_up()/down() prototypes just do nothing if NFS4 is not
enabled.  Also make the down function void type since we can't really do
anything if it fails.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

5ae1fbce

NFS: Disambiguate nfs_stat_to_errno() · 0a8ea437

由 David Howells 提交于 8月 22, 2006

Rename the NFS4 version of nfs_stat_to_errno() so that it doesn't conflict with
the common one used by NFS2 and NFS3.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

0a8ea437

NFS: Fix up split of fs/nfs/inode.c · 7d4e2747

由 David Howells 提交于 8月 22, 2006

Fix ups for the splitting of the superblock stuff out of fs/nfs/inode.c,
including:

 (*) Move the callback tcpport module param into callback.c.

 (*) Move the idmap cache timeout module param into idmap.c.

 (*) Changes to internal.h:

     (*) namespace-nfs4.c was renamed to nfs4namespace.c.

     (*) nfs_stat_to_errno() is in nfs2xdr.c, not nfs4xdr.c.

     (*) nfs4xdr.c is contingent on CONFIG_NFS_V4.

     (*) nfs4_path() is only uses if CONFIG_NFS_V4 is set.

Plus also:

 (*) The sec_flavours[] table should really be const.
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

7d4e2747

NFS: Add an ACCESS cache memory shrinker · 979df72e

由 Trond Myklebust 提交于 7月 25, 2006

A pinned inode may in theory end up filling memory with cached ACCESS
calls. This patch ensures that the VM may shrink away the cache in these
particular cases.
The shrinker works by iterating through the list of inodes on the global
nfs_access_lru_list, and removing the least recently used access
cache entry until it is done (or until the entire cache is empty).
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

979df72e

NFS: Add a global LRU list for the ACCESS cache · cfcea3e8

由 Trond Myklebust 提交于 7月 25, 2006

...in order to allow the addition of a memory shrinker.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

cfcea3e8

NFS: Add a new ACCESS rpc call cache to the linux nfs client · 1c3c07e9

由 Trond Myklebust 提交于 7月 25, 2006

The current access cache only allows one entry at a time to be cached for each
inode. Add a per-inode red-black tree in order to allow more than one to
be cached at a time.

Should significantly cut down the time spent in path traversal for shared
directories such as ${PATH}, /usr/share, etc.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

1c3c07e9

19 9月, 2006 3 次提交
- T
  NFS: Fix nfs_page use after free issues in fs/nfs/write.c · 5c2d97cb
  由 Trond Myklebust 提交于 9月 18, 2006
```
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
```
  5c2d97cb
- T
  NFSv4: Fix incorrect semaphore release in _nfs4_do_open() · 76723de0
  由 Trond Myklebust 提交于 9月 15, 2006
```
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
```
  76723de0
- T
  NFS: Fix Oopsable condition in nfs_readpage_sync() · 7a524111
  由 Trond Myklebust 提交于 9月 15, 2006
```
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
```
  7a524111
09 9月, 2006 1 次提交

[PATCH] NFS: large non-page-aligned direct I/O clobbers memory · e9f7bee1

由 Trond Myklebust 提交于 9月 08, 2006

The logic in nfs_direct_read_schedule and nfs_direct_write_schedule can
allow data->npages to be one larger than rpages.  This causes a page
pointer to be written beyond the end of the pagevec in nfs_read_data (or
nfs_write_data).

Fix this by making nfs_(read|write)_alloc() calculate the size of the
pagevec array, and initialise data->npages.

Also get rid of the redundant argument to nfs_commit_alloc().
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
Cc: Chuck Lever <chuck.lever@oracle.com>
Signed-off-by: NAndrew Morton <akpm@osdl.org>
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>

e9f7bee1

25 8月, 2006 6 次提交

NFSv4: Add v4 exception handling for the ACL functions. · 16b4289c

由 Trond Myklebust 提交于 8月 24, 2006

This is needed in order to handle any NFS4ERR_DELAY errors that might be
returned by the server. It also ensures that we map the NFSv4 errors before
they are returned to userland.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from 71c12b3f0abc7501f6ed231a6d17bc9c05a238dc commit)

16b4289c

NFS: Check lengths more thoroughly in NFS4 readdir XDR decode · e8896495

由 David Howells 提交于 8月 24, 2006

Check the bounds of length specifiers more thoroughly in the XDR decoding of
NFS4 readdir reply data.

Currently, if the server returns a bitmap or attr length that causes the
current decode point pointer to wrap, this could go undetected (consider a
small "negative" length on a 32-bit machine).

Also add a check into the main XDR decode handler to make sure that the amount
of data is a multiple of four bytes (as specified by RFC-1014).  This makes
sure that we can do u32* pointer subtraction in the NFS client without risking
an undefined result (the result is undefined if the pointers are not correctly
aligned with respect to one another).
Signed-Off-By: NDavid Howells <dhowells@redhat.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from 5861fddd64a7eaf7e8b1a9997455a24e7f688092 commit)

e8896495

NFS: Fix issue with EIO on NFS read · 79558f36

由 Trond Myklebust 提交于 8月 22, 2006

The problem is that we may be caching writes that would extend the file and
create a hole in the region that we are reading. In this case, we need to
detect the eof from the server, ensure that we zero out the pages that
are part of the hole and mark them as up to date.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from 856b603b01b99146918c093969b6cb1b1b0f1c01 commit)

79558f36

SUNRPC: Fix dentry refcounting issues with users of rpc_pipefs · 8f8e7a50

由 Trond Myklebust 提交于 8月 14, 2006

rpc_unlink() and rpc_rmdir() will dput the dentry reference for you.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from a05a57effa71a1f67ccbfc52335c10c8b85f3f6a commit)

8f8e7a50

SUNRPC: make rpc_unlink() take a dentry argument instead of a path · 5d67476f

由 Trond Myklebust 提交于 7月 31, 2006

Signe-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from 88bf6d811b01a4be7fd507d18bf5f1c527989089 commit)

5d67476f

NFS: Fix a potential deadlock in nfs_release_page · ddeff520

由 Nikita Danilov 提交于 8月 09, 2006

nfs_wb_page() waits on request completion and, as a result, is not safe to be
called from nfs_release_page() invoked by VM scanner as part of GFP_NOFS
allocation. Fix possible deadlock by analyzing gfp mask and refusing to
release page if __GFP_FS is not set.
Signed-off-by: NNikita Danilov <danilov@gmail.com>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from 374d969debfb290bafcb41d28918dc6f7e43ce31 commit)

ddeff520

04 8月, 2006 2 次提交

NFS: make 2 functions static · e4e20512

由 Adrian Bunk 提交于 8月 03, 2006

nfs_writedata_free() and nfs_readdata_free() can now become static.
Signed-off-by: NAdrian Bunk <bunk@stusta.de>
Cc: Trond Myklebust <trond.myklebust@fys.uio.no>
Signed-off-by: NAndrew Morton <akpm@osdl.org>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from 5e1ce40f0c3c8f67591aff17756930d7a18ceb1a commit)

e4e20512

NFS: Release dcache_lock in an error path of nfs_path · ce510193

由 Josh Triplett 提交于 7月 24, 2006

In one of the error paths of nfs_path, it may return with dcache_lock still
held; fix this by adding and using a new error path Elong_unlock which unlocks
dcache_lock.
Signed-off-by: NJosh Triplett <josh@freedesktop.org>
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
(cherry picked from f4b90b43677fb23297c56802c3056fc304f988d9 commit)

ce510193

06 7月, 2006 5 次提交

NFS: Optimise away an excessive GETATTR call when a file is symlinked · 4e0641a7

由 Trond Myklebust 提交于 7月 05, 2006

In the case when compiling via a symlink tree, we want to ensure that the
close-to-open GETATTR call is applied only to the final file, and not to
the symlink.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

4e0641a7

NFS: Fix NFS page_state usage · 83715ad5

由 Trond Myklebust 提交于 7月 05, 2006

The introduction of the FLUSH_INVALIDATE argument to nfs_sync_inode_wait()
does not clear the nr_unstable page state counter for pages that are being
released.

Also fix a longstanding similar bug when nfs_commit_list() fails.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

83715ad5

NLM,NFSv4: Wait on local locks before we put RPC calls on the wire · 01c3b861

由 Trond Myklebust 提交于 6月 29, 2006

Use FL_ACCESS flag to test and/or wait for local locks before we try
requesting a lock from the server
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

01c3b861

T
NFSv4: Ensure nfs4_lock_expired() caches delegated locks · 42a2d13e
由 Trond Myklebust 提交于 6月 29, 2006
```
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>
```
42a2d13e

NLM,NFSv4: Don't put UNLOCK requests on the wire unless we hold a lock · 9b073574

由 Trond Myklebust 提交于 6月 29, 2006

Use the new behaviour of {flock,posix}_file_lock(F_UNLCK) to determine if
we held a lock, and only send the RPC request to the server if this was the
case.
Signed-off-by: NTrond Myklebust <Trond.Myklebust@netapp.com>

9b073574

03 7月, 2006 1 次提交

[PATCH] nfs: non-procfs build fix · 4ebd9ab3

由 Dominik Hackl 提交于 7月 02, 2006

This fixes a bug in fs/nfs which makes it impossible to build nfs
without having procfs enabled.
Signed-off-by: NDominik Hackl <dominik@hackl.dhs.org>
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>

4ebd9ab3

01 7月, 2006 2 次提交

[PATCH] zoned vm counters: conversion of nr_unstable to per zone counter · fd39fc85

由 Christoph Lameter 提交于 6月 30, 2006

Conversion of nr_unstable to a per zone counter

We need to do some special modifications to the nfs code since there are
multiple cases of disposition and we need to have a page ref for proper
accounting.

This converts the last critical page state of the VM and therefore we need to
remove several functions that were depending on GET_PAGE_STATE_LAST in order
to make the kernel compile again.  We are only left with event type counters
in page state.

[akpm@osdl.org: bugfixes]
Signed-off-by: NChristoph Lameter <clameter@sgi.com>
Cc: Trond Myklebust <trond.myklebust@fys.uio.no>
Signed-off-by: NAndrew Morton <akpm@osdl.org>
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>

fd39fc85

[PATCH] zoned vm counters: conversion of nr_dirty to per zone counter · b1e7a8fd

由 Christoph Lameter 提交于 6月 30, 2006

This makes nr_dirty a per zone counter.  Looping over all processors is
avoided during writeback state determination.

The counter aggregation for nr_dirty had to be undone in the NFS layer since
we summed up the page counts from multiple zones.  Someone more familiar with
NFS should probably review what I have done.

[akpm@osdl.org: bugfix]
Signed-off-by: NChristoph Lameter <clameter@sgi.com>
Cc: Trond Myklebust <trond.myklebust@fys.uio.no>
Signed-off-by: NAndrew Morton <akpm@osdl.org>
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>

b1e7a8fd