提交 · 31459fe4b24c1e09712eff0d82a5276f4fd0e3cf · openeuler / Kernel

18 5月, 2010 1 次提交

ceph: use __page_cache_alloc and add_to_page_cache_lru · 31459fe4

由 Yehuda Sadeh 提交于 3月 17, 2010

Following Nick Piggin patches in btrfs, pagecache pages should be
allocated with __page_cache_alloc, so they obey pagecache memory
policies.

Also, using add_to_page_cache_lru instead of using a private
pagevec where applicable.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

31459fe4

12 5月, 2010 2 次提交

ceph: preserve seq # on requeued messages after transient transport errors · e84346b7

由 Sage Weil 提交于 5月 11, 2010

If the tcp connection drops and we reconnect to reestablish a stateful
session (with the mds), we need to resend previously sent (and possibly
received) messages with the _same_ seq # so that they can be dropped on
the other end if needed.  Only assign a new seq once after the message is
queued.
Signed-off-by: NSage Weil <sage@newdream.net>

e84346b7

ceph: zero unused message header, footer fields · 45c6ceb5

由 Sage Weil 提交于 5月 11, 2010

We shouldn't leak any prior memory contents to other parties.  And random
data, particularly in the 'version' field, can cause problems down the
line.
Signed-off-by: NSage Weil <sage@newdream.net>

45c6ceb5

04 5月, 2010 2 次提交

ceph: discard incoming messages with bad seq # · ae18756b

由 Sage Weil 提交于 4月 22, 2010

We can get old message seq #'s after a tcp reconnect for stateful sessions
(i.e., the MDS).  If we get a higher seq #, that is an error, and we
shouldn't see any bad seq #'s for stateless (mon, osd) connections.
Signed-off-by: NSage Weil <sage@newdream.net>

ae18756b

ceph: fix seq counting for skipped messages · 684be25c

由 Sage Weil 提交于 4月 21, 2010

Increment in_seq even when the message is skipped for some reason.
Signed-off-by: NSage Weil <sage@newdream.net>

684be25c

14 4月, 2010 1 次提交

ceph: use separate class for ceph sockets' sk_lock · a6a5349d

由 Sage Weil 提交于 4月 13, 2010

Use a separate class for ceph sockets to prevent lockdep confusion.
Because ceph sockets only get passed kernel pointers, there is no
dependency from sk_lock -> mmap_sem.  If we share the same class as other
sockets, lockdep detects a circular dependency from

	mmap_sem (page fault) -> fs mutex -> sk_lock -> mmap_sem

because dependencies are noted from both ceph and user contexts.  Using
a separate class prevents the sk_lock(ceph) -> mmap_sem dependency and
makes lockdep happy.
Signed-off-by: NSage Weil <sage@newdream.net>

a6a5349d

03 4月, 2010 1 次提交

ceph: fix ack counter reset on connection reset · 0e0d5e0c

由 Sage Weil 提交于 4月 02, 2010

If in_seq_acked isn't reset along with in_seq, we don't ack received
messages until we reach the old count, consuming gobs memory on the other
end of the connection and introducing a large delay when those messages
are eventually deleted.
Signed-off-by: NSage Weil <sage@newdream.net>

0e0d5e0c

30 3月, 2010 1 次提交

include cleanup: Update gfp.h and slab.h includes to prepare for breaking... · 5a0e3ad6

由 Tejun Heo 提交于 3月 24, 2010

include cleanup: Update gfp.h and slab.h includes to prepare for breaking implicit slab.h inclusion from percpu.h

percpu.h is included by sched.h and module.h and thus ends up being
included when building most .c files.  percpu.h includes slab.h which
in turn includes gfp.h making everything defined by the two files
universally available and complicating inclusion dependencies.

percpu.h -> slab.h dependency is about to be removed.  Prepare for
this change by updating users of gfp and slab facilities include those
headers directly instead of assuming availability.  As this conversion
needs to touch large number of source files, the following script is
used as the basis of conversion.

  http://userweb.kernel.org/~tj/misc/slabh-sweep.py

The script does the followings.

* Scan files for gfp and slab usages and update includes such that
  only the necessary includes are there.  ie. if only gfp is used,
  gfp.h, if slab is used, slab.h.

* When the script inserts a new include, it looks at the include
  blocks and try to put the new include such that its order conforms
  to its surrounding.  It's put in the include block which contains
  core kernel includes, in the same order that the rest are ordered -
  alphabetical, Christmas tree, rev-Xmas-tree or at the end if there
  doesn't seem to be any matching order.

* If the script can't find a place to put a new include (mostly
  because the file doesn't have fitting include block), it prints out
  an error message indicating which .h file needs to be added to the
  file.

The conversion was done in the following steps.

1. The initial automatic conversion of all .c files updated slightly
   over 4000 files, deleting around 700 includes and adding ~480 gfp.h
   and ~3000 slab.h inclusions.  The script emitted errors for ~400
   files.

2. Each error was manually checked.  Some didn't need the inclusion,
   some needed manual addition while adding it to implementation .h or
   embedding .c file was more appropriate for others.  This step added
   inclusions to around 150 files.

3. The script was run again and the output was compared to the edits
   from #2 to make sure no file was left behind.

4. Several build tests were done and a couple of problems were fixed.
   e.g. lib/decompress_*.c used malloc/free() wrappers around slab
   APIs requiring slab.h to be added manually.

5. The script was run on all .h files but without automatically
   editing them as sprinkling gfp.h and slab.h inclusions around .h
   files could easily lead to inclusion dependency hell.  Most gfp.h
   inclusion directives were ignored as stuff from gfp.h was usually
   wildly available and often used in preprocessor macros.  Each
   slab.h inclusion directive was examined and added manually as
   necessary.

6. percpu.h was updated not to include slab.h.

7. Build test were done on the following configurations and failures
   were fixed.  CONFIG_GCOV_KERNEL was turned off for all tests (as my
   distributed build env didn't work with gcov compiles) and a few
   more options had to be turned off depending on archs to make things
   build (like ipr on powerpc/64 which failed due to missing writeq).

   * x86 and x86_64 UP and SMP allmodconfig and a custom test config.
   * powerpc and powerpc64 SMP allmodconfig
   * sparc and sparc64 SMP allmodconfig
   * ia64 SMP allmodconfig
   * s390 SMP allmodconfig
   * alpha SMP allmodconfig
   * um on x86_64 SMP allmodconfig

8. percpu.h modifications were reverted so that it could be applied as
   a separate patch and serve as bisection point.

Given the fact that I had only a couple of failures from tests on step
6, I'm fairly confident about the coverage of this conversion patch.
If there is a breakage, it's likely to be something in one of the arch
headers which should be easily discoverable easily on most builds of
the specific arch.
Signed-off-by: NTejun Heo <tj@kernel.org>
Guess-its-ok-by: NChristoph Lameter <cl@linux-foundation.org>
Cc: Ingo Molnar <mingo@redhat.com>
Cc: Lee Schermerhorn <Lee.Schermerhorn@hp.com>

5a0e3ad6

23 3月, 2010 2 次提交

ceph: avoid reopening osd connections when address hasn't changed · 87b315a5

由 Sage Weil 提交于 3月 22, 2010

We get a fault callback on _every_ tcp connection fault.  Normally, we
want to reopen the connection when that happens.  If the address we have
is bad, however, and connection attempts always result in a connection
refused or similar error, explicitly closing and reopening the msgr
connection just prevents the messenger's backoff logic from kicking in.
The result can be a console full of

[ 3974.417106] ceph: osd11 10.3.14.138:6800 connection failed
[ 3974.423295] ceph: osd11 10.3.14.138:6800 connection failed
[ 3974.429709] ceph: osd11 10.3.14.138:6800 connection failed

Instead, if we get a fault, and have outstanding requests, but the osd
address hasn't changed and the connection never successfully connected in
the first place, do nothing to the osd connection.  The messenger layer
will back off and retry periodically, because we never connected and thus
the lossy bit is not set.

Instead, touch each request's r_stamp so that handle_timeout can tell the
request is still alive and kicking.
Signed-off-by: NSage Weil <sage@newdream.net>

87b315a5

ceph: fix connection fault con_work reentrancy problem · 3c3f2e32

由 Sage Weil 提交于 3月 18, 2010

The messenger fault was clearing the BUSY bit, for reasons unclear. This
made it possible for the con->ops->fault function to reopen the connection,
and requeue work in the workqueue--even though the current thread was
already in con_work.

This avoids a problem where the client busy loops with connection failures
on an unreachable OSD, but doesn't address the root cause of that problem.
Signed-off-by: NSage Weil <sage@newdream.net>

3c3f2e32

21 3月, 2010 1 次提交

ceph: fix authenticator timeout · 63733a0f

由 Sage Weil 提交于 3月 15, 2010

We were failing to reconnect to services due to an old authenticator, even
though we had the new ticket, because we weren't properly retrying the
connect handshake, because we were calling an old/incorrect helper that
left in_base_pos incorrect.  The result was a failure to reconnect to the
OSD or MDS (with an authentication error) if the MDS restarted after the
service had been up a few hours (long enough for the original authenticator
to be invalid).  This was only a problem if the AUTH_X authentication was
enabled.

Now that the 'negotiate' and 'connect' stages are fully separated, we
should use the prepare_read_connect() helper instead, and remove the
obsolete one.
Signed-off-by: NSage Weil <sage@newdream.net>

63733a0f

02 3月, 2010 2 次提交

ceph: reset front len on return to msgpool; BUG on mismatched front iov · 3ca02ef9

由 Sage Weil 提交于 3月 01, 2010

Reset msg front len when a message is returned to the pool: the caller
may have changed it.

BUG if we try to send a message with a hdr.front_len that doesn't match
the front iov.
Signed-off-by: NSage Weil <sage@newdream.net>

3ca02ef9

ceph: reset bits on connection close · 1679f876

由 Sage Weil 提交于 2月 26, 2010

Clear LOSSYTX bit, so that if/when we reconnect, said reconnect
will retry on failure.

Clear _PENDING bits too, to avoid polluting subsequent
connection state.

Drop unused REGISTERED bit.
Signed-off-by: NSage Weil <sage@newdream.net>

1679f876

26 2月, 2010 2 次提交

ceph: fix connection fault STANDBY check · e80a52d1

由 Sage Weil 提交于 2月 25, 2010

Move any out_sent messages to out_queue _before_ checking if
out_queue is empty and going to STANDBY, or else we may drop
something that was never acked.

And clean up the code a bit (less goto).
Signed-off-by: NSage Weil <sage@newdream.net>

e80a52d1

ceph: invalidate_authorizer without con->mutex held · 161fd65a

由 Sage Weil 提交于 2月 25, 2010

This fixes lock ABBA inversion, as the ->invalidate_authorizer()
op may need to take a lock (or even call back into the
messenger).
Signed-off-by: NSage Weil <sage@newdream.net>

161fd65a

24 2月, 2010 1 次提交

ceph: fix up unexpected message handling · 5b3a4db3

由 Sage Weil 提交于 2月 19, 2010

Fix skipping of unexpected message types from osd, mon.

Clean up pr_info and debug output.
Signed-off-by: NSage Weil <sage@newdream.net>

5b3a4db3

17 2月, 2010 2 次提交

ceph: cancel delayed work when closing connection · 91e45ce3

由 Sage Weil 提交于 2月 15, 2010

This ensures that if/when we reopen the connection, we can requeue work on
the connection immediately, without waiting for an old timer to expire.
Queue new delayed work inside con->mutex to avoid any race.

This fixes problems with clients failing to reconnect to the MDS due to
the client_reconnect message arriving too late (due to waiting for an old
delayed work timeout to expire).
Signed-off-by: NSage Weil <sage@newdream.net>

91e45ce3

ceph: allow connection to be reopened by fault callback · e2663ab6

由 Sage Weil 提交于 2月 16, 2010

Fix the messenger to allow a ceph_con_open() during the fault callback.
Previously the work wasn't getting queued on the connection because the
fault path avoids requeued work (normally spurious).  Loop on reopening by
checking for the OPENING state bit.

This fixes OSD reconnects when a TCP connection drops.
Signed-off-by: NSage Weil <sage@newdream.net>

e2663ab6

14 2月, 2010 1 次提交

ceph: fix msgr to keep sent messages until acked · 6c5d1a49

由 Sage Weil 提交于 2月 13, 2010

The test was backwards from commit b3d1dbbd: keep the message if the
connection _isn't_ lossy.  This allows the client to continue when the
TCP connection drops for some reason (network glitch) but both ends
survive.
Signed-off-by: NSage Weil <sage@newdream.net>

6c5d1a49

11 2月, 2010 1 次提交

ceph: allow renewal of auth credentials · 9bd2e6f8

由 Sage Weil 提交于 2月 02, 2010

Add infrastructure to allow the mon_client to periodically renew its auth
credentials.  Also add a messenger callback that will force such a renewal
if a peer rejects our authenticator.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

9bd2e6f8

30 1月, 2010 1 次提交

ceph: include type in ceph_entity_addr, filepath · ac8839d7

由 Sage Weil 提交于 1月 27, 2010

Include a type/version in ceph_entity_addr and filepath.  Include extra
byte in filepath encoding as necessary.
Signed-off-by: NSage Weil <sage@newdream.net>

ac8839d7

26 1月, 2010 4 次提交

ceph: keep reserved replies on the request structure · 0d59ab81

由 Yehuda Sadeh 提交于 1月 13, 2010

This includes treating all the data preallocation and revokation
at the same place, not having to have a special case for
the reserved pages.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>

0d59ab81

ceph: alloc message data pages and check if tid exists · 0547a9b3

由 Yehuda Sadeh 提交于 1月 11, 2010

Now doing it in the same callback that is also responsible for
allocating the 'front' part of the message. If we get a message
that we haven't got a corresponding tid for, mark it for skipping.

Moving the mutex unlock/lock from the osd alloc_msg callback
to the calling function in the messenger.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>

0547a9b3

Y
ceph: refactor messages data section allocation · 9d7f0f13
由 Yehuda Sadeh 提交于 1月 11, 2010
```
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
```
9d7f0f13

ceph: allocate middle of message before stating to read · 2450418c

由 Yehuda Sadeh 提交于 1月 08, 2010

Both front and middle parts of the message are now being
allocated at the ceph_alloc_msg().
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>

2450418c

15 1月, 2010 1 次提交

ceph: remove unused erank field · 103e2d3a

由 Sage Weil 提交于 1月 07, 2010

The ceph_entity_addr erank field is obsolete; remove it.  Get rid of
trivial addr comparison helpers while we're at it.
Signed-off-by: NSage Weil <sage@newdream.net>

103e2d3a

24 12月, 2009 4 次提交

ceph: support ceph_pagelist for message payload · 58bb3b37

由 Sage Weil 提交于 12月 23, 2009

The ceph_pagelist is a simple list of whole pages, strung together via
their lru list_head.  It facilitates encoding to a "buffer" of unknown
size.  Allow its use in place of the ceph_msg page vector.

This will be used to fix the huge buffer preallocation woes of MDS
reconnection.
Signed-off-by: NSage Weil <sage@newdream.net>

58bb3b37

ceph: add feature bits to connection handshake (protocol change) · 04a419f9

由 Sage Weil 提交于 12月 23, 2009

Define supported and required feature set.  Fail connection if the server
requires features we do not support (TAG_FEATURES), or if the server does
not support features we require.
Signed-off-by: NSage Weil <sage@newdream.net>

04a419f9

ceph: control access to page vector for incoming data · 350b1c32

由 Sage Weil 提交于 12月 22, 2009

When we issue an OSD read, we specify a vector of pages that the data is to
be read into. The request may be sent multiple times, to multiple OSDs, if
the osdmap changes, which means we can get more than one reply.

Only read data into the page vector if the reply is coming from the
OSD we last sent the request to. Keep track of which connection is using
the vector by taking a reference. If another connection was already
using the vector before and a new reply comes in on the right connection,
revoke the pages from the other connection.
Signed-off-by: NSage Weil <sage@newdream.net>

350b1c32

ceph: use connection mutex to protect read and write stages · ec302645

由 Sage Weil 提交于 12月 22, 2009

Use a single mutex (previously out_mutex) to protect both read and write
activity from concurrent ceph_con_* calls.  Drop the mutex when doing
callbacks to avoid nested locking (the callback may need to call something
like ceph_con_close).
Signed-off-by: NSage Weil <sage@newdream.net>

ec302645

22 12月, 2009 7 次提交

Y
ceph: remove unaccessible code · 169e16ce
由 Yehuda Sadeh 提交于 12月 16, 2009
```
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
```
169e16ce

ceph: plug leak of incoming message during connection fault/close · cf3e5c40

由 Sage Weil 提交于 12月 11, 2009

If we explicitly close a connection, or there is a socket error, we need
to drop any partially received message.
Signed-off-by: NSage Weil <sage@newdream.net>

cf3e5c40

S
ceph: hex dump corrupt server data to KERN_DEBUG · 9ec7cab1
由 Sage Weil 提交于 12月 14, 2009
```
Also, print fsid using standard format, NOT hex dump.
Signed-off-by: NSage Weil <sage@newdream.net>
```
9ec7cab1

ceph: don't save sent messages on lossy connections · b3d1dbbd

由 Sage Weil 提交于 12月 14, 2009

For lossy connections we drop all state on socket errors, so there is no
reason to keep sent ceph_msg's around.
Signed-off-by: NSage Weil <sage@newdream.net>

b3d1dbbd

ceph: detect lossy state of connection · 92ac41d0

由 Sage Weil 提交于 12月 14, 2009

The server indicates whether a connection is lossy; set our LOSSYTX bit
appropriately.  Do not set lossy bit on outgoing connections.
Signed-off-by: NSage Weil <sage@newdream.net>

92ac41d0

S
ceph: plug msg leak in con_fault · 5e095e8b
由 Sage Weil 提交于 12月 14, 2009
```
Signed-off-by: NSage Weil <sage@newdream.net>
```
5e095e8b

ceph: carry explicit msg reference for currently sending message · c86a2930

由 Sage Weil 提交于 12月 14, 2009

Carry a ceph_msg reference for connection->out_msg.  This will allow us to
make out_sent optional.
Signed-off-by: NSage Weil <sage@newdream.net>

c86a2930

08 12月, 2009 2 次提交

S
ceph: use kref for ceph_msg · c2e552e7
由 Sage Weil 提交于 12月 07, 2009
```
Signed-off-by: NSage Weil <sage@newdream.net>
```
c2e552e7

ceph: simplify ceph_buffer interface · b6c1d5b8

由 Sage Weil 提交于 12月 07, 2009

We never allocate the ceph_buffer and buffer separtely, so use a single
constructor.

Disallow put on NULL buffer; make the caller check.
Signed-off-by: NSage Weil <sage@newdream.net>

b6c1d5b8

21 11月, 2009 1 次提交

ceph: reset msgr backoff during open, not after successful handshake · 03c677e1

由 Sage Weil 提交于 11月 20, 2009

Reset the backoff delay when we reopen the connection, so that the delays
for any initial connection problems are reasonable. We were resetting only
after a successful handshake, which was of limited utility.
Signed-off-by: NSage Weil <sage@newdream.net>

03c677e1

openeuler / Kernel 1 年多 前同步成功

openeuler / Kernel
1 年多前同步成功