提交 · 1679f876a641d209e7b22e43ebda0693c71003cf · openanolis / cloud-kernel

02 3月, 2010 1 次提交

ceph: reset bits on connection close · 1679f876

由 Sage Weil 提交于 2月 26, 2010

Clear LOSSYTX bit, so that if/when we reconnect, said reconnect
will retry on failure.

Clear _PENDING bits too, to avoid polluting subsequent
connection state.

Drop unused REGISTERED bit.
Signed-off-by: NSage Weil <sage@newdream.net>

1679f876

27 2月, 2010 2 次提交

ceph: remove bogus mds forward warning · 080af17e

由 Sage Weil 提交于 2月 25, 2010

The must_resend flag is always true, not false.  In any case, we can
just ignore it anyway.
Signed-off-by: NSage Weil <sage@newdream.net>

080af17e

ceph: remove fragile __map_osds optimization · c99eb1c7

由 Sage Weil 提交于 2月 26, 2010

We used to try to avoid freeing and then reallocating the osd
struct.  This is a bit fragile due to potential interactions with
other references (beyond o_requests), and may be the cause of
this crash:

[120633.442358] BUG: unable to handle kernel NULL pointer dereference at (null)
[120633.443292] IP: [<ffffffff812549b6>] rb_erase+0x11d/0x277
[120633.443292] PGD f7ff3067 PUD f7f53067 PMD 0
[120633.443292] Oops: 0000 [#1] PREEMPT SMP
[120633.443292] last sysfs file: /sys/kernel/uevent_seqnum
[120633.443292] CPU 1
[120633.443292] Modules linked in: ceph fan ac battery psmouse ehci_hcd ide_pci_generic ohci_hcd thermal processor button
[120633.443292] Pid: 3023, comm: ceph-msgr/1 Not tainted 2.6.32-rc2 #12 H8SSL
[120633.443292] RIP: 0010:[<ffffffff812549b6>]  [<ffffffff812549b6>] rb_erase+0x11d/0x277
[120633.443292] RSP: 0018:ffff8800f7b13a50  EFLAGS: 00010246
[120633.443292] RAX: ffff880022907819 RBX: ffff880022907818 RCX: 0000000000000000
[120633.443292] RDX: ffff8800f7b13a80 RSI: ffff8800f587eb48 RDI: 0000000000000000
[120633.443292] RBP: ffff8800f7b13a60 R08: 0000000000000000 R09: 0000000000000004
[120633.443292] R10: 0000000000000000 R11: ffff8800c4441000 R12: ffff8800f587eb48
[120633.443292] R13: ffff8800f58eaa00 R14: ffff8800f413c000 R15: 0000000000000001
[120633.443292] FS:  00007fbef6e226e0(0000) GS:ffff880009200000(0000) knlGS:0000000000000000
[120633.443292] CS:  0010 DS: 0018 ES: 0018 CR0: 000000008005003b
[120633.443292] CR2: 0000000000000000 CR3: 00000000f7c53000 CR4: 00000000000006e0
[120633.443292] DR0: 0000000000000000 DR1: 0000000000000000 DR2: 0000000000000000
[120633.443292] DR3: 0000000000000000 DR6: 00000000ffff0ff0 DR7: 0000000000000400
[120633.443292] Process ceph-msgr/1 (pid: 3023, threadinfo ffff8800f7b12000, task ffff8800f5858b40)
[120633.443292] Stack:
[120633.443292]  ffff8800f413c000 ffff8800f587e9c0 ffff8800f7b13a80 ffffffffa0098a86
[120633.443292] <0> 00000000000006f1 0000000000000000 ffff8800f7b13af0 ffffffffa009959b
[120633.443292] <0> ffff8800f413c000 ffff880022a68400 ffff880022a68400 ffff8800f587e9c0
[120633.443292] Call Trace:
[120633.443292]  [<ffffffffa0098a86>] __remove_osd+0x4d/0xbc [ceph]
[120633.443292]  [<ffffffffa009959b>] __map_osds+0x199/0x4fa [ceph]
[120633.443292]  [<ffffffffa00999f4>] ? __send_request+0xf8/0x186 [ceph]
[120633.443292]  [<ffffffffa0099beb>] kick_requests+0x169/0x3cb [ceph]
[120633.443292]  [<ffffffffa009a8c1>] ceph_osdc_handle_map+0x370/0x522 [ceph]

Since we're probably screwed anyway if a small kmalloc is
failing, don't bother with trying to be clever here.
Signed-off-by: NSage Weil <sage@newdream.net>

c99eb1c7

26 2月, 2010 2 次提交

ceph: fix connection fault STANDBY check · e80a52d1

由 Sage Weil 提交于 2月 25, 2010

Move any out_sent messages to out_queue _before_ checking if
out_queue is empty and going to STANDBY, or else we may drop
something that was never acked.

And clean up the code a bit (less goto).
Signed-off-by: NSage Weil <sage@newdream.net>

e80a52d1

ceph: invalidate_authorizer without con->mutex held · 161fd65a

由 Sage Weil 提交于 2月 25, 2010

This fixes lock ABBA inversion, as the ->invalidate_authorizer()
op may need to take a lock (or even call back into the
messenger).
Signed-off-by: NSage Weil <sage@newdream.net>

161fd65a

24 2月, 2010 6 次提交

Y
ceph: don't clobber write return value when using O_SYNC · 88d892a3
由 Yehuda Sadeh 提交于 2月 23, 2010
```
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>
```
88d892a3

ceph: fix client_request_forward decoding · a1ea787c

由 Sage Weil 提交于 2月 23, 2010

The tid is in the message header, not body.  Broken since 6df058c0.

No need to look at next mds session; just mark the request and be done.
(The old error path was broken too, but now it's gone.)
Signed-off-by: NSage Weil <sage@newdream.net>

a1ea787c

ceph: drop messages on unregistered mds sessions; cleanup · 2600d2dd

由 Sage Weil 提交于 2月 22, 2010

Verify the mds session is currently registered before handling
incoming messages.  Clean up message handlers to pull mds out
of session->s_mds instead of less trustworthy src field.

Clean up con_{get,put} debug output.
Signed-off-by: NSage Weil <sage@newdream.net>

2600d2dd

ceph: fix comments, locking in destroy_inode · a6369741

由 Sage Weil 提交于 2月 22, 2010

The destroy_inode path needs no inode locks since there are no
inode references.  Update __ceph_remove_cap comment to reflect
that it is called without cap->session->s_mutex in this case.
Signed-off-by: NSage Weil <sage@newdream.net>

a6369741

ceph: move dereference after NULL test · 4ce1e9ad

由 Alexander Beregalov 提交于 2月 22, 2010

Signed-off-by: NAlexander Beregalov <a.beregalov@gmail.com>
Signed-off-by: NSage Weil <sage@newdream.net>

4ce1e9ad

ceph: fix up unexpected message handling · 5b3a4db3

由 Sage Weil 提交于 2月 19, 2010

Fix skipping of unexpected message types from osd, mon.

Clean up pr_info and debug output.
Signed-off-by: NSage Weil <sage@newdream.net>

5b3a4db3

20 2月, 2010 4 次提交

ceph: cleanup redundant code in handle_cap_grant · bcd2cbd1

由 Yehuda Sadeh 提交于 2月 19, 2010

There is no state in local vars that requires us to loop after temporarily
dropping i_lock.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

bcd2cbd1

ceph: don't truncate dirty pages in invalidate work thread · c9af9fb6

由 Yehuda Sadeh 提交于 2月 19, 2010

Instead of truncating the whole range of pages, we skip those
pages that are dirty or in the middle of writeback. Those pages
will be cleared later when the writeback completes.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

c9af9fb6

ceph: remove page upon writeback completion if lost cache cap · e63dc5c7

由 Yehuda Sadeh 提交于 2月 19, 2010

This page should have been removed earlier when the cache cap was
revoked, but a writeback was in flight, so it was skipped. We truncate
it here just as the writeback finishes, while it's still locked.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

e63dc5c7

ceph: fix check for invalidate_mapping_pages success · 5ecad6fd

由 Sage Weil 提交于 2月 17, 2010

We need to know whether there was any page left behind, and not the
return value (the total number of pages invalidated).  Look at the mapping
to see if we were successful or not.

Move it all into a helper to simplify the two callers.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

5ecad6fd

18 2月, 2010 6 次提交

S
ceph: fix typo in ceph_queue_writeback debug output · 2c27c9a5
由 Sage Weil 提交于 2月 17, 2010
```
Signed-off-by: NSage Weil <sage@newdream.net>
```
2c27c9a5
S
ceph: v0.19 release · a17d6473
由 Sage Weil 提交于 2月 17, 2010
```
Signed-off-by: NSage Weil <sage@newdream.net>
```
a17d6473

ceph: use rbtree for pg pools; decode new osdmap format · 4fc51be8

由 Sage Weil 提交于 2月 16, 2010

Since we can now create and destroy pg pools, the pool ids will be sparse,
and an array no longer makes sense for looking up by pool id.  Use an
rbtree instead.

The OSDMap encoding also no longer has a max pool count (previously used to
allocate the array).  There is a new pool_max, that is the largest pool id
we've ever used, although we don't actually need it in the client.
Signed-off-by: NSage Weil <sage@newdream.net>

4fc51be8

S
ceph: fix memory leak when destroying osdmap with pg_temp mappings · 9794b146
由 Sage Weil 提交于 2月 16, 2010
```
Also move _lookup_pg_mapping into a helper.
Signed-off-by: NSage Weil <sage@newdream.net>
```
9794b146

ceph: fix iterate_caps removal race · 7c1332b8

由 Sage Weil 提交于 2月 16, 2010

We need to be able to iterate over all caps on a session with a
possibly slow callback on each cap.  To allow this, we used to
prevent cap reordering while we were iterating.  However, we were
not safe from races with removal: removing the 'next' cap would
make the next pointer from list_for_each_entry_safe be invalid,
and cause a lock up or similar badness.

Instead, we keep an iterator pointer in the session pointing to
the current cap.  As before, we avoid reordering.  For removal,
if the cap isn't the current cap we are iterating over, we are
fine.  If it is, we clear cap->ci (to mark the cap as pending
removal) but leave it in the session list.  In iterate_caps, we
can safely finish removal and get the next cap pointer.

While we're at it, clean up put_cap to not take a cap reservation
context, as it was never used.
Signed-off-by: NSage Weil <sage@newdream.net>

7c1332b8

ceph: clean up readdir caps reservation · 85ccce43

由 Sage Weil 提交于 2月 17, 2010

Use a global counter for the minimum number of allocated caps instead of
hard coding a check against readdir_max.  This takes into account multiple
client instances, and avoids examining the superblock mount options when a
cap is dropped.
Signed-off-by: NSage Weil <sage@newdream.net>

85ccce43

17 2月, 2010 6 次提交

ceph: fix authentication races, auth_none oops · 5ce6e9db

由 Sage Weil 提交于 2月 15, 2010

Call __validate_auth() under monc->mutex, and use helper for
initial hello so that the pending_auth flag is set.  This fixes
possible races in which we have an authentication request (hello
or otherwise) pending and send another one.  In particular, with
auth_none, we _never_ want to call ceph_build_auth() from
__validate_auth(), since the ->build_request() method is NULL.
Signed-off-by: NSage Weil <sage@newdream.net>

5ce6e9db

ceph: use rbtree for mon statfs requests · 85ff03f6

由 Sage Weil 提交于 2月 15, 2010

An rbtree is lighter weight, particularly given we will generally have
very few in-flight statfs requests.
Signed-off-by: NSage Weil <sage@newdream.net>

85ff03f6

ceph: use rbtree for snap_realms · a105f00c

由 Sage Weil 提交于 2月 15, 2010

Switch from radix tree to rbtree for snap realms.  This is much more
appropriate given that realm keys are few and far between.
Signed-off-by: NSage Weil <sage@newdream.net>

a105f00c

ceph: use rbtree for mds requests · 44ca18f2

由 Sage Weil 提交于 2月 15, 2010

The rbtree is a more appropriate data structure than a radix_tree.  It
avoids extra memory usage and simplifies the code.

It also fixes a bug where the debugfs 'mdsc' file wasn't including the
most recent mds request.
Signed-off-by: NSage Weil <sage@newdream.net>

44ca18f2

ceph: cancel delayed work when closing connection · 91e45ce3

由 Sage Weil 提交于 2月 15, 2010

This ensures that if/when we reopen the connection, we can requeue work on
the connection immediately, without waiting for an old timer to expire.
Queue new delayed work inside con->mutex to avoid any race.

This fixes problems with clients failing to reconnect to the MDS due to
the client_reconnect message arriving too late (due to waiting for an old
delayed work timeout to expire).
Signed-off-by: NSage Weil <sage@newdream.net>

91e45ce3

ceph: allow connection to be reopened by fault callback · e2663ab6

由 Sage Weil 提交于 2月 16, 2010

Fix the messenger to allow a ceph_con_open() during the fault callback.
Previously the work wasn't getting queued on the connection because the
fault path avoids requeued work (normally spurious).  Loop on reopening by
checking for the OPENING state bit.

This fixes OSD reconnects when a TCP connection drops.
Signed-off-by: NSage Weil <sage@newdream.net>

e2663ab6

16 2月, 2010 1 次提交

ceph: reset osd connections after fault · 153a008b

由 Sage Weil 提交于 2月 15, 2010

A single osd connection fault (e.g. tcp disconnect) wasn't
reopening the connection, which causes all current and future
requests for that osd to hang.
Signed-off-by: NSage Weil <sage@newdream.net>

153a008b

14 2月, 2010 1 次提交

ceph: fix msgr to keep sent messages until acked · 6c5d1a49

由 Sage Weil 提交于 2月 13, 2010

The test was backwards from commit b3d1dbbd: keep the message if the
connection _isn't_ lossy.  This allows the client to continue when the
TCP connection drops for some reason (network glitch) but both ends
survive.
Signed-off-by: NSage Weil <sage@newdream.net>

6c5d1a49

12 2月, 2010 11 次提交

ceph: remove bogus invalidate_mapping_pages · 80310491

由 Sage Weil 提交于 2月 09, 2010

We were invalidating mapping pages when dropping FILE_CACHE in
__send_cap().  But ceph_check_caps attempts to invalidate already, and
also checks for success, so we should never get to this point.
Signed-off-by: NSage Weil <sage@newdream.net>

80310491

ceph: invalidate pages even if truncate is pending · 0840d8af

由 Sage Weil 提交于 2月 09, 2010

There is no reason not to invalidate pages when a truncate is pending.
Both throw out page cache pages.
Signed-off-by: NSage Weil <sage@newdream.net>

0840d8af

ceph: cleanup async writeback, truncation, invalidate helpers · 3c6f6b79

由 Sage Weil 提交于 2月 09, 2010

Grab inode ref in helper.  Make work functions static, with consistent
naming.
Signed-off-by: NSage Weil <sage@newdream.net>

3c6f6b79

ceph: fix sync read eof check deadlock · 6a026589

由 Sage Weil 提交于 2月 09, 2010

If a sync read gets a short result from the OSD, it may need to do a
getattr to see if it is short due to reaching end-of-file. The getattr
was being done while holding a reference to FILE_RD, which can lead to
a deadlock if the MDS is revoking that capability bit and can't process
the getattr until it does.

We fix this by setting a flag if EOF size validation is needed, and doing
the getattr in ceph_aio_read, after the RD cap ref is dropped. If the
read needs to be continued, we loop and continue traversing the file.
Signed-off-by: NSage Weil <sage@newdream.net>

6a026589

ceph: do not retain caps that are being revoked · 68c28323

由 Sage Weil 提交于 2月 09, 2010

Never retain caps in __send_cap() that are being revoked.
Signed-off-by: NSage Weil <sage@newdream.net>

68c28323

ceph: cap revocation fixes · cbd03635

由 Sage Weil 提交于 2月 09, 2010

Try to invalidate pages in ceph_check_caps() if FILE_CACHE is being
revoked. If we fail, queue an immediate async invalidate if FILE_CACHE
is being revoked. (If it's not being revoked, we just queue the caps
for later evaluation later, as per the old behavior.)
Signed-off-by: NSage Weil <sage@newdream.net>

cbd03635

ceph: sync read/write considers page cache · 29065a51

由 Yehuda Sadeh 提交于 2月 09, 2010

In the cases where we either do a sync read or a write, we
need to make sure that everything in the page cache is flushed.
In the case of a sync write we invalidate the relevant pages,
so that subsequent read/write reflects the new data written.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

29065a51

ceph: fix truncation when not holding caps · 3d497d85

由 Yehuda Sadeh 提交于 2月 09, 2010

A truncation should occur when either we have the
specified caps for the file, or (in cases where we are
not the only ones referencing the file) when it is mapped
or when it is opened. The latter two cases were not
handled.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

3d497d85

ceph: refactor ceph_write_begin, fix ceph_page_mkwrite · 4af6b225

由 Yehuda Sadeh 提交于 2月 09, 2010

Originally ceph_page_mkwrite called ceph_write_begin, hoping that
the returned locked page would be the page that it was requested
to mkwrite. Factored out relevant part of ceph_page_mkwrite and
we lock the right page anyway.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

4af6b225

ceph: fix short synchronous reads · 972f0d3a

由 Yehuda Sadeh 提交于 2月 04, 2010

Zeroing of holes was not done correctly: page_off was miscalculated and
zeroing the tail didn't not adjust the 'read' value to include the zeroed
portion.
Signed-off-by: NYehuda Sadeh <yehuda@hq.newdream.net>
Signed-off-by: NSage Weil <sage@newdream.net>

972f0d3a

S
ceph: add uid field to ceph_pg_pool · 02f90c61
由 Sage Weil 提交于 2月 04, 2010
```
Also verify encoding version as we go.
Signed-off-by: NSage Weil <sage@newdream.net>
```
02f90c61

openanolis / cloud-kernel 接近 2 年 前同步成功

openanolis / cloud-kernel
接近 2 年前同步成功