提交 · ab75144be08cfc1d80f49e9c37970fcadb1215a2 · openeuler / Kernel

07 7月, 2017 9 次提交

libceph: kill __{insert,lookup,remove}_pg_mapping() · ab75144b

由 Ilya Dryomov 提交于 6月 21, 2017

Switch to DEFINE_RB_FUNCS2-generated {insert,lookup,erase}_pg_mapping().
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

ab75144b

I
libceph: introduce and switch to decode_pg_mapping() · a303bb0e
由 Ilya Dryomov 提交于 6月 21, 2017
```
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
```
a303bb0e

libceph: don't pass pgid by value · 33333d10

由 Ilya Dryomov 提交于 6月 21, 2017

Make __{lookup,remove}_pg_mapping() look like their ceph_spg_mapping
counterparts: take const struct ceph_pg *.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

33333d10

I
libceph: respect RADOS_BACKOFF backoffs · a02a946d
由 Ilya Dryomov 提交于 6月 19, 2017
```
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
```
a02a946d
I
libceph: avoid unnecessary pi lookups in calc_target() · df28152d
由 Ilya Dryomov 提交于 6月 15, 2017
```
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
```
df28152d

libceph: resend on PG splits if OSD has RESEND_ON_SPLIT · 7de030d6

由 Ilya Dryomov 提交于 6月 15, 2017

Note that ceph_osd_request_target fields are updated regardless of
RESEND_ON_SPLIT.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

7de030d6

libceph: introduce ceph_spg, ceph_pg_to_primary_shard() · dc98ff72

由 Ilya Dryomov 提交于 6月 15, 2017

Store both raw pgid and actual spgid in ceph_osd_request_target.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

dc98ff72

libceph: new pi->last_force_request_resend · 8e48cf00

由 Ilya Dryomov 提交于 6月 05, 2017

The old (v15) pi->last_force_request_resend has been repurposed to
make pre-RESEND_ON_SPLIT clients that don't check for PG splits but do
obey pi->last_force_request_resend resend on splits.  See ceph.git
commit 189ca7ec6420 ("mon/OSDMonitor: make pre-luminous clients resend
ops on split").
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

8e48cf00

I
libceph: handle non-empty dest in ceph_{oloc,oid}_copy() · ca35ffea
由 Ilya Dryomov 提交于 6月 05, 2017
```
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
```
ca35ffea

24 5月, 2017 1 次提交

libceph: NULL deref on crush_decode() error path · 293dffaa

由 Dan Carpenter 提交于 5月 23, 2017

If there is not enough space then ceph_decode_32_safe() does a goto bad.
We need to return an error code in that situation.  The current code
returns ERR_PTR(0) which is NULL.  The callers are not expecting that
and it results in a NULL dereference.

Fixes: f24e9980 ("ceph: OSD client")
Signed-off-by: NDan Carpenter <dan.carpenter@oracle.com>
Reviewed-by: NIlya Dryomov <idryomov@gmail.com>
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

293dffaa

07 3月, 2017 2 次提交

libceph: don't set weight to IN when OSD is destroyed · b581a585

由 Ilya Dryomov 提交于 3月 01, 2017

Since ceph.git commit 4e28f9e63644 ("osd/OSDMap: clear osd_info,
osd_xinfo on osd deletion"), weight is set to IN when OSD is deleted.
This changes the result of applying an incremental for clients, not
just OSDs. Because CRUSH computations are obviously affected,
pre-4e28f9e63644 servers disagree with post-4e28f9e63644 clients on
object placement, resulting in misdirected requests.

Mirrors ceph.git commit a6009d1039a55e2c77f431662b3d6cc5a8e8e63f.

Fixes: 930c5328 ("libceph: apply new_state before new_up_client on incrementals")
Link: http://tracker.ceph.com/issues/19122Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
Reviewed-by: NSage Weil <sage@redhat.com>

b581a585

libceph: fix crush_decode() for older maps · 9afd30db

由 Ilya Dryomov 提交于 2月 28, 2017

Older (shorter) CRUSH maps too need to be finalized.

Fixes: 66a0e2d5 ("crush: remove mutable part of CRUSH map")
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

9afd30db

20 2月, 2017 4 次提交

libceph: don't go through with the mapping if the PG is too wide · ef9324bb

由 Ilya Dryomov 提交于 2月 08, 2017

With EC overwrites maturing, the kernel client will be getting exposed
to potentially very wide EC pools. While "min(pi->size, X)" works fine
when the cluster is stable and happy, truncating OSD sets interferes
with resend logic (ceph_is_new_interval(), etc). Abort the mapping if
the pool is too wide, assigning the request to the homeless session.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
Reviewed-by: NSage Weil <sage@redhat.com>

ef9324bb

crush: merge working data and scratch · 743efcff

由 Ilya Dryomov 提交于 1月 31, 2017

Much like Arlo Guthrie, I decided that one big pile is better than two
little piles.

Reflects ceph.git commit 95c2df6c7e0b22d2ea9d91db500cf8b9441c73ba.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

743efcff

crush: remove mutable part of CRUSH map · 66a0e2d5

由 Ilya Dryomov 提交于 1月 31, 2017

Then add it to the working state. It would be very nice if we didn't
have to take a lock to calculate a crush placement. By moving the
permutation array into the working data, we can treat the CRUSH map as
immutable.

Reflects ceph.git commit cbcd039651c0569551cb90d26ce27e1432671f2a.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

66a0e2d5

libceph: add osdmap_set_crush() helper · 1b6a78b5

由 Ilya Dryomov 提交于 1月 31, 2017

Simplify osdmap_decode() and osdmap_apply_incremental() a bit.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

1b6a78b5

28 7月, 2016 2 次提交

libceph: rados pool namespace support · 30c156d9

由 Yan, Zheng 提交于 2月 14, 2016

Add pool namesapce pointer to struct ceph_file_layout and struct
ceph_object_locator. Pool namespace is used by when mapping object
to PG, it's also used when composing OSD request.

The namespace pointer in struct ceph_file_layout is RCU protected.
So libceph can read namespace without taking lock.
Signed-off-by: NYan, Zheng <zyan@redhat.com>
[idryomov@gmail.com: ceph_oloc_destroy(), misc minor changes]
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

30c156d9

libceph: define new ceph_file_layout structure · 7627151e

由 Yan, Zheng 提交于 2月 03, 2016

Define new ceph_file_layout structure and rename old ceph_file_layout
to ceph_file_layout_legacy. This is preparation for adding namespace
to ceph_file_layout structure.
Signed-off-by: NYan, Zheng <zyan@redhat.com>

7627151e

22 7月, 2016 1 次提交

libceph: apply new_state before new_up_client on incrementals · 930c5328

由 Ilya Dryomov 提交于 7月 19, 2016

Currently, osd_weight and osd_state fields are updated in the encoding
order.  This is wrong, because an incremental map may look like e.g.

    new_up_client: { osd=6, addr=... } # set osd_state and addr
    new_state: { osd=6, xorstate=EXISTS } # clear osd_state

Suppose osd6's current osd_state is EXISTS (i.e. osd6 is down).  After
applying new_up_client, osd_state is changed to EXISTS | UP.  Carrying
on with the new_state update, we flip EXISTS and leave osd6 in a weird
"!EXISTS but UP" state.  A non-existent OSD is considered down by the
mapping code

2087    for (i = 0; i < pg->pg_temp.len; i++) {
2088            if (ceph_osd_is_down(osdmap, pg->pg_temp.osds[i])) {
2089                    if (ceph_can_shift_osds(pi))
2090                            continue;
2091
2092                    temp->osds[temp->size++] = CRUSH_ITEM_NONE;

and so requests get directed to the second OSD in the set instead of
the first, resulting in OSD-side errors like:

[WRN] : client.4239 192.168.122.21:0/2444980242 misdirected client.4239.1:2827 pg 2.5df899f2 to osd.4 not [1,4,6] in e680/680

and hung rbds on the client:

[  493.566367] rbd: rbd0: write 400000 at 11cc00000 (0)
[  493.566805] rbd: rbd0:   result -6 xferred 400000
[  493.567011] blk_update_request: I/O error, dev rbd0, sector 9330688

The fix is to decouple application from the decoding and:
- apply new_weight first
- apply new_state before new_up_client
- twiddle osd_state flags if marking in
- clear out some of the state if osd is destroyed

Fixes: http://tracker.ceph.com/issues/14901

Cc: stable@vger.kernel.org # 3.15+: 6dd74e44: libceph: set 'exists' flag for newly up osd
Cc: stable@vger.kernel.org # 3.15+
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>
Reviewed-by: NJosh Durgin <jdurgin@redhat.com>

930c5328

31 5月, 2016 1 次提交

libceph: use %s instead of %pE in dout()s · 4a3262b1

由 Ilya Dryomov 提交于 5月 30, 2016

Commit d30291b9 ("libceph: variable-sized ceph_object_id") changed
dout()s in what is now encode_request() and ceph_object_locator_to_pg()
to use %pE, mostly to document that, although all rbd and cephfs object
names are NULL-terminated strings, ceph_object_id will handle any RADOS
object name, including the one containing NULs, just fine.

However, it turns out that vbin_printf() can't handle anything but ints
and %s - all %p suffixes are ignored. The buffer %p** points to isn't
recorded, resulting in trash in the messages if the buffer had been
reused by the time bstr_printf() got to it.
Signed-off-by: NIlya Dryomov <idryomov@gmail.com>

4a3262b1

26 5月, 2016 9 次提交

libceph: allocate dummy osdmap in ceph_osdc_init() · e5253a7b