提交 · 40dae7ec537c5619fc93ad602c62f37be786d161 · Greenplum / Gpdb

19 3月, 2014 5 次提交

Make the handling of interrupted B-tree page splits more robust. · 40dae7ec

由 Heikki Linnakangas 提交于 3月 18, 2014

Splitting a page consists of two separate steps: splitting the child page,
and inserting the downlink for the new right page to the parent. Previously,
we handled the case that you crash in between those steps with a cleanup
routine after the WAL recovery had finished, which finished the incomplete
split. However, that doesn't help if the page split is interrupted but the
database doesn't crash, so that you don't perform WAL recovery. That could
happen for example if you run out of disk space.

Remove the end-of-recovery cleanup step. Instead, when a page is split, the
left page is marked with a new INCOMPLETE_SPLIT flag, and when the downlink
is inserted to the parent, the flag is cleared again. If an insertion sees
a page with the flag set, it knows that the split was interrupted for some
reason, and inserts the missing downlink before proceeding.

I used the same approach to fix GIN and GiST split algorithms earlier. This
was the last WAL cleanup routine, so we could get rid of that whole
machinery now, but I'll leave that for a separate patch.

Reviewed by Peter Geoghegan.

40dae7ec

T
Fix some remaining int64 vestiges in contrib/test_shm_mq. · b6ec7c92
由 Tom Lane 提交于 3月 18, 2014
```
Andres Freund and Tom Lane
```
b6ec7c92

test_shm_mq: Use Size rather than uint64. · c676ac0f

由 Robert Haas 提交于 3月 18, 2014

Commit 3bd261ca updated the API but
neglected to make the corresponding edits here.

Per Tom Lane and the buildfarm.

c676ac0f

R
Documentation for logical decoding. · 49c0864d
由 Robert Haas 提交于 3月 18, 2014
```
Craig Ringer, Andres Freund, Christian Kruse, with edits by me.
```
49c0864d

Add pg_recvlogical, a tool to receive data logical decoding data. · 8bdd12bb

由 Robert Haas 提交于 3月 18, 2014

This is fairly basic at the moment, but it's at least useful for
testing and debugging, and possibly more.

Andres Freund

8bdd12bb

18 3月, 2014 8 次提交

Rewrite comment for shm_mq_receive_bytes. · 250f8a7b

由 Robert Haas 提交于 3月 18, 2014

The comment and the code diverged at some point before the initial
commit of this feature, and I failed to notice.

Noted by Tom Lane.

250f8a7b

Fix relcache reference leak in refresh_by_match_merge(). · f7271c44

由 Tom Lane 提交于 3月 18, 2014

One path through the loop over indexes forgot to do index_close(). Rather
than adding a fourth call, restructure slightly so that there's only one.

In passing, get rid of an unnecessary syscache lookup: the pg_index struct
for the index is already available from its relcache entry.

Per report from YAMAMOTO Takashi, though this is a bit different from his
suggested patch. This is new code in HEAD, so no need for back-patch.

f7271c44

Improve shm_mq portability around MAXIMUM_ALIGNOF and sizeof(Size). · 3bd261ca

由 Robert Haas 提交于 3月 18, 2014

Revise the original decision to expose a uint64-based interface and
use Size everywhere possible.  Avoid assuming that MAXIMUM_ALIGNOF is
8, or making any assumption about the relationship between that value
and sizeof(Size).  If MAXIMUM_ALIGNOF is bigger, we'll now insert
padding after the length word; if it's smaller, we are now prepared
to read and write the length word in chunks.

Per discussion with Tom Lane.

3bd261ca

T
Fix pg_dumpall option parsing: -i doesn't take an argument. · 19f2d6cd
由 Tom Lane 提交于 3月 18, 2014
```
This used to work properly, but got fat-fingered in commit
3dee636e.  Per bug #9620 from
Nicolas Payart.
```
19f2d6cd
F
Fix help message and document in pg_receivexlog. · e726e59d
由 Fujii Masao 提交于 3月 18, 2014
```
Add SLOTNAME placeholder to --slot option in help message and
document.
```
e726e59d

Make it easy to detach completely from shared memory. · 79a4d24f

由 Robert Haas 提交于 3月 18, 2014

The new function dsm_detach_all() can be used either by postmaster
children that don't wish to take any risk of accidentally corrupting
shared memory; or by forked children of regular backends with
the same need.  This patch also updates the postmaster children that
already do PGSharedMemoryDetach() to do dsm_detach_all() as well.

Per discussion with Tom Lane.

79a4d24f

T

Release notes for 9.3.4, 9.2.8, 9.1.13, 9.0.17, 8.4.21. · 551fb5ac
由 Tom Lane 提交于 3月 17, 2014

551fb5ac

During index build, check and elog (not just Assert) for broken HOT chain. · d70cf811

由 Tom Lane 提交于 3月 17, 2014

The recently-fixed bug in WAL replay could result in not finding a parent
tuple for a heap-only tuple.  The existing code would either Assert or
generate an invalid index entry, neither of which is desirable.  Throw a
regular error instead.

d70cf811

17 3月, 2014 9 次提交

H
Fix thinko: have trueTriConsistentFn return GIN_TRUE. · d663d439
由 Heikki Linnakangas 提交于 3月 17, 2014
```
While we're at it, also improve comments in ginlogic.c.
```
d663d439
F
Fix typos in comments. · 2bccced1
由 Fujii Masao 提交于 3月 17, 2014
```
Thom Brown
```
2bccced1

Fix bug in clean shutdown of walsender that pg_receiving is connecting to. · 5c6d9fc4

由 Fujii Masao 提交于 3月 17, 2014

On clean shutdown, walsender waits for all WAL to be replicated to a standby,
and exits. It determined whether that replication had been completed by
checking whether its sent location had been equal to a standby's flush
location. Unfortunately this condition never becomes true when the standby
such as pg_receivexlog which always returns an invalid flush location is
connecting to walsender, and then walsender waits forever.

This commit changes walsender so that it just checks a standby's write
location if a flush location is invalid.

Back-patch to 9.1 where enough infrastructure for this exists.

5c6d9fc4

M
Fix small typo in comment · 02703ff2
由 Magnus Hagander 提交于 3月 17, 2014
```
Michael Paquier
```
02703ff2

plperl: Fix memory leak in hek2cstr · bd1154ed

由 Alvaro Herrera 提交于 3月 16, 2014

Backpatch all the way back to 9.1, where it was introduced by commit
50d89d42.

Reported by Sergey Burladyan in #9223
Author: Alex Hunsaker

bd1154ed

Fix unportable shell-script syntax in pg_upgrade's test.sh. · 0268d21e

由 Tom Lane 提交于 3月 16, 2014

I discovered the hard way that on some old shells, the locution
    FOO=""   unset FOO
does not behave the same as
    FOO="";  unset FOO
and in fact leaves FOO set to an empty string.  test.sh was inconsistently
spelling it different ways on adjacent lines.

This got broken relatively recently, in commit c737a2e5, so the lack of
field reports to date doesn't represent a lot of evidence that the problem
is rare.

0268d21e

P

Make punctuation consistent · 2861e8e9
由 Peter Eisentraut 提交于 3月 16, 2014

2861e8e9
P

Fix whitespace · e2b95947
由 Peter Eisentraut 提交于 3月 16, 2014

e2b95947

Fix advertised dispsize for libpq's sslmode connection parameter. · f4051e36

由 Tom Lane 提交于 3月 16, 2014

"8" was correct back when "disable" was the longest allowed value, but
since "verify-full" was added, it should be "12". Given the lack of
complaints, I wouldn't be surprised if nobody is actually using these
values ... but still, if they're in the API, they should be right.

Noticed while pursuing a different problem. It's been wrong for quite
a long time, so back-patch to all supported branches.

f4051e36

16 3月, 2014 3 次提交

Cleanups from the remove-native-krb5 patch · 0294023a

由 Magnus Hagander 提交于 3月 16, 2014

krb_srvname is actually not available anymore as a parameter server-side, since
with gssapi we accept all principals in our keytab. It's still used in libpq for
client side specification.

In passing remove declaration of krb_server_hostname, where all the functionality
was already removed.

Noted by Stephen Frost, though a different solution than his suggestion

0294023a

First-draft release notes for 9.3.4. · e3c9f232

由 Tom Lane 提交于 3月 15, 2014

As usual, the release notes for older branches will be made by cutting
these down, but put them up for community review first.

e3c9f232

T
Update time zone data files to tzdata release 2014a. · aba7f567
由 Tom Lane 提交于 3月 15, 2014
```
DST law changes in Fiji, Turkey; historical changes in Israel, Ukraine.
```
aba7f567

14 3月, 2014 4 次提交

Fix race condition in B-tree page deletion. · efada2b8

由 Heikki Linnakangas 提交于 3月 14, 2014

In short, we don't allow a page to be deleted if it's the rightmost child
of its parent, but that situation can change after we check for it.

Problem
-------

We check that the page to be deleted is not the rightmost child of its
parent, and then lock its left sibling, the page itself, its right sibling,
and the parent, in that order. However, if the parent page is split after
the check but before acquiring the locks, the target page might become the
rightmost child, if the split happens at the right place. That leads to an
error in vacuum (I reproduced this by setting a breakpoint in debugger):

ERROR: failed to delete rightmost child 41 of block 3 in index "foo_pkey"

We currently re-check that the page is still the rightmost child, and throw
the above error if it's not. We could easily just give up rather than throw
an error, but that approach doesn't scale to half-dead pages. To recap,
although we don't normally allow deleting the rightmost child, if the page
is the *only* child of its parent, we delete the child page and mark the
parent page as half-dead in one atomic operation. But before we do that, we
check that the parent can later be deleted, by checking that it in turn is
not the rightmost child of the grandparent (potentially recursing all the
way up to the root). But the same situation can arise there - the
grandparent can be split while we're not holding the locks. We end up with
a half-dead page that we cannot delete.

To make things worse, the keyspace of the deleted page has already been
transferred to its right sibling. As the README points out, the keyspace at
the grandparent level is "out-of-whack" until the half-dead page is deleted,
and if enough tuples with keys in the transferred keyspace are inserted, the
page might get split and a downlink might be inserted into the grandparent
that is out-of-order. That might not cause any serious problem if it's
transient (as the README ponders), but is surely bad if it stays that way.

Solution
--------

This patch changes the page deletion algorithm to avoid that problem. After
checking that the topmost page in the chain of to-be-deleted pages is not
the rightmost child of its parent, and then deleting the pages from bottom
up, unlink the pages from top to bottom. This way, the intermediate stages
are similar to the intermediate stages in page splitting, and there is no
transient stage where the keyspace is "out-of-whack". The topmost page in
the to-be-deleted chain doesn't have a downlink pointing to it, like a page
split before the downlink has been inserted.

This also allows us to get rid of the cleanup step after WAL recovery, if we
crash during page deletion. The deletion will be continued at next VACUUM,
but the tree is consistent for searches and insertions at every step.

This bug is old, all supported versions are affected, but this patch is too
big to back-patch (and changes the WAL record formats of related records).
We have not heard any reports of the bug from users, so clearly it's not
easy to bump into. Maybe backpatch later, after this has had some field
testing.

Reviewed by Kevin Grittner and Peter Geoghegan.

efada2b8

Prevent interrupts while reporting non-ERROR elog messages. · 6c461cb9

由 Tom Lane 提交于 3月 13, 2014

This should eliminate the risk of recursive entry to syslog(3), which
appears to be the cause of the hang reported in bug #9551 from James
Morton.

Arguably, the real problem here is auth.c's willingness to turn on
ImmediateInterruptOK while executing fairly wide swaths of backend code.
We may well need to work at narrowing the code ranges in which the
authentication_timeout interrupt is enabled.  For the moment, though,
this is a cheap and reasonably noninvasive fix for a field-reported
failure; the other approach would be complex and not necessarily
bug-free itself.

Back-patch to all supported branches.

6c461cb9

Allow psql to print COPY command status in more cases. · f70a78bc

由 Tom Lane 提交于 3月 13, 2014

Previously, psql would print the "COPY nnn" command status only for COPY
commands executed server-side. Now it will print that for frontend copies
too (including \copy). However, we continue to suppress the command status
for COPY TO STDOUT, since in that case the copy data has been routed to the
same place that the command status would go, and there is a risk of the
status line being mistaken for another line of COPY data. Doing that would
break existing scripts, and it doesn't seem worth the benefit --- this case
seems fairly analogous to SELECT, for which we also suppress the command
status.

Kumar Rajeev Rastogi, with substantial review by Amit Khandekar

f70a78bc

Avoid transaction-commit race condition while receiving a NOTIFY message. · 7bae0284

由 Tom Lane 提交于 3月 13, 2014

Use TransactionIdIsInProgress, then TransactionIdDidCommit, to distinguish
whether a NOTIFY message's originating transaction is in progress,
committed, or aborted. The previous coding could accept a message from a
transaction that was still in-progress according to the PGPROC array;
if the client were fast enough at starting a new transaction, it might fail
to see table rows added/updated by the message-sending transaction. Which
of course would usually be the point of receiving the message. We noted
this type of race condition long ago in tqual.c, but async.c overlooked it.

The race condition probably cannot occur unless there are multiple NOTIFY
senders in action, since an individual backend doesn't send NOTIFY signals
until well after it's done committing. But if two senders commit in close
succession, it's certainly possible that we could see the second sender's
message within the race condition window while responding to the signal
from the first one.

Per bug #9557 from Marko Tiikkaja. This patch is slightly more invasive
than what he proposed, since it removes the now-redundant
TransactionIdDidAbort call.

Back-patch to 9.0, where the current NOTIFY implementation was introduced.

7bae0284

13 3月, 2014 9 次提交
- H
  Fix a couple of typos in docs. · 16ff08b7
  由 Heikki Linnakangas 提交于 3月 13, 2014
```
Thom Brown
```
  16ff08b7
- B
  C comments: remove odd blank lines after #ifdef WIN32 lines · 242c2737
  由 Bruce Momjian 提交于 3月 13, 2014
```
A few more
```
  242c2737
- B
  
  C comments: remove odd blank lines after #ifdef WIN32 lines · 886c0be3
  由 Bruce Momjian 提交于 3月 13, 2014
  
  886c0be3
- H
  Only WAL-log the modified portion in an UPDATE, if possible. · a3115f0d
  由 Heikki Linnakangas 提交于 3月 12, 2014
```
When a row is updated, and the new tuple version is put on the same page as
the old one, only WAL-log the part of the new tuple that's not identical to
the old. This saves significantly on the amount of WAL that needs to be
written, in the common case that most fields are not modified.

Amit Kapila, with a lot of back and forth with me, Robert Haas, and others.
```
  a3115f0d
- H
  Items on GIN data pages are no longer always 6 bytes; update gincostestimate. · 17d787a3
  由 Heikki Linnakangas 提交于 3月 12, 2014
```
Also improve the comments a bit.
```
  17d787a3
- F
  Show PIDs of lock holders and waiters in log_lock_waits log message. · 588fb507
  由 Fujii Masao 提交于 3月 13, 2014
```
Christian Kruse, reviewed by Kumar Rajeev Rastogi.
```
  588fb507
- R
  test_decoding: Documentation fix. · a0b4c355
  由 Robert Haas 提交于 3月 12, 2014
```
Andres Freund
```
  a0b4c355
- R
  Fix incorrect assertion about historical snapshots. · 336a578b
  由 Robert Haas 提交于 3月 12, 2014
```
Also fix some nearby comments.

Andres Freund
```
  336a578b
- R
  Comment fixes related to logical decoding. · 890194f1
  由 Robert Haas 提交于 3月 12, 2014
```
Andres Freund, per complaints by Peter Eisentraut.
```
  890194f1
12 3月, 2014 2 次提交

Allow opclasses to provide tri-valued GIN consistent functions. · c5608ea2

由 Heikki Linnakangas 提交于 3月 12, 2014

With the GIN "fast scan" feature, GIN can skip items without fetching all
the keys for them, if it can prove that they don't match regardless of
those keys. So far, it has done the proving by calling the boolean
consistent function with all combinations of TRUE/FALSE for the unfetched
keys, but since that's O(n^2), it becomes unfeasible with more than a few
keys. We can avoid calling consistent with all the combinations, if we can
tell the operator class implementation directly which keys are unknown.

This commit includes a triConsistent function for the built-in array and
tsvector opclasses.

Alexander Korotkov, with some changes by me.

c5608ea2

In WAL replay, restore GIN metapage unconditionally to avoid torn page. · fecfc2b9

由 Heikki Linnakangas 提交于 3月 12, 2014

We don't take a full-page image of the GIN metapage; instead, the WAL record
contains all the information required to reconstruct it from scratch. But
to avoid torn page hazards, we must re-initialize it from the WAL record
every time, even if it already has a greater LSN, similar to how normal full
page images are restored.

This was highly unlikely to cause any problems in practice, because the GIN
metapage is small. We rely on an update smaller than a 512 byte disk sector
to be atomic elsewhere, at least in pg_control. But better safe than sorry,
and this would be easy to overlook if more fields are added to the metapage
so that it's no longer small.

Reported by Noah Misch. Backpatch to all supported versions.

fecfc2b9