提交 · 53dda6ff6263a3f350514d9edae600468c946ed4 · 李少辉-开发者 / git

27 9月, 2006 1 次提交

introduce delta objects with offset to base · eb32d236

由 Nicolas Pitre 提交于 9月 21, 2006

This adds a new object, namely OBJ_OFS_DELTA, renames OBJ_DELTA to
OBJ_REF_DELTA to better make the distinction between those two delta
objects, and adds support for the handling of those new delta objects
in sha1_file.c only.

The OBJ_OFS_DELTA contains a relative offset from the delta object's
position in a pack instead of the 20-byte SHA1 reference to identify
the base object.  Since the base is likely to be not so far away, the
relative offset is more likely to have a smaller encoding on average
than an absolute offset.  And for those delta objects the base must
always be stored first because there is no way to know the distance of
later objects when streaming a pack.  Hence this relative offset is
always meant to be negative.

The offset encoding is slightly denser than the one used for object
size -- credits to <linux@horizon.com> (whoever this is) for bringing
it to my attention.

This allows for pack size reduction between 3.2% (Linux-2.6) to over 5%
(linux-historic).  Runtime pack access should be faster too since delta
replay does skip a search in the pack index for each delta in a chain.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

eb32d236

23 9月, 2006 1 次提交

many cleanups to sha1_file.c · 43057304

由 Nicolas Pitre 提交于 9月 21, 2006

Those cleanups are mainly to set the table for the support of deltas
with base objects referenced by offsets instead of sha1.  This means
that many pack lookup functions are converted to take a pack/offset
tuple instead of a sha1.

This eliminates many struct pack_entry usages since this structure
carried redundent information in many cases, and it increased stack
footprint needlessly for a couple recursively called functions that used
to declare a local copy of it for every recursion loop.

In the process, packed_object_info_detail() has been reorganized as well
so to look much saner and more amenable to deltas with offset support.

Finally the appropriate adjustments have been made to functions that
depend on the above changes.  But there is no functionality changes yet
simply some code refactoring at this point.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

43057304

13 9月, 2006 1 次提交
- J
  pack-objects: document --revs, --unpacked and --all. · 4321134c
  由 Junio C Hamano 提交于 9月 12, 2006
```
Signed-off-by: NJunio C Hamano <junkio@cox.net>
```
  4321134c
07 9月, 2006 2 次提交

pack-objects: further work on internal rev-list logic. · 8d1d8f83

由 Junio C Hamano 提交于 9月 06, 2006

This teaches the internal rev-list logic to understand options
that are needed for pack handling: --all, --unpacked, and --thin.

It also moves two functions from builtin-rev-list to list-objects
so that the two programs can share more code.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

8d1d8f83

pack-objects: run rev-list equivalent internally. · b5d97e6b

由 Junio C Hamano 提交于 9月 04, 2006

Instead of piping the rev-list output from its standard input,
you can say:

	pack-objects --all --unpacked --revs pack

and feed the rev parameters you would otherwise give the
rev-list on its command line from the standard input.
In other words:

	echo 'master..next' | pack-objects --revs pack

and

	rev-list --objects master..next | pack-objects pack

are equivalent.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

b5d97e6b

04 9月, 2006 2 次提交

more lightweight revalidation while reusing deflated stream in packing · 72518e9c

由 Junio C Hamano 提交于 9月 03, 2006

When copying from an existing pack and when copying from a loose
object with new style header, the code makes sure that the piece
we are going to copy out inflates well and inflate() consumes
the data in full while doing so.

The check to see if the xdelta really apply is quite expensive
as you described, because you would need to have the image of
the base object which can be represented as a delta against
something else.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

72518e9c

pack-objects: fix thinko in revalidate code · 7042dbf7

由 Junio C Hamano 提交于 9月 03, 2006

When revalidating an entry from an existing pack entry->size and
entry->type are not necessarily the size of the final object
when the entry is deltified, but for base objects they must
match.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

7042dbf7

03 9月, 2006 1 次提交

pack-objects: re-validate data we copy from elsewhere. · df6d6101

由 Junio C Hamano 提交于 9月 01, 2006

When reusing data from an existing pack and from a new style
loose objects, we used to just copy it staight into the
resulting pack.  Instead make sure they are not corrupt, but
do so only when we are not streaming to stdout, in which case
the receiving end will do the validation either by unpacking
the stream or by constructing the .idx file.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

df6d6101

24 8月, 2006 1 次提交

Convert memcpy(a,b,20) to hashcpy(a,b). · e702496e

由 Shawn Pearce 提交于 8月 23, 2006

This abstracts away the size of the hash values when copying them
from memory location to memory location, much as the introduction
of hashcmp abstracted away hash value comparsion.

A few call sites were using char* rather than unsigned char* so
I added the cast rather than open hashcpy to be void*.  This is a
reasonable tradeoff as most call sites already use unsigned char*
and the existing hashcmp is also declared to be unsigned char*.

[jc: Splitted the patch to "master" part, to be followed by a
 patch for merge-recursive.c which is not in "master" yet.

 Fixed the cast in the latter hunk to combine-diff.c which was
 wrong in the original.

 Also converted ones left-over in combine-diff.c, diff-lib.c and
 upload-pack.c ]
Signed-off-by: NShawn O. Pearce <spearce@spearce.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

e702496e

18 8月, 2006 1 次提交

Do not use memcmp(sha1_1, sha1_2, 20) with hardcoded length. · a89fccd2

由 David Rientjes 提交于 8月 17, 2006

Introduces global inline:

	hashcmp(const unsigned char *sha1, const unsigned char *sha2)

Uses memcmp for comparison and returns the result based on the length of
the hash name (a future runtime decision).
Acked-by: NAlex Riesen <raa.lkml@gmail.com>
Signed-off-by: NDavid Rientjes <rientjes@google.com>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

a89fccd2

16 8月, 2006 1 次提交

remove unnecessary initializations · 96f1e58f

由 David Rientjes 提交于 8月 15, 2006

[jc: I needed to hand merge the changes to the updated codebase,
 so the result needs to be checked.]
Signed-off-by: NDavid Rientjes <rientjes@google.com>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

96f1e58f

04 8月, 2006 1 次提交

Make git-pack-objects a builtin · 5d4a6003

由 Matthias Kestenholz 提交于 8月 03, 2006

Signed-off-by: NMatthias Kestenholz <matthias@spinlock.ch>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

5d4a6003

26 7月, 2006 1 次提交

pack-objects: reuse deflated data from new-style loose objects. · ceec1361

由 Junio C Hamano 提交于 7月 17, 2006

When packing an object without deltifying, if the data is stored in
a loose object that is encoded with a new style header, copy it without
inflating and deflating.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

ceec1361

24 7月, 2006 1 次提交

pack-objects: check pack.window for default window size · 4812a93a

由 Jeff King 提交于 7月 23, 2006

For some repositories, deltas simply don't make sense. One can disable
them for git-repack by adding --window, but git-push insists on making
the deltas which can be very CPU-intensive for little benefit.
Signed-off-by: NJeff King <peff@peff.net>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

4812a93a

10 7月, 2006 1 次提交

Fix more typos, primarily in the code · 82e5a82f

由 Pavel Roskin 提交于 7月 10, 2006

The only visible change is that git-blame doesn't understand
"--compability" anymore, but it does accept "--compatibility" instead,
which is already documented.
Signed-off-by: NPavel Roskin <proski@gnu.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

82e5a82f

01 7月, 2006 1 次提交

don't load objects needlessly when repacking · 560b25a8

由 Nicolas Pitre 提交于 6月 30, 2006

If no delta is attempted on some objects then it is useless to load them
in memory, neither create any delta index for them.  The best thing to
do is therefore to load and index them only when really needed.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

560b25a8

30 6月, 2006 2 次提交

consider previous pack undeltified object state only when reusing delta data · 8dbbd14e

由 Nicolas Pitre 提交于 6月 29, 2006

Without this there would never be a chance to improve packing for
previously undeltified objects.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

8dbbd14e

Do not try futile object pairs when repacking. · 51d1e83f

由 Linus Torvalds 提交于 6月 29, 2006

In the repacking window, if both objects we are looking at already came
from the same (old) pack-file, don't bother delta'ing them against each
other.

That means that we'll still always check for better deltas for (and
against!) _unpacked_ objects, but assuming incremental repacks, you'll
avoid the delta creation 99% of the time.
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

51d1e83f

21 6月, 2006 1 次提交

upload-pack: prepare for sideband message support. · 363b7817

由 Junio C Hamano 提交于 6月 20, 2006

This does not implement sideband for propagating the status to
the downloader yet, but add code to capture the standard error
output from the pack-objects process in preparation for sending
it off to the client when the protocol extension allows us to do
so.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

363b7817

20 6月, 2006 1 次提交

Remove all void-pointer arithmetic. · 1d7f171c

由 Florian Forster 提交于 6月 18, 2006

ANSI C99 doesn't allow void-pointer arithmetic. This patch fixes this in
various ways. Usually the strategy that required the least changes was used.
Signed-off-by: NFlorian Forster <octo@verplant.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

1d7f171c

06 6月, 2006 1 次提交

pack-objects: improve path grouping heuristics. · ce0bd642

由 Linus Torvalds 提交于 6月 05, 2006

This trivial patch not only simplifies the name hashing, it actually
improves packing for both git and the kernel.

The git archive pack shrinks from 6824090->6622627 bytes (a 3%
improvement), and the kernel pack shrinks from 108756213 to 108219021 (a
mere 0.5% improvement, but still, it's an improvement from making the
hashing much simpler!)

We just create a 32-bit hash, where we "age" previous characters by two
bits, so the last characters in a filename count most. So when we then
compare the hashes in the sort routine, filenames that end the same way
sort the same way.

It takes the subdirectory into account (unless the filename is > 16
characters), but files with the same name within the same subdirectory
will obviously sort closer than files in different subdirectories.

And, incidentally (which is why I tried the hash change in the first
place, of course) builtin-rev-list.c will sort fairly close to rev-list.c.

And no, it's not a "good hash" in the sense of being secure or unique, but
that's not what we're looking for. The whole "hash" thing is misnamed
here. It's not so much a hash as a "sorting number".

[jc: rolled in simplification for computing the sorting number
 computation for thin pack base objects]
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

ce0bd642

31 5月, 2006 1 次提交

tree_entry(): new tree-walking helper function · 4c068a98

由 Linus Torvalds 提交于 5月 30, 2006

This adds a "tree_entry()" function that combines the common operation of
doing a "tree_entry_extract()" + "update_tree_entry()".

It also has a simplified calling convention, designed for simple loops
that traverse over a whole tree: the arguments are pointers to the tree
descriptor and a name_entry structure to fill in, and it returns a boolean
"true" if there was an entry left to be gotten in the tree.

This allows tree traversal with

	struct tree_desc desc;
	struct name_entry entry;

	desc.buf = tree->buffer;
	desc.size = tree->size;
	while (tree_entry(&desc, &entry) {
		... use "entry.{path, sha1, mode, pathlen}" ...
	}

which is not only shorter than writing it out in full, it's hopefully less
error prone too.

[ It's actually a tad faster too - we don't need to recalculate the entry
  pathlength in both extract and update, but need to do it only once.
  Also, some callers can avoid doing a "strlen()" on the result, since
  it's returned as part of the name_entry structure.

  However, by now we're talking just 1% speedup on "git-rev-list --objects
  --all", and we're definitely at the point where tree walking is no
  longer the issue any more. ]

NOTE! Not everybody wants to use this new helper function, since some of
the tree walkers very much on purpose do the descriptor update separately
from the entry extraction. So the "extract + update" sequence still
remains as the core sequence, this is just a simplified interface.

We should probably add a silly two-line inline helper function for
initializing the descriptor from the "struct tree" too, just to cut down
on the noise from that common "desc" initializer.
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

4c068a98

17 5月, 2006 1 次提交

improve depth heuristic for maximum delta size · c3b06a69

由 Nicolas Pitre 提交于 5月 16, 2006

This provides a linear decrement on the penalty related to delta depth
instead of being an 1/x function. With this another 5% reduction is
observed on packs for both the GIT repo and the Linux kernel repo, as
well as fixing a pack size regression in another sample repo I have.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

c3b06a69

16 5月, 2006 3 次提交

Fix pack-index issue on 64-bit platforms a bit more portably. · 1b9bc5a7

由 Junio C Hamano 提交于 5月 15, 2006

Apparently <stdint.h> is not enough for uint32_t on OpenBSD; use
"unsigned int" -- hopefully that would stay 32-bit on every
platform we care about, at least until we update the pack-index
file format.

Our sha1 routines optimized for architectures use uint32_t and
expects '#include <stdint.h>' to be enough, so OpenBSD on arm or
ppc might have similar issues down the road, I dunno.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

1b9bc5a7

pack-object: slightly more efficient · ff45715c

由 Nicolas Pitre 提交于 5月 15, 2006

Avoid creating a delta index for objects with maximum depth since they
are not going to be used as delta base anyway.  This also reduce peak
memory usage slightly as the current object's delta index is not useful
until the next object in the loop is considered for deltification. This
saves a bit more than 1% on CPU usage.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

ff45715c

simple euristic for further free packing improvements · 4e8da195

由 Nicolas Pitre 提交于 5月 15, 2006

Given that the early eviction of objects with maximum delta depth
may exhibit bad packing on its own, why not considering a bias against
deep base objects in try_delta() to mitigate that bad behavior.

This patch adjust the MAX_size allowed for a delta based on the depth of
the base object as well as enabling the early eviction of max depth
objects from the object window. When used separately, those two things
produce slightly better and much worse results respectively. But their
combined effect is a surprising significant packing improvement.

With this really simple patch the GIT repo gets nearly 15% smaller, and
the Linux kernel repo about 5% smaller, with no significantly measurable
CPU usage difference.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

4e8da195

15 5月, 2006 1 次提交
- B
  include header to define uint32_t, necessary on Mac OS X · d9635e9c
  由 Ben Clifford 提交于 5月 14, 2006
```
Signed-off-by: NJunio C Hamano <junkio@cox.net>
```
  d9635e9c
14 5月, 2006 1 次提交

Fix git-pack-objects for 64-bit platforms · 66561f5a

由 Dennis Stosberg 提交于 5月 11, 2006

The offset of an object in the pack is recorded as a 4-byte integer
in the index file.  When reading the offset from the mmap'ed index
in prepare_pack_revindex(), the address is dereferenced as a long*.
This works fine as long as the long type is four bytes wide.  On
NetBSD/sparc64, however, a long is 8 bytes wide and so dereferencing
the offset produces garbage.

[jc: taking suggestion by Linus to use uint32_t]
Signed-off-by: NDennis Stosberg <dennis@stosberg.net>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

66561f5a

06 5月, 2006 1 次提交

pack-object: squelch eye-candy on non-tty · 86118bcb

由 Junio C Hamano 提交于 5月 05, 2006

One of my post-update scripts runs a git-fetch into a separate
repository and sends the results back to me (2>&1); I end up
getting this in the mail:

    Generating pack...
    Done counting 180 objects.
    Result has 131 objects.
    Deltifying 131 objects.
       0% (0/131) done^M   1% (2/131) done^M...

This defaults not to do the progress report when not on a tty.

You could give --progress to force the progress report, but
let's not bother even documenting it nor mentioning it in the
usage string.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

86118bcb

28 4月, 2006 1 次提交

pack-objects: update size heuristucs. · 9a8b6a0a

由 Junio C Hamano 提交于 4月 27, 2006

We used to omit delta base candidates that is much bigger than
the target, but delta size does not grow when we delete more, so
that was not a very good heuristics.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

9a8b6a0a

27 4月, 2006 1 次提交

use delta index data when finding best delta matches · f6c7081a

由 Nicolas Pitre 提交于 4月 26, 2006

This patch allows for computing the delta index for each base object
only once and reuse it when trying to find the best delta match.

This should set the mark and pave the way for possibly better delta
generator algorithms.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

f6c7081a

21 4月, 2006 2 次提交

fix pack-object buffer size · 0dec30b9

由 Nicolas Pitre 提交于 4月 20, 2006

The input line has 40 _chars_ of sha1 and no 20 _bytes_. It should also
account for the space before the pathname, and the terminating \n and \0.
Signed-off-by: NNicolas Pitre <nico@cam.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

0dec30b9

pack-objects: do not stop at object that is "too small" · f527cb8c

由 Junio C Hamano 提交于 4月 20, 2006

Because we sort the delta window by name-hash and then size,
hitting an object that is too small to consider as a delta base
for the current object does not mean we do not have better
candidate in the window beyond it.

Noticed by Shawn Pearce, analyzed by Nico, Linus and me.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

f527cb8c

17 4月, 2006 1 次提交

Try using Geert similarity code in pack-objects. · ca9de6ca

由 Junio C Hamano 提交于 4月 16, 2006

It appears the fingerprinting itself is too expensive to be worth doing
for this purpose.  A failed experiment.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

ca9de6ca

07 4月, 2006 1 次提交

Thin pack generation: optimization. · 5379a5c5

由 Junio C Hamano 提交于 4月 05, 2006

Jens Axboe noticed that recent "git push" has become very slow
since we made --thin transfer the default.

Thin pack generation to push a handful revisions that touch
relatively small number of paths out of huge tree was stupid; it
registered _everything_ from the excluded revisions.  As a
result, "Counting objects" phase was unnecessarily expensive.

This changes the logic to register the blobs and trees from
excluded revisions only for paths we are actually going to send
to the other end.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

5379a5c5

04 4月, 2006 2 次提交

Use blob_, commit_, tag_, and tree_type throughout. · 8e440259

由 Peter Eriksen 提交于 4月 02, 2006

This replaces occurences of "blob", "commit", "tag", and "tree",
where they're really used as type specifiers, which we already
have defined global constants for.
Signed-off-by: NPeter Eriksen <s022018@student.dtu.dk>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

8e440259

safe_fgets() - even more anal fgets() · 687dd75c

由 Junio C Hamano 提交于 4月 03, 2006

This is from Linus -- the previous round forgot to clear error
after EINTR case.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

687dd75c

03 4月, 2006 2 次提交

pack-objects: be incredibly anal about stdio semantics · da93d12b

由 Linus Torvalds 提交于 4月 02, 2006

This is the "letter of the law" version of using fgets() properly in the
face of incredibly broken stdio implementations.  We can work around the
Solaris breakage with SA_RESTART, but in case anybody else is ever that
stupid, here's the "safe" (read: "insanely anal") way to use fgets.

It probably goes without saying that I'm not terribly impressed by
Solaris libc.
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

da93d12b

Fix Solaris stdio signal handling stupidities · fb7a6531

由 Linus Torvalds 提交于 4月 02, 2006

This uses sigaction() to install the SIGALRM handler with SA_RESTART, so
that Solaris stdio doesn't break completely when a signal interrupts a
read.

Thanks to Jason Riedy for confirming the silly Solaris signal behaviour.
Signed-off-by: NLinus Torvalds <torvalds@osdl.org>
Signed-off-by: NJunio C Hamano <junkio@cox.net>

fb7a6531

30 3月, 2006 1 次提交

tree/diff header cleanup. · 1b0c7174

由 Junio C Hamano 提交于 3月 29, 2006

Introduce tree-walk.[ch] and move "struct tree_desc" and
associated functions from various places.

Rename DIFF_FILE_CANON_MODE(mode) macro to canon_mode(mode) and
move it to cache.h.  This macro returns the canonicalized
st_mode value in the host byte order for files, symlinks and
directories -- to be compared with a tree_desc entry.
create_ce_mode(mode) in cache.h is similar but is intended to be
used for index entries (so it does not work for directories) and
returns the value in the network byte order.
Signed-off-by: NJunio C Hamano <junkio@cox.net>

1b0c7174

李少辉-开发者 / git 与 Fork 源项目一致

李少辉-开发者 / git
与 Fork 源项目一致