提交 · 9d5bd411be284c36d2758cf7d1ec731a783d029e · kvdb / rocksdb

06 1月, 2015 5 次提交

benchmark.sh won't run through all tests properly if one specifies wal_dir to... · 9d5bd411

由 Leonidas Galanis 提交于 1月 05, 2015

benchmark.sh won't run through all tests properly if one specifies wal_dir to be different than db directory.

Summary:
A command line like this to run all the tests:
source benchmark.config.sh && nohup ./benchmark.sh 'bulkload,fillseq,overwrite,filluniquerandom,readrandom,readwhilewriting'
where
benchmark.config.sh is:
export DB_DIR=/data/mysql/rocksdata
export WAL_DIR=/txlogs/rockswal
export OUTPUT_DIR=/root/rocks_benchmarking/output

Will fail for the tests that need a new DB .

Also 1) set disable_data_sync=0 and 2) add debug mode to run through all the tests more quickly

Test Plan: run ./benchmark.sh 'debug,bulkload,fillseq,overwrite,filluniquerandom,readrandom,readwhilewriting' and verify that there are no complaints about WAL dir not being empty.

Reviewers: sdong, yhchiang, rven, igor

Reviewed By: igor

Subscribers: dhruba

Differential Revision: https://reviews.facebook.net/D30909

9d5bd411

Deprecating skip_log_error_on_recovery · 62ad0a9b

由 Igor Canadi 提交于 1月 05, 2015

Summary:
Since https://reviews.facebook.net/D16119, we ignore partial tailing writes. Because of that, we no longer need skip_log_error_on_recovery.

The documentation says "Skip log corruption error on recovery (If client is ok with losing most recent changes)", while the option actually ignores any corruption of the WAL (not only just the most recent changes). This is very dangerous and can lead to DB inconsistencies. This was originally set up to ignore partial tailing writes, which we now do automatically (after D16119). I have digged up old task t2416297 which confirms my findings.

Test Plan: There was actually no tests that verified correct behavior of skip_log_error_on_recovery.

Reviewers: yhchiang, rven, dhruba, sdong

Reviewed By: sdong

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30603

62ad0a9b

I

Fix corruption_test -- if status is not OK, return status -- during recovery · fa0b126c
由 Igor Canadi 提交于 1月 05, 2015

fa0b126c

Fail DB::Open() on WAL corruption · d7b4bb62

由 Igor Canadi 提交于 1月 05, 2015

Summary:
This is a serious bug. If paranod_check == true and WAL is corrupted, we don't fail DB::Open(). I tried going into history and it seems we've been doing this for a long long time.

I found this when investigating t5852041.

Test Plan: Added unit test to verify correct behavior.

Reviewers: yhchiang, rven, sdong

Reviewed By: sdong

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30597

d7b4bb62

I
Merge pull request #449 from robertabcd/improve-backupable · 9619081d
由 Igor Canadi 提交于 1月 05, 2015
```
Improve backupable db performance on loading BackupMeta
```
9619081d

05 1月, 2015 1 次提交
- R
  
  Fix errors when using -Wshorten-64-to-32. · 49376bfe
  由 Robert 提交于 1月 05, 2015
  
  49376bfe
04 1月, 2015 2 次提交
- R
  
  Do not issue extra GetFileSize() calls when loading BackupMeta. · a8c5564a
  由 Robert 提交于 1月 04, 2015
  
  a8c5564a
- R
  Improve performance when loading BackupMeta. · caa1fd0e
  由 Robert 提交于 1月 04, 2015
```
* Use strtoul() and strtoull() instead of sscanf().
  glibc's sscanf() will do a implicit strlen().

* Move implicit construction of Slice("crc32 ") out of loop.
```
  caa1fd0e
31 12月, 2014 2 次提交

Fix CLANG build for db_bench · e9ca3581

由 sdong 提交于 12月 30, 2014

Summary: CLANG was broken for a recent change in db_ench. Fix it.

Test Plan: Build db_bench using CLANG.

Reviewers: rven, igor, yhchiang

Reviewed By: yhchiang

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30801

e9ca3581

Add structures for exposing thread events and operations. · bf287b76

由 Yueh-Hsuan Chiang 提交于 12月 30, 2014

Summary:
Add structures for exposing events and operations.  Event describes
high-level action about a thread such as doing compaciton or
doing flush, while an operation describes lower-level action
of a thread such as reading / writing a SST table, waiting for
mutex.  Events and operations are designed to be independent.
One thread would typically involve in one event and one operation.

Code instrument will be in a separate diff.

Test Plan:
Add unit-tests in thread_list_test
make dbg -j32
./thread_list_test
export ROCKSDB_TESTS=ThreadList
./db_test

Reviewers: ljin, igor, sdong

Reviewed By: sdong

Subscribers: rven, jonahcohen, dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D29781

bf287b76

25 12月, 2014 1 次提交

db_bench --num_hot_column_families to be default off · a801c1fb

由 sdong 提交于 12月 24, 2014

Summary: Having --num_hot_column_families default on fails some existing regression tests. By default turn it off

Test Plan: Run db_bench to make sure it is default off.

Reviewers: yhchiang, rven, igor

Reviewed By: igor

Subscribers: leveldb, dhruba

Differential Revision: https://reviews.facebook.net/D30705

a801c1fb

24 12月, 2014 6 次提交

Dump routine to BlockBasedTableReader (valgrind) · 2067058a

由 Manish Patil 提交于 11月 30, 2014

Summary: Fixed valgrind issue

Test Plan: valgrind check done

Reviewers: rven, sdong

Reviewed By: sdong

Subscribers: sdong, dhruba

Differential Revision: https://reviews.facebook.net/D30699

2067058a

db_bench to add an option as number of hot column families to add to · ddc81440

由 sdong 提交于 12月 22, 2014

Summary:
Add option --num_hot_column_families in db_bench. If it is set, write options will first write to that number of column families, and then move on to next set of hot column families. The working set of column families can be smaller than total number of CFs.

It is to test how RocksDB can handle cold column families

Test Plan: Run db_bench with --num_hot_column_families set and not set.

Reviewers: yhchiang, rven, igor

Reviewed By: igor

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30663

ddc81440

Y

Fixed a compile error in db/db_impl.cc on ROCKSDB_LITE · a944afd3
由 Yueh-Hsuan Chiang 提交于 12月 23, 2014

a944afd3

Dump routine to BlockBasedTableReader · 7ea7bdf0

由 Manish Patil 提交于 12月 23, 2014

Summary: Added necessary routines for dumping block based SST with block filter

Test Plan: Added "raw" mode to utility sst_dump

Reviewers: sdong, rven

Reviewed By: rven

Subscribers: dhruba

Differential Revision: https://reviews.facebook.net/D29679

7ea7bdf0

I

Clean up compile for c_simple_example · ae508df9
由 Igor Canadi 提交于 12月 23, 2014

ae508df9
I

Fix compile of compact_file_example · b6230096
由 Igor Canadi 提交于 12月 23, 2014

b6230096

23 12月, 2014 6 次提交

I
Merge pull request #444 from adamretter/java-api-fix · ded26605
由 Igor Canadi 提交于 12月 23, 2014
```
Fix the Java API build on Mac OS X
```
ded26605
A

Fix the build on Mac OS X · 98490bcc
由 Adam Retter 提交于 12月 23, 2014

98490bcc
Y
Merge pull request #443 from behanna/master · 4d997297
由 Yueh-Hsuan Chiang 提交于 12月 22, 2014
```
Fix the build with -DNDEBUG.
```
4d997297

add support for nested BlockBasedTableOptions in config string · 5045c439

由 Lei Jin 提交于 12月 22, 2014

Summary:
Add support to allow nested config for block-based table factory. The format looks like this:

"write_buffer_size=1024;block_based_table_factory={block_size=4k};max_write_buffer_num=2"

Test Plan: unit test

Reviewers: yhchiang, rven, igor, ljin, jonahcohen

Reviewed By: jonahcohen

Subscribers: jonahcohen, dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D29223

5045c439

Fix the build with -DNDEBUG. · d232cb15

由 Chris BeHanna 提交于 12月 22, 2014

Dike out the body of VerifyCompactionResult.  With assert() compiled out, the
loop index variable in the inner loop was unused, breaking the build when
-Werror is enabled.

d232cb15

Move GetThreadList() feature under Env. · 45bab305

由 Yueh-Hsuan Chiang 提交于 12月 22, 2014

Summary:
GetThreadList() feature depends on the thread creation and destruction, which is currently handled under Env.
This patch moves GetThreadList() feature under Env to better manage the dependency of GetThreadList() feature
on thread creation and destruction.

Renamed ThreadStatusImpl to ThreadStatusUpdater.  Add ThreadStatusUtil, which is a static class contains
utility functions for ThreadStatusUpdater.

Test Plan: run db_test, thread_list_test and db_bench and verify the life cycle of Env and ThreadStatusUpdater is properly managed.

Reviewers: igor, sdong

Reviewed By: sdong

Subscribers: ljin, dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30057

45bab305

22 12月, 2014 4 次提交

Only execute flush from compaction if max_background_flushes = 0 · 4fd26f28

由 Igor Canadi 提交于 12月 22, 2014

Summary: As title. We shouldn't need to execute flush from compaction if there are dedicated threads doing flushes.

Test Plan: make check

Reviewers: rven, yhchiang, sdong

Reviewed By: sdong

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30579

4fd26f28

Speed up FindObsoleteFiles() · 0acc7388

由 Igor Canadi 提交于 12月 22, 2014

Summary:
There are two versions of FindObsoleteFiles():
* full scan, which is executed every 6 hours (and it's terribly slow)
* no full scan, which is executed every time a background process finishes and iterator is deleted

This diff is optimizing the second case (no full scan). Here's what we do before the diff:
* Get the list of obsolete files (files with ref==0). Some files in obsolete_files set might actually be live.
* Get the list of live files to avoid deleting files that are live.
* Delete files that are in obsolete_files and not in live_files.

After this diff:
* The only files with ref==0 that are still live are files that have been part of move compaction. Don't include moved files in obsolete_files.
* Get the list of obsolete files (which exclude moved files).
* No need to get the list of live files, since all files in obsolete_files need to be deleted.

I'll post the benchmark results, but you can get the feel of it here: https://reviews.facebook.net/D30123

This depends on D30123.

P.S. We should do full scan only in failure scenarios, not every 6 hours. I'll do this in a follow-up diff.

Test Plan:
One new unit test. Made sure that unit test fails if we don't have a `if (!f->moved)` safeguard in ~Version.

make check

Big number of compactions and flushes:

  ./db_stress --threads=30 --ops_per_thread=20000000 --max_key=10000 --column_families=20 --clear_column_family_one_in=10000000 --verify_before_write=0  --reopen=15 --max_background_compactions=10 --max_background_flushes=10 --db=/fast-rocksdb-tmp/db_stress --prefixpercent=0 --iterpercent=0 --writepercent=75 --db_write_buffer_size=2000000

Reviewers: yhchiang, rven, sdong

Reviewed By: sdong

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30249

0acc7388

I
Merge pull request #442 from alabid/alabid/fix-example-typo · d8c4ce6b
由 Igor Canadi 提交于 12月 22, 2014
```
fix really trivial typo in column families example
```
d8c4ce6b
A

fix really trivial typo · 949bd71f
由 alabid 提交于 12月 22, 2014

949bd71f

21 12月, 2014 1 次提交

Fix a SIGSEGV in BackgroundFlush · f8999fcf

由 Igor Canadi 提交于 12月 21, 2014

Summary:
This one wasn't easy to find :)

What happens is we go through all cfds on flush_queue_ and find no cfds to flush, *but* the cfd is set to the last CF we looped through and following code assumes we want it flushed.

BTW @sdong do you think we should also make BackgroundFlush() only check a single cfd for flushing instead of doing this `while (!flush_queue_.empty())`?

Test Plan: regression test no longer fails

Reviewers: sdong, rven, yhchiang

Reviewed By: yhchiang

Subscribers: sdong, dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30591

f8999fcf

20 12月, 2014 3 次提交

MultiGet for DBWithTTL · ade4034a

由 Igor Canadi 提交于 12月 20, 2014

Summary: This is a feature request from rocksdb's user. I didn't even realize we don't support multigets on TTL DB :)

Test Plan: added a unit test

Reviewers: yhchiang, rven, sdong

Reviewed By: sdong

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30561

ade4034a

Rewritten system for scheduling background work · fdb6be4e

由 Igor Canadi 提交于 12月 19, 2014

Summary:
When scaling to higher number of column families, the worst bottleneck was MaybeScheduleFlushOrCompaction(), which did a for loop over all column families while holding a mutex. This patch addresses the issue.

The approach is similar to our earlier efforts: instead of a pull-model, where we do something for every column family, we can do a push-based model -- when we detect that column family is ready to be flushed/compacted, we add it to the flush_queue_/compaction_queue_. That way we don't need to loop over every column family in MaybeScheduleFlushOrCompaction.

Here are the performance results:

Command:

    ./db_bench --write_buffer_size=268435456 --db_write_buffer_size=268435456 --db=/fast-rocksdb-tmp/rocks_lots_of_cf --use_existing_db=0 --open_files=55000 --statistics=1 --histogram=1 --disable_data_sync=1 --max_write_buffer_number=2 --sync=0 --benchmarks=fillrandom --threads=16 --num_column_families=5000  --disable_wal=1 --max_background_flushes=16 --max_background_compactions=16 --level0_file_num_compaction_trigger=2 --level0_slowdown_writes_trigger=2 --level0_stop_writes_trigger=3 --hard_rate_limit=1 --num=33333333 --writes=33333333

Before the patch:

     fillrandom   :      26.950 micros/op 37105 ops/sec;    4.1 MB/s

After the patch:

      fillrandom   :      17.404 micros/op 57456 ops/sec;    6.4 MB/s

Next bottleneck is VersionSet::AddLiveFiles, which is painfully slow when we have a lot of files. This is coming in the next patch, but when I removed that code, here's what I got:

      fillrandom   :       7.590 micros/op 131758 ops/sec;   14.6 MB/s

Test Plan:
make check

two stress tests:

Big number of compactions and flushes:

    ./db_stress --threads=30 --ops_per_thread=20000000 --max_key=10000 --column_families=20 --clear_column_family_one_in=10000000 --verify_before_write=0  --reopen=15 --max_background_compactions=10 --max_background_flushes=10 --db=/fast-rocksdb-tmp/db_stress --prefixpercent=0 --iterpercent=0 --writepercent=75 --db_write_buffer_size=2000000

max_background_flushes=0, to verify that this case also works correctly

    ./db_stress --threads=30 --ops_per_thread=2000000 --max_key=10000 --column_families=20 --clear_column_family_one_in=10000000 --verify_before_write=0  --reopen=3 --max_background_compactions=3 --max_background_flushes=0 --db=/fast-rocksdb-tmp/db_stress --prefixpercent=0 --iterpercent=0 --writepercent=75 --db_write_buffer_size=2000000

Reviewers: ljin, rven, yhchiang, sdong

Reviewed By: sdong

Subscribers: dhruba, leveldb

Differential Revision: https://reviews.facebook.net/D30123

fdb6be4e

I

Remove -mtune=native because it's redundant · a3001b1d
由 Igor Canadi 提交于 12月 19, 2014

a3001b1d

19 12月, 2014 8 次提交
- Y
  Merge pull request #437 from fyrz/RocksJava-SliceTests-Fixes · e27c8452
  由 Yueh-Hsuan Chiang 提交于 12月 18, 2014
```
[RocksJava] Slice / DirectSlice improvements
```
  e27c8452
- F
  
  [RocksJava] Incorporated changes D30081 · 1fed1282
  由 fyrz 提交于 12月 18, 2014
  
  1fed1282
- F
  
  [RocksJava] JavaDoc correction · 5b9ceef0
  由 fyrz 提交于 12月 18, 2014
  
  5b9ceef0
- F
  
  [RocksJava] Incorporated changes D30081 · 5fbba60b
  由 fyrz 提交于 12月 18, 2014
  
  5fbba60b
- F
  
  [RocksJava] Incorporate additions for D30081 · b0230d7e
  由 fyrz 提交于 12月 14, 2014
  
  b0230d7e
- F
  [RocksJava] Slice / DirectSlice improvements · b015ed0c
  由 fyrz 提交于 12月 10, 2014
```
Summary:
- AssertionError when initialized with Non-Direct Buffer
- Tests + coverage for DirectSlice
- Slice sigsegv fixes when initializing from String and byte arrays
- Slice Tests

Test Plan: Run tests without source modifications.

Reviewers: yhchiang, adamretter, ankgup87

Subscribers: dhruba

Differential Revision: https://reviews.facebook.net/D30081
```
  b015ed0c
- Y
  Merge pull request #430 from adamretter/increase-parallelism · 4d422db0
  由 Yueh-Hsuan Chiang 提交于 12月 18, 2014
```
Added setIncreaseParallelism() to Java API Options
```
  4d422db0
- Y
  Merge pull request #411 from fyrz/RocksJava-RangeCompaction · 04c4e496
  由 Yueh-Hsuan Chiang 提交于 12月 18, 2014
```
[RocksJava] Range compaction
```
  04c4e496
18 12月, 2014 1 次提交
- I
  Merge pull request #427 from haneefmubarak/c-examples · 62d19b7b
  由 Igor Canadi 提交于 12月 18, 2014
```
C example
```
  62d19b7b

kvdb / rocksdb 12 个月 前同步成功

kvdb / rocksdb
12 个月前同步成功