odh2: second on-disk-header format (insert/update timestamps, larger blobs) - #6102
odh2: second on-disk-header format (insert/update timestamps, larger blobs)#6102markhannum wants to merge 13 commits into
Conversation
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
cdb2jdbc [failed with core dumped]
sp_twofiles_generated [failed with core dumped]
sc_redo_step [failed with core dumped]
sp_queueodh_generated [failed with core dumped]
sp_snapshot_generated [failed with core dumped]
sp [failed with core dumped]
sc_redo_logicalsc_generated [failed with core dumped]
sc_partial_datacopy_logicalsc_generated [failed with core dumped] **quarantined**
scindex_logicalsc_generated
sc_constraints_logicalsc_generated
72c8062 to
a86d8b2
Compare
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_redo_step [failed with core dumped]
sc_partial_datacopy_logicalsc_generated [failed with core dumped] **quarantined**
sc_datacopy_logicalsc_generated [failed with core dumped] **quarantined**
scindex_logicalsc_generated
sc_constraints_logicalsc_generated
sc_inserts_deletes_logicalsc_generated
sc_redo_logicalsc_generated
sc_inserts_logicalsc_generated
sc_blob_update_logicalsc_generated
comdb2sys_queueodh_generated
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
scindex_logicalsc_generated
sc_constraints_logicalsc_generated
sc_inserts_deletes_logicalsc_generated
sc_inserts_logicalsc_generated
sc_blob_update_logicalsc_generated
comdb2sys_queueodh_generated
comdb2sys_pagesize_generated
comdb2sys **quarantined**
noresetgen
sc_tableversion_logicalsc_generated
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
scindex_logicalsc_generated
sc_constraints_logicalsc_generated
sc_inserts_deletes_logicalsc_generated
sc_inserts_logicalsc_generated
sc_blob_update_logicalsc_generated
comdb2sys_queueodh_generated
comdb2sys_pagesize_generated
comdb2sys **quarantined**
sc_tableversion_logicalsc_generated
sc_parallel_logicalsc_generated
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
noresetgen
incoherent_startup
load_cache_dumpmax_generated
diskspace_nollmeta
diskspace_nollmeta_nostripe_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
tunables
cdb2dump
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_truncate [db unavailable at finish]
load_cache_dumpmax_generated
diskspace_nollmeta_nostripe_generated
diskspace_nollmeta
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
tunables
cdb2dump
sc_downgrade [timeout] **quarantined**
a58220a to
9c2a3e6
Compare
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_resume
reco-ddlk-sql **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
cdb2dump
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
ssl_san
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
ssl_set_cmd
ssl_dbname
ssl_prefer
sc_downgrade [timeout] **quarantined**
logical_ops [timeout]
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
truncatesc_offline_generated **quarantined**
sc_resume_logicalsc_generated **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_truncate [db unavailable at finish]
queuedb_rollover_noroll1_generated **quarantined**
queuedb_rollover **quarantined**
noresetgen
reco-ddlk-sql **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_truncate [db unavailable at finish]
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
sc_downgrade [timeout] **quarantined**
reco-ddlk-sql [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_resume_logicalsc_generated **quarantined**
ssl_san
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
ssl_set_cmd
ssl_dbname
ssl_prefer
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
sc_truncate_multiddl_generated [db unavailable at finish] **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
sc_downgrade [timeout] **quarantined**
truncatesc_offline_generated [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
incoherent_slow **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
sc_downgrade [timeout] **quarantined**
reco-ddlk-sql [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
auth_queueodh_generated [setup failed with core dumped]
sc_truncate_multiddl_generated [db unavailable at finish] **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
blob_size_limit
logical_ops
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
ssl_san
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
logical_ops
ssl_set_cmd
ssl_prefer
ssl_dbname
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
logical_ops
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
sc_downgrade [timeout] **quarantined**
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
noresetgen
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
sc_downgrade [timeout] **quarantined**
…length)
Introduce a second on-disk record header format ("odh2") alongside the
existing 7-byte odh1 header. odh2 is a strict superset of odh1: the flags,
csc2vers, and updateid bytes are encoded identically, so odh1 readers and
writers are unaffected and the two formats coexist in the same table.
odh2 (16 bytes) is discriminated by ODH2_FLAG (bit 7 of the flags byte),
which odh1 never sets. It replaces odh1's packed 28-bit length with a clean
32-bit length field and appends two 32-bit unsigned timestamps: insert_secs
and update_secs (seconds since the 1970 epoch, stored unsigned to survive the
2038 signed-time_t rollover). The wider length lifts the odh1 256MB ceiling.
This commit adds only the codec and plumbing primitives:
- struct odh gains insert_secs / update_secs
- ODH2_SIZE, ODH2_FLAG, odh_size_from_flags(); ODH_SIZE_RESERVE bumped to
the max header size so buffers fit either format
- write_odh / read_odh branch on ODH2_FLAG; unpack peeks the flags byte to
learn the real header size before validating
- IPU paths (bdb_update_updateid, bdb_cposition) read into ODH2_SIZE buffers
and rewrite exactly odh_size_from_flags() bytes
- poke_update_secs() helper for stamping update time in place
Nothing sets ODH2_FLAG yet, so no odh2 record is produced; this is the
codec baseline. Also rewrites the odh.c header diagram to document both
layouts and removes the stale note-to-self.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
…rcing)
Make odh2 a real, persisted, per-table attribute and produce odh2 records
at write time. Mirrors the instant_schema_change plumbing end-to-end.
bdb layer:
- bdb_state_type gains an 'odh2' flag; bdb_set_odh2() sets it (gated on
ondisk_header, like inplace_updates)
- init_odh() now sets ODH2_FLAG and stamps insert_secs/update_secs with the
current epoch when the table opts into odh2 OR the database is in genid48
format. The genid48 forcing enforces the project invariant that a genid48
record is never written as odh1 (odh1 relies on the genid carrying the
insert time, which genid48 does not).
config / persistence (clone of instant_schema_change):
- dbtable gains 'odh2'; META_ODH2 llmeta key; get/put_db_odh2 accessors
- set_bdb_option_flags() takes an odh2 argument; all callers updated
- new-table create (init_odh_lrl) seeds from gbl_init_with_odh2 and persists;
restart (init_odh_llmeta) reads it back
- gbl_init_with_odh2 (default OFF -- opt-in) plus init_with_odh2 /
dont_init_with_odh2 lrl tunables
- alter/fastinit preserve the existing table's odh2 setting and re-persist it
(odh2 is cleared if the ondisk header is turned off)
Timestamps stamp correctly for inserts. Two follow-ups remain: (1) preserve
insert_secs across updates (currently an update would reset it to now); this
needs the old record's insert time threaded down the update path. (2) a SQL
'OPTIONS odh2 {on,off}' clause -- for now odh2 is enabled via init_with_odh2,
genid48, or preserved across schema changes. Read-side columns and tests
follow in later commits.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
An update funnels through init_odh, which stamps both timestamps with "now" --
correct for an insert, but it would reset a record's original insert time on
every update. Carry the original insert time forward instead.
- bdb_prepare_put_pack_updateid() gains a preserve_insert_secs argument; when
non-zero and the record is odh2 it overrides insert_secs after init_odh
(update_secs stays "now"). Inserts and non-odh2 tables pass 0.
- ll_dta_upd_int() computes it from the old record, which is in scope at the
single pack site shared by the in-place and new-genid (delete+add) update
paths: the old odh2 record's insert_secs, or bdb_genid_timestamp(oldgenid)
for an odh1 record being upgraded (odh1 => time-based genid, so the genid
carries the insert time). peek_odh2_insert_secs() reads it off the raw,
still-packed old header (the ODH is plaintext; only the payload is
compressed).
- bdb_update_updateid() (the genid-only / blob-optimization path that rewrites
just the header) now refreshes update_secs for odh2 records; insert_secs is
preserved from the record it read.
Add path and fresh inserts are unchanged (preserve_insert_secs = 0 => both
timestamps are "now").
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
… columns)
Surface a row's odh2 timestamps to SQL and make comdb2_rowtimestamp keep
working after a genid48 conversion.
Backend plumbing (mirrors how `ver` flows through the real/shadow merge):
- bdb_cursor_impl and the real-stream berkdb tag (u.rl) gain insert_secs /
update_secs; process_bulk_odh stashes them from the decoded odh
- new berkdb accessor bdb_berkdb_odh2_times (real+odh stream only; returns 0
otherwise so callers fall back to the genid time)
- bdb_cursor_ifn gains insert_secs()/update_secs() methods; the merge sets the
cursor's values from the winning real stream, and 0 for synthetic/shadow
rows
- BtCursor gains insert_secs/update_secs, snapshotted at the two data-record
get_found_data sites; sqlite3BtreeInsertTimestamp/UpdateTimestamp expose them
SQL surface (clones the comdb2_rowtimestamp pseudo-column mechanics):
- new magic columns comdb2_insert_timestamp (iColumn -4) and
comdb2_update_timestamp (iColumn -5); name match, resolver, span/name and
columnType/columnName cases, OP_Rowid P3 codes 3/4, and getRowid branches
- getRowid for all three columns: use the odh2 header time when present, else
fall back to the insert time in the (time-based) genid -- valid because an
odh1 record predates any genid48 conversion
- comdb2_rowtimestamp is no longer gated on comdb2genidcontainstime(): it now
resolves regardless of genid format (odh2 rows read the header, odh1 rows
read the genid), which is what lets it survive the genid48 switch
Known limitation: rows read from the shadow/addcur (own uncommitted writes in a
transaction) path report 0 and fall back to the genid time; only committed
(real-stream) reads carry the odh2 header timestamps. Sufficient for normal
SELECTs; can be extended to the shadow path later.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
comdb2_insert_timestamp / comdb2_update_timestamp / comdb2_rowtimestamp return
the odh times snapshotted on the cursor. For a data-file scan that is the data
record's odh -- correct. But when a row is reached via a secondary index the
cursor sits on the *index entry*, whose odh timestamps are independent of the row
and, crucially, do not advance on a non-key update (the index entry is not
rewritten). So `select comdb2_update_timestamp ... where <indexed-col>=?`
returned the index entry's stale time instead of the row's last-update time.
Read the DATA record's odh times by genid instead:
- bdb_fetch_args_t gains insert_secs/update_secs, populated in bdb_fetch_int_ll
from the decoded odh (alongside ver) on the data-record unpack.
- get_ondisk_timestamps_by_genid() reads them via bdb_fetch_by_rrn_and_genid;
the data record (dtafile 0) is bounded by the row size, so it is a small
fetch, not the (up to 2GB) blob.
- the SQL timestamp accessors call it for a real (non-synthetic) index cursor
and return the data record's times; data cursors and synthetic rows are
unchanged.
Cost: one small data-record fetch per row, only when a timestamp column is
selected on an index scan.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
The genid format lives only on the parent (env) bdb_state; every genid helper normalises `bdb_state = bdb_state->parent` before reading it. init_odh was testing the *table* (child) handle's genid_format, which is never set (always 0 == LLMETA_GENID_ORIGINAL), so the genid48 forcing never fired and a genid48 database would still have written odh1 records -- losing insert timestamps, the exact failure this project exists to prevent. Use genid_contains_time() instead, which normalises to the parent. This also reads better: force odh2 precisely when the genid no longer carries an insert time. The forcing is independent of the per-table odh2 attribute, so the invariant "an odh1 record is never written under genid48" holds for every odh-enabled table regardless of its odh2 setting. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
odh2's 32-bit length field lifts the odh1 256MB ceiling. Enforce the limit
per-table: odh2 tables allow up to INT_MAX (MAXBLOBLENGTH2), everything else
stays at odh1's 28-bit MAXBLOBLENGTH (odh1 physically can't store more).
- MAXBLOBLENGTH2 (INT_MAX) added; MAXBLOBLENGTH kept as the odh1 limit
- max_blob_length_for_table() returns the per-table cap (odh2 when the table
opts in or the genid lacks time, matching init_odh); used by the reject
checks in toblock.c (blob receive) and record.c (check_blob_sizes)
- mem_to_ondisk() enforces the per-table limit on the *uncompressed* sqlite
value, before it is packed/compressed. This is essential: the value is
compressed by the time check_blob_sizes() runs, so that downstream check
sees only the compressed size and cannot catch a >256MB blob that
compresses small. The per-table limit is threaded in via
mem_info.max_blob_length (set by sqlite3MakeRecordForComdb2 from the
cursor's table); a 0 value falls back to the coarse MAXBLOBLENGTH2 bound
for the callers (index keys etc.) that do not supply it.
- bdb_pack() refuses (EINVAL) to emit an odh1 header for a record longer
than 28 bits instead of silently truncating the length -- defence in depth
for any write path that bypasses the db-layer check.
- SQLITE_MAX_LENGTH raised above the stock 1e9 but capped at INT_MAX/2: a
value travels in one newsql message whose length is a signed 32-bit field,
so a value near INT_MAX plus framing could not be sent; INT_MAX/2 leaves
headroom. Truly ~2GB values need a wider wire length (future work).
- comdb2_limits.max_blob_length reports the odh2 storage ceiling
- LZ4 path stores payloads above LZ4_MAX_INPUT_SIZE uncompressed (LZ4 takes
int sizes)
- cdb2api rejects (CDB2ERR_REJECTED, non-retryable) a query whose packed
size exceeds the signed-int wire length instead of overflowing it
The core client paths already allocate dynamically (blob receive mallocs to the
blob length; bdb_unpack mallocs to odh->length), so they scale without change.
Schema change sizes its reconstruct buffer to the table's max, so altering a
table with >256MB blobs works.
Known limitation: the logicalops systable (logical replication reader) keeps a
fixed 256MB scratch buffer; records larger than that now error cleanly instead
of asserting/overflowing. Full 2GB support there (grow-on-demand) is a
follow-up. rowlocks debug-print buffers are likewise bounded at 256MB.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
Nothing today exercises both header formats living in the same table -- the
per-record flag decode in read_odh() that most needs coverage, and the shape
produced in the field when an odh1/time-based database is converted to genid48.
A database is otherwise uniform per creation: genid48 forces odh2 everywhere,
time-based-genid tables are odh1.
Add a general, default-off tunable "randomize_odh2". When enabled, for any
record that would otherwise be odh1, init_odh() flips a coin and writes odh2
instead. The random branch is only reachable when genids are time-based (where
odh1 is legal), so the genid48-never-odh1 invariant is untouched -- genid48
databases still force odh2 as before. With the tunable on and a database
created under time-based genids (init_with_time_based_genids + dont_init_with_odh2),
every table ends up with a random mix of both formats.
The coin is flipped per write, so a record's format can change from one update
to the next. That is intentional and harmless: reads decode each record from
its own flag byte, and it broadens coverage to both the odh1->odh2 and
odh2->odh1 pack transitions. (An earlier draft enforced upgrade-only/never-
downgrade, which required threading the old record's format through the write
path; the plain coin flip is simpler and a better fuzzer.)
No existing test selects the new comdb2_insert_timestamp / comdb2_update_timestamp
columns on a base table, so the randomization is invisible to expected output
and the suite passes as-is. A read-only "odh2_random_upgrades" counter (visible
via the comdb2_tunables system table) lets a run confirm coexistence occurred.
To run the suite as a coexistence fuzzer, point CUSTOMLRLPATH at a fragment
containing:
init_with_time_based_genids
dont_init_with_odh2
randomize_odh2 1
and run `make -kjN CUSTOMLRLPATH=<frag>` from tests/. Normal operation is
unaffected (tunable defaults off); not for production use.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
Several paths still allocated scratch buffers sized for the odh1 header or the odh1 256MB ceiling and would truncate, reject, or overrun an odh2 record even though the write and recovery paths already handle it: - schemachange/sc_records.c: the logical-redo thread (live_sc_logical_redo_thd and the live_sc_redo_add/update helpers) sized its four packed-record scratch buffers -- and the reconstruct capacity hints handed to bdb_reconstruct_add / bdb_reconstruct_update / bdb_reconstruct_inplace_update -- as lrl + ODH_SIZE (the 7-byte odh1 header). An odh2 record carries a 16-byte header, so a full-size reconstructed record overran the buffer by up to 9 bytes, corrupting the heap; the damage only surfaced later as an abort inside mspace_free during convert_record_data_cleanup (seen by the sc_redo_step test). Size these to lrl + ODH_SIZE_RESERVE (the maximum header, == ODH2_SIZE) so either header fits, matching the blob path already fixed here. The delete path already bounds its buffer by the record's exact logged length and is unchanged. - sqlite/ext/comdb2/logicalops.c: the comdb2_logicalops systable cursor kept fixed 256MB packed/packedprev buffers and errored on bigger records. Grow them on demand to the table's max_blob_length_for_table() (256MB for odh1, up to ~2GB for odh2), tracked by new packedcap/packedprevcap fields and freed in logicalopsClose(). These use plain realloc()/free() rather than sqlite3's allocator, which caps a single allocation below the odh2 ~2GB ceiling and would otherwise fail to allocate a large odh2 record. odh1 cursors are unchanged; only odh2 tables allocate more, lazily, per active cursor. The delete path sizes to the deleted record's exact logged length. Also normalise the log rectype (normalize_rectype) before the sanity assert in unpack_logical_record -- the on-disk rectype can carry the rowlocks/utxnid variant offset, which otherwise trips the assert. - bdb/rowlocks.c: the four db_printlog-only buffers (case DB_TXN_PRINT, reached only by the standalone comdb2_db_printlog tool -- not the recovery replay path) are sized to MAXBLOBLENGTH2 so a large odh2 record isn't truncated in the log dump. This only ever allocates inside that debug tool. Known ceiling left in place: the newsql wire header length is a signed int, so a single ~2GB message still can't round-trip; the SQL-insertable limit is SQLITE_MAX_LENGTH (INT_MAX/2) and MAXBLOBLENGTH2 is INT_MAX by design. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
New test odh2_timestamps.test: on an odh2 table (init_with_odh2, time-based genids, randomizer off) it checks that comdb2_insert_timestamp / comdb2_update_timestamp / comdb2_rowtimestamp are real datetimes, that all three match on a fresh insert, and that after an update the insert timestamp is preserved while the update timestamp advances. Assertions reduce to booleans/typeof/server-side comparisons so no raw datetime is diffed as text. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
New test odh2_bigblob.test: mirrors blob_size_limit.test but on an init_with_odh2 table, driving comdb2_blobtest at sizes through and beyond 256MB (up to 512MB). Where the odh1 test expects failures/zero length at >=256MB, the odh2 table stores them successfully (full length), and verify decodes every large record. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
New test odh2_coexist.test: creates a table under time-based genids with odh2 off (odh1 rows), converts the db to genid48 with "put genid48 enable", then inserts new rows and updates some original ones so the table holds both odh1 and odh2 records at once. verify drives read_odh() over the mixed table, row counts and updated values are checked, and every row reports a non-null rowtimestamp (genid-derived for odh1, header-derived for odh2). This mirrors the real migration the feature exists for. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
New test odh2_randomize.test: with randomize_odh2 on under time-based genids, inserts many rows so the per-record coin produces both formats, asserts the odh2_random_upgrades counter (via comdb2_tunables) is greater than zero, then verifies the mixed table and checks all rowtimestamps are non-null. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Signed-off-by: Mark Hannum <mhannum@bloomberg.net>
roborivers
left a comment
There was a problem hiding this comment.
Cbuild submission: Error ⚠.
Regression testing: Success ✓.
The first 10 failing tests are:
cldeadlock [setup failed]
sc_resume_logicalsc_generated **quarantined**
load_cache_dumpmax_generated
consumer_non_atomic_default_consumer_generated **quarantined**
sc_downgrade [timeout] **quarantined**
Adds "odh2", a second on-disk-header format that coexists with the legacy odh1 header (told apart per-record by bit 7 of the flags byte). odh2 records store the row's insert-time and last-update-time as unsigned 32-bit epoch seconds (valid past the 2038 rollover) and use a clean 32-bit length, lifting the odh1 256MB blob ceiling toward ~2GB.
Motivation: converting a database to genid48. Time-based genids encode the insert epoch in the genid; genid48 does not. odh2 keeps the times in the header so the conversion doesn't lose them. Invariant preserved: a genid48 record is never written as odh1 (genid48 => odh2; odh2 does not imply genid48).
Highlights:
Tests: tests/odh2_timestamps, tests/odh2_bigblob, tests/odh2_coexist, tests/odh2_randomize (all pass).
Known follow-ups: no SQL
OPTIONS odh2 {on,off}grammar yet; the newsql wire header length is a signed int so a single ~2GB message can't round-trip (MAXBLOBLENGTH2 = INT_MAX is chosen to stay within it).🤖 Generated with Claude Code