Date: 2026-09-14
Stream: //p4-sdp/dev
Job: SDP-1400 (Bug)
Fix changelist: 33774
A customer-reported log showed live_checkpoint.sh performing two full
offline_db rebuilds back-to-back on an edge server (each replay taking
~4.2 hours), roughly doubling the total run time of the live checkpoint
process.
checkpoint() in Server/Unix/p4/common/bin/backup_functions.sh has two
branches:
SERVERID == P4MASTER_ID): dumps a live
checkpoint from P4ROOT (-jc/-jcp/-jcpm). Does not touch
offline_db at all.else branch): calls
request_replica_checkpoint.sh ... -now to ask the commit server for a
checkpoint, waits (polling every 5s) for the corresponding *.md5 file to
appear -- which, per the customer log, took ~5 hours -- then re-logs in
(ticket may have expired during the wait) and, prior to this fix, called
recreate_offline_db_files directly.Separately, live_checkpoint.sh's main flow always calls checkpoint() and
then, unless -skip_odb was passed, unconditionally calls
recreate_offline_db_files again:
checkpoint
if [[ "$RebuildOfflineDB" -eq 1 ]]; then
recreate_offline_db_files
else
log "Skipping rebuild of offline_db due to '-skip_odb'."
fi
Because checkpoint() already did the rebuild internally on edge/replica
servers, this second unconditional call redundantly rebuilt offline_db a
second time from the exact same checkpoint file, which is exactly what the
customer log shows: two consecutive Recovering from serial checkpoint file: .../p4_1.edge_kef.ckp.13865.gz / -jr replays, each ~4.2 hours,
sandwiched around the "Seeking highest journal counter..." lookup being
run twice.
The double rebuild only ever occurred on edge/replica/standby servers,
never on the commit/master server. On the commit server
(isCommitServer == 1 branch), checkpoint() only performs the live
checkpoint dump from P4ROOT -- it never calls recreate_offline_db_files
internally. The single call to recreate_offline_db_files made afterward
by live_checkpoint.sh's main flow (gated by RebuildOfflineDB/
-skip_odb) was, and remains, the only rebuild performed on a commit
server. This code path was not touched by the fix.
The fix is scoped entirely to the else (non-commit-server) branch of
checkpoint(), so commit server behavior is provably unchanged.
Removed the internal recreate_offline_db_files call from the edge/
replica branch of checkpoint(), while keeping the re-p4login (needed
because the wait for the MD5 file can span hours, long enough for the
ticket to expire). The single rebuild is now performed exactly once, by
live_checkpoint.sh's existing -skip_odb-gated call, for both commit
and non-commit servers alike.
Server/Unix/p4/common/bin/backup_functions.sh -- the actual fix;
removed duplicate recreate_offline_db_files call from checkpoint().test/bsw/test_script-live-checkpoint-off-commit.sh -- added a
regression check: counts occurrences of Recovering from .* checkpoint
in the edge/standby checkpoint.log after a live checkpoint run and
fails if more than one rebuild occurred.doc/ReleaseNotes.adoc -- added a SDP-1400 change-history entry under
"SDP 2026.1 Patch 1".bash -n on all modified shell scripts -- clean.shellcheck on backup_functions.sh -- no new warnings introduced.checkpoint() in the repo; confirmed
live_checkpoint.sh is the only caller, so no other script relies on
checkpoint() performing its own offline_db rebuild.p4-04/p4-02); the added assertions are ready for the
next BSW lab run.# Session Log: SDP-1400 -- live_checkpoint.sh double offline_db rebuild Date: 2026-09-14 Stream: `//p4-sdp/dev` Job: [SDP-1400](https://perforce.atlassian.net/browse/SDP-1400) (Bug) Fix changelist: 33774 ## Summary A customer-reported log showed `live_checkpoint.sh` performing two full offline_db rebuilds back-to-back on an edge server (each replay taking ~4.2 hours), roughly doubling the total run time of the live checkpoint process. ## Root cause `checkpoint()` in `Server/Unix/p4/common/bin/backup_functions.sh` has two branches: * **Commit/master server** (`SERVERID == P4MASTER_ID`): dumps a live checkpoint from `P4ROOT` (`-jc`/`-jcp`/`-jcpm`). Does **not** touch `offline_db` at all. * **Edge/replica/standby server** (the `else` branch): calls `request_replica_checkpoint.sh ... -now` to ask the commit server for a checkpoint, waits (polling every 5s) for the corresponding `*.md5` file to appear -- which, per the customer log, took ~5 hours -- then re-logs in (ticket may have expired during the wait) and, prior to this fix, called `recreate_offline_db_files` directly. Separately, `live_checkpoint.sh`'s main flow always calls `checkpoint()` and then, unless `-skip_odb` was passed, unconditionally calls `recreate_offline_db_files` again: ```bash checkpoint if [[ "$RebuildOfflineDB" -eq 1 ]]; then recreate_offline_db_files else log "Skipping rebuild of offline_db due to '-skip_odb'." fi ``` Because `checkpoint()` already did the rebuild internally on edge/replica servers, this second unconditional call redundantly rebuilt `offline_db` a second time from the exact same checkpoint file, which is exactly what the customer log shows: two consecutive `Recovering from serial checkpoint file: .../p4_1.edge_kef.ckp.13865.gz` / `-jr` replays, each ~4.2 hours, sandwiched around the "Seeking highest journal counter..." lookup being run twice. ## Scope confirmation: commit server vs. edge/replica **The double rebuild only ever occurred on edge/replica/standby servers, never on the commit/master server.** On the commit server (`isCommitServer == 1` branch), `checkpoint()` only performs the live checkpoint dump from `P4ROOT` -- it never calls `recreate_offline_db_files` internally. The single call to `recreate_offline_db_files` made afterward by `live_checkpoint.sh`'s main flow (gated by `RebuildOfflineDB`/ `-skip_odb`) was, and remains, the *only* rebuild performed on a commit server. This code path was not touched by the fix. The fix is scoped entirely to the `else` (non-commit-server) branch of `checkpoint()`, so commit server behavior is provably unchanged. ## Fix Removed the internal `recreate_offline_db_files` call from the edge/ replica branch of `checkpoint()`, while keeping the re-`p4login` (needed because the wait for the MD5 file can span hours, long enough for the ticket to expire). The single rebuild is now performed exactly once, by `live_checkpoint.sh`'s existing `-skip_odb`-gated call, for both commit and non-commit servers alike. ## Files changed (changelist 33774) * `Server/Unix/p4/common/bin/backup_functions.sh` -- the actual fix; removed duplicate `recreate_offline_db_files` call from `checkpoint()`. * `test/bsw/test_script-live-checkpoint-off-commit.sh` -- added a regression check: counts occurrences of `Recovering from .* checkpoint` in the edge/standby `checkpoint.log` after a live checkpoint run and fails if more than one rebuild occurred. * `doc/ReleaseNotes.adoc` -- added a SDP-1400 change-history entry under "SDP 2026.1 Patch 1". ## Verification performed * `bash -n` on all modified shell scripts -- clean. * `shellcheck` on `backup_functions.sh` -- no new warnings introduced. * Traced all call sites of `checkpoint()` in the repo; confirmed `live_checkpoint.sh` is the only caller, so no other script relies on `checkpoint()` performing its own offline_db rebuild. * Did not run the BSW test itself in this session (requires live lab hosts `p4-04`/`p4-02`); the added assertions are ready for the next BSW lab run.
| # | Change | User | Description | Committed | |
|---|---|---|---|---|---|
| #1 | 33778 | C. Thomas Tyler |
SDP-1400: Add session log to ai_dev_support (dev-process artifact, not shipped) Session log documenting investigation and fix for SDP-1400 (live_checkpoint.sh double offline_db rebuild on edge/replica servers). Kept separate from the product fix changelist per ai_dev_support convention. |