# Session Log: SDP-1400 -- live_checkpoint.sh double offline_db rebuild Date: 2026-09-14 Stream: `//p4-sdp/dev` Job: [SDP-1400](https://perforce.atlassian.net/browse/SDP-1400) (Bug) Fix changelist: 33774 ## Summary A customer-reported log showed `live_checkpoint.sh` performing two full offline_db rebuilds back-to-back on an edge server (each replay taking ~4.2 hours), roughly doubling the total run time of the live checkpoint process. ## Root cause `checkpoint()` in `Server/Unix/p4/common/bin/backup_functions.sh` has two branches: * **Commit/master server** (`SERVERID == P4MASTER_ID`): dumps a live checkpoint from `P4ROOT` (`-jc`/`-jcp`/`-jcpm`). Does **not** touch `offline_db` at all. * **Edge/replica/standby server** (the `else` branch): calls `request_replica_checkpoint.sh ... -now` to ask the commit server for a checkpoint, waits (polling every 5s) for the corresponding `*.md5` file to appear -- which, per the customer log, took ~5 hours -- then re-logs in (ticket may have expired during the wait) and, prior to this fix, called `recreate_offline_db_files` directly. Separately, `live_checkpoint.sh`'s main flow always calls `checkpoint()` and then, unless `-skip_odb` was passed, unconditionally calls `recreate_offline_db_files` again: ```bash checkpoint if [[ "$RebuildOfflineDB" -eq 1 ]]; then recreate_offline_db_files else log "Skipping rebuild of offline_db due to '-skip_odb'." fi ``` Because `checkpoint()` already did the rebuild internally on edge/replica servers, this second unconditional call redundantly rebuilt `offline_db` a second time from the exact same checkpoint file, which is exactly what the customer log shows: two consecutive `Recovering from serial checkpoint file: .../p4_1.edge_kef.ckp.13865.gz` / `-jr` replays, each ~4.2 hours, sandwiched around the "Seeking highest journal counter..." lookup being run twice. ## Scope confirmation: commit server vs. edge/replica **The double rebuild only ever occurred on edge/replica/standby servers, never on the commit/master server.** On the commit server (`isCommitServer == 1` branch), `checkpoint()` only performs the live checkpoint dump from `P4ROOT` -- it never calls `recreate_offline_db_files` internally. The single call to `recreate_offline_db_files` made afterward by `live_checkpoint.sh`'s main flow (gated by `RebuildOfflineDB`/ `-skip_odb`) was, and remains, the *only* rebuild performed on a commit server. This code path was not touched by the fix. The fix is scoped entirely to the `else` (non-commit-server) branch of `checkpoint()`, so commit server behavior is provably unchanged. ## Fix Removed the internal `recreate_offline_db_files` call from the edge/ replica branch of `checkpoint()`, while keeping the re-`p4login` (needed because the wait for the MD5 file can span hours, long enough for the ticket to expire). The single rebuild is now performed exactly once, by `live_checkpoint.sh`'s existing `-skip_odb`-gated call, for both commit and non-commit servers alike. ## Files changed (changelist 33774) * `Server/Unix/p4/common/bin/backup_functions.sh` -- the actual fix; removed duplicate `recreate_offline_db_files` call from `checkpoint()`. * `test/bsw/test_script-live-checkpoint-off-commit.sh` -- added a regression check: counts occurrences of `Recovering from .* checkpoint` in the edge/standby `checkpoint.log` after a live checkpoint run and fails if more than one rebuild occurred. * `doc/ReleaseNotes.adoc` -- added a SDP-1400 change-history entry under "SDP 2026.1 Patch 1". ## Verification performed * `bash -n` on all modified shell scripts -- clean. * `shellcheck` on `backup_functions.sh` -- no new warnings introduced. * Traced all call sites of `checkpoint()` in the repo; confirmed `live_checkpoint.sh` is the only caller, so no other script relies on `checkpoint()` performing its own offline_db rebuild. * Did not run the BSW test itself in this session (requires live lab hosts `p4-04`/`p4-02`); the added assertions are ready for the next BSW lab run.