2026-09-14_SDP-1400_session_log.md #1

  • //
  • p4-sdp/
  • dev/
  • ai_dev_support/
  • session_logs/
  • 2026-09-14_SDP-1400_session_log.md
  • Markdown
  • View
  • Commits
  • Open Download .zip Download (4 KB)

Session Log: SDP-1400 -- live_checkpoint.sh double offline_db rebuild

Date: 2026-09-14 Stream: //p4-sdp/dev Job: SDP-1400 (Bug) Fix changelist: 33774

Summary

A customer-reported log showed live_checkpoint.sh performing two full offline_db rebuilds back-to-back on an edge server (each replay taking ~4.2 hours), roughly doubling the total run time of the live checkpoint process.

Root cause

checkpoint() in Server/Unix/p4/common/bin/backup_functions.sh has two branches:

  • Commit/master server (SERVERID == P4MASTER_ID): dumps a live checkpoint from P4ROOT (-jc/-jcp/-jcpm). Does not touch offline_db at all.
  • Edge/replica/standby server (the else branch): calls request_replica_checkpoint.sh ... -now to ask the commit server for a checkpoint, waits (polling every 5s) for the corresponding *.md5 file to appear -- which, per the customer log, took ~5 hours -- then re-logs in (ticket may have expired during the wait) and, prior to this fix, called recreate_offline_db_files directly.

Separately, live_checkpoint.sh's main flow always calls checkpoint() and then, unless -skip_odb was passed, unconditionally calls recreate_offline_db_files again:

checkpoint
if [[ "$RebuildOfflineDB" -eq 1 ]]; then
   recreate_offline_db_files
else
   log "Skipping rebuild of offline_db due to '-skip_odb'."
fi

Because checkpoint() already did the rebuild internally on edge/replica servers, this second unconditional call redundantly rebuilt offline_db a second time from the exact same checkpoint file, which is exactly what the customer log shows: two consecutive Recovering from serial checkpoint file: .../p4_1.edge_kef.ckp.13865.gz / -jr replays, each ~4.2 hours, sandwiched around the "Seeking highest journal counter..." lookup being run twice.

Scope confirmation: commit server vs. edge/replica

The double rebuild only ever occurred on edge/replica/standby servers, never on the commit/master server. On the commit server (isCommitServer == 1 branch), checkpoint() only performs the live checkpoint dump from P4ROOT -- it never calls recreate_offline_db_files internally. The single call to recreate_offline_db_files made afterward by live_checkpoint.sh's main flow (gated by RebuildOfflineDB/ -skip_odb) was, and remains, the only rebuild performed on a commit server. This code path was not touched by the fix.

The fix is scoped entirely to the else (non-commit-server) branch of checkpoint(), so commit server behavior is provably unchanged.

Fix

Removed the internal recreate_offline_db_files call from the edge/ replica branch of checkpoint(), while keeping the re-p4login (needed because the wait for the MD5 file can span hours, long enough for the ticket to expire). The single rebuild is now performed exactly once, by live_checkpoint.sh's existing -skip_odb-gated call, for both commit and non-commit servers alike.

Files changed (changelist 33774)

  • Server/Unix/p4/common/bin/backup_functions.sh -- the actual fix; removed duplicate recreate_offline_db_files call from checkpoint().
  • test/bsw/test_script-live-checkpoint-off-commit.sh -- added a regression check: counts occurrences of Recovering from .* checkpoint in the edge/standby checkpoint.log after a live checkpoint run and fails if more than one rebuild occurred.
  • doc/ReleaseNotes.adoc -- added a SDP-1400 change-history entry under "SDP 2026.1 Patch 1".

Verification performed

  • bash -n on all modified shell scripts -- clean.
  • shellcheck on backup_functions.sh -- no new warnings introduced.
  • Traced all call sites of checkpoint() in the repo; confirmed live_checkpoint.sh is the only caller, so no other script relies on checkpoint() performing its own offline_db rebuild.
  • Did not run the BSW test itself in this session (requires live lab hosts p4-04/p4-02); the added assertions are ready for the next BSW lab run.
# Session Log: SDP-1400 -- live_checkpoint.sh double offline_db rebuild

Date: 2026-09-14
Stream: `//p4-sdp/dev`
Job: [SDP-1400](https://perforce.atlassian.net/browse/SDP-1400) (Bug)
Fix changelist: 33774

## Summary

A customer-reported log showed `live_checkpoint.sh` performing two full
offline_db rebuilds back-to-back on an edge server (each replay taking
~4.2 hours), roughly doubling the total run time of the live checkpoint
process.

## Root cause

`checkpoint()` in `Server/Unix/p4/common/bin/backup_functions.sh` has two
branches:

* **Commit/master server** (`SERVERID == P4MASTER_ID`): dumps a live
  checkpoint from `P4ROOT` (`-jc`/`-jcp`/`-jcpm`). Does **not** touch
  `offline_db` at all.
* **Edge/replica/standby server** (the `else` branch): calls
  `request_replica_checkpoint.sh ... -now` to ask the commit server for a
  checkpoint, waits (polling every 5s) for the corresponding `*.md5` file to
  appear -- which, per the customer log, took ~5 hours -- then re-logs in
  (ticket may have expired during the wait) and, prior to this fix, called
  `recreate_offline_db_files` directly.

Separately, `live_checkpoint.sh`'s main flow always calls `checkpoint()` and
then, unless `-skip_odb` was passed, unconditionally calls
`recreate_offline_db_files` again:

```bash
checkpoint
if [[ "$RebuildOfflineDB" -eq 1 ]]; then
   recreate_offline_db_files
else
   log "Skipping rebuild of offline_db due to '-skip_odb'."
fi
```

Because `checkpoint()` already did the rebuild internally on edge/replica
servers, this second unconditional call redundantly rebuilt `offline_db` a
second time from the exact same checkpoint file, which is exactly what the
customer log shows: two consecutive `Recovering from serial checkpoint
file: .../p4_1.edge_kef.ckp.13865.gz` / `-jr` replays, each ~4.2 hours,
sandwiched around the "Seeking highest journal counter..." lookup being
run twice.

## Scope confirmation: commit server vs. edge/replica

**The double rebuild only ever occurred on edge/replica/standby servers,
never on the commit/master server.** On the commit server
(`isCommitServer == 1` branch), `checkpoint()` only performs the live
checkpoint dump from `P4ROOT` -- it never calls `recreate_offline_db_files`
internally. The single call to `recreate_offline_db_files` made afterward
by `live_checkpoint.sh`'s main flow (gated by `RebuildOfflineDB`/
`-skip_odb`) was, and remains, the *only* rebuild performed on a commit
server. This code path was not touched by the fix.

The fix is scoped entirely to the `else` (non-commit-server) branch of
`checkpoint()`, so commit server behavior is provably unchanged.

## Fix

Removed the internal `recreate_offline_db_files` call from the edge/
replica branch of `checkpoint()`, while keeping the re-`p4login` (needed
because the wait for the MD5 file can span hours, long enough for the
ticket to expire). The single rebuild is now performed exactly once, by
`live_checkpoint.sh`'s existing `-skip_odb`-gated call, for both commit
and non-commit servers alike.

## Files changed (changelist 33774)

* `Server/Unix/p4/common/bin/backup_functions.sh` -- the actual fix;
  removed duplicate `recreate_offline_db_files` call from `checkpoint()`.
* `test/bsw/test_script-live-checkpoint-off-commit.sh` -- added a
  regression check: counts occurrences of `Recovering from .* checkpoint`
  in the edge/standby `checkpoint.log` after a live checkpoint run and
  fails if more than one rebuild occurred.
* `doc/ReleaseNotes.adoc` -- added a SDP-1400 change-history entry under
  "SDP 2026.1 Patch 1".

## Verification performed

* `bash -n` on all modified shell scripts -- clean.
* `shellcheck` on `backup_functions.sh` -- no new warnings introduced.
* Traced all call sites of `checkpoint()` in the repo; confirmed
  `live_checkpoint.sh` is the only caller, so no other script relies on
  `checkpoint()` performing its own offline_db rebuild.
* Did not run the BSW test itself in this session (requires live
  lab hosts `p4-04`/`p4-02`); the added assertions are ready for the
  next BSW lab run.
# Change User Description Committed
#1 33778 C. Thomas Tyler SDP-1400: Add session log to ai_dev_support (dev-process artifact, not shipped)

Session log documenting investigation and fix for SDP-1400 (live_checkpoint.sh
double offline_db rebuild on edge/replica servers). Kept separate from the
product fix changelist per ai_dev_support convention.