Repository navigation
Conversation
| BaseQuery: db.BaseQuery{IDs: []string{jobID}}, | ||
| }, | ||
| false, 0, 1) | ||
| if err != nil { |
There was a problem hiding this comment.
is continuing on DB failure acceptable even though ownership cannot currently be verified?
How long it will continue if the error persists ?
There was a problem hiding this comment.
Yes, deliberately. It's the same contract the progress write already had before this PR: a failed UpdateProgressCounts is logged and the job keeps going.
How long it continues: until the DB is reachable again (the next successful check aborts the job if the epoch moved in the meantime), or until the job reaches finalization, whose write is epoch-fenced and can't commit for a non-owner. So any extra work is bounded by the outage. A non-owner never gets to commit terminal state.
Why I didn't make it abort after N failures: aborting here goes through the neutral "lost ownership" path, which makes no terminal write. If we are actually still the owner, which is the common case for a transient error, the job would be left in_progress with nobody working on it until something reclaims it. Two more points:
- If the DB is down for everyone, nobody can reclaim either, so continuing is correct.
- Processor readiness doesn't depend on DB health today, so a processor that can't reach the DB isn't reclaimed through Reconciler treats pod readiness as processor liveness #712's readiness path.
Duplicate work therefore needs both a reclaim and this processor being cut off from the DB at the same time.
The principled bound is the lease from #712: a processor that can't renew its lease stops at lease expiry, which is exactly when a reclaimer may take over. I'd rather add the bound there than invent a second timeout here. If it would help in the meantime, I can raise the log level or add a counter for consecutive failed checks, so a persistent failure is visible.
| return u.fenced(u.inner.UpdateProgressCounts(ctx, u.jobID, u.epoch, counts)) | ||
| } | ||
|
|
||
| // CheckJobStatus runs on every progress interval, so a job with no new results |
There was a problem hiding this comment.
nit on the comment style:
// CheckJobStatus calls onFencedOut when epoch no longer owns the job and
// onCancelling when the job is cancelling. Other errors are returned.
wseaton
left a comment
There was a problem hiding this comment.
LGTM. Ran the consistency sim harness (rebased past #676: https://github.com/wseaton/llm-d-batch-gateway/tree/sim-post-676) against main and against this PR on compose, with a 2s progress interval. Two new scenarios hit the quiet interval directly.
| scenario | main | this PR |
|---|---|---|
quiet_interval_cancel_lost: cancelling written, event never inserted, every request still generating |
cancelling ignored for 40s, then cancelling -> completed |
cancelled 3s after the cancelling write |
quiet_interval_reclaim: epoch bumped with 5 requests in flight and 3 waiting |
old owner started all 3 waiting requests | 0 started after the bump |
cancel_event_lost (existing) |
cancelling -> completed |
cancelled |
cancel_racing_completion (existing) |
holds | holds |
Not for this PR: the finalizing write is fenced on epoch only, so a cancelling that lands after the last tick but before finalization can still get overwritten. The window is one progress interval now instead of the rest of the job.
f39d8a4 to
c6a3b0e
Compare
|
@zdtsw I think we can safely merge this. |
c6a3b0e to
8872316
Compare
The epoch fence on a running job only ran through ProgressTracker, and the tracker called the updater only when counts were dirty. A job whose requests run long without completing could go indefinitely without checking its epoch or cancelling status. A failed progress write also cleared the dirty flag until the next result arrived. The tracker now calls CheckJobStatus on every tick, whether or not counts changed. The worker implements it as a read of the job's row by primary key, fenced like the progress write, so quiet intervals add no writes. A missing row, a different epoch or a terminal status aborts the job as lost ownership, the same as a fenced-out progress write. A cancelling status aborts it as a user cancel, which backstops a missed cancel event. Read errors are logged and the job keeps running, the same as progress write errors. A failed progress write now leaves the counts dirty so the next tick retries it. Fixes llm-d#713 Signed-off-by: Jiazhou Gao <gjz140103@gmail.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jiazhou Gao <gjz140103@gmail.com>
…ick case Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Jiazhou Gao <gjz140103@gmail.com>
|
Rebased onto main to resolve the conflict with #725 (initial push in
|
Head branch was pushed to by a user without write access
8872316 to
381eef4
Compare
Why is this PR needed?
The epoch fence on a running job only runs through
ProgressTracker, and the tracker calls the updater only when counts are dirty. A job whose requests run long without completing can go indefinitely without checking its epoch orcancellingstatus: a reclaimed job keeps running on the old owner until its next completion, and a missed cancel event isn't noticed until finalization. A failed progress write also clears the dirty flag until another result arrives.This addresses @lioraron's review point on #676. cc @wseaton
What does this PR do?
ProgressTrackercalls a newCheckJobStatuson every tick, after the push when counts were dirty (skipped if that push already stopped the job).DBUpdateProgress:cancellingaborts it as a user cancel, a backstop for a missed cancel event.DBGet. There are no storage interface changes, and a quiet interval adds no row version or WAL (a fenced no-opUPDATEwould add one per job per interval).Open questions:
cancellingquery would need a registry of running jobs. Is per-job OK?DBGetlogs "DBGet: succeeded" at V(1) and the chart's default verbosity is 2, so this adds one log line per running job per interval. A dedicated fenced-read method onBatchProgressDBClientwould avoid that at the cost of an interface change. Happy to switch if you prefer.How was this tested?
Unit tests added/updated/verified
Integration/e2e tests added/updated/verified (PostgreSQL-backed test; e2e not run)
Manual testing performed
TestExecuteJob_QuietIntervalStatusChange: a request runs with no completions; after the first progress write the row's epoch is bumped (or set tocancelling) and the job must abort with the matching cause. Both cases time out on main.TestJobProgressUpdater_CheckJobStatus: owned, epoch bumped, missing row, terminal, cancelling, cancelling under a newer epoch, transient read error.TestProgressTracker_Tick: quiet ticks check without writing; a failed push is retried on the next tick with no new results.TestProgressTrackerQuietIntervalPostgres(make test-postgres): the same scenarios against PostgreSQL, plus the row'sxminstays unchanged across quiet ticks.make test(with-race),make lint,make test-regression,make test-integration,make test-postgresandmake pre-commitpass locally.Checklist
git commit -s) per DCOmake ci)make test-e2e)Related Issues
Fixes #713
Related to #676. Lease-based liveness (#712) is out of scope.