Repository navigation
docs(spec): one document that cannot be serialized fails its whole reindex batch (#37269) - #37837
fabrizzio-dotCMS wants to merge 7 commits into
Conversation
|
Claude finished @fabrizzio-dotCMS's task in 57s —— View job Spec review — #37269 bulk batch discardThis PR carries only New Issues
Observations (non-blocking, no change required)
The spec is otherwise coherent and the acceptance criteria are testable (Red-first integration test required at Phases 0/1/3, AC-006 transport-vs-content unit test, AC-007 Postman + no-field-values unit test). Only the |
…37269) Address review: on the OpenSearch path a serialization error throws at add time and is caught per entry in appendBulkRequestToProcessor; flush()'s catch only covers errors from the client.bulk() send. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
What "fails individually" means when one reindex entry maps to several documents (AC-001) A reindex entry is per identifier, but it can turn into several index documents.
Suggestion: state in the spec that the check runs before anything for the entry is queued. In practice that means inside |
…al row (#37269) Address review: one journal row covers every language and working/live document of an identifier, and any sibling's bulk success deletes it. Move the check to the mapping step, before anything of the entry is queued, on the final document map (catchall included) against Jackson's effective string limit. Add the sibling acceptance scenarios. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@dario-daza Good catch — confirmed in the code, and it already applies to the OS path today. Updated the spec (285e4f4):
|
|
Wording correction, no change in behaviour or acceptance criteria: The spec says OpenSearch write failures are fire-and-forget "in Phases 1 and 2", and scopes the shadow fix to "Phase 1/2". In the reindex path OS is a shadow only in Phase 1 ( I'd like to change those two sentences to "Phase 1". OK to update the spec, or do you prefer it stays as approved? |
Two sentences said OpenSearch write failures are fire-and-forget in Phases 1 and 2. In the reindex path OS is a shadow only in Phase 1 (createBulkProcessor: !isReadEnabled(), since #37276); from Phase 2 on its failures propagate. Wording only: no change to scope or acceptance criteria (AC-004 already says Phase 1). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@dario-daza I went ahead with the wording change in f3f2454: "Phases 1 and 2" → "Phase 1" in the two sentences (the second also now says failures propagate from Phase 2 on). No change to scope or acceptance criteria. The push dismissed your approval — could you re-approve when you get a chance? Implementation PR: #37924. |
) Found while verifying the fix by hand: GET /api/v1/esindex/failed embeds each failed record's full contentlet, so failed records of oversized content made a 168 MB response and the Maintenance "Download Failed Records" button crashed the browser tab. Add GET /api/v1/index/failed (typed, no field values) used by the button; the old endpoint keeps its response and is marked deprecated. New AC-007. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@dario-daza one more addition before you re-approve, found while testing the fix by hand: |
Root carries the document limits and retry policy; each record carries the pending operation, failed attempts and, for a document-limit failure, the violation (limit, field, actual, allowed, unit). Explicit nulls. AC-007 updated. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Automated rollback-safety check (and note: an earlier placeholder/test string briefly appeared here while verifying tool permissions — this comment replaces it). Verdict: Safe to roll back. This PR's diff ( |
|
@dario-daza re-validation request — the spec changed in three places since your approval (all in this PR, and the implementation PR #37924 carries the identical spec.md):
Could you take another look and re-approve if it reads right? |
Spec PR (1 of 2) for #37269. It carries
spec.mdonly; the implementation PR follows once this is approved.Problem
When one document in a reindex group can't be serialized (e.g. a field over Jackson's 20,000,000-character read limit), the whole engine request is abandoned and every contentlet in the group is recorded as failed with that one document's message. On a real dataset, 10 oversized documents caused 380 failures, 363 of them healthy contentlets under 100 KB.
The trigger fails on Elasticsearch because the 7.10.2 client re-parses every document while building the bulk body (
RequestConverters.bulk()). OpenSearch parses each document as it is added, so there only the bad document fails. But OpenSearch'sflush()still discards the whole group on any other client-side error, and Phase 3 inherits that.What the spec asks for
Out of scope
data:URIs: separate issue.Verification
A new integration test (one field over 20M characters plus N healthy contentlets in the same fetched group), registered in a MainSuite and run at Phases 0, 1 and 3. It must fail on current code first.
🤖 Generated with Claude Code