Repository navigation
fix(opensearch): readiness report degrades instead of failing when one engine is unreachable (#37636) - #37724
Conversation
…e engine is unreachable (#37636) GET /api/v1/index/migration/readiness returned only the connection error whenever Elasticsearch or OpenSearch could not be reached, because both reconcilers called getIndicesStats()/listIndices() on each engine with no failure handling. That includes retiring the old cluster at phase 3, the runbook's last step. An engine that cannot be read is now recorded in unreachableEngines, its side of every row is EngineCopy.unavailable(...) with the reason, and the row's verdict is UNMEASURED rather than MISSING_COUNTERPART, so the report never prescribes a reindex or re-crawl over an unknown. The verdict: safeToRollback is false while either engine is unread; safeToAdvance is false in phases 0-2; at phase 3 an unreachable Elasticsearch is reported in the summary but does not block, an unreachable OpenSearch does. The Elasticsearch client's node-failure log line now says what happened instead of printing only the host. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…"no indices" (#37636) OSIndexAPIImpl.getIndicesStats() and listIndices() answer any failure with an empty result, so with OpenSearch down the reconcilers read "OpenSearch holds no indices": a false "safe to advance" at phase 0 and a "no OpenSearch copy, reindex" blocker at phase 3. Add getIndicesStatsOrThrow()/listIndicesOrThrow() to IndexAPI and listIndicesOrThrow() to SiteSearchAPI as default methods that delegate to the existing ones, override them on the OpenSearch side to propagate, and use them from the two reconcilers only. Every other caller keeps today's behaviour. Verified against a running stack (ES 7.10.2 + OS 3.4.0) at phases 0, 1 and 3 with each engine stopped: HTTP 200, the stopped engine listed in unreachableEngines with its reason, rows UNMEASURED. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The previous commit rewrote the file with LF endings, turning a 10-line addition into a whole-file diff. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Claude finished @fabrizzio-dotCMS's task in 2m 43s —— View job PR ReviewI reviewed the diff ( One confirmed regression, matching @dario-daza's earlier comment. New Issues
Notes (non-blocking)
The |
…37636) The both-down message read "Neither Elasticsearch nor OpenSearch could not be reached", a double negative that states the opposite. It now says both could not be reached and gives each engine's reason. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Closes #37636.
The problem
GET /api/v1/index/migration/readinessgoes dark whenever one search engine cannot be reached: the response is only the connection error ({"message":"elasticsearch: Name or service not known"}), with no phase, no verdict and no rows for the engine that is up. That includes retiring the old Elasticsearch cluster at phase 3, the runbook's last step.Two causes:
ContentIndexMirrorReconcilercallsgetIndicesStats(), andSiteSearchMirrorReconcilercallslistIndices(). Elasticsearch throws, and the wholeevaluate()goes with it.OSIndexAPIImpl.getIndicesStats()/listIndices()do the opposite and answer any failure with an empty result, so a stopped OpenSearch read as "OpenSearch has no indices": a false "safe to advance" at phase 0, and a "no OpenSearch copy — reindex" blocker at phase 3.The change
unreachableEnginesfield on the report (engine → reason, omitted when both answered). Its side of every row becomesEngineCopy.unavailable(...)(exists=false,docCount=-1, newunavailableReason), it is asked nothing else, and the row's verdict is a newUNMEASUREDrather thanMISSING_COUNTERPART, so the report never prescribes a reindex or re-crawl over an unknown.IndexAPI.getIndicesStatsOrThrow()/listIndicesOrThrow()andSiteSearchAPI.listIndicesOrThrow()are new default methods that delegate to the existing ones; the OpenSearch implementations override them to propagate. Only the two reconcilers use them — every other caller (router merge, index-stats JSPs) keeps today's behaviour.[host=...].With both engines reachable the report is unchanged (the new fields are omitted).
Verdict rules — the decision to review
The issue does not say what the verdict should be over a partial view. This PR sets:
safeToAdvanceisfalsein phases 0–2 whenever either engine could not be read (every next phase writes to both), with one blocker per engine that names it and the reason.safeToRollbackisfalsewhenever either engine could not be read. This follows the existing rule that an unmeasurable count on either side makes a downgrade unsafe. It is conservative: with only OpenSearch down, Elasticsearch's completeness could still be judged against the database, so this could be relaxed to "false only when Elasticsearch is unread". Left conservative here; happy to relax it.About the logging half of the issue
The issue also reports the deprecated
$estool.esSearch()path failing with a single unsearchable log line. That does not reproduce onmain: at phase 3 the path no longer reaches Elasticsearch (#37667 fails it on the phase with a descriptive message), and in phases 0–2, reproduced with Elasticsearch stopped, the bare listener line is always followed by a WARN namingesSearch, theDotStateExceptionand the cause — on page render, the Page API and/api/vtl/dynamic. Only the listener wording is changed here.Verified
Unit: 13 new tests (content and Site Search reconcilers, readiness service per phase), all failing first against the old code; 84 tests across the five affected classes plus 44 in the neighbouring OpenSearch classes pass.
Live, Elasticsearch 7.10.2 + OpenSearch 3.4.0, stopping one engine at a time — every case HTTP 200, the stopped engine in
unreachableEngineswith its reason, rowsUNMEASURED:The live runs needed #37723: on current
mainthe Docker image cannot reach OpenSearch at all (#37722), so they ran on an image withjdk.netadded to the runtime.Out of scope, noted: with OpenSearch down at phase 3 the summary says "an outOfSyncCount of 0 here" while the field reads 2 — that wording predates this PR and belongs to #37638.
🤖 Generated with Claude Code
This PR fixes: #37636