Repository navigation
Retiring the source engine: deprecated search fails with an unsearchable log line, and the readiness report stops responding #37636
Description
Activity
github-actions commented
on Sep 23, 2026 on Sep 23, 2026 – with GitHub ActionsContributorMore actionsQA Note — how to test this fix
The fix is in #37724. Use the migration test stack from the tester guide (
docs/backend/OPENSEARCH_MIGRATION_TESTER_GUIDE.md): dotCMS on8082, the source engine (Elasticsearch) on9200, the target engine (OpenSearch) on9201.⚠️ Prerequisite — the image must include #37723Any dotCMS image built from
mainsince 2026-09-22 cannot talk to OpenSearch at all: every OpenSearch request hangs (#37722). #37723 fixes that, but only once thedotcms/java-baseimage has been republished. Before testing, confirm the image has it:docker exec <dotcms-container> /java/bin/java --list-modules | grep jdk.net
If that prints nothing, stop: the readiness report will simply hang, and nothing below can be tested.
What the bug was
The migration readiness report (
GET /api/v1/index/migration/readiness) stopped answering whenever one of the two search engines could not be reached. Instead of the report, it returned only the connection error, for example{"message":"elasticsearch: Name or service not known"}: no phase, no verdict, nothing about the engine that was still up. This happens in the runbook's last step, switching the old Elasticsearch cluster off at Phase 3, and during any outage in any phase.With OpenSearch down it was worse: the report did answer, but read the outage as "OpenSearch has no indices". In Phase 0 it said it was safe to advance, and in Phase 3 it told the operator to reindex.
After the fix: the report always answers. It names the engine that could not be reached and why, shows the side that did answer, and never draws a conclusion from the engine it could not read.
Setup
- Start at Phase 1 with content in the site, run a full reindex (System → Maintenance → Index) and wait for it to finish, so both engines hold the same content.
- Take a baseline report:
(The user needs CMS Admin plus the role in
curl -s -u admin:admin http://localhost:8082/api/v1/index/migration/readiness | jq . > baseline.json
DOT_OS_MIGRATION_INDEX_VISIBILITY_ROLE_KEY, as the tester guide explains.) Both content entries should sayIN_SYNC. - Change phases the way the tester guide does it (
DOT_FEATURE_FLAG_OPEN_SEARCH_PHASEindocker-compose.yml, then restart dotCMS). - After starting an engine again, wait until it is healthy before taking the next report. If you check too soon, it is correctly reported as unreachable.
Test 1 — Phase 3, the old Elasticsearch cluster is switched off (the main case)
- Set Phase 3, restart dotCMS, then stop Elasticsearch:
docker compose stop elasticsearch. - Take the report.
- Expected:
- The report answers with the full body, not just an error message.
unreachableEngineslistsElasticsearchwith the reason.- Each content entry shows the OpenSearch side with real document counts, and the Elasticsearch side with
unavailableReasonset. The entry's verdict isUNMEASURED. safeToAdvanceistrueand there are no blockers: nothing in Phase 3 depends on Elasticsearch.safeToRollbackisfalse, and the summary says that rollback safety cannot be judged because Elasticsearch could not be reached.
- Before the fix this step returned only
{"message":"elasticsearch: Name or service not known"}.
Test 2 — Phase 3, OpenSearch is down
- Start Elasticsearch again, wait until it is healthy, then stop OpenSearch.
- Expected: the report answers.
unreachableEngineslistsOpenSearch,safeToAdvanceandsafeToRollbackare bothfalse, and there is one blocker that names OpenSearch and says not to reindex because of it. No blocker should say an index "has no OpenSearch copy".
Test 3 — dual-write phase (1 or 2), each engine down in turn
- Set Phase 1 or 2. Stop Elasticsearch, take the report, start it again and wait. Then do the same with OpenSearch.
- Expected, each time: the report answers, the stopped engine is in
unreachableEngines,safeToAdvanceandsafeToRollbackare bothfalse, and there is one blocker naming the stopped engine, with no advice to reindex or re-crawl.
Test 4 — Phase 0, OpenSearch is down
- Set Phase 0 and stop OpenSearch.
- Expected:
safeToAdvanceisfalse, with one blocker saying OpenSearch could not be reached. Before the fix this said it was safe to advance to Phase 1.
Test 5 — both engines up again
- Start everything, wait until both are healthy, and take the report in the phase you started from.
- Expected: the same as
baseline.json. There is nounreachableEnginesfield at all.
Optional — the log line when Elasticsearch is down
With Elasticsearch stopped, open a page that uses the old
$estool.esSearch(...)call in Phase 0, 1 or 2. The log now saysElasticsearch node failed a request and was marked dead by the client, followed by the host, instead of only the host in brackets. The WARN line namingesSearchright after it was already there before this fix.Known limits (do not raise as bugs)
- With OpenSearch down in Phase 3, the summary text says "an outOfSyncCount of 0 here" even when the field reads 2. That wording predates this fix and is tracked in Readiness report prose contradicts its own data: healthy indexes called leftovers at phase 0, wrong drift direction at phase 3 #37638.
- The rollback verdict stays
falsewhenever either engine is down, including when only OpenSearch is. That is deliberate and conservative.
- added a commit that references this issue
on Sep 24, 2026 QA Result: ✅ PASSED (with notes)
Tested on the local OpenSearch migration stack (
docker/docker-compose-examples/single-node-os-migration, trunk image built 2026-09-29): Elasticsearch side = OpenSearch 1.3.20, target = OpenSearch 3.8.0. Content: 1,506 live / 1,508 working. Baseline after a full reindex at Phase 1: both entriesIN_SYNC, nounreachableEnginesfield.✅ QA note tests
Test Phase Engine stopped Result 1 (main case) 3 Elasticsearch ✅ Full body; unreachableEngineslists Elasticsearch with the reason; OpenSearch counts 1,506 / 1,508; verdictUNMEASURED;safeToAdvance: true, no blockers;safeToRollback: false, "cannot be judged"2 3 OpenSearch ✅ One blocker naming OpenSearch ("do not reindex because of it"); both flags false; no "has no OpenSearch copy" blocker3 1 each in turn ✅ One blocker naming the stopped engine; both flags false; no reindex advice4 0 OpenSearch ✅ safeToAdvance: false, one blocker "OpenSearch could not be reached"5 1 none ✅ Identical to the baseline, no unreachableEnginesfield✅ Extra cases (not in the QA note)
Case Result Both engines down at once (Phase 3) ✅ Report answers (20 s), lists both engines, all UNMEASURED, both flagsfalseEngine hangs instead of being down ( docker pause)✅ Report answers and lists the engine as unreachable (see note 1 for timing) Retired end state: Phase 3, old engine stopped, then dotCMS restarted ✅ Startup passes ("ES is decommissioned and ES_ENDPOINTS is not required"); content search, Content Drive and the report work Engine started but not yet healthy, then healthy ✅ Reported as unreachable while starting; once healthy, matches the baseline without restarting dotCMS Admin without the os_migration_qarole, engine down✅ HTTP 403 with the usual refusal, no engine details leaked $estool.esSearch()at Phase 3✅ Clear DotStateException: the deprecated path "is not available once the OpenSearch migration reaches its final phase", with the call to migrate toNotes (not blocking)
- A hung engine makes the report slow. With Elasticsearch paused the report took 91 s (about three 30 s timeouts in a row); with OpenSearch paused, 41 s. It answers correctly, but an operator checking it during a network problem waits a long time.
- The deprecated-path log still has no stack trace. At Phase 1,
$estool.esSearch()with Elasticsearch down now logsElasticsearch node failed a request and was marked dead by the client…(greppable) and a WARN namingesSearch,DotStateExceptionand the cause. The issue's first criterion also asks for the stack trace and call-site context; neither line has a stack trace. (Called through/api/vtl/dynamic, not a page.) - After a restart with the old engine gone, the reason is only a host name.
unavailableReasonfor Elasticsearch reads justopensearch1(fromjava.io.IOException: opensearch1); before the restart it readopensearch1: Name or service not known. - Something still calls Elasticsearch at Phase 3 startup. An
Elasticsearch node failed a request…ERROR appears during startup with no stack trace or caller, so the log doesn't say what made the call.
Video
video.mov
Drafted with Claude Code, reviewed and posted by @rjvelazco.
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
Found during a lab run of the ES→OpenSearch 3 migration against
docker/docker-compose-examples/single-node-os-migration, by Jamie Mauro.Description
Switching the source engine off is the event that actually breaks the deprecated search path — not any
phase transition, since
esSearch()/esRaw()reachAPILocator.getEsSearchAPI()in every phase andnever touch the phase router. It is also the intended end state: the runbook has the operator retire
the old cluster after a cooling-off period at phase 3.
Two separate defects surface at that moment, and together they leave an operator with no way to
diagnose what happened.
1. The deprecated path fails with one context-free, unsearchable log line
With the source engine stopped, the entire log output for a failing
esSearch()request was:No message. No exception class. No stack trace. No URL, template, or call site. A hostname in
brackets.
Compare the same template failing through the router, which produced a descriptive
ERROR PhaseRouterline, a fullDotStateExceptionwith stack trace, and aWARN ESContentToolcarryingthe request URL, language, IP and user. The rich logging lives in the router, and this path bypasses
it.
It is also literally unsearchable. A targeted grep for
esSearch|NoNodeAvailable|ConnectException|Connection refusedacross 200 lines of output returnednothing, because the line contains none of those words — nor any word an operator would think to
search for. It was found only by tailing the window blind.
Practical effect: retire the old cluster and every deprecated call site breaks at once, with
essentially no diagnostic signal pointing at why.
2. The readiness report cannot answer at all
With
opensearch1stopped, the endpoint returns:{"message":"opensearch1: Name or service not known"}No phase, no verdict, no content — the whole report is gone. The stack trace shows why:
evaluate()calls the reconciler, which callsESIndexAPI.getIndicesStats()unconditionally. Oneunreachable engine takes down the entire endpoint rather than degrading to "OpenSearch state is X,
Elasticsearch unavailable".
Why this matters
The operator documentation calls this report the only reliable view of migration state, and has the
operator reading it at phase 3 — in the closing checks and throughout troubleshooting. Phase 3 is also
the state in which the source engine is meant to be decommissioned.
So the primary instrument stops working at exactly the point the migration is supposed to end, and the
one other signal available — the log — carries nothing usable. There is a window where the report
still works: phase 3 with the old cluster left running. The failure arrives whenever someone finally
switches that cluster off, by which time nobody connects it to the migration.
Acceptance Criteria
enough context to locate the call site — at minimum matching what the router path already emits.
that engine's side marked unavailable, the other side populated, and a verdict that reflects the
partial view.
ESIndexAPI.getIndicesStats()is not called unconditionally from the reconciler at phase 3.Additional Context
Same lab run as #37635, which is the other way the readiness report stops telling the truth at phase 3
— there it answers, but with nothing in it. The two compound: after a phase-3 reindex the report goes
empty, and once the old cluster is off it stops answering entirely.