Skip to content

Retiring the source engine: deprecated search fails with an unsearchable log line, and the readiness report stops responding #37636

Description

@fabrizzio-dotCMS

Found during a lab run of the ES→OpenSearch 3 migration against
docker/docker-compose-examples/single-node-os-migration, by Jamie Mauro.

Description

Switching the source engine off is the event that actually breaks the deprecated search path — not any
phase transition, since esSearch() / esRaw() reach APILocator.getEsSearchAPI() in every phase and
never touch the phase router. It is also the intended end state: the runbook has the operator retire
the old cluster after a cooling-off period at phase 3.

Two separate defects surface at that moment, and together they leave an operator with no way to
diagnose what happened.

1. The deprecated path fails with one context-free, unsearchable log line

With the source engine stopped, the entire log output for a failing esSearch() request was:

14:27:19.609  ERROR util.DotRestHighLevelClientProvider$1 - [host=https://opensearch1:9200]

No message. No exception class. No stack trace. No URL, template, or call site. A hostname in
brackets.

Compare the same template failing through the router, which produced a descriptive ERROR PhaseRouter line, a full DotStateException with stack trace, and a WARN ESContentTool carrying
the request URL, language, IP and user. The rich logging lives in the router, and this path bypasses
it.

It is also literally unsearchable. A targeted grep for
esSearch|NoNodeAvailable|ConnectException|Connection refused across 200 lines of output returned
nothing, because the line contains none of those words — nor any word an operator would think to
search for. It was found only by tailing the window blind.

Practical effect: retire the old cluster and every deprecated call site breaks at once, with
essentially no diagnostic signal pointing at why.

2. The readiness report cannot answer at all

With opensearch1 stopped, the endpoint returns:

{"message":"opensearch1: Name or service not known"}

No phase, no verdict, no content — the whole report is gone. The stack trace shows why:

ESIndexAPI.getIndicesStats(ESIndexAPI.java:135)
  ContentIndexMirrorReconciler.statuses(ContentIndexMirrorReconciler.java:101)
  MigrationReadinessService.evaluate(MigrationReadinessService.java:56)
  MigrationReadinessResource.readiness(MigrationReadinessResource.java:109)

evaluate() calls the reconciler, which calls ESIndexAPI.getIndicesStats() unconditionally. One
unreachable engine takes down the entire endpoint rather than degrading to "OpenSearch state is X,
Elasticsearch unavailable".

Why this matters

The operator documentation calls this report the only reliable view of migration state, and has the
operator reading it at phase 3 — in the closing checks and throughout troubleshooting. Phase 3 is also
the state in which the source engine is meant to be decommissioned.

So the primary instrument stops working at exactly the point the migration is supposed to end, and the
one other signal available — the log — carries nothing usable. There is a window where the report
still works: phase 3 with the old cluster left running. The failure arrives whenever someone finally
switches that cluster off, by which time nobody connects it to the migration.

Acceptance Criteria

  • A failure on the deprecated search path logs the exception class, message and stack trace, plus
    enough context to locate the call site — at minimum matching what the router path already emits.
  • That log line contains at least one term an operator would plausibly grep for.
  • The readiness report degrades rather than failing: an unreachable engine yields a report with
    that engine's side marked unavailable, the other side populated, and a verdict that reflects the
    partial view.
  • ESIndexAPI.getIndicesStats() is not called unconditionally from the reconciler at phase 3.
  • Regression test: stop the source engine at phase 3 and assert the endpoint returns a usable body.

Additional Context

Same lab run as #37635, which is the other way the readiness report stops telling the truth at phase 3
— there it answers, but with nothing in it. The two compound: after a phase-3 reindex the report goes
empty, and once the old cluster is off it stops answering entirely.

Activity

  1. fabrizzio-dotCMS commented on Sep 23, 2026

    @fabrizzio-dotCMS
    MemberAuthor

    QA Note — how to test this fix

    The fix is in #37724. Use the migration test stack from the tester guide (docs/backend/OPENSEARCH_MIGRATION_TESTER_GUIDE.md): dotCMS on 8082, the source engine (Elasticsearch) on 9200, the target engine (OpenSearch) on 9201.

    ⚠️ Prerequisite — the image must include #37723

    Any dotCMS image built from main since 2026-09-22 cannot talk to OpenSearch at all: every OpenSearch request hangs (#37722). #37723 fixes that, but only once the dotcms/java-base image has been republished. Before testing, confirm the image has it:

    docker exec <dotcms-container> /java/bin/java --list-modules | grep jdk.net

    If that prints nothing, stop: the readiness report will simply hang, and nothing below can be tested.

    What the bug was

    The migration readiness report (GET /api/v1/index/migration/readiness) stopped answering whenever one of the two search engines could not be reached. Instead of the report, it returned only the connection error, for example {"message":"elasticsearch: Name or service not known"}: no phase, no verdict, nothing about the engine that was still up. This happens in the runbook's last step, switching the old Elasticsearch cluster off at Phase 3, and during any outage in any phase.

    With OpenSearch down it was worse: the report did answer, but read the outage as "OpenSearch has no indices". In Phase 0 it said it was safe to advance, and in Phase 3 it told the operator to reindex.

    After the fix: the report always answers. It names the engine that could not be reached and why, shows the side that did answer, and never draws a conclusion from the engine it could not read.

    Setup

    1. Start at Phase 1 with content in the site, run a full reindex (System → Maintenance → Index) and wait for it to finish, so both engines hold the same content.
    2. Take a baseline report:
      curl -s -u admin:admin http://localhost:8082/api/v1/index/migration/readiness | jq . > baseline.json
      (The user needs CMS Admin plus the role in DOT_OS_MIGRATION_INDEX_VISIBILITY_ROLE_KEY, as the tester guide explains.) Both content entries should say IN_SYNC.
    3. Change phases the way the tester guide does it (DOT_FEATURE_FLAG_OPEN_SEARCH_PHASE in docker-compose.yml, then restart dotCMS).
    4. After starting an engine again, wait until it is healthy before taking the next report. If you check too soon, it is correctly reported as unreachable.

    Test 1 — Phase 3, the old Elasticsearch cluster is switched off (the main case)

    1. Set Phase 3, restart dotCMS, then stop Elasticsearch: docker compose stop elasticsearch.
    2. Take the report.
    3. Expected:
      • The report answers with the full body, not just an error message.
      • unreachableEngines lists Elasticsearch with the reason.
      • Each content entry shows the OpenSearch side with real document counts, and the Elasticsearch side with unavailableReason set. The entry's verdict is UNMEASURED.
      • safeToAdvance is true and there are no blockers: nothing in Phase 3 depends on Elasticsearch.
      • safeToRollback is false, and the summary says that rollback safety cannot be judged because Elasticsearch could not be reached.
    4. Before the fix this step returned only {"message":"elasticsearch: Name or service not known"}.

    Test 2 — Phase 3, OpenSearch is down

    1. Start Elasticsearch again, wait until it is healthy, then stop OpenSearch.
    2. Expected: the report answers. unreachableEngines lists OpenSearch, safeToAdvance and safeToRollback are both false, and there is one blocker that names OpenSearch and says not to reindex because of it. No blocker should say an index "has no OpenSearch copy".

    Test 3 — dual-write phase (1 or 2), each engine down in turn

    1. Set Phase 1 or 2. Stop Elasticsearch, take the report, start it again and wait. Then do the same with OpenSearch.
    2. Expected, each time: the report answers, the stopped engine is in unreachableEngines, safeToAdvance and safeToRollback are both false, and there is one blocker naming the stopped engine, with no advice to reindex or re-crawl.

    Test 4 — Phase 0, OpenSearch is down

    1. Set Phase 0 and stop OpenSearch.
    2. Expected: safeToAdvance is false, with one blocker saying OpenSearch could not be reached. Before the fix this said it was safe to advance to Phase 1.

    Test 5 — both engines up again

    1. Start everything, wait until both are healthy, and take the report in the phase you started from.
    2. Expected: the same as baseline.json. There is no unreachableEngines field at all.

    Optional — the log line when Elasticsearch is down

    With Elasticsearch stopped, open a page that uses the old $estool.esSearch(...) call in Phase 0, 1 or 2. The log now says Elasticsearch node failed a request and was marked dead by the client, followed by the host, instead of only the host in brackets. The WARN line naming esSearch right after it was already there before this fix.

    Known limits (do not raise as bugs)

  2. self-assigned this
    on Oct 1, 2026
  3. rjvelazco commented on Oct 2, 2026

    @rjvelazco
    Member

    QA Result: ✅ PASSED (with notes)

    Tested on the local OpenSearch migration stack (docker/docker-compose-examples/single-node-os-migration, trunk image built 2026-09-29): Elasticsearch side = OpenSearch 1.3.20, target = OpenSearch 3.8.0. Content: 1,506 live / 1,508 working. Baseline after a full reindex at Phase 1: both entries IN_SYNC, no unreachableEngines field.

    ✅ QA note tests

    Test Phase Engine stopped Result
    1 (main case) 3 Elasticsearch ✅ Full body; unreachableEngines lists Elasticsearch with the reason; OpenSearch counts 1,506 / 1,508; verdict UNMEASURED; safeToAdvance: true, no blockers; safeToRollback: false, "cannot be judged"
    2 3 OpenSearch ✅ One blocker naming OpenSearch ("do not reindex because of it"); both flags false; no "has no OpenSearch copy" blocker
    3 1 each in turn ✅ One blocker naming the stopped engine; both flags false; no reindex advice
    4 0 OpenSearch ✅ safeToAdvance: false, one blocker "OpenSearch could not be reached"
    5 1 none ✅ Identical to the baseline, no unreachableEngines field

    ✅ Extra cases (not in the QA note)

    Case Result
    Both engines down at once (Phase 3) ✅ Report answers (20 s), lists both engines, all UNMEASURED, both flags false
    Engine hangs instead of being down (docker pause) ✅ Report answers and lists the engine as unreachable (see note 1 for timing)
    Retired end state: Phase 3, old engine stopped, then dotCMS restarted ✅ Startup passes ("ES is decommissioned and ES_ENDPOINTS is not required"); content search, Content Drive and the report work
    Engine started but not yet healthy, then healthy ✅ Reported as unreachable while starting; once healthy, matches the baseline without restarting dotCMS
    Admin without the os_migration_qa role, engine down ✅ HTTP 403 with the usual refusal, no engine details leaked
    $estool.esSearch() at Phase 3 ✅ Clear DotStateException: the deprecated path "is not available once the OpenSearch migration reaches its final phase", with the call to migrate to

    Notes (not blocking)

    1. A hung engine makes the report slow. With Elasticsearch paused the report took 91 s (about three 30 s timeouts in a row); with OpenSearch paused, 41 s. It answers correctly, but an operator checking it during a network problem waits a long time.
    2. The deprecated-path log still has no stack trace. At Phase 1, $estool.esSearch() with Elasticsearch down now logs Elasticsearch node failed a request and was marked dead by the client… (greppable) and a WARN naming esSearch, DotStateException and the cause. The issue's first criterion also asks for the stack trace and call-site context; neither line has a stack trace. (Called through /api/vtl/dynamic, not a page.)
    3. After a restart with the old engine gone, the reason is only a host name. unavailableReason for Elasticsearch reads just opensearch1 (from java.io.IOException: opensearch1); before the restart it read opensearch1: Name or service not known.
    4. Something still calls Elasticsearch at Phase 3 startup. An Elasticsearch node failed a request… ERROR appears during startup with no stack trace or caller, so the log doesn't say what made the call.

    Video

    video.mov

    Drafted with Claude Code, reviewed and posted by @rjvelazco.

  4. removed their assignment
    on Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions