Skip to content

PhaseRouter reports a malformed query as an OpenSearch availability failure and falls back to Elasticsearch #37637

Description

@fabrizzio-dotCMS

Found during a lab run of the ES→OpenSearch 3 migration, by Jamie Mauro.

Description

A template sent a query that is not valid JSON. PhaseRouter reported it as an OpenSearch
availability problem and fell back to Elasticsearch:

ERROR index.PhaseRouter - OS read failed in Phase 2 — falling back to ES.
      OS index may be stale or unavailable. Cause: Unable to parse the given query.

The actual cause is client-side and visible in the message itself —
OSSearchAPIImpl.searchRaw does new JSONObject(query) and throws DotStateException("Unable to parse the given query.") when the input is not JSON. Nothing was wrong with the cluster.

Two problems follow from that.

Wrong diagnosis in the log. An operator reading "OS index may be stale or unavailable" goes
looking for a cluster fault. There isn't one; a template sent a bad query. During a migration, when
people are already primed to suspect the new engine, this is an expensive misdirection.

The fallback fires for the wrong class of error. Phase-2 fallback exists so that a degraded
OpenSearch does not take the site down. A 4xx-class client error is not that: the same malformed query
fails identically against the other engine, as it did here — the request still aborted, after doing
the work twice.

Suggested direction

PhaseRouter should separate transport and availability failures from request-validity failures.
Only the former should trigger the fallback and the "stale or unavailable" wording; the latter should
propagate with the underlying message intact.

A reasonable dividing line: DotStateException raised by query parsing (and anything else that means
"the caller sent something invalid") is not a fallback condition.

Acceptance Criteria

  • A malformed query at phase 2 does not trigger the Elasticsearch fallback.
  • Its log line names the real cause and does not suggest the index is stale or unavailable.
  • A genuine connectivity or availability failure still falls back, and still logs as it does today.
  • Tests cover both branches, so the classification cannot silently collapse back into one.

Additional Context

Surfaced alongside #37635 and #37636 in the same lab run. Lower severity than those two — it wastes an
operator's time rather than hiding a problem — but it is in the code path the migration leans on most.

Activity

  1. github-actions commented on Sep 25, 2026

    @github-actions
    Contributor
  2. fabrizzio-dotCMS commented on Sep 25, 2026

    @fabrizzio-dotCMS
    MemberAuthor

    QA Note — how to test this fix

    The fix is in #37747. Use the migration test stack from the tester guide (docs/backend/OPENSEARCH_MIGRATION_TESTER_GUIDE.md): dotCMS on 8082, Elasticsearch on 9200, OpenSearch on 9201. All tests run at Phase 2, the only phase where reads fall back from OpenSearch to Elasticsearch.

    ⚠️ Prerequisite — the image must include #37723

    Images built from main between 2026-09-22 and the base-image republish cannot talk to OpenSearch at all (#37722). Check first:

    docker exec <dotcms-container> /java/bin/java --list-modules | grep jdk.net

    If that prints nothing, stop.

    What the bug was

    At Phase 2, when a page or an API call sent a search that was wrong, dotCMS treated it as if OpenSearch were broken. The log said OS read failed in Phase 2 — falling back to ES. OS index may be stale or unavailable, and anyone reading the log went looking for a cluster problem that did not exist.

    After the fix, a wrong search is reported for what it is:

    • A query that is not even JSON is reported as a bad request and is not retried on Elasticsearch, where it would fail the same way.
    • A query OpenSearch refuses as malformed is still retried on Elasticsearch, because Elasticsearch 7 accepts some older syntax that OpenSearch 3 dropped. The log now warns that the query will fail at Phase 3 unless it is rewritten, instead of blaming the index.
    • A real OpenSearch outage still falls back to Elasticsearch, exactly as before.

    Setup

    1. Set Phase 2 (as the tester guide explains) and restart dotCMS.
    2. Keep a terminal following the log:
      docker logs -f <dotcms-container> 2>&1 | grep PhaseRouter

    Test 1 — a query that is not JSON

    curl -s -u admin:admin -H 'Content-Type: application/json' -X POST http://localhost:8082/api/es/raw -d 'this is not json'

    Expected:

    • The request fails with Unable to parse the given query., as it did before the fix.
    • The log shows one WARN line: OS read rejected in Phase 2 — the request is invalid, not the index; not retrying on ES…, followed by the cause.
    • There is no ERROR line that says falling back to ES or stale or unavailable.

    Test 2 — a query only Elasticsearch accepts

    curl -s -u admin:admin -H 'Content-Type: application/json' -X POST http://localhost:8082/api/es/raw -d '{"query":{"type":{"value":"_doc"}},"size":1}'

    Expected:

    • The request succeeds and returns results, served by Elasticsearch, as it did before the fix.
    • The log shows one WARN line: OS rejected a read in Phase 2 as malformed; retrying on ES, which may accept syntax OpenSearch does not. This query will fail at Phase 3 unless it is rewritten for OpenSearch., followed by OpenSearch's reason (unknown query [type]).
    • There is no line saying stale or unavailable.

    Test 3 — a valid query still works

    curl -s -u admin:admin -H 'Content-Type: application/json' -X POST http://localhost:8082/api/es/raw -d '{"query":{"query_string":{"query":"+baseType:1"}},"size":1}'

    Expected: a normal result, and no PhaseRouter line in the log at all.

    Test 4 — a real OpenSearch outage still falls back

    1. Stop OpenSearch: docker compose stop opensearch.
    2. Run the valid query from Test 3 again.
    3. Expected:
      • The request succeeds and returns results, now served by Elasticsearch.
      • The log shows the ERROR line OS read failed in Phase 2 — falling back to ES. OS index may be stale or unavailable, with a connection error as its cause. This is the fallback working as designed.
    4. Start OpenSearch again and wait until it is healthy.

    Known limits (do not raise as bugs)

    • A query that is not JSON still returns HTTP 500, as it always has. This fix changes only the fallback and the log line, not the response code.
    • A query that neither engine accepts (for example {"query":{"bogus_clause":{}}}) is retried on Elasticsearch and still fails. The log shows the Phase 3 WARN for it, which is expected.
    • Pages that use the older Lucene-style queries (for example $dotcontent.pull(...) with a broken query) are unchanged. There, Elasticsearch quietly returns no results for a broken query, so falling back keeps pages behaving as they did before the migration.
  3. rjvelazco commented on Oct 1, 2026

    @rjvelazco
    Member

    Passed QA

    • dotCMS Docker Image: [dotcms/dotcms:trunk_42cb984]

    Video

    video.mov
  4. rjvelazco commented on Oct 1, 2026

    @rjvelazco
    Member

    After QA Note:

    I noticed that this log message:

    WARN index.PhaseRouter: OS read rejected in Phase 2 â€" the request is invalid, not the index; not retrying on ES

    Contains the symbol â€"; this looks like a UTF-8 encoding issue where an em dash (—) is being rendered incorrectly. Not sure if this is intentional or a bug worth fixing.

    cc: @nollymar @fabrizzio-dotCMS @fmontes

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions