Skip to content

feat(sql): add native query support for system tables - #20183

Open
FrankChen021 wants to merge 12 commits into
apache:masterfrom
FrankChen021:codex/native-sys-tasks-mvp
Open

FrankChen021 wants to merge 12 commits into
apache:masterfrom
FrankChen021:codex/native-sys-tasks-mvp

Conversation

@FrankChen021

@FrankChen021 FrankChen021 commented Aug 28, 2026 •

Copy link
Copy Markdown
Member

Follow-up sys.tasks implementation: FrankChen021/druid#198 (stacked on this PR).

Description

This PR introduces opt-in native-query execution for supported sys.* tables and enables it for sys.server_properties. Existing Bindable execution remains the default for compatibility.

The native path is selected only when:

  • The resolved SQL engine is native.
  • The query context contains "useNativeQueryForSystemTables": true.
  • The table implements NativeSystemTable.

The context parameter defaults to false. Unsupported system tables continue to use their existing Bindable path, and other SQL engines are unaffected.

Motivation

System tables traditionally use Calcite's Bindable execution path and are not represented as native Druid datasources. Consequently, native operators and aggregations are unavailable for these tables. For example:

SELECT COUNT(DISTINCT server) FROM sys.server_properties;

This PR represents supported system tables as SystemTableDataSource instances. Components provide their local rows through native Scan queries, after which the Broker executes the original query using Druid's native query engine.

Scope This PR
Shared native system-table planning and execution Yes
sys.server_properties Yes; rows are collected from discovered Druid server nodes
sys.tasks Follow-up FrankChen021/druid#198
Execution path
POST /druid/v2/sql
        |
        v
Router -- normal SQL routing --> Broker SQL planner
                                      |
                            native SystemTableDataSource
                                      |
                       discover contributing server nodes
                                      |
                    local call or POST /druid/v2 with
                    X-Druid-Native-Query-Route: local
                                      |
                                      v
                         component-local Scan query
                                      |
                                      v
                              ScanResultValue rows
                                      |
                                      v
                 Broker residual filtering and aggregation

The Router forwards SQL requests to a Broker normally. For component fanout, the Broker uses the standard /druid/v2 endpoint and response format rather than introducing a system-table-specific protocol.

The local-route header is a routing instruction, not an authentication credential. The receiving process still applies its normal authentication and system-table authorization. It executes only the local system-table Scan, which avoids recursive Router-to-Broker or Broker-to-Broker fanout.

If the Broker itself contributes rows, it calls the raw local handler in-process with the escalator's authentication result rather than performing an HTTP loopback.

Implementation details and extension points

SQL planning

NativeSystemTable supplies the native DruidTable representation backed by SystemTableDataSource. Both planner strategies are supported:

  • In COUPLED mode, DruidTableScanRule converts the system table directly.
  • In DECOUPLED mode, DruidBindableTableScanRule reconstructs the native project/filter/scan plan when Calcite has already embedded filters or projections in a BindableTableScan.

The DECOUPLED conversion is needed for sys.server_properties, whose traditional table implements ProjectableFilterableTable. It runs only after native system-table planning has been selected; the Bindable path is unchanged.

Registering another native system table

A new table uses the generic infrastructure and does not require a new HTTP resource, RPC client, or Broker query branch:

  1. Define and register a SystemTableDescriptor containing the table name, contributing node roles, row signature, routing mode, and Broker-side row authorizer.
  2. Implement and register SystemTableDataProvider on the nodes that own the table's rows.
  3. Make the Calcite table implement NativeSystemTable and return a table backed by SystemTableDataSource.

The existing DataSourceQueryHandler registration then handles discovery, fanout, authorization, and component-local execution.

Filter pushdown in this PR

The PR includes the generic SystemTablePushdownFilter extraction and column-mapping mechanism. Eligible filters from the original query are copied into each component Scan, and the component handler passes provider-supported filters to SystemTableDataProvider.

ServerPropertiesTableDataProvider advertises equality pushdown for server and service_name. A node that does not match returns no rows before enumerating its local properties. The original filter remains on the Broker as a residual filter, so pushdown is an optimization rather than the source of final query correctness.

This PR does not generate metadata SQL. Task metadata-store pushdown is part of the stacked sys.tasks follow-up.

Native query module

NativeQueryEngineModule is a shared facade installed by Druid server components. It combines the queryable, query-runner, segment-wrangler, joinable-factory, system-table, and query-resource modules.

Its builder supports:

  • scanOnly() for management components that need component-local Scan execution without merge buffers or the full aggregation stack.
  • withOverrideModule(...) for role-specific execution bindings such as the Coordinator's segment schema cache.
  • withQueryResourceModule(...) for Broker- and Router-specific query resources.
Compatibility, authorization, and current limitations

Remote requests use Druid's existing escalated HTTP client. Components authenticate them through the normal authenticator chain, and the table descriptor applies row authorization for the original user after rows return to the Broker.

useNativeQueryForSystemTables defaults to false for rolling-upgrade compatibility. It can be enabled per query or globally on Brokers after relevant components are upgraded:

druid.query.default.context.useNativeQueryForSystemTables=true

Current limitations:

  • Components execute Scan queries only. Aggregations, expressions, residual filters, sorting, limits, and window processing execute on the Broker.
  • Component-side aggregation is intentionally deferred because management components do not otherwise require merge buffers, and distributed partial-aggregation results need a well-defined merge contract.
  • Component cancellation closes result consumption and the HTTP request but cannot guarantee interruption of provider work already in progress.
  • Component-local scans are not governed by the normal QueryScheduler lane and capacity limits.
  • Component results use row-oriented ScanResultValue transport rather than frame-based exchange.
  • Most results stream lazily, but window queries require Broker-side materialization.
  • Tables without a native representation retain the Bindable execution path.
  • Pushdown is provider-dependent; unsupported filters remain Broker-side residual filters.

Validation

Coverage includes:

  • Native and Bindable planning in COUPLED and DECOUPLED modes.
  • Native aggregations including COUNT(DISTINCT ...).
  • Component-local Scan execution and Broker-side result processing.
  • Provider registration, node discovery, and failure handling.
  • Router local-header routing and Broker in-process execution without forwarding loops.
  • Authentication and row authorization.
  • Filter extraction and sys.server_properties node pruning.
  • Embedded end-to-end sys.server_properties queries across Druid component roles.

Release note

SQL queries against supported system tables can opt into Druid's native query engine with the useNativeQueryForSystemTables query context parameter. The initial implementation supports sys.server_properties, uses the standard /druid/v2 endpoint for component fanout, and retains the existing Bindable path by default for rolling-upgrade compatibility.

This PR has:

  • been self-reviewed.
  • added documentation for the query context and native system-table behavior.
  • included a release note in the PR description.
  • added Javadocs for the principal new classes and interfaces.
  • added unit and embedded end-to-end tests.
  • been tested in a local Druid cluster.

Copilot AI lite review requested due to automatic review settings August 28, 2026 15:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@FrankChen021

Copy link
Copy Markdown
Member Author

@clintropolis @gianm please help review the changes. I hope this can be merged as soon as possible so that we can move on migrating other system tables into native execution path, and push forward for #18087 and #19855

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review complete: no high-confidence correctness, security, or reliability issues found in the current changes.

Reviewed 100 of 100 changed files.

Validation: focused git diff --no-ext-diff --check passed. Builds and tests were not run.

The dedicated reviewer timed out; this review was completed by the main agent from the prepared worktree.


This is an automated review by Codex GPT-5.6-Luna(max)

@abhishekrb19 abhishekrb19 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for these changes @FrankChen021! The sys.* tables have historically been slow and less performant, especially sys.segments and sys.tasks.

I wonder if we can benchmark sys.tasks with thousands of tasks to see how these changes perform at scale with the native support + filter pushdown.

Also, regarding the scope of these changes, I feel it would be helpful to break this into a few more patches for ease of review. Perhaps something like:

  • Separate PRs for sys.tasks and sys.server_properties, isolating the appropriate wirings, filter pushdown mechanisms and tests to make it work for each table.
  • Handle DECOUPLED planner support in a separate change, since COUPLED is the default planner strategy today (and DECOUPLED is currently undocumented).

Comment on lines -95 to -98
new QueryableModule(),
new QueryRunnerFactoryModule(),
new SegmentWranglerModule(),
new JoinableFactoryModule(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are these not needed? Wondering if this is general cleanup, outside the scope of this PR

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not out of the scope. it's a tiny refactoring of how native query related module are installed. otherwise, for overlord, coordinator, middle manager, we will repeate these code. Now we only need the NativeQueryEngineModule installed, all dependencies are enclosed in this new module.

`X-Druid-Native-Query-Route: local`. Local execution uses the authenticated request identity and applies the table's
authorization rules. The Broker uses the same header for remote node fan-out requests. If the Broker itself is
one of the selected nodes, it executes that node Scan in-process without an HTTP request. The header
controls routing and doesn't grant additional permissions.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hmm, is this header required only for the sys.server_properties table? I wonder if there's a better way to achieve this without passing in additional headers.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

it's needed as the broker will send native query to router or itself to retrieve data. for router, we need to tell whether it acts as a proxy or execute the native query.

another way is by using query context, but I don't think the query context is better way because this value here is just an internal flag, not user-facing context parameter.

Comment thread docs/querying/sql-metadata-tables.md Outdated
@FrankChen021

FrankChen021 commented Sep 9, 2026 •

Copy link
Copy Markdown
Member Author

Thanks for these changes @FrankChen021! The sys.* tables have historically been slow and less performant, especially sys.segments and sys.tasks.

I wonder if we can benchmark sys.tasks with thousands of tasks to see how these changes perform at scale with the native support + filter pushdown.

Also, regarding the scope of these changes, I feel it would be helpful to break this into a few more patches for ease of review. Perhaps something like:

  • Separate PRs for sys.tasks and sys.server_properties, isolating the appropriate wirings, filter pushdown mechanisms and tests to make it work for each table.
  • Handle DECOUPLED planner support in a separate change, since COUPLED is the default planner strategy today (and DECOUPLED is currently undocumented).

Thanks for reviewing.

I deliberately chose the sys.tasks and sys.server_properties included in this PR to demonstrate how the native query for system tables are supported.

These two system tables have different fan-out paths and push down policies:

  1. for sys.server_properties, we need to fan out sub queries to all nodes, while for sys.tasks we need to fan out the query to leader overlord only. This is also why the header X-Druid-Native-Query-Route is introduced to serve the purpose
  2. for sys.tasks, it also demonstrates how filters are pushed down, which is one of the most important goal we need to achieve

If we move any of these system table out of this PR, we can't have a full picture of the change, and understand how the small framework for the system tables work.

As for DECOUPLED mode, only a few files are involved, I think it's better to included in this PR.

To address your concern, I can split the changes for sys.tasks out of this one. But We still need to have a full picture of the core change.

@gianm

gianm commented Sep 9, 2026

Copy link
Copy Markdown
Contributor

@clintropolis @gianm please help review the changes. I hope this can be merged as soon as possible so that we can move on migrating other system tables into native execution path, and push forward for #18087 and #19855

Very cool. I will try to have a look soon.

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Severity Findings
P0 0
P1 0
P2 1
P3 0
Total 1
Severity Findings
P0 0
P1 0
P2 1
P3 0
Total 1

Reviewed 90 of 90 changed files.


This is an automated review by Codex GPT-5.6-Luna(max)

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

Rechecked the updated head against the prior reviewed SHA, beginning with the incremental patch. The prior P2 concerning nested or composite system-table datasources is resolved: local execution now rejects a non-system-table root with BadQueryContextException, and the updated tests cover both broker and node-local handlers. Composite queries without the local route continue through the recursive system-table client path.

The full current diff and surrounding code were reviewed across all 90 changed files, including SQL planning and authorization, native datasource conversion, node discovery/fanout/recovery, request routing, lifecycle and cancellation, service-module wiring, and regression tests. No additional high-confidence correctness, security, compatibility, data-loss, concurrency, lifecycle, integration, or missing-test finding was identified.

Validation performed: git diff --check 8c05f05adbe8f7ccf647c135caa948cc083fae8b -- passed. No builds, broad tests, dependency installs, formatters, prepare, fetch, or cleanup were run, per the review scope.


This is an automated review by Codex GPT-5.6-Luna(max)

@FrankChen021 FrankChen021 added this to the 39.0.0 milestone Sep 19, 2026
@FrankChen021 FrankChen021 mentioned this pull request Sep 19, 2026
1 of 10 tasks
…asks-mvp

# Conflicts:
#	server/src/test/java/org/apache/druid/client/DirectDruidClientTest.java

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

Rechecked the current head beginning with the incremental patch. The incremental patch is the merge of current master into the previously reviewed PR head; the PR logic is unchanged since the prior review. The prior P2 about composite or nested system-table datasources remains resolved: local routing rejects a non-system-table root, while non-local Broker execution recursively resolves system-table leaves.

Reviewed 90 of 90 changed files. The full current diff and relevant surrounding code were reviewed across SQL planner selection and authorization, native datasource conversion, node discovery and leader routing, request forwarding, local and remote execution, failure recovery, cancellation and lifecycle cleanup, service-module wiring, and regression tests. No additional high-confidence correctness, security, compatibility, data-loss, concurrency, lifecycle, integration, or missing-test issue was found.

Validation: git diff --check cf84e5f1b09ce3b05c2f5c015e5259bf66b7a7e6 -- passed. No builds, broad tests, dependency installs, formatters, prepare, fetch, or cleanup were run.


This is an automated review by Codex GPT-5.6-Luna(max)

…asks-mvp

# Conflicts:
#	processing/src/main/java/org/apache/druid/query/rowsandcols/LazilyDecoratedRowsAndColumns.java
#	services/src/main/java/org/apache/druid/cli/CliPeon.java

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

Rechecked the incremental patch first. It is the merge of current master into the previously reviewed PR head, so it contains upstream churn rather than additional PR logic; I then revalidated the full current PR diff and the prior composite/nested system-table routing finding at the current head.

Reviewed 90 of 90 changed files. Coverage included SQL planner native/bindable selection and authorization, native datasource conversion, system-table descriptors and providers, node discovery and leader routing, Router/Broker/local forwarding, local and remote execution, failure recovery, cancellation and lifecycle cleanup, service-module wiring, and regression and embedded tests.

Validation: git diff --check 631f05de5d1d05edfbbab3adc8dbbb99745678ee 0e7033010c84ac45dc909e36d843046fd18bb919 -- passed. No broad builds or test suites were run per review scope.

No actionable PR-caused correctness, security, compatibility, data-loss, concurrency, lifecycle, integration, or high-confidence missing-test issue found. The prior composite/nested system-table routing finding remains resolved by the current root-datasource checks and recursive non-local resolution.


This is an automated review by Codex GPT-5.6-Luna(max)

Merge tests that exercise the same code path into parameterized tests and
drop tests fully covered by stronger equivalents.
Derive native system table names from the logical plan instead of building
the native query, and collapse identical branches in RecoveringNodeIterator.

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Fix virtual-column handling in provider filter pushdown before merging: filters on virtual columns named server or service_name can currently discard matching rows by comparing against physical node metadata. The incremental authorization-plan traversal and recovery simplification introduce no additional confirmed issue.

Reviewed 90 of 90 changed files, starting with the 6-file incremental diff and then reviewing the full current diff and relevant surrounding code. Rechecked the prior composite/nested system-table routing finding; it remains resolved by the local root checks and recursive Broker resolution. Coverage included SQL planning and authorization, native datasource conversion, node discovery and fanout, request routing, recovery, cancellation and cleanup, module wiring, and regression tests.

Validation: git diff --check 631f05de5d1d05edfbbab3adc8dbbb99745678ee HEAD -- passed. This was a static review; builds and test suites were not run.

Severity Findings
P0 0
P1 0
P2 1
P3 0
Total 1

This is an automated review by Codex GPT-5.6-Luna(max)

After addressing the findings or replying to the comments, you can request another review from me to trigger a new automated review.

A virtual column shadows the physical column of the same name, so a filter on it
must be evaluated by the native scan instead of being compared against the
provider's physical column.

@FrankChen021 FrankChen021 left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

SystemTableQueryClient can throw an NPE when constructing a filtered WindowOperatorQuery whose scan leaf has no virtual columns. Fix this before merging.

Reviewed 90 of 90 changed files. The prior virtual-column shadowing issue is resolved.

Validation: git diff --check passed; no tests or builds were run.

Severity Findings
P0 0
P1 1
P2 0
P3 0
Total 1

This is an automated review by Codex GPT-5.6-Luna(max)

After addressing the findings or replying to the comments, you can request another review from me to trigger a new automated review.

if (query instanceof WindowOperatorQuery) {
for (final OperatorFactory operator : ((WindowOperatorQuery) query).getLeafOperators()) {
if (operator instanceof ScanOperatorFactory) {
virtualColumns.addAll(List.of(((ScanOperatorFactory) operator).getVirtualColumns().getVirtualColumns()));

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Handle null window-leaf virtual columns

Finding: WindowOperatorQuery intentionally stores null in ScanOperatorFactory when a leaf has no virtual columns. When the owning window query has a pushed-down filter, nodeFilter is non-null, so this expression dereferences that null value for any empty leaf and fails with an NPE before the node query executes.

Suggestion: Treat null leaf virtual columns as empty before adding them, and add regression coverage for a filtered window query with a leaf that has no virtual columns.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in d4ddcf9e31.

Reproduced first: the new SystemTableQueryClientTest#testWindowLeafFilterWithoutVirtualColumns builds the window query the way production does (a filtered Scan subquery with leafOperators left for the WindowOperatorQuery constructor, which stores null virtual columns) and failed with the reported NPE at nodeVirtualColumns. nodeVirtualColumns now skips leaves whose virtual columns are null; the test asserts the filter is still pushed to the node scan with no virtual columns.

This path is reachable through a native windowOperator query without leafOperators; SQL planning does not produce it today. While adding SQL-level coverage (NativeSysServerPropertiesQueryTest#testNativeWindowFunction, both planners), a separate DECOUPLED bug surfaced: the window query is planned directly over filter(systemTable) and failed with Segment [filtered->ArrayListSegment] cannot shapeshift, because FilteredSegment never provided a CloseableShapeshifter. Fixed generally in 24dee021a5: FilteredSegment now returns CursorFactoryRowsAndColumns over its filtered cursor factory, the same way HashJoinSegment does, with a new FilteredSegmentTest. The window, DECOUPLED and Drill window SQL suites still pass.

…ries

WindowOperatorQuery stores null rather than empty virtual columns for a leaf
scan without any, which caused an NPE when building the node query for a
filtered window query.
FilteredSegment did not provide a CloseableShapeshifter, so window queries over
a FilteredDataSource failed with 'cannot shapeshift'. DECOUPLED planning
produces this shape for filtered window queries on native system tables.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants