Repository navigation
Fix metrics lost after metric service restart and conflicting metric series - #18748
Conversation
…series Changing dn_metric_level through a hot reload, or restarting the metric service over JMX, unbinds and rebinds every metric set. Many metric sets dropped their registrations when unbinding, kept metric objects that the restart replaced, or created their metrics only once, so their metrics disappeared for good after the restart, e.g. the thread pool, pipe, file and Ratis metrics. - Keep the registrations when unbinding, bind them again, and remove them only when the owner deregisters: thread pools, pipe metrics on the DataNode and ConfigNode, pipe event commit, subscription queues and IoTConsensusV2 sync lag. - Refresh the cached metric objects when binding: the Ratis metric registry bridge and the pipe sink compression timers. - Create again the metrics created by one-shot paths and restore the recorded values: file metrics, per-table schema gauges, writing thresholds and active counts, load TsFile memory, active loading counters, system disk metrics and JVM GC pause timers. Stop the GC listeners of the previous binding. - Keep binding the other metric sets when one fails, register the partition cache metrics, and give the RPC connection and decoding metrics a single owner each. Some metrics were also never exported or reported wrong values: - A metric name must keep the same tag keys, otherwise the metrics registered later are silently dropped. Move the per-table device number gauges to their own names, and add the region tag to the pipe-based subscription queue metrics, which share their names with the consensus-based queues. - The database memory and chunk metadata memory gauges sum up all data regions and TsFile processors of the database, instead of each of them replacing and removing the shared gauge. - The IoTConsensus stage timers of a region are shared by the dispatcher threads of all its peers, so only the last one to unbind removes them. - Record the active memtable count under the region tag of the counter created with the region. - Count the latest sample in the GC time window, and reset the ring buffer when the monitor binds again.
|
Reviewed commit
Validation: a clean reactor test build of the 28 selected classes listed in the PR description passed all 99 tests. Three additional targeted tests reproduced the findings above. The parent comparisons replaced only the affected production classes (and the old pipe message constants), using the same JVM, test classes, and dependency classpath; they were not full parent-revision builds. No distributed end-to-end test was run for this review. |
- A pipe or subscription registration may be removed while the metric service restart unbinds it, which failed the unbinding and skipped binding the whole metric set again. Synchronize binding and unbinding with registration and deregistration in the metric sets that unbind a snapshot of their registrations. - The dispatchers recovered at startup register their metrics before the metric service starts, which binds them again. Track the owners of the shared IoTConsensus stage timers instead of counting the bindings, so that the last owner still removes the timers.
Description
Changing
dn_metric_levelwithSET CONFIGURATION+LOAD CONFIGURATION, or callingrestartService()of the metric service over JMX, runsMetricService.restartService(): it drops all metrics, then unbinds and binds every metric set again. Many metric sets could not survive this, so their metric families disappeared until the node restarted, e.g. allthread_pool_*series (#18681 fixed onlyclient_manager). Some other metric series were never exported or reported wrong values, because two owners registered the same series, or the same metric name with different tag keys.Metrics lost after the metric service restarts
The metric sets followed three broken patterns:
unbindFromdropped the registrations thatbindTobinds from, so nothing was bound again. NowunbindFromonly removes the metrics and keeps the registrations, which are removed when their owner deregisters:ThreadPoolMetrics, the pipe metric sets on the DataNode and the ConfigNode,PipeEventCommitMetrics, the subscription queue metrics, and the IoTConsensusV2 sync lag, which is now released when the server stops instead of when its metrics unbind. Thread pools are also unregistered by instance, so a new pool with the same name is kept when the old one shuts down.TsFileMetrics(file_global_*,file_level_*), the per-table schema gauges, theWritingMetricsthresholds, WAL queue size and active memtable / time partition counts, the load TsFile memory, the active loading counters, the system disk metrics (also when the level was OFF before), and the JVM GC pause timers.JvmGcMetricsalso stops the GC listeners of the previous binding, which kept running and reset the timers.Besides, a metric set failing to rebind no longer stops the others,
CacheMetrics(partition cache hits) is registered at all, andRPCServiceThriftHandlerMetricsno longer has two instances owning the same series.The ConfigNode starts the metric service after the consensus layer, so its Ratis metrics (
ConfigRegion...) were missing from the start for the same reason. They are exported now.Metric series with conflicting owners
tabletag toschema_engine/schema_region, so they were never exported. They now have their own namesschema_engine_tableandschema_region_table, with the same tags.regiontag ([Subscription] Fix consensus subscription metrics across Regions #18277), so whichever registered later lost its metrics. The pipe-based queues now have theregiontag too, with an empty value as they are not bound to a region. Renaming the metrics of the consensus-based queues instead would drop the dashboard compatibility that [Subscription] Fix consensus subscription metrics across Regions #18277 kept on purpose.mem{name="database_<db>"}was registered by each data region of the database, andmem{name="chunkMetaData_<db>"}by each TsFileProcessor, so the gauge reported one of them and disappeared when any of them was closed. They now sum up all data regions / processors of the database, and the last one removes the gauge.iot_send_logstage timers of a region were created and removed by the dispatcher thread of each peer: a peer leaving removed the timers the others still used, and after a restart only one peer recorded into the exported timers. The peers of a region now share the timers, and the last one to unbind removes them. Making the timers per peer instead would change thenametag.WritingMetricsrecorded the active memtable count underregion="<N>", while it creates and removes the counter underregion="DataRegion[<N>]", so each region had two series and one of them leaked when the region was deleted.JvmGcMonitorMetrics(jvm_gc_accumulated_time_percentage) left the latest sample out of the observation window, and kept the ring buffer indices of the previous binding, which counted the samples before a restart again. As the latest sample now counts, the first one is taken after a whole interval (3 s instead of 50 ms), so the percentage is never computed over a tiny window.Compatibility
schema_engine_tableandschema_region_table. These gauges were never exported before.subscription_uncommitted_event_count,subscription_current_commit_idandsubscription_event_transferof pipe-based queues getregion="". Prometheus treats an empty label like a missing one; the IoTDB reporter paths of these metrics get aregion=node.active_memtable_countis only exported withregion="DataRegion[<N>]".mem{name="database_<db>"}andmem{name="chunkMetaData_<db>"}report the totals of the database.Validation
MetricService, as a mocked one cannot show the tag key conflicts.restartService(), the table gauge not registered, 2 of 3 subscription queues exported, 2 counters for one region, 1 of 2 timer records counted, 200 instead of 300 for the database memory.JvmGcMonitorMetricsTestand the load TsFile memory and cache cases ofDataNodeMetricsRestartTestcall methods added by this PR, so they cannot run on the old code.ALL, with tree model data, table model data and a pipe: scraped, setdn_metric_leveltoIMPORTANTand back toALLwithLOAD CONFIGURATION ON LOCAL, and scraped again, twice.ConfigRegionRatis series (master: 112 families, no Ratis series).mem{name="database_*"}equaled the sum ofdata_region_mem_cost(4255 = 4255; master: 2770 vs 4255).mvn test -pl iotdb-core/node-commons,iotdb-core/consensus,iotdb-core/confignode,iotdb-core/datanode -am -Dtest=ThreadPoolMetricsTest,MetricServiceRestartTest,JvmGcMonitorMetricsTest,PipeEventCommitMetricsTest,ClientManagerMetricsTest,MetricReporterSwitchTest,IoTDBThreadPoolFactoryTest,MetricManagerLifecycleTest,PrometheusReporterTest,RatisMetricSetTest,IoTConsensusV2ServerMetricsTest,RatisConsensusTest,LogDispatcherThreadMetricsTest,PipeConfigNodeMetricsRestartTest,PipeMetricsRestartTest,PipeSchemaRegionListenerMetricsTest,PipeSchemaRegionSinkMetricsTest,TsFileMetricsTest,DataNodeMetricsRestartTest,SchemaMemMetricTableTest,SubscriptionMetricsRestartTest,ConsensusSubscriptionPrefetchingQueueMetricsTest,WritingMetricsTest,MetricServiceTest,DatabaseMemMetricsTest,ActiveLoadListeningDirConfigTest,ActiveLoadDirScannerTest,ActiveLoadTsFileLoaderTest -Dsurefire.failIfNoSpecifiedTests=false -DfailIfNoTests=false mvn test-compile -DskipTests mvn test-compile -P with-zh-locale -DskipTestsThis PR has:
Key changed/added classes in this PR
MetricService,ThreadPoolMetrics,PipeEventCommitMetrics,JvmGcMonitorMetrics,MetricRatisMetricSet,MetricRegistryManager,IoTDBMetricRegistry,CounterProxy,TimerProxyIoTConsensusV2ServerImpl,IoTConsensusV2ServerMetrics,LogDispatcherThreadMetricsorg.apache.iotdb.db.pipe.metricandorg.apache.iotdb.confignode.manager.pipe.metricSubscriptionPrefetchingQueueMetrics,ConsensusSubscriptionPrefetchingQueueMetricsTsFileMetrics,WritingMetrics,SchemaEngineMemMetric,SchemaRegionMemMetric,DataRegionMetrics,TsFileProcessorInfoMetricsLoadTsFileMemMetricSet,ActiveLoadingFilesMetricsSet,ActiveLoadingFilesNumberMetricsSetRPCServiceThriftHandlerMetrics,CacheMetricsSystemMetrics,JvmGcMetrics🤖 Generated with Claude Code