A Rust implementation of nginx-module-vts for virtual host traffic status monitoring, built on top of the ngx-rust framework.
Status: experimental, but the cross-worker aggregation path, the
vts_zone directive, and the /status endpoint are working end-to-end
with nginx 1.31.
┌─────────────────────────────────────────────────────────────────────┐
│ nginx master │
│ │
│ vts_zone main 1m; ─► ngx_shared_memory_add ─► shm_zone │
│ │ │
│ ▼ │
│ vts_init_shm_zone (Rust) │
│ │ │
│ ▼ │
│ ┌──────────────────────────────────────────────┐ │
│ │ slab pool (SlabPool: Allocator) │ │
│ │ ┌─ VtsShared │ │
│ │ │ ├─ RwLock< RbTreeMap< │ │
│ │ │ │ NgxString<SlabPool>, │ │
│ │ │ │ ServerCounters, │ │
│ │ │ │ SlabPool> > (servers) │ │
│ │ │ └─ RwLock< RbTreeMap< │ │
│ │ │ NgxString<SlabPool>, │ │
│ │ │ UpstreamCounters, │ │
│ │ │ SlabPool> > (upstreams) │ │
│ │ └─ shpool->data = &VtsShared (reload-safe) │ │
│ └──────────────────────────────────────────────┘ │
└──────────────────────────────┬──────────────────────────────────────┘
│ fork
┌────────────────┴────────────────┐
▼ ▼
┌──────────────────────┐ ┌──────────────────────┐
│ worker 1 │ │ worker 2 │
│ LOG_PHASE handler │ │ LOG_PHASE handler │
│ └─► record_server │ │ └─► record_server │
│ └─► record_upstr. │ ──┬────► │ └─► record_upstr. │
│ /status handler │ │ │ /status handler │
│ └─► snapshot_* │ │ │ └─► snapshot_* │
└──────────────────────┘ │ └──────────────────────┘
│
ngx::sync::RwLock guards each map independently:
- writers (record_*) take .write()
- readers (/status snapshot_*) take .read()
The shared state is two RbTreeMaps allocated from the slab pool — one
keyed by server_name, one keyed by "upstream\0server" — wrapped in
ngx::sync::RwLock for concurrent worker access. Capacity scales with
the configured vts_zone size rather than being capped at compile time,
and /status reads no longer block concurrent writers thanks to the
reader-writer lock.
Keys are derived from nginx configuration (the matched server block's
first server_name, the upstream block name) — never from the raw
Host header — so attacker-controlled values cannot expand the key
space.
When vts_zone is not declared the FFI transparently falls back to
a process-local manager — this is how the unit tests exercise the data
model without nginx.
- Cross-worker aggregation — every worker writes to the same slab
table;
/statusreturns totals across the whole nginx instance. vts_zonedirective — declares a real shared-memory zone (ngx_shared_memory_add) whoseinitcallback creates twoRbTreeMaps inside the slab pool from Rust.- Server-zone metrics keyed by the matched server block's first
server_name(not the rawHostheader), so the table can't be blown up by adversarial Host values. - Metric names compatible with nginx-module-vts — the Prometheus families, label names and label values follow the original module's output, so dashboards and alerts written against it keep working. See Compatibility with nginx-module-vts for what still differs.
- Upstream metrics per
(upstream, backend)peer — status-code class counters, bytes in/out, request and upstream response times. - Per-attempt upstream tracking —
r->upstream_statesis iterated so each retry attempt (e.g.502from peer A followed by200from peer B) contributes its own sample to the upstream counters, not just the final state. - Upstream and server-zone request-time histograms — classic
Prometheus
_bucket{le=...}/_sum/_countover a shared fixed 11-bucket layout (client_golang defaults), exposed asnginx_vts_upstream_response_duration_seconds_bucket{...}andnginx_vts_server_request_duration_seconds_bucket{...}. Both feedhistogram_quantile(0.99, ...)for p50/p90/p99 panels per upstream peer and per vhost. - Cache hit/miss metrics per cache zone (
proxy_cache_path keys_zone=NAME:SIZE) — counts ofHIT,MISS,BYPASS,EXPIRED,STALE,UPDATING,REVALIDATED,SCARCEaggregated across workers, exposed asnginx_vts_cache_requests_total. - Cache size gauges per cache zone —
proxy_cache_path max_size=…and current on-disk usage (sh->size × bsize) exposed asnginx_vts_cache_usage_bytes{cache_size="max"}and{cache_size="used"}. - Shared zone accounting —
nginx_vts_main_shm_usage_byteswithshared="max_size","used_size","used_node"and"free_size", so a full zone can be told from an idle one.free_sizecomes from the slab's own page count andused_sizeis the rest of the zone, rather than the sum of node sizes the original reports, because the slab spends a whole page or slot per node and that sum reads low right up to the point where inserts start failing. - Accurate connection counters via the global
ngx_stat_*atomics when nginx is built with--with-http_stub_status_module;reading/writing/waitingmatch whatstub_statuswould report. Without that build flag the module falls back to a cycle-table walk and only theactivetotal stays meaningful. - Subrequest- and
/status-aware counting — the LOG_PHASE handler skips internal subrequests (auth_request,mirror,addition, …) and the module's own/statusscrapes, so neither double-counts the per-vhost counters. - Prometheus text format at
/statuswith thetext/plain; version=0.0.4Content-Type that Prometheus 3.x requires. - Reload-safe —
nginx -s reloadreuses the existing shared table, so counters survive a config reload.
- Rust 1.85 or later (ngx-rust 0.5 uses edition 2024).
- nginx source tree (any 1.24+ release; CI is pinned to 1.28.0).
- A C compiler (
cc/clang). - pcre2 and zlib headers for the nginx build.
export NGINX_SOURCE_DIR=/path/to/nginx-source # ngx-rust looks here
cargo build --releaseOutput: target/release/libngx_vts_rust.{so,dylib}.
cd /path/to/nginx-source
auto/configure --prefix=/tmp/nginx-vts-test \
--with-compat \
--add-dynamic-module=/path/to/ngx_vts
makeThis produces:
objs/nginx— the nginx binary (only needed if you don't already have one built from the same source).objs/ngx_http_vts_module.so— the dynamic module you load fromnginx.confviaload_module.
The repository's config script picks .dylib on macOS and .so on
Linux automatically.
Minimal nginx.conf that proxies through an upstream and exposes
/status:
load_module modules/ngx_http_vts_module.so;
events {
worker_connections 64;
}
http {
vts_zone main 1m;
upstream backend {
server 127.0.0.1:18091;
server 127.0.0.1:18092;
}
# Two local servers acting as the upstream peers.
server { listen 18091; location / { return 200 "peer1\n"; } }
server { listen 18092; location / { return 200 "peer2\n"; } }
server {
listen 18080;
server_name example.test;
location / { proxy_pass http://backend; }
location /status { vts_status; allow 127.0.0.1; deny all; }
}
}Run it:
mkdir -p /tmp/nginx-vts-test/{conf,logs,modules}
cp objs/ngx_http_vts_module.so /tmp/nginx-vts-test/modules/
cp nginx.conf /tmp/nginx-vts-test/conf/
objs/nginx -p /tmp/nginx-vts-test -c conf/nginx.confDrive traffic and read the metrics:
$ seq 1 100 | xargs -P 8 -I{} curl -sS -o /dev/null http://127.0.0.1:18080/
$ curl -sS http://127.0.0.1:18080/statusTrimmed where marked ….
# Prometheus Metrics:
# HELP nginx_vts_info Nginx info
# TYPE nginx_vts_info gauge
nginx_vts_info{hostname="…",module_version="0.1.0",version="1.31.6"} 1
# HELP nginx_vts_start_time_seconds Nginx start time
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791080808
…
# HELP nginx_vts_main_connections Nginx connections
# TYPE nginx_vts_main_connections gauge
nginx_vts_main_connections{status="accepted"} 10
nginx_vts_main_connections{status="active"} 10
…
# HELP nginx_vts_main_shm_usage_bytes Shared memory zone usage
# TYPE nginx_vts_main_shm_usage_bytes gauge
nginx_vts_main_shm_usage_bytes{shared="max_size"} 1048576
nginx_vts_main_shm_usage_bytes{shared="used_size"} 98304
nginx_vts_main_shm_usage_bytes{shared="used_node"} 4
nginx_vts_main_shm_usage_bytes{shared="free_size"} 950272
# HELP nginx_vts_server_bytes_total The request/response bytes
# TYPE nginx_vts_server_bytes_total counter
nginx_vts_server_bytes_total{host="example.test",direction="in"} 8190
nginx_vts_server_bytes_total{host="example.test",direction="out"} 16065
…
# HELP nginx_vts_server_requests_total The requests counter
# TYPE nginx_vts_server_requests_total counter
nginx_vts_server_requests_total{host="example.test",code="2xx"} 105
…
nginx_vts_server_requests_total{host="_",code="2xx"} 105
…
nginx_vts_server_requests_total{host="*",code="2xx"} 210
…
# HELP nginx_vts_upstream_requests_total The upstream requests counter
# TYPE nginx_vts_upstream_requests_total counter
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18091",code="2xx"} 53
…
nginx_vts_upstream_requests_total{upstream="backend",backend="127.0.0.1:18092",code="2xx"} 52
…
# HELP nginx_vts_upstream_server_up Upstream server status (1=up, 0=down)
# TYPE nginx_vts_upstream_server_up gauge
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18091"} 1
nginx_vts_upstream_server_up{upstream="backend",backend="127.0.0.1:18092"} 1
Note that peer1 (53) + peer2 (52) = 105: both workers feed the same
table, so /status shows the totals regardless of which worker
happened to handle the request. host="_" is the two peer servers,
which have no server_name and live in the same nginx here, so the
host="*" row that sums every zone counts each request twice.
The Prometheus output uses the original module's metric names, label
names (host, code, upstream, backend, cache_zone, …) and
label values, including the host="*" row that sums every server
zone. What still differs:
- Additions —
process_start_time_seconds(see Persistence) andnginx_vts_upstream_server_up. - Values —
used_sizeis what the slab has spent, not the sum of node sizes.nginx_vts_start_time_secondsis when the counters started from zero, which a reload does not change. The*_request_secondsand*_response_secondsaverages are cumulative (sum / count), not the original's moving average. - Histograms —
nginx_vts_server_request_duration_secondsandnginx_vts_upstream_response_duration_secondsare always emitted, with fixed buckets, andleis written1rather than1.000. The original emits them only whenhistogram_bucketsis set. - Not emitted yet —
nginx_vts_server_cache_total,nginx_vts_cache_bytes_total,nginx_vts_upstream_request_duration_seconds,nginx_vts_status_code_requests_totaland thenginx_vts_filter_*families. - Upstreams without a group — a
proxy_passstraight to an address is reported under its own name rather than the original'supstream="::nogroups".
| Directive | Context | Args | Description |
|---|---|---|---|
vts_zone |
http |
name size |
Declare the shared-memory zone backing all counters. Minimum size is 1 MB; without this directive the module silently falls back to process-local counters (mainly useful for tests). |
vts_status |
location |
— | Render the Prometheus text response at this location. |
vts_upstream_stats |
http, server, location |
on | off |
Accepted for backward compatibility; currently a no-op (upstream stats are always collected when vts_zone is set). |
The shared state is two RbTreeMaps — one keyed by server_name, one
keyed by the (upstream, server) pair — allocated inside the slab pool
that backs the vts_zone. There is no compile-time slot cap: how many
distinct keys you can track is bounded only by the slab pool size you
configure with vts_zone <name> <size>.
Rough sizing rule of thumb: a 1m zone comfortably holds a few thousand
server-zone keys plus a few thousand upstream pairs. Each entry is on the
order of ~200 bytes for the counters plus the key length plus rbtree
node overhead. Bump the size if you genuinely have more virtual hosts.
When a new key cannot be allocated (the slab pool is full), it is dropped
silently and existing counters keep updating. There is also a defensive
upper bound on key length (VTS_MAX_KEY_BYTES = 256) to keep
misconfigured server_name directives from chewing up the pool.
Keys are derived from nginx configuration (the matched server block's
first server_name, the upstream block name) — never from the raw Host
header — so attacker-controlled values cannot expand the key space.
Counters live only in shared memory; nothing is written to disk.
| Operation | Counters |
|---|---|
nginx -s reload |
Kept — nginx reuses the zone |
Reload after changing the vts_zone size |
Reset to zero — a new zone is allocated |
| Stop / start, container restart | Reset to zero |
Binary upgrade (USR2) |
Reset to zero — the new master does not inherit the zone |
All *_total metrics are Prometheus counters, so query them with
rate() / increase(), which detect and absorb resets. What a restart
loses is only the increment since the last scrape.
When a zone is configured, /status also reports when the counters
last started from zero, under the original module's name and under the
name other collectors look for:
# TYPE nginx_vts_start_time_seconds gauge
nginx_vts_start_time_seconds 1791039391
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1791039391
It is the time the zone was built, not the time a process started, so
it stays put across a reload and moves forward exactly when the
counters reset. process_start_time_seconds is unprefixed because
that is the name other collectors look for to tell a reset from a
series they have only just started watching.
OpenTelemetry Collector — scrape /status with the prometheus
receiver and let the metric_start_time processor take the start time
from this metric, before any batching:
receivers:
prometheus:
config:
scrape_configs:
- job_name: nginx-vts
metrics_path: /status
static_configs:
- targets: ["localhost:80"]
processors:
metric_start_time:
strategy: start_time_metric
service:
pipelines:
metrics:
receivers: [prometheus]
processors: [metric_start_time]
exporters: [otlp] # Datadog, Mackerel, or any OTLP backendDatadog Agent — the openmetrics check reads the same metric with
use_process_start_time, so counters that started after the Agent did
are counted from zero on the first scrape instead of being dropped:
instances:
- openmetrics_endpoint: http://localhost/status
namespace: nginx_vts
metrics: [".*"]
use_process_start_time: trueMackerel — send OTLP through the Collector configuration above;
Mackerel stores the counters as cumulative sums and its PromQL
rate() / increase() absorb resets. Scraping with
mackerel-plugin-prometheus-exporter instead posts every value as is,
so a restart shows up as a drop to zero on the graph.
NGINX_SOURCE_DIR=/path/to/nginx-source cargo test --lib~75 unit tests cover the shared-table data model, the upstream
tracker, the Prometheus formatter (per metric family), the cache
statistics helpers, the LOG_PHASE-level FFI, and the rendered
/status output via the process-local VTS_MANAGER fallback.
NGINX_SOURCE_DIR=/path/to/nginx-source cargo clippy --all-targets -- -D warnings
cargo fmt --all -- --checkThe list below tracks known gaps relative to the original
nginx-module-vts. None of them block normal traffic monitoring.
- JSON / HTML / JSONP output formats — only Prometheus text is emitted.
/controlAPI for reset/delete.vts_dumpdirective (periodic on-disk dump for counter recovery across restarts). Not planned while output is Prometheus-only; see Persistence.
- Filter zones (
vhost_traffic_status_filter_by_set_key,_filter_by_host,_filter_max_node) — no dynamic key-based grouping yet. - Traffic limiting (
vhost_traffic_status_limit_traffic,_limit_traffic_by_set_key) — the module is observation-only; it cannot rate-limit responses.
- Some of the original's families are not emitted yet; see Compatibility with nginx-module-vts.
- Upstream peer state (
down,weight,max_fails,fail_timeout,backup) is not yet read from the nginx upstream configuration. - Per-status-code counters
(
vhost_traffic_status_measure_status_codes) — only the1xx/2xx/3xx/4xx/5xxclass buckets are exposed. - Histogram bucket layout is fixed at the Prometheus
client_golangdefaults (5ms..10s, 11 buckets). There is novts_histogram_buckets-style directive to customise the bounds. - Average method (
vhost_traffic_status_average_methodAMM / WMA) — averages are plain cumulativesum / count. - Embedded
$vts_*variables for use inlog_format/if— upstream module exposes ~20; we expose none.
Licensed under either of
- Apache License, Version 2.0 (LICENSE-APACHE or http://www.apache.org/licenses/LICENSE-2.0)
- MIT license (LICENSE-MIT or http://opensource.org/licenses/MIT)
at your option.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above, without any additional terms or conditions.