Repository navigation
Add vLLM server intropsection endpoints - #2
Conversation
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
elevran
left a comment
There was a problem hiding this comment.
Solid, well-tested first PR for this repo. The endpoint-plugin architecture matches the stated design, naming was renamed consistently everywhere (no leftover vllm_server_introspection references), and CI passes. Found one startup-crash risk that contradicts its own docstring, a secrets-exposure gap the README already warns about, and a CODEOWNERS entry GitHub can't validate. A complexity pass also found ~180 lines of avoidable duplication (logger boilerplate x4, a response-handler pattern duplicated between devices_plugin.py and kv_cache_plugin.py, and repetitive per-kind test builders in test_kv_cache_plugin.py / test_schemas.py) - not blocking, just worth a follow-up cleanup pass.
The devices plugin no longer takes down vllm serve when the worker extension isn't installed. It logs a warning and the endpoint returns 503 with a message pointing at --worker-extension-cls. The RPC method name is now a shared constant so the plugin and the worker extension can't drift apart. kv_connector_extra_config values are now redacted by default with only the key names returned. Set LLM_D_INTROSPECTION_EXPOSE_KV_EXTRA_CONFIG=1 to get the real values back. Also: - fixed the config cache test which compared values to themselves - added unit tests for DeviceInfoWorkerExtension, no GPU needed - set read-only contents permissions on the CI workflows - fixed typos, broken README link and the e2e test path Review comment: - #2 (review) Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
Thanks @elevran for the review. I'll do the duplication cleanup in a follow up PR. |
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
Adds to the repo the vLLM introspection endpoints which are ported from the draft at hickeyma/vllm-server-introspection.
Background is in llm-d/llm-d#2088. The ask is for vLLM to provide endpoints that llm-d could use for routing and scheduling. vLLM community pushed back on growing the core API. So the endpoints ship as endpoint plugins instead which are installed next to the server and loaded at startup. No fork, no patch and no vLLM core changes.
The endpoints
Three of them each a separate plugin so you can turn on only what you need:
GET /plugins/llm-d-server-introspection/config: how the server was launched. Needs no engine, so it works on the CPU only render serverGET /plugins/llm-d-server-introspection/devices: per-rank hardware. Needs the worker extension via--worker-extension-clsGET /plugins/llm-d-server-introspection/kv-cache: KV cache capacity and attention group structure, post-profilingThey stay off unless you name them in
VLLM_PLUGINS, so installing the package doesn't change how an existing server behaves.Layout
Top level is inference server agnostic, as agreed on the issue. vLLM code sits under
vllm/as one installable package (llm-d-api-extensions-vllm) with each extension family as a subpackage. If SGLang grows a similar hook later, it lands as a sibling directory.What got renamed
Everything picked up llm-d naming on the way over:
vllm_server_introspectionllm_d_api_extensions_vllm.server_introspectionVLLM_PLUGINSnamesvllm_server_introspection_*llm_d_server_introspection_*/plugins/vllm-server-introspection/*/plugins/llm-d-server-introspection/*Dropping
vllmfrom a path that vLLM itself serves seemed like the right call and it namespaces llm-d against anyone else's plugins.Community files
Adds standard files like
CONTRIBUTING.md,MAINTAINERS.md,OWNERS,SECURITY.mdand a Python.gitignore. AlsoCODE_OF_CONDUCT.md,CODEOWNERS, issue and PR templates, dependabot config, a Makefile and two workflows (ruff, unit tests). This follows whatllm-d-benchmarkandllm-d-kv-cache-manageralready do.Closes llm-d/llm-d#2088
Closes #1