Repository navigation
feat: report lame duck mode and let the application move off the server - #9
Draft
MattiasAng wants to merge 1 commit into
Draft
MattiasAng wants to merge 1 commit into
MattiasAng wants to merge 1 commit into
Conversation
A server being taken out of service announces it with the ldm flag on an INFO, then closes its clients gradually. The client records the server as draining and reports it to lameDuckModeHandler. Like the go, python and javascript clients it only reports: the connection is kept until the server closes it, and whether to migrate is the application's decision. None of those clients reconnects on its own, and nats.py documents calling force_reconnect() from inside the callback. The notification is edge triggered, since the flag is repeated on every later update, and is suppressed for the first connection so that starting up against a server that is already draining is not announced as a transition. The draining flag is set when the notification is delivered rather than when the INFO is read: a connection still being confirmed has not marked its server connected yet, and doing so clears the flag. forceReconnect() is how the application acts on it. It is deferred to the next read or write, because the way to leave a draining server is to call it from the lame duck handler, which runs inside the read that delivered the notification, and reconnecting from there would tear down the call stack still processing that message. It is honoured whatever the reconnect option says, that option governing whether the client recovers by itself rather than whether it does as it is told, and it is not routed through the failure path, so it logs no error and reports the real connection error when nothing answers. It makes one pass over the servers however the budget is set: with the default unlimited budget a sweep until something answered would never return when nothing did, and here the answer to "move me" can be "no". Asked for before any connection exists it is a no-op, since the next connection picks a server regardless. The end to end test signals a node of the cluster with SIGUSR2 and restarts it afterwards, so it owns that node for the duration. It is skipped when docker is not available. Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
MattiasAng
added this pull request to stack #10
October 7, 2026 12:08
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A server being taken out of service announces it with the ldm flag on an INFO,
then closes its clients gradually. The client records the server as draining and
reports it to lameDuckModeHandler. Like the go, python and javascript clients it
only reports: the connection is kept until the server closes it, and whether to
migrate is the application's decision. None of those clients reconnects on its
own, and nats.py documents calling force_reconnect() from inside the callback.
The notification is edge triggered, since the flag is repeated on every later
update, and is suppressed for the first connection so that starting up against a
server that is already draining is not announced as a transition. The draining
flag is set when the notification is delivered rather than when the INFO is
read: a connection still being confirmed has not marked its server connected yet,
and doing so clears the flag.
forceReconnect() is how the application acts on it. It is deferred to the next
read or write, because the way to leave a draining server is to call it from the
lame duck handler, which runs inside the read that delivered the notification,
and reconnecting from there would tear down the call stack still processing that
message. It is honoured whatever the reconnect option says, that option governing
whether the client recovers by itself rather than whether it does as it is told,
and it is not routed through the failure path, so it logs no error and reports
the real connection error when nothing answers. It makes one pass over the
servers however the budget is set: with the default unlimited budget a sweep
until something answered would never return when nothing did, and here the
answer to "move me" can be "no". Asked for before any connection exists it is a
no-op, since the next connection picks a server regardless.
The end to end test signals a node of the cluster with SIGUSR2 and restarts it
afterwards, so it owns that node for the duration. It is skipped when docker is
not available.
Co-Authored-By: Claude Sonnet 5.5 noreply@anthropic.com
Stack created with GitHub Stacks CLI • Give Feedback 💬
🤖 Generated with Claude Code