Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
16 commits
Select commit Hold shift + click to select a range
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -550,6 +550,10 @@ no legacy spelling of either to warn anybody about.
| `GITLAB_MCP_DRAIN_DELAY` | No | HTTP mode: after `SIGTERM`, keep the listener open and answer `/health` with `503 draining` (`Cache-Control: no-store`) for this long before closing it, so a balancer that polls `/health` removes the instance before the close (`0` default, max 5m). Set it to at least one probe interval. `/health` also carries `build` (closest release plus short commit, `.dirty` when the tree had changes; `cmd/server/health.go`) and `config_digest` (twelve hex over tool surface, capability surface, meta parameter schema, tier and whether it was pinned or is detected, scope detection, read-only, safe mode and excluded tools; a comparison fingerprint, not a secret) so the instances behind one balancer can be compared. Flag `--drain-delay` |
| `GITLAB_MCP_RATE_LIMIT_RPS` | No | Per-credential rate limit, in req/s, on every call that reaches GitLab (`tools/call`, `resources/read`, `resources/subscribe`, `subscriptions/listen`, `prompts/get`) (`0` = disabled). The bucket lives on the pool entry, so each token and URL pair is limited on its own, whatever server shape serves it. `tools/list` is charged too, on a bucket of its own refilled a tenth as fast, for the opposite reason to the rest: it reaches no GitLab, and instead spends the processor every tenant of the process is waiting for (about 3.2 MB marshalled per listing on the individual surface). Its burst is the configured one, undivided, so a fleet of clients sharing one credential can still all discover at once; the refill is what bounds a client listing in a loop. Keeping the buckets apart is what lets a drained tool-call bucket still answer a client's discovery |
| `GITLAB_MCP_RATE_LIMIT_BURST` | No | Token-bucket burst size when RPS > 0 (`40` default) |
| `GITLAB_MCP_AUTH_FAILURE_LIMIT` | No | Failed authentications one address may produce inside the failure window before it is blocked (`10` default, max 100000). `0` turns the budget off, and that is the one value it must not be read literally on: the limiter blocks once a record reaches its limit, so a zero would block an address after a single failure, which is the harshest setting there is reached by typing the figure every other budget here reads as "none". `cmd/server.authFailureLimiter` returns nil instead, and every consulting site already tolerates that |
| `GITLAB_MCP_AUTH_FAILURE_WINDOW` | No | Window the failure budget counts in (`1m` default, max 24h), and the step the distinct-credential escalation is built from: one window, then ten, then sixty, reset after sixty of silence. Deriving the ladder rather than giving it four more settings is what lets a test set this to a second and watch the whole thing in seconds |
| `GITLAB_MCP_AUTH_DISTINCT_TOKEN_LIMIT` | No | Distinct credentials one address may have refused inside the distinct window before it is blocked, for longer each time (`50` default, max 100000, `0` disables). It asks a different question from the failure budget, and that is what makes it safe to escalate on: ten failures in a minute is a stuck client as much as an attack, while fifty _distinct_ invalid tokens is only an attack, since a person has one token and a fleet behind a NAT has one each. Credentials are held as truncated SHA-256 digests, never in the clear |
| `GITLAB_MCP_AUTH_DISTINCT_TOKEN_WINDOW` | No | Window the distinct-credential budget counts in (`10m` default, max 24h) |
| `GITLAB_MCP_CLIENT_COMPAT` | No | Per-client response compatibility (`auto` default): Codex sessions get float `priority` in annotations rounded to 0/1; `off` disables. Read from the process environment in both stdio and HTTP modes; the `--client-compat` flag writes the same variable. See `internal/clientcompat` and `docs/guides/client-compatibility.md` |
| `GITLAB_MCP_DESCRIPTION_SUBSTITUTIONS` | No | Rewrite listed catalog text for strict MCP gateway validators: comma-separated `old=new` pairs applied in order to every listed description and title (backslash escapes `\,` `\=` `\\`). Covers tools (all surfaces, schema-embedded descriptions included), prompts, resources, and resource templates; never names, URIs, `pattern`, `const`, enum values, or tool-call payloads. Malformed values refuse startup. Flag `--description-substitutions`; both modes. The served surface itself stays pure ASCII with no semicolons (`make check-gateway-chars`); verify a config with `go run ./cmd/audit_gateway_chars/ -apply -check`. See `internal/gatewaycompat` and `docs/guides/client-compatibility.md` |
| `GITLAB_MCP_TELEMETRY` | No | Export OpenTelemetry traces, metrics and logs over OTLP (`false` default). Off for privacy: telemetry goes to a collector the operator configures, never to the maintainer; the exporters' only default is their own `localhost`. Endpoint, credentials, sampling and batching all come from the standard `OTEL_EXPORTER_OTLP_*` variables the exporters read themselves. `OTEL_SDK_DISABLED=true` vetoes it regardless. Flag `--telemetry`. See `docs/guides/telemetry.md` |
Expand Down Expand Up @@ -610,6 +614,10 @@ In **HTTP mode**, configuration is resolved in three layers: an explicitly passe
| `--trusted-proxy-header` | _(empty)_ | HTTP header with real client IP for rate limiting behind proxies (e.g. `CF-Connecting-IP`, `X-Forwarded-For`); believed only on a connection from an address in `--trusted-proxies`, which it requires |
| `--rate-limit-rps` | `10` | Per-credential rate limit, in req/s, on every call that reaches GitLab (`tools/call`, `resources/read`, `resources/subscribe`, `subscriptions/listen`, `prompts/get`) (`0` disables it), plus `tools/list` on a separate bucket refilled a tenth as fast, with the same burst, charged because it spends the shared processor rather than because it reaches GitLab. The bucket lives on the pool entry, not on the shared server, so each token and URL pair is limited on its own. On by default in HTTP mode; the `GITLAB_MCP_RATE_LIMIT_RPS` env var used by stdio still defaults to `0` |
| `--rate-limit-burst` | `40` | Token-bucket burst size when --rate-limit-rps > 0 |
| `--auth-failure-limit` | `10` | Failed authentications one address may produce inside `--auth-failure-window` before it is blocked (`0` disables the budget rather than blocking on the first failure) |
| `--auth-failure-window` | `1m` | Window the failure budget counts in, and the step the distinct-credential escalation is built from. A 429 announces in `Retry-After` the **longest** block actually holding the request rather than this window, since one failure can raise the minute-long lockout and an hour-long escalated block together and announcing the shorter one would tell a client to come back while it is still refused (`longestAuthBlock` in `cmd/server/auth_gate.go`) |
| `--auth-distinct-token-limit` | `50` | Distinct credentials one address may have refused inside `--auth-distinct-token-window` before it is blocked, for longer each time (`0` disables). The admitted-credential exemption applies here too: a credential the pool already holds is still served from a blocked address |
| `--auth-distinct-token-window` | `10m` | Window the distinct-credential budget counts in |
| `--telemetry` | `false` | Export OpenTelemetry traces, metrics and logs over OTLP. Applies to both transports. Endpoint and credentials come from the standard `OTEL_EXPORTER_OTLP_*` environment |
| `--telemetry-identity` | `none` | What telemetry records about the caller: `none`, `pseudonymous` or `full` |
| `--telemetry-tool-name` | `auto` | Whether `gen_ai.tool.name` is a metric dimension: `auto`, `on` or `off` |
Expand Down
38 changes: 19 additions & 19 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -449,39 +449,39 @@ and is never logged. Full details: [PRIVACY.md](PRIVACY.md).

| Category | Files | Lines |
| ------------------------ | --------: | ----------: |
| Source (`.go`, non-test) | 1,296 | 287,327 |
| Unit tests (`_test.go`) | 878 | 565,726 |
| End-to-end tests | 492 | 94,277 |
| **Total** | **2,666** | **947,330** |
| Source (`.go`, non-test) | 1,300 | 288,328 |
| Unit tests (`_test.go`) | 882 | 566,988 |
| End-to-end tests | 492 | 94,381 |
| **Total** | **2,674** | **949,697** |

### Functions

| Category | Count |
| ------------------------------- | -----: |
| Source functions | 10,064 |
| . Exported (public) | 3,179 |
| . Unexported (private) | 6,885 |
| Unit test functions (`TestXxx`) | 17,143 |
| Subtests (`t.Run(...)`) | 6,019 |
| End-to-end test functions | 1,199 |
| Source functions | 10,087 |
| . Exported (public) | 3,186 |
| . Unexported (private) | 6,901 |
| Unit test functions (`TestXxx`) | 17,183 |
| Subtests (`t.Run(...)`) | 6,034 |
| End-to-end test functions | 1,201 |

### Ratios worth noting

| Observation | Value |
| ---------------------------------- | -------------------------: |
| Test lines vs source lines | 1.97× more tests than code |
| Average source file length | ~222 lines |
| Average test file length | ~644 lines |
| Comment lines in source | 65,408 (~22.8% of source) |
| Average test file length | ~643 lines |
| Comment lines in source | 65,795 (~22.8% of source) |
| Test functions per source function | 1.7× |

### Code patterns

| Pattern | Count |
| ---------------------------------- | ----: |
| `if err != nil` checks | 9,186 |
| `defer` statements | 1,110 |
| `struct` types defined | 3,299 |
| `if err != nil` checks | 9,198 |
| `defer` statements | 1,114 |
| `struct` types defined | 3,304 |
| `//nolint` suppressions | 226 |
| `TODO` / `FIXME` / `HACK` comments | 1 |

Expand All @@ -497,15 +497,15 @@ and is never logged. Full details: [PRIVACY.md](PRIVACY.md).

| Record | File |
| ------------------- | --------------------------------------- |
| Longest source file | `cmd/server/main.go`. 4,746 lines |
| Longest test file | `cmd/server/main_test.go`. 11,210 lines |
| Longest source file | `cmd/server/main.go`. 4,826 lines |
| Longest test file | `cmd/server/main_test.go`. 11,326 lines |

### Because why not

| Fact | Value |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------- |
| Source code printed at 55 lines/page | ~5,224 pages of A4 |
| Source lines mentioning `"gitlab"` | 13,967 (impossible to avoid) |
| Source code printed at 55 lines/page | ~5,242 pages of A4 |
| Source lines mentioning `"gitlab"` | 13,979 (impossible to avoid) |
| Longest function name in source | `assertDynamicCompatibilityPolicyOwnedByActionCompat` (51 chars) |
| Longest test function name | `TestNewOperationIndex_TwoRoutesMountedAtOnePath_KeepTheFirstAnswerAndMergeThePagination` (87 chars) |

Expand Down
162 changes: 162 additions & 0 deletions cmd/server/auth_blocks.go
Original file line number Diff line number Diff line change
@@ -0,0 +1,162 @@
package main

import (
"fmt"
"log/slog"
"sync/atomic"
"time"

"github.com/jmrplens/gitlab-mcp-server/v3/internal/config"
"github.com/jmrplens/gitlab-mcp-server/v3/internal/mcpotel"
"github.com/jmrplens/gitlab-mcp-server/v3/internal/serverpool"
)

// authBlockCounters counts the requests each authentication budget refused,
// which is what telemetry publishes through [mcpotel.ObserveAuthBlocks].
//
// It is shared by the two guards rather than kept per guard, because they share
// the budgets themselves: a caller refused by the bearer guard and one refused
// by the gate behind it were refused by the same rule, and two series would
// have to be added back together to answer any question anybody asks.
//
// What is counted is the refusal, not the block. A block is raised once and
// then refuses every request that arrives while it lasts, so counting the
// raising would measure how often a sprayer starts and counting the refusals
// measures what the deployment is actually turning away. The second is the one
// an operator acts on.
//
// A nil *authBlockCounters counts nothing and is usable, which is what a guard
// built without telemetry wiring gets.
type authBlockCounters struct {
failureLockout atomic.Int64
transportSource atomic.Int64
distinctTokens atomic.Int64
}

// record counts one refusal under the given reason, which is one of the
// mcpotel.AuthBlock* values. A reason this does not know is not counted, so a
// new budget that forgets to add its counter here is silent rather than
// attributed to the wrong one.
func (c *authBlockCounters) record(reason string) {
if c == nil {
return
}
switch reason {
case mcpotel.AuthBlockFailureLockout:
c.failureLockout.Add(1)
case mcpotel.AuthBlockTransportSource:
c.transportSource.Add(1)
case mcpotel.AuthBlockDistinctTokens:
c.distinctTokens.Add(1)
}
}

// validateAuthBudgetBounds holds the four authentication-budget flags to the
// same bounds their environment spellings are held to.
//
// HTTP mode never runs (*config.Config).validate, so a bound enforced only
// there is a bound the flags escape, which is the class of defect
// [validateHTTPPoolAndRateBounds] exists to have fixed once. A negative value
// is refused rather than read as "off", because zero already says that and a
// minus sign is a typo.
func validateAuthBudgetBounds(cfg *config.Config) error {
for _, b := range []struct {
flag string
value int
max int
}{
{"--auth-failure-limit", cfg.AuthFailureLimit, config.MaxAuthFailureLimit},
{"--auth-distinct-token-limit", cfg.AuthDistinctTokenLimit, config.MaxAuthDistinctTokenLimit},
} {
if b.value < 0 {
return fmt.Errorf("%s must not be negative, got %d (0 disables the budget)", b.flag, b.value)
}
if b.value > b.max {
return fmt.Errorf("%s %d exceeds maximum of %d", b.flag, b.value, b.max)
}
}
for _, w := range []struct {
flag string
value time.Duration
max time.Duration
}{
{"--auth-failure-window", cfg.AuthFailureWindow, config.MaxAuthFailureWindow},
{"--auth-distinct-token-window", cfg.AuthDistinctWindow, config.MaxAuthDistinctWindow},
} {
if w.value < 0 {
return fmt.Errorf("%s must not be negative, got %s (0 disables the budget)", w.flag, w.value)
}
if w.value > w.max {
return fmt.Errorf("%s %s exceeds maximum of %s", w.flag, w.value, w.max)
}
}
return nil
}

// authFailureLimiter builds the fast per-address budget, or returns nil when
// the deployment turned it off.
//
// The nil is what "off" has to be. [serverpool.AuthRateLimiter] blocks once a
// record's count reaches its limit, so a limit of zero blocks an address after
// a single failure: the most aggressive setting there is, reached by typing the
// figure that every other budget here reads as "no budget". Every consulting
// site already tolerates a nil limiter, so refusing to build one is both the
// smallest change and the only one that cannot be read two ways.
func authFailureLimiter(cfg *config.Config) *serverpool.AuthRateLimiter {
if cfg.AuthFailureLimit <= 0 || cfg.AuthFailureWindow <= 0 {
return nil
}
return serverpool.NewAuthRateLimiter(cfg.AuthFailureLimit, cfg.AuthFailureWindow)
}

// authSprayBudget builds the distinct-token budget from the configuration, or
// returns nil when the deployment turned it off.
//
// The escalation step is the fast window rather than a setting of its own, so
// the ladder is one window, then ten, then sixty, and a deployment that
// shortens the window to watch the behavior shortens the whole ladder with
// it. At the defaults that is a minute, ten minutes and an hour.
func authSprayBudget(cfg *config.Config) *serverpool.DistinctTokenBudget {
return serverpool.NewDistinctTokenBudget(cfg.AuthDistinctTokenLimit, cfg.AuthDistinctWindow, cfg.AuthFailureWindow)
}

// observeAuthBlocks registers the refusal counters with OpenTelemetry.
//
// A failure is logged and not returned: telemetry that cannot be registered is
// a reason to say so, never a reason to refuse to serve. With telemetry off the
// global meter is a no-op and this costs one registration at startup.
func observeAuthBlocks(counts *authBlockCounters) {
if _, err := mcpotel.ObserveAuthBlocks(counts.counts); err != nil {
slog.Warn("authentication block metrics are not being exported", "error", err)
}
}

// retryAfterSeconds renders a block's remaining time for the Retry-After
// header, which RFC 9110 defines in whole seconds.
//
// It rounds up and never answers zero: a block with 400 milliseconds left
// truncates to 0, and "come back in no time at all" invites the client to
// retry immediately into the same refusal. One second is the smallest honest
// answer.
func retryAfterSeconds(d time.Duration) int {
if d <= 0 {
return 1
}
// Rounding up cannot produce less than one here: the smallest d this
// reaches is a nanosecond, and a nanosecond rounded up is a second. A
// floor check would be a second guard on that arithmetic, reachable by no
// input.
return int((d + time.Second - 1) / time.Second)
}

// counts reads the three counters, in the shape mcpotel publishes them.
func (c *authBlockCounters) counts() mcpotel.AuthBlockCounts {
if c == nil {
return mcpotel.AuthBlockCounts{}
}
return mcpotel.AuthBlockCounts{
FailureLockout: c.failureLockout.Load(),
TransportSource: c.transportSource.Load(),
DistinctTokens: c.distinctTokens.Load(),
}
}
Loading
Loading