Skip to content

Commit e2879fa

Browse files
committed
Rate limit ADR: new measurements
- implementation on #5361 was improved and reaches more throughput now - all measurements run for 300s - more detailed explanation how to determine max throughput - some more arguments for replacing nginx by a dedicated CAPI router
1 parent 19c213c commit e2879fa

1 file changed

Lines changed: 23 additions & 22 deletions

File tree

‎decisions/0016-rate-limiting.md‎

Lines changed: 23 additions & 22 deletions
Original file line numberDiff line numberDiff line change
@@ -29,11 +29,11 @@ CF API rate limiting should be moved from CCNG (Ruby process) into a dedicated r
2929

3030
### Short-term
3131

32-
A single user can't overload the CF API anymore by running more parallel requests than the CC api VMs can process. The CF API is protected up to ~750 req/s (~1.8k with token caching) per CC VM (max 429 response rate).
32+
A single user can't overload the CF API anymore by running more parallel requests than the CC api VMs can process. The CF API is protected up to ~2.2k req/s per CC VM (max 429 response rate).
3333

3434
### Long-term
3535

36-
The protection level can be increased at least by factor 4.
36+
The protection level can be substantially increased (we expect at least factor 2).
3737

3838
CCNG is offloaded by token decoding/validation and rate limiting middleware. We may be able to remove Redis/Valkey from CCNG which was introduced because of rate limiting.
3939

@@ -47,14 +47,15 @@ Downside is the implementation and testing effort but we also gain full control
4747

4848
[ccng #5361](https://github.com/cloudfoundry/cloud_controller_ng/pull/5361)
4949

50-
Improves also also the middleware order to increase throughput when rate limits are hit.
50+
The PR also improves the middleware order to increase throughput when rate limits are hit.
5151

5252
Pros:
5353
- simple
5454
- big improvement compared to no concurrency rate limiting
5555

5656
Cons:
57-
- limited protection (max ~750 req/s per api VM, max 1.8k with token caching)
57+
- less protection compared to a rate limiter outside of CCNG (max 2.2k req/s per api VM)
58+
- rate-limited requests consume CCNG resources
5859

5960
### Rate limiting in nginx using OpenResty
6061

@@ -92,7 +93,7 @@ Cons:
9293

9394
### Rate limiting in nginx using a golang server for token decoding
9495

95-
Aborted POC implementation.
96+
POC implementation started but then aborted.
9697

9798
Pros:
9899
- no Redis/Valkey needed for rate limiters
@@ -108,7 +109,7 @@ Cons:
108109

109110
### Rate limiting in nginx using token decoding module in C
110111

111-
Aborted POC implementation.
112+
POC implementation started but then aborted.
112113

113114
Pros:
114115
- no Redis/Valkey needed for rate limiters
@@ -122,13 +123,13 @@ Cons:
122123

123124
### Replacing nginx by a CAPI specific go implementation
124125

125-
No POC yet. Ideas is to replace nginx by a CAPI specific router implementation in golang.
126-
The new capi router has to implement routing rules, mTLS, rate limiting and file upload.
126+
No POC yet. Idea is to replace nginx by a CAPI specific router implementation in golang.
127+
A new capi router has to implement routing rules, mTLS, rate limiting and file upload.
127128

128129
Pros
129130
- CAPI specific implementation without hacks
130131
- implementation in a well-understood and performant language (no Lua or JavaScript)
131-
- should perform even better than nginx because of in-process token decoding (to be validated)
132+
- should perform better than nginx because of in-process token decoding and validation (to be validated)
132133
- no Redis/Valkey needed for rate limiters
133134
- offloads CCNG (token decoding/validation and rate limiting)
134135
- no dependency to nginx and nginx modules, removes unmaintained nginx upload module
@@ -150,39 +151,39 @@ Test setup:
150151
- vegeta on external VM
151152
- 300s attack time
152153
- all requests by a single user
153-
- assumed an API response time of 100ms (i.e. a rather slow endpoint), simulated by adding a 100ms delay to the /v3/info endpoint
154+
- assumed an API response time of ~106ms (i.e. a rather slow endpoint), simulated by adding a 100ms delay to the /v3/info endpoint
154155
- rate limiting (per VM)
155156
- fixed-window rate limiter with very high limit so that it has no effect
156157
- concurrent request rate limiter with max 10 parallel requests per user
157158

158-
Table shows highest achievable request rate w/o 5xx responses.
159+
The max measured 200 response throughput for this test setup is ~189 req/s (= 56,700 200 responses in 300s).
159160

160-
- TODO: Update data. Some tests were run until reaching 40k 200, should rerun for 300s with fixed-window rate limiter set to very high limit so that we get comparable results.
161+
The table shows the highest achievable attack request rate for which the CF API is still working stable:
162+
- no 5xx responses
163+
- achieves at least 90% of the max 200 responses (~51k 200 responses in 300s, ~170 req/s).
161164

162165
| Implementation | Rate | Duration | 200s | 429s | 5xx |
163166
|---|---|---|---|---|---|
164167
| Baseline (1) |600 req/s | 300s | 180,000 | 0 | 0 |
165-
| CCNG middleware |1200 req/s | 183s | 40,000 | 179,100 | 0 |
166-
| CCNG middleware w. token caching |3500 req/s | 300s | 37,534 | 1,012,415 | 0 |
167-
| nginx OpenResty |5250 req/s| 238s | 40,000 | 1,178,881 | 0 |
168-
| nginx njs |5250 req/s| 246s | 39,295 | 1,253,393 | 0 |
168+
| CCNG middleware |4000 req/s | 300s | 52,140 | 1,147,873 | 0 |
169+
| nginx njs |4750 req/s| 300s | 51,771 | 1,373,232 | 0 |
170+
| nginx OpenResty |5250 req/s| 300s | 53,143 | 1,521,858 | 0 |
169171

170172
(1) Baseline = current implementation w/o concurrent request rate limiter.
171173

172-
Additional measurement of rate limiting throughput, i.e. all requests are rejected with 429. Concurrent request rate limiter set to 0:
174+
Additional measurement of max rate limiting throughput, i.e. all requests are rejected with 429. Concurrent request rate limiter set to 0:
173175

174176
| Implementation | Rate | Duration | 200s | 429s | 5xx |
175177
|---|---|---|---|---|---|
176178
| Baseline (2) | 1050 req/s | 300s | - | 315,000 | 0 |
177-
| CCNG middleware | 1500 req/s | 300s | - | 450,000 | 0 |
178-
| CCNG middleware w. token caching |3650 req/s | 300s | - | 1,095,002 | 0 |
179-
| nginx OpenResty | 6000 req/s | 300s | - | 1,800,007 | 0 |
179+
| CCNG middleware | 4500 req/s | 300s | - | 1,353,000 | 0 |
180180
| nginx njs | 5750 req/s | 300s | - | 1,725,001 | 0 |
181+
| nginx OpenResty | 6000 req/s | 300s | - | 1,800,007 | 0 |
181182

182183
(2) Baseline uses fixed-window rate limit with limit 0 instead of concurrent request rate limiter
183184

184185
## Additional Information
185186

186-
The rate limiter protection performance (i.e. which max load can be responded with 429) can be further improved by caching tokens that have been validated and decoded. This applies to all implementation options. POC was only done for CCNG middleware.
187+
The rate limiter protection performance (i.e. which max load can be responded with 429) can be further improved by caching tokens that have been validated and decoded. This applies to all implementation options.
187188

188-
Max nginx response rate is ~8..10k req/s per CC api VM (static response by nginx w/o token decoding and rate limiting). It might be possible to increase it further by tuning nginx configuration.
189+
Max nginx response rate is ~8..10k req/s per CC api VM (static response by nginx w/o token decoding and rate limiting). It might be possible to increase it further by tuning nginx configuration. This indicates that a CAPI specific and properly optimized rate limiting implementation outside of CCNG can provide a higher protection level than an implementation within CCNG.

0 commit comments

Comments
 (0)