Commit Graph
1 Commits
Author SHA1 Message Date
Максименко Никита ВладимировичandClaude Sonnet 5 832738891c feat: self-tuning concurrency limit instead of a fixed Semaphore
Under 0.5-CPU containers, a static Semaphore(2000) never tripped —
latency ballooned to 1.5-2s instead of the service answering 429.
Runtime.availableProcessors() can't help pick a number either: it
ignores the cgroups --cpus quota and reports full host cores.

AdaptiveConcurrencyLimiter reacts to observed latency instead of
guessing capacity: starts at min-concurrent, grows by one per
adjustment window when latency stays under target, halves it the
moment it doesn't. Adjustment is gated by wall-clock time, not by
request count — an earlier per-request version let the limit race to
the ceiling in milliseconds under high RPS, before any real overload
had a chance to show up in the samples.

Verified under load (native image, 250MB/0.5 CPU): p50 latency at 3x
overload dropped from ~1.3s to under 4ms; normal-load p95 unaffected.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 21:29:57 +03:00