Cut p99 serving latency in half
ML designHardCorroborated · 4 reportsLast seen Aug 27, 2026
The question
An inference service meets its p50 target but misses p99 by 3x. Walk through how you would diagnose it, then which of batching, caching, quantisation, or admission control you would reach for and why.
No write-up yet
We have the question but not yet a full breakdown. If you were asked this, the fastest way to improve the page is to say what the interviewer pushed on.
Add what you remember →4 people reported this question.
I was asked this tooAlso reported at Anthropic, Netflix, Perplexity, xAI.