Engineering

Latency Budgets: How We Keep Every Interaction Under 200ms

Perceived speed is the single most important predictor of whether a user will trust an AI assistant enough to integrate it into their daily workflow. Research consistently shows that interactions exceeding 400ms feel sluggish, while anything under 200ms feels instant. At Maple, we enforce a strict 200ms latency budget across every user-facing interaction path.

"You cannot optimize what you do not measure. Every Maple interaction emits a structured latency trace with microsecond-resolution timestamps for each phase."

This post explains the engineering systems we built to guarantee that budget, the tradeoffs we navigated, and the instrumentation that keeps us honest.

The latency budget framework

A latency budget works like a financial budget: you have a fixed total (200ms), and every subsystem gets an allocation. We decompose our total budget into five sequential phases: Hotkey Capture (10ms), Overlay Render (16ms), Context Assembly (80ms), Model Routing (30ms), and First Token Streaming (64ms).