Latency (AI)

From ALT-TEXT
Revision as of 10:36, 7 September 2026 by imported>ALT-TEXT (Import: AI terminology and people glossary)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigation Jump to search

Latency (AI)

The delay between sending a request to an AI model and receiving its response, chiefly determined by model size, compute availability, and network conditions. Latency is a key practical constraint on where AI can realistically be used: real-time voice assistants and self-driving systems need very low latency, which is one reason smaller, locally run models are attractive for such applications. (See also: Compute, Small language model)