Latency (AI)

From ALT-TEXT
Jump to navigation Jump to search

Latency (AI)

The delay between sending a request to an AI model and receiving its response, chiefly determined by model size, compute availability, and network conditions. Latency is a key practical constraint on where AI can realistically be used: real-time voice assistants and self-driving systems need very low latency, which is one reason smaller, locally run models are attractive for such applications. (See also: Compute, Small language model)