The time delay between sending a request to an AI model and receiving a response.
Latency matters significantly for user-facing AI applications — a technically accurate response that takes 30 seconds may be a worse product experience than a slightly less thorough one that returns in 2.