METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅇTechnical words in the news

End-to-End Latency

The total time it takes from the moment a request is sent until the complete response is received.

In plain words

End-to-end latency is the total time from the moment you send a request until you fully receive the result you wanted. It's like measuring the whole time from placing an order at a restaurant, through the kitchen cooking it, to the food finally being served in front of you. No matter how fast the kitchen is, if serving is slow, the customer still feels like it took a long time.

In AI services, this time breaks down into several stages: the time it takes for the request to reach the server (or device) and start being processed, the time to generate the answer, and the time to deliver that answer to the screen — all added together in sequence. So making just one stage faster doesn't necessarily make the whole thing faster; you need to measure the entire process from start to finish to know the speed users actually experience.

This becomes especially important when running an AI model directly on a phone or laptop. Since processing has to be completed on the device itself without going through an internet server, end-to-end latency can vary greatly depending on which compression method is used and which runtime the same model is loaded onto.

How it shows up in the news

Articles explain this concept by measuring "how many seconds it takes to input a 1,024-token prompt and receive 256 tokens back." A common misconception here is thinking that only the moment the answer starts matters — but end-to-end latency refers to the time until the entire answer is complete, not just the first character of the response.

Try it yourself

Type any question into a chatbot app and time how long it takes from the moment you hit send until the response completely stops. Try shortening or lengthening the same question and timing it again — you'll get a feel for how the length of the request and the response affect the total time.

See also

Stories using this term

Browse every entry