AI GlossaryㅊTechnical words in the news
Time to First Token
A metric that measures how long it takes for the first character of an AI's response to appear after you send a query.
In plain words
Time to First Token (TTFT) refers to the time between the moment you send a query to an AI service and the moment the first character of its answer appears on screen. It's similar to timing how long it takes for the first side dish to arrive after you place an order at a restaurant. It doesn't measure how long it takes for the whole meal to come out, just the moment you first sense that something is happening.
The shorter this time is, the more responsive the AI feels to users. Even if generating the full answer takes a while, if the first character appears quickly, the person waiting feels less frustrated. That's why, in real-time interactive services like chatbots or coding assistants, this initial response speed is treated as being just as important as overall response speed.
Throughput and end-to-end latency, on the other hand, are different metrics that look at how fast and how much of an answer is generated overall. Time to First Token isolates specifically the part that corresponds to the user's first impression.
How it shows up in the news
The article mentions measuring "throughput, TTFT (Time to First Token), and end-to-end latency." Here, it's important not to confuse TTFT with how fast the service completes the entire response — it specifically refers only to the time it takes for the first response to appear on screen after a user submits a query.
Try it yourself
Try asking a chatbot a short question, and use a stopwatch to measure the time from the moment you hit enter to the moment the first character or cursor movement appears on screen. Then measure how long it takes for the entire answer to finish, and compare the two numbers — you'll get a concrete sense that initial response speed and overall completion speed are two different things.
See also
Stories using this term
- AWS integrates LLM inference optimization into SageMaker SDKAI · 2026.08.09
- F1 cuts data onboarding time from 8 weeks to 40 minutes with agentic AI on AWSAI · 2026.08.09
- 2026 Comparison of the Top 4 AI Video Generation APIsAI · 2026.08.05
- Gemini 3.7 Flash Benchmark: Score 56, 1.7 Minutes per TaskAI · 2026.08.14
- AI agent memory needs different prescriptions by model size to boost performanceAI · 2026.08.19
- GLM-5.3 API released, Terminal-Bench score jumps from 4.6 to 28.3AI · 2026.08.19
