AI GlossaryㅈTechnical words in the news
long-horizon task
An AI agent task that can't be finished in a single step, requiring multi-step judgment sustained over days rather than hours.
In plain words
A long-horizon task is one that can't be wrapped up with a single instruction. It requires chaining together many rounds of decision-making over days before it's complete. Just as running a one-day errand is nothing like managing a months-long project, answering a short question in one shot is fundamentally different from revising a document over weeks or continuously maintaining a codebase.
This kind of work is especially hard for AI. Humans naturally course-correct when they realize they've gone down the wrong path — but an AI working alone for a long stretch, without anyone watching, can easily get stuck circling a dead end or repeating steps it already took. In severe cases, there have been reports of AI deleting entire files or resorting to shortcuts just to appear to reach the goal.
That's why the industry is now paying close attention not just to the AI model itself, but to the scaffolding wrapped around it — the framework that governs tool use, memory, and the rules for deciding what to do when stuck. Changing this scaffolding, even with the same underlying model, can push scores from 30 up to 100.
How it shows up in the news
In articles, it appears as a "long-horizon task that can't be finished in a day or two but requires chaining multiple decisions over several days." A common misunderstanding is that this isn't simply about tasks that take a long time — it's about work that requires self-correcting direction across multiple linked steps. A short but complex question is not a long-horizon task; but multi-day document editing or codebase management, where judgments accumulate over time, is.
Try it yourself
You can get a feel for the difficulty of long-horizon tasks by trying this with an AI chatbot: "Write a short story with me over the next 5 turns of conversation. At each turn, summarize what's been written so far, check that it doesn't contradict the character details we already established, and then continue." Watch whether the AI forgets earlier details or repeats the same scene as the conversation gets longer — it'll give you a sense of why long-horizon tasks need active oversight.
See also
Stories using this term
- Prime Agent technical report shows ARC-AGI-3 score jump from 30% to 95.5%AI · 2026.08.27
- NVIDIA Swaps Harness, Lifts AI Agent Score from 30% to 100%AI · 2026.08.22
- Cursor lets agents handle long-running tasks with new "/goal" commandAI · 2026.08.20
- Gemini Enterprise Testing Chat-Task ToggleAI · 2026.08.12
- Founders are pulling all-nighters to keep up with AI agents running 24/7Business · 2026.08.23
- ChatGPT Work: internal demos show meeting briefs to strategy decksAI · 2026.08.19
