METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

AI GlossaryㅌInfrastructure and chips

topology-aware scheduler

A scheduling method that assigns jobs to the most efficient locations by understanding the physical layout and connections between GPUs and servers

In plain words

A topology-aware scheduler doesn't just look for an empty slot when deciding where to place a computing job—it also considers how servers and GPUs are actually connected to each other. It's similar to seating team members working on the same project close together in an office, so they spend less time walking to meetings. Because physically nearby GPUs can exchange data much faster, grouping the GPUs used for the same job as close together as possible boosts overall processing speed.

A typical basic resource manager just checks for free slots and drops jobs in. A topology-aware scheduler, by contrast, factors in the GPU connection structure, pre-set quotas for each team, and real-time available resources when deciding placement. This means that even when multiple teams share several physical servers, each team's jobs avoid interfering with one another and can efficiently receive their fair share of resources.

How it shows up in the news

The article explains that "KAI Scheduler is a topology-aware scheduler designed for AI workloads, supporting per-team quotas and dynamic GPU allocation." A common misunderstanding here is assuming it completely replaces Kubernetes' default scheduler—in reality, it operates alongside the default scheduler, handling only jobs specifically designated to use it while the default scheduler continues to handle everything else.

See also

Stories using this term

Browse every entry