Skip to main content
L14.8

GPU Scheduling Basics

Goal

Place inference work on GPUs using memory, health, queue, and data-location limits instead of treating a GPU as unlimited.

Fit the request before optimizing the queue​

Imagine assigning students to worktables. One table has no line but only a tiny empty corner. Another has two students waiting but plenty of space for a large poster. If the next student needs a large poster area, choosing the shortest line first would be a mistake: the first table simply cannot fit the work.

GPU scheduling begins with the same kind of hard constraint. Suppose GPU A has 2 GB of safe free memory and GPU B has 8 GB. A request is estimated to need 4 GB of additional state. GPU A is not a valid target even if its queue is empty. Only after removing targets that cannot safely fit the request does it make sense to compare queue length, locality, or priority.

request needs 4 GB additional safe memory
GPU A: 2 GB free → not feasible
GPU B: 8 GB free → feasible
only then compare queue length and priority

A GPU has its own memory and compute resources. Moving data to or from it also costs time. Even when a serving runtime handles low-level kernels, the scheduler still needs a simple resource model.

Memory and compute pressure differ​

Start with memory. A replica must fit the model, runtime workspace, and request state. If a long request needs more KV-cache memory than a replica can safely provide, idle compute on that GPU does not make the request fit.

Compute pressure and memory pressure are different. A service can run out of memory while compute is idle, or have spare memory while compute is busy. Watch both kinds of pressure.

Placement changes transfer and failure costs​

With several GPUs, placement matters. Different parallel strategies split or copy work in different ways. You do not need to choose one universal strategy here. Remember that placement changes communication cost, memory use, and what fails together.

Locality means keeping data near the compute that uses it. Moving data between host and GPU, or between GPUs, consumes time and bandwidth. Avoid needless transfers on the request's critical path.

Shared GPUs also need resource reservations. The scheduler must know how much capacity is actually available, and the serving process still needs its own admission limits.

Health is a scheduling constraint​

Health is a hard scheduling input. A replica with repeated out-of-memory errors, kernel failures, or failed readiness should stop receiving new work until it recovers.

Placement also changes failure risk. Sending everything to one fast replica creates a larger impact when that process fails. More replicas can improve resilience and make rollouts easier.

Choose among feasible targets​

For each request, ask: Which replica can fit it? Which is healthy? Which already has queued work? Then choose among the safe options.

Suppose three requests arrive together. One is short, one has a large expected KV cache, and one belongs to a latency-sensitive endpoint. GPU A has the shortest queue but little free memory; GPU B has more headroom but is already busy. A safe scheduler first removes targets that cannot fit each request. It then considers health, queue delay, locality, and service priority among the feasible targets. This order matters. Optimizing queue length before checking memory can send work to a replica that is fast only until it crashes. Scheduling is therefore a constrained choice, not simply “pick the least busy GPU.”

Predict

GPU A has lower queue length but insufficient safe memory for a long request; GPU B has more queue work but enough memory. Which should admission prefer?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.

Run:

python3 labs/notebooks/level-14/l14-08-gpu-scheduling.py

The Lab checks memory feasibility before using queue length as a preference.

  1. Run it unchanged with request = {"cache_gb": 5.0}. gpu-a lacks enough headroom, so the request should be placed on gpu-b.
  2. Before editing, predict what the scheduler should do if the same request now needs 9 GB, more than either replica can safely hold.
  3. Change only "cache_gb": 5.0 to "cache_gb": 9.0, then rerun.
  4. Confirm no replica is selected and the Lab reports that the request must wait or be shed. A shorter queue never overrides an infeasible memory boundary.

Loading lab…

Quick Check

1. What must a simple GPU scheduler know before placing work?
2. Why does data locality matter?
3. How should an unhealthy replica affect scheduling?

0 of 3 questions answered.

Explain it back​

Given two GPUs with different free memory, queue depth, and health, describe an admission order for three requests with different cache estimates.

Key Takeaways

  • GPU scheduling starts with resource feasibility.
  • Memory and compute pressure are separate signals.
  • Placement can change communication and failure scope.
  • Health must influence routing.
  • Queue optimization comes after hard safety constraints.

Next Lesson

Next, L14.9 — Containers and Reproducible Deployments turns the serving environment into a versioned artifact.

References

Lesson actions

Completion is stored locally on this device.

View progress