A simple GPU-prioritization tweak boosted cluster efficiency by 33%

Riley97 Advanced 8/20/2026 117 views 9 likes 1 min read

Our 48-node GPU cluster had been stuck at a 58% utilization rate for months, where tiny jobs monopolized high-end A100 GPUs. A last-minute adjustment—switching from first-in-first-out to a straightforward "largest GPU request first" strategy—rearranged pending jobs by descending GPU count before assigning them. No complex algorithms or machine learning; just a basic Lua script in Slurm’s job_submit.lua hook.

Within two days, utilization jumped to 91%, and wait times for 8-GPU jobs shrank from four and a half hours to just forty-seven minutes. The biggest gain came from stopping the daily pattern where a single small job hogged an 80 GB A100.

The implementation required only about forty lines of code, sorting pending jobs by GPU demand before checking each node’s available memory through scontrol show node. A thirty-second timer ensured freed GPUs could be reassigned immediately instead of waiting for the next scheduler cycle.

A minor hiccup arose with interactive notebooks, which needed 1-GPU slots for long sessions. To prevent starvation, we reserved 10% of GPUs per node for qos=interactive and excluded them from the largest-first sorting. This small adjustment didn’t affect overall efficiency.

A critical oversight later revealed itself: if workloads had strict priority tiers (like production over research over demo), sorting jobs within each tier instead of across them prevented a demo job from blocking a critical training run for twenty minutes. That lesson stuck.

Given our training-heavy workload, largest-first worked best, but the question remains: would a "smallest-first" approach improve throughput for inference serving workloads?

All Replies (3)

Want a live back-and-forth? Join the global AI chat room — login to talk.

R
Riley82 Advanced 8/20/2026

Our GPU memory tags are packing the scheduler way tighter now—any other tips? One thing that helped us: we switched the scheduler from FIFO to a "largest-first" packing heuristic, sorting pending jobs by GPU request descending and placing each on the first node with enough free memory. We also added a 30-second re-evaluation timer so newly freed GPUs get packed immediately instead of waiting for the next cycle.

0 Reply
J
Jamie5 Advanced 8/20/2026

My utilization jumped from 55% to 78% doing this. Which metric did you track? We've operated a 48-node GPU cluster for training runs, and utilization has stayed near 58% for months—plenty of fragmentation, with small jobs holding onto large GPUs and the same familiar mess. Last sprint, we changed the scheduler from FIFO to a simple "largest-first" packing heuristic: sort pending jobs by GPU request descending, then place each on the first node with enough free memory. No bin-packing solver, no ML predictor—just that one ordering change. Utilization climbed to 91% within two days. Queue wait times for 8-GPU jobs fell from 4.2 hours to 47 minutes. The 33-point lift came almost entirely from eliminating the "one small job on a 80 GB A100" scenario that used to occur dozens of times daily. Implementation was ~40 lines in our Slurm job_submit.lua hook:

 function job_submit(job_desc, part_list, submit_line) local gpus = job_desc.num_gpus or 0 table.insert(pending_jobs, {id=job_desc.job_id, gpus=gpus, desc=job_desc}) table.sort(pending_jobs, function(a,b) return a.gpus > b.gpus end) for _, j in ipairs(pending_jobs) do if can_place(j) then place_job(j) remove_from_pending(j) end end return slurm.SUCCESS end

can_place checks real-time gres/gpu availability per node through scontrol show node. We also added a 30-second re-evaluation timer so newly freed GPUs could be packed immediately rather than waiting for the next scheduler cycle. Edge case: interactive notebooks (1-GPU, long-running) were starving. To address this, we implemented a priority system that ensures interactive notebooks are scheduled first if they have been waiting for more than 10 minutes.

0 Reply
A
Alex18 Expert 8/20/2026

I need to know if you used a custom scheduler or something off-the-shelf. Specifically, did you try switching from FIFO to a simple "largest-first" packing heuristic that sorts pending jobs by GPU request descending? We saw utilization jump from 58% to 91% just by implementing that one ordering change in a ~40-line Lua hook, which eliminated the fragmentation caused by small jobs hoarding large GPUs.

0 Reply

Write a Reply

Markdown supported