Supercomputers and their workloads have grown steadily more heterogeneous, and job scheduling has moved from static heuristics to Deep Reinforcement Learning (DRL) policies, though no consensus exists on which algorithm suits the task. Comparisons to date benchmark DRL policies only against heuristics, leaving open a prior question: how much of the reported performance comes from the learned policy, and how much from the action masking that makes such policies runnable at all.
This paper compares six DRL treatments (PPO, A2C, and DQN, each with and without action masking) against three Slurm heuristics and, critically, a seed-paired masked uniform-random control. The comparison runs across ten seeds on homogeneous and heterogeneous Slurm traces under a non-parametric, rank-based framework.
Papers, presentations, and recordings for each submission milestone.
April 2026
Opening submission, setting out the problem statement, the literature review findings, the proposed methodology, and the progress so far.
May 2026
Progress update covering the DRL training infrastructure, the first training results, and a preliminary statistical analysis.
August 2026
Results and the full paper: the complete statistical analysis, the critical difference diagrams, and the evaluation of all ten treatments across both traces.
September 2026
The final paper, alongside a cleaned-up project repository and a wiki covering how to run the pipeline end to end.
Researchers and supervisors contributing to this project.
Ten treatments, two real Slurm traces, ten fixed seeds, and a rank-based framework, chosen because normality does not hold at n = 10.
stable-baselines3 and
sb3-contrib; MaskableA2C and MaskableDQN are custom. Three Slurm heuristics:
SJF, LCFS, and UNICEP. One seed-paired masked uniform-random control.
best_fit allocator and backfill stay
fixed as controlled factors, so the comparison reflects scheduling policy rather than allocation strategy.
nixos-25.05 makes every dependency content-addressed, Apptainer
containerizes the environment for cluster execution, and Snakemake 9.4.3 runs the pipeline. Every output
carries the commit hash and a sha256 manifest fingerprint in its metadata.
What separates the two workloads is headroom: the gap between the best heuristic and the random control. Headroom is how much room a learned policy has to show a benefit.
The best heuristic, LCFS, reaches 2,090 s against the random control's 2,166 s. That leaves almost no room to optimize, and no DRL treatment beat the control.
The best heuristic, UNICEP, reaches 846 s against the random control's 1,170 s. That is roughly ten times more room, and MaskablePPO beat the control on every primary metric.
| Treatment | Physical (s) | Deeplearn (s) |
|---|---|---|
| MaskablePPO | 2,243 ± 41 | 846 ± 9 |
| MaskableA2C | 2,223 ± 42 | diverged a |
| MaskableDQN | 4,474 ± 1,684 | 45,729 ± 40,335 |
| SJF | 2,176 | 851 |
| LCFS | 2,090 | 912 |
| UNICEP | 2,499 | 846 |
| Random (masked) | 2,166 ± 76 | 1,170 ± 26 |
Mean average waiting time in seconds (± standard deviation across ten seeds) on the dev split, with the heuristics run without backfill. Lower is better.
a MaskableA2C diverged on 2 of 10 seeds on the heterogeneous trace, which puts its mean at 24,520,863 s. On the median it reaches 837 s and outperforms MaskablePPO. It is capable of the better result, but not reliably. This table omits the three unmasked treatments, because none completed an evaluation episode and truncation biases their figures.
MaskablePPO is the best-ranked treatment that completes an evaluation episode on both traces. The differences between the algorithms, however, are not significant in this data.
Three questions the study set out to answer, and what the data supports.
Deployability, not schedule quality.
No unmasked policy terminated an evaluation episode on either trace. Without a mask, the policy can select an action that names no allocatable job. The simulator then freezes on an identical state, and a hang guard stops the run. Masking removes that failure mode, but it does not remove divergence: MaskableDQN still averaged 4,474 ± 1,684 s of waiting on the physical trace, with a maximum slowdown above 190,000.
0 of 6 unmasked runs completed an episode · masking removes one failure mode but leaves the others
Not on schedule quality.
A PPO-family treatment topped both traces but separated from the other leading treatments on neither, so the data supports no claim of better schedule quality. Cost is the axis where the family does hold an advantage: both PPO treatments trained 9% to 22% faster than the A2C variants. That advantage follows from an eightyfold lower gradient-update count under inherited defaults, which makes it a budget artifact rather than algorithmic efficiency.
Nemenyi critical difference = 2.39 rank units (n = 10, k = 6) · no treatment separates
On heterogeneous workloads, and only for tail fairness.
On the homogeneous trace the agents could not beat a uniform draw over the same valid actions, so the heuristics win outright. On the heterogeneous trace MaskablePPO holds parity with the heuristics on average waiting and improves maximum slowdown by up to 22%. It does not schedule the average job better; it handles the awkward ones better. That costs 13 to 34 times the evaluation time of the fastest heuristic.
Parity on the average job · −22.0% on max slowdown vs. LCFS · 13 to 34 times slower to run
A masked random control, paired by seed, is inexpensive and decisive. Without it, this study would have reported the parity its masked treatments reach with production heuristics on the physical trace as a success
This literature should adopt it as a required baseline.