Honours Research Project · 2026

deep reinforcement learning for HPC job scheduling: a statistical evaluation

Masking buys deployability, not schedule quality.

Supercomputers and their workloads have grown steadily more heterogeneous, and job scheduling has moved from static heuristics to Deep Reinforcement Learning (DRL) policies, though no consensus exists on which algorithm suits the task. Comparisons to date benchmark DRL policies only against heuristics, leaving open a prior question: how much of the reported performance comes from the learned policy, and how much from the action masking that makes such policies runnable at all.

This paper compares six DRL treatments (PPO, A2C, and DQN, each with and without action masking) against three Slurm heuristics and, critically, a seed-paired masked uniform-random control. The comparison runs across ten seeds on homogeneous and heterogeneous Slurm traces under a non-parametric, rank-based framework.

Institution University of the Western Cape
Programme BSc (Hons) Computer Science
Supervisor Prof. Clement Nyirenda
Year 2026
Computation UWC High Performance Computing and Research Cloud
10 Treatments compared
120 Evaluated DRL runs
10 Seeds per treatment
153K Slurm jobs replayed

Research Deliverables

Milestones

Papers, presentations, and recordings for each submission milestone.

Submission 1

Available

April 2026

Opening submission, setting out the problem statement, the literature review findings, the proposed methodology, and the progress so far.

Submission 2

Available

May 2026

Progress update covering the DRL training infrastructure, the first training results, and a preliminary statistical analysis.

Submission 3

Available

August 2026

Results and the full paper: the complete statistical analysis, the critical difference diagrams, and the evaluation of all ten treatments across both traces.

Paper (PDF) Slides (PDF)

Final Submission

In progress

September 2026

The final paper, alongside a cleaned-up project repository and a wiki covering how to run the pipeline end to end.

Paper (PDF) Slides (PDF) Recording Demo

The Team

People

Researchers and supervisors contributing to this project.

Justin M. Cheney

Honours Student

BSc Computer Science
Department of Computer Science
University of the Western Cape

Prof. Clement Nyirenda

Supervisor

PhD Computational Intelligence and Systems Science
Department of Computer Science
University of the Western Cape

Method

How it was run

Ten treatments, two real Slurm traces, ten fixed seeds, and a rank-based framework, chosen because normality does not hold at n = 10.

Treatments
Six DRL algorithms: PPO, A2C, and DQN, each with and without action masking. The three unmasked algorithms and MaskablePPO come from stable-baselines3 and sb3-contrib; MaskableA2C and MaskableDQN are custom. Three Slurm heuristics: SJF, LCFS, and UNICEP. One seed-paired masked uniform-random control.
Environment
HPCSim, a scheduling environment built on Gymnasium. The best_fit allocator and backfill stay fixed as controlled factors, so the comparison reflects scheduling policy rather than allocation strategy.
Traces
Physical, homogeneous: 84,135 jobs over 37 days. Deeplearn, heterogeneous: 68,720 jobs over 382 days. A 70/30 split by submission time divides both traces, so no model sees the holdout during training.
Statistics
A Friedman omnibus test per metric per trace, reporting Kendall's W alongside every p-value. Nemenyi post-hoc comparisons with critical difference plots. Wilcoxon confidence intervals for effect size, TOST equivalence against a ±10% margin, and Holm correction within each comparison family.
Reproducibility
A Nix flake pinned to nixos-25.05 makes every dependency content-addressed, Apptainer containerizes the environment for cluster execution, and Snakemake 9.4.3 runs the pipeline. Every output carries the commit hash and a sha256 manifest fingerprint in its metadata.

Results

Two traces

What separates the two workloads is headroom: the gap between the best heuristic and the random control. Headroom is how much room a learned policy has to show a benefit.

3.7% Physical · homogeneous

The best heuristic, LCFS, reaches 2,090 s against the random control's 2,166 s. That leaves almost no room to optimize, and no DRL treatment beat the control.

38.3% Deeplearn · heterogeneous

The best heuristic, UNICEP, reaches 846 s against the random control's 1,170 s. That is roughly ten times more room, and MaskablePPO beat the control on every primary metric.

Average waiting time by treatment

Treatment Physical (s) Deeplearn (s)
MaskablePPO 2,243 ± 41 846 ± 9
MaskableA2C 2,223 ± 42 diverged a
MaskableDQN 4,474 ± 1,684 45,729 ± 40,335
SJF 2,176 851
LCFS 2,090 912
UNICEP 2,499 846
Random (masked) 2,166 ± 76 1,170 ± 26

Mean average waiting time in seconds (± standard deviation across ten seeds) on the dev split, with the heuristics run without backfill. Lower is better.

a MaskableA2C diverged on 2 of 10 seeds on the heterogeneous trace, which puts its mean at 24,520,863 s. On the median it reaches 837 s and outperforms MaskablePPO. It is capable of the better result, but not reliably. This table omits the three unmasked treatments, because none completed an evaluation episode and truncation biases their figures.

Critical difference diagrams

Critical difference diagram for average waiting time on the physical trace. MaskablePPO and MaskableA2C hold the top two mean ranks and are joined by a bar indicating no significant difference between them.
Physical trace MaskablePPO (1.7) and MaskableA2C (1.9) take the top two mean ranks, joined by a crossbar. Nemenyi returns p = 1.000 between them, so they cannot be ordered at n = 10.
Critical difference diagram for average waiting time on the deeplearn trace, showing the treatments' mean ranks and the groups that are not significantly different.
Deeplearn trace MaskablePPO is the best-ranked treatment that completes an evaluation episode. No other treatment separates from it at n = 10.

MaskablePPO is the best-ranked treatment that completes an evaluation episode on both traces. The differences between the algorithms, however, are not significant in this data.

Findings

Research questions

Three questions the study set out to answer, and what the data supports.

  1. RQ1

    what does action masking contribute?

    Deployability, not schedule quality.

    No unmasked policy terminated an evaluation episode on either trace. Without a mask, the policy can select an action that names no allocatable job. The simulator then freezes on an identical state, and a hang guard stops the run. Masking removes that failure mode, but it does not remove divergence: MaskableDQN still averaged 4,474 ± 1,684 s of waiting on the physical trace, with a maximum slowdown above 190,000.

    0 of 6 unmasked runs completed an episode · masking removes one failure mode but leaves the others

  2. RQ2

    is the field's concentration on PPO justified?

    Not on schedule quality.

    A PPO-family treatment topped both traces but separated from the other leading treatments on neither, so the data supports no claim of better schedule quality. Cost is the axis where the family does hold an advantage: both PPO treatments trained 9% to 22% faster than the A2C variants. That advantage follows from an eightyfold lower gradient-update count under inherited defaults, which makes it a budget artifact rather than algorithmic efficiency.

    Nemenyi critical difference = 2.39 rank units (n = 10, k = 6) · no treatment separates

  3. RQ3

    when is a DRL scheduler worth its inference cost?

    On heterogeneous workloads, and only for tail fairness.

    On the homogeneous trace the agents could not beat a uniform draw over the same valid actions, so the heuristics win outright. On the heterogeneous trace MaskablePPO holds parity with the heuristics on average waiting and improves maximum slowdown by up to 22%. It does not schedule the average job better; it handles the awkward ones better. That costs 13 to 34 times the evaluation time of the fastest heuristic.

    Parity on the average job · −22.0% on max slowdown vs. LCFS · 13 to 34 times slower to run

The strongest claim is methodological

A masked random control, paired by seed, is inexpensive and decisive. Without it, this study would have reported the parity its masked treatments reach with production heuristics on the physical trace as a success

This literature should adopt it as a required baseline.