8 October 2026 to 18 February 2027
Europe/Brussels timezone

Stop and Restart program: (how to beat walltime on HPC)

19 Nov 2026, 09:30
1h 30m
CYCL09a

CYCL09a

chemin du cyclotron 2, 1348 Louvain-La-Neuve

Speaker

Dr Olivier Mattelaer (UCLouvain/CISM)

Description

Checkpointing and restarting, or the art of stopping computations and resuming them later, is a powerful technique for overcoming job time limits, surviving hardware or software failures, and improving the robustness of long-running computations on HPC clusters. This session introduces the main checkpointing approaches and teaches how to use them effectively on CÉCI clusters.

Contents Information
• Uses and challenges of checkpointing
• The different checkpointing approaches
• Checkpointing support in Slurm
• Using DMTCP for application-level checkpointing and restart
• Best practices for long-running jobs
Prerequisite:
• Being able to use SSH with private keys
• Being familiar with a text editor
• Mastering the Linux command line and GNU utilities (mkdir, cp, scp, etc.)
• Passive knowledge of either C, Fortran, Octave, Python, or R

Type: Hands-on
Target audience: Everyone
Must: This session is a must-have for anyone feeling oppressed by time limits.

Presentation materials

There are no materials yet.