Speaker
Dr
Olivier Mattelaer
(UCLouvain/CISM)
Description
Checkpointing and restarting, or the art of stopping computations and resuming them later, is a powerful technique for overcoming job time limits, surviving hardware or software failures, and improving the robustness of long-running computations on HPC clusters. This session introduces the main checkpointing approaches and teaches how to use them effectively on CÉCI clusters.
| Contents | Information |
|---|---|
| • Uses and challenges of checkpointing • The different checkpointing approaches • Checkpointing support in Slurm • Using DMTCP for application-level checkpointing and restart • Best practices for long-running jobs |
Prerequisite: • Being able to use SSH with private keys • Being familiar with a text editor • Mastering the Linux command line and GNU utilities (mkdir, cp, scp, etc.) • Passive knowledge of either C, Fortran, Octave, Python, or R Type: Hands-on Target audience: Everyone Must: This session is a must-have for anyone feeling oppressed by time limits. |