Speaker
Damien François
(UCLouvain/CISM)
Description
DataLad is a data management and distribution tool built on top of Git and git-annex. It enables researchers to track datasets, code, and computational workflows in a reproducible manner while efficiently handling large files that would be impractical to store directly in Git. This session introduces DataLad as a solution for reproducible research, collaborative projects, and large-scale scientific data management on HPC systems.
| Contents | Information |
|---|---|
| • Motivation for reproducible data management • Introduction to Git, git-annex, and DataLad • Creating and managing DataLad datasets • Tracking large datasets efficiently • Cloning, sharing, and publishing datasets • Organizing code, data, and results in a reproducible workflow • Working with remote storage and HPC filesystems • Data provenance and dataset versioning • Collaborative workflows with DataLad • Practical examples for scientific projects |
Prerequisite: • Being able to use SSH with private keys • Familiarity with the Linux command line • Basic knowledge of Git concepts (commits, branches, repositories) Type: Hands-on Target audience: Researchers, data scientists, and software developers managing datasets on HPC systems Must: Strongly recommended for anyone interested in reproducible research, collaborative data management, or versioning large scientific datasets. |