Tuesday 14 April 2026 9:30am
Research Informatics Training Room, Craik-Marshall Building
About
Dates: Tue 14 and Wed 15 April 2026
Time: 09:30 - 17:30
Location: Research Informatics Training Room, Craik-Marshall Building
♿ The training room is located on the first floor and there is currently no wheelchair or level access.
Participants can make use of the computers in the training room, unless otherwise advised. Instructions on specific system requirements and downloads will be provided when a booking is secured.
Please ensure you meet the target audience criteria and prerequisites before registering for a course.
DescriptionHigh-throughput data analyses usually involve many data processing steps, including the use of a range of command line tools and scripts to transform, filter, aggregate and visualise data. Each tool may require a specific set of inputs and options to be defined and, as we chain multiple tools together, this can become challenging to manage.
The Nextflow workflow management system is a tool to create reproducible and scalable data analyses workflows. Workflows can be seamlessly scaled to server, cluster, and cloud environments, without the need to modify the workflow definition. These workflows can also include a description of the required software, which will be automatically deployed to any execution environment.
This course will cover the principles for building workflows using Nextflow, as well as advanced strategies to fully customise, automate and scale your analysis. It will also cover how to take advantage of modules and workflow design principles from the nf-core initiative, allowing you to build robust and high-quality pipelines following current best practices.
How to book Target audience- This course is aimed at researchers who perform computational data analysis and wish to build reproducible, scalable pipelines using Nextflow.
- It is suitable for participants who have experience running command line tools and scripting analyses, but who are new to workflow management systems or have limited experience with Nextflow.
- The course is not designed for complete beginners in programming or the UNIX shell
Essential
- Command line competence: Navigate directories, manage files, inspect file contents, run command line tools with options. Prior experience writing basic shell scripts is also advantageous.
- Scripting experience: Familiarity with at least one scripting language (e.g. Python, R, or Bash). Participants should be comfortable editing text scripts, using variables, and understanding simple control flow.
Desirable
- Data formats: Familiarity with tabular file formats (e.g. CSV, TSV/TXT). Familiarity with common bioinformatic file formats (e.g. FASTQ, BAM, VCF) is advantageous but not required.
- HPC environments: Experience using job schedulers (e.g. SLURM, PBS, LSF), and running jobs on shared compute resources.
- Software packaging: Familiarity with Conda, Mamba, or similar package managers.
- Container systems: Basic understanding of Docker or Singularity/Apptainer concepts (images, containers), as these are commonly used for software provisioning in Nextflow.
- Reproducible analysis practices: Experience structuring computational projects or pipelines, even informally (e.g. a collection of interdependent bash scripts).
Fees must be paid at registration.
Free for registered University of Cambridge students
£ 65/full day for all University of Cambridge staff, including postdocs.
£ 65/full day for all academic participants from external Institutions and charitable organizations.
£ 130/full day for all Industry participants.
For further information about the courses, please email the Research Informatics Training Team.