Courses

Structured material for building a skill deliberately rather than picking it up on the job. Some of it ramps colleagues up ahead of a project, while others are targeted at people taking their very first steps into data and AI.

Agent Skills for Data Practitioners

A half-day workshop on what Agent Skills are, how other data professionals use them, and how to build your own.

AI-Assisted Coding Workshop

Level up your engineering workflow with AI coding tools.

AI Discovery Workshop

A one-day workshop that turns scattered AI experiments into a prioritized, feasibility-checked roadmap leadership can act on.

Building Agentic AI Solutions Workshop

A hands-on, one-day workshop that takes engineers from prompting AI tools to building production-grade agentic AI systems: RAG, MCP, agents, and AI security.

Capstone Project

Apply all your knowledge in a capstone project similar to an actual data product in a modern data platform.

CI/CD with Azure DevOps

Build end-to-end CI/CD pipelines on Azure DevOps. From repos and code reviews to YAML pipelines, multi-environment deployments, and automated testing.

CI/CD with GitHub Actions

Set up CI/CD pipelines with GitHub Actions: workflows, triggers, secrets, environments, and container registries.

Containerization with Docker

Learn how to get your application working on any machine. Reach production faster.

Data Product Workshop

A one-day, cross-functional workshop that gets business and technical teams on the same page about data products, with a shared vocabulary, a scoping framework, quality principles, and clear roles.

DataFrame of Mind

Build data pipelines with Polars and PySpark. Learn when to use single-node processing and when to scale out.

Developer Productivity

Get more out of your dev environment: Python tooling, debugging, code quality automation, and agentic coding harnesses.

Effective PySpark

Build production-grade PySpark pipelines with sound design patterns, performance tuning, data quality checks, and testing.

Infrastructure as Code with Terraform

Learn how to create and maintain infrastructure as code, using Terraform.

Introduction to Cloud Providers (with AWS)

Learn when and how to use some of Amazon Web Services' most commonly used resources.

Introduction to Git for Version Control

Become proficient with the most widely used version control software. Stop sending code as attachments. Collaborate fearlessly.

Introduction to Linux & Bash

Learn bash and Linux basics to make you productive on any data project.

Kafka & Spark Structured Streaming

Build real-time streaming data pipelines by integrating Apache Kafka with Spark Structured Streaming.

Kubernetes

Deploy, scale, and manage containerized applications on Kubernetes, from core components to operators and networking.

Modernize SQL Analytics with DBT

A core component of any good analytics solution is still a relational database. Learn about SQL and build pipelines using DBT.

Principles of Modern Data Platforms

Prepare your data platform to support multiple use cases, at scale.

PySpark Fundamentals

Learn distributed data processing with PySpark. Understand Spark's architecture and DataFrame API, then apply them to build clean, modular, and testable data transformation pipelines.

Python for Data Engineering

Learn how to process data with Python and maintain a good code base.

Snowflake, a Modern Data Warehouse

Learn how to maintain and run analytics in this virtually infinitely scalable warehouse.

Sovereign Data Workshop

A one-day workshop that assesses whether and how to extend your data platform with sovereign EU cloud capabilities, from regulatory exposure to a migration roadmap.

Spark on Kubernetes

Submit, configure, and troubleshoot Spark jobs running on a Kubernetes cluster.

Workflow Orchestration with Airflow

Manage multiple workloads with the orchestration framework of Apache Airflow.