Effective PySpark
Build production-grade PySpark pipelines with sound design patterns, performance tuning, data quality checks, and testing.
Data Platform
advanced
8h
Overview
A one-day advanced course on what makes a PySpark pipeline production-grade: design patterns, performance tuning with caching and partitioning, data quality checks, dimensional modeling, and testing.
What you'll cover
- ▸Pipeline design patterns
- ▸Lazy evaluation, caching, and partitioning
- ▸Data validation and quality checks
- ▸Dataset catalogs
- ▸Dimensional modeling: fact and dimension tables
- ▸Testing strategy for pipelines
What you'll be able to do
- ✓Structure a PySpark pipeline for production
- ✓Tune performance with caching and partitioning
- ✓Add validation and quality checks to a pipeline
- ✓Model data into fact and dimension tables
- ✓Write unit and integration tests for a pipeline
Tags
Pipeline DesignPerformance TuningData QualityPySparkPythonData EngineeringAdvanced
Prerequisites
- Getting Started with PySpark or equivalent experience
