# Apache Spark Shuffle Joins Day 1 Lecture

> In this lecture, Zach discusses how we handle two petabytes of data daily at Netflix, focusing on different sampling techniques to optimize processing. He shares insights on the importance of precision in data analysis and how we managed to reduce processing time and costs significantly by using a…

- Web page: https://www.dataexpert.io/lesson/apache-spark-shuffle-joins-day-1-lecture-may2025-p736-l1142
- Program: [DataExpert](https://www.dataexpert.io/program/data-expert)
- Module: Week 2: Databricks & Advanced Spark
- Access: Requires enrollment in DataExpert
- Length: 40 min video
- Skills: Data Modeling, ETL/ELT, Apache Spark
- Academy: DataExpert.io Academy

## About this lesson

In this lecture, Zach discusses how we handle two petabytes of data daily at Netflix, focusing on different sampling techniques to optimize processing. He shares insights on the importance of precision in data analysis and how we managed to reduce processing time and costs significantly by using a 0.1% sample. He also touches on the challenges of dynamic IP addresses in our cloud environment and the need for collaboration with application owners to implement effective logging.

## Navigation

- Previous lesson: [Apache Spark Core Day 3 Lab](https://www.dataexpert.io/lesson/apache-spark-core-day-3-lab-jan2025-p736-l780.md)
- Next lesson: [Apache Spark Memory Turning, Partitioning Day 2 Lecture](https://www.dataexpert.io/lesson/apache-spark-memory-turning-partitioning-day-2-lecture-may2025-p736-l1145.md)
