# Spark Batch Processing - Managing Spark Jobs and Notebooks: Challenges, Caching, and Performance Optimization (Day 2 Lecture)

> In this lecture, Zach discusses the intricacies of managing Spark jobs and notebooks, highlighting the challenges of achieving code modularity and the nuances of jar submissions for Scala Spark and PySpark. Exploring the advantages and distinctions between notebooks and spark-submit for production…

- Web page: https://www.dataexpert.io/lesson/spark-batch-day-2-lecture-v3
- Program: [Data Engineering Boot Camp V2 Combined Track](https://www.dataexpert.io/program/data-engineering-boot-camp-v2-combined-track)
- Module: Week 4: Batch Pipelines with Apache Spark
- Access: Requires enrollment in Data Engineering Boot Camp V2 Combined Track
- Length: 39 min video
- Skills: Python, Apache Spark
- Academy: DataExpert.io Academy

## About this lesson

In this lecture, Zach discusses the intricacies of managing Spark jobs and notebooks, highlighting the challenges of achieving code modularity and the nuances of jar submissions for Scala Spark and PySpark. Exploring the advantages and distinctions between notebooks and spark-submit for production jobs, Zach also sheds light on caching, the auto broadcast join threshold, and the effective use of UDFs in PySpark. Closing with expert insights, he offers valuable tips on performance optimization and the optimal language selection for Spark jobs. [Recorded on Dec 7 2023]

## Navigation

- Previous lesson: [Spark Batch Processing - Hands-On Techniques for Broadcast Hash Join (Day 1 Lab) ](https://www.dataexpert.io/lesson/spark-batch-day-1-lab-v3.md)
- Next lesson: [Spark Batch Processing - User-Defined Functions (UDFs) and Broadcast Join (Day 2 Lab)](https://www.dataexpert.io/lesson/spark-batch-day-2-lab-v3.md)
