# Exploring UDFs and SQL Benchmarks in Spark Streaming

> In this video, Zach dives into the differences between UDFs and built-in SQL functions in Spark Streaming, using a benchmark that processes 5 million random numbers. He explores how UDFs can be slower, with results showing about a 10% speed-up when using built-in functions, but this can vary based…

- Web page: https://www.dataexpert.io/lesson/exploringudfsandsqlbenchmarksinsparkstreaming-feb26-p736-l2113
- Program: [DataExpert](https://www.dataexpert.io/program/data-expert)
- Module: Week 4: Structured Streaming with Spark & Kafka
- Access: Requires enrollment in DataExpert
- Length: 6 min video
- Skills: SQL, Python, Apache Spark
- Academy: DataExpert.io Academy

## About this lesson

In this video, Zach dives into the differences between UDFs and built-in SQL functions in Spark Streaming, using a benchmark that processes 5 million random numbers. He explores how UDFs can be slower, with results showing about a 10% speed-up when using built-in functions, but this can vary based on caching and the complexity of operations. Zach also experiments with increasing the row count to 100 million and even 1 billion to see how performance changes. He encourages everyone to run similar benchmarks and observe the results for themselves, as there are nuances that can affect performance.

## Navigation

- Previous lesson: [Structured Streaming Kafka to Delta Live Table Day2 Lab](https://www.dataexpert.io/lesson/structuredstreamingkafkatodeltalivetableday2lab-feb26-p736-l2032.md)
- Next lesson: [Advanced Spark Optimization Techniques Day 1 Lecture](https://www.dataexpert.io/lesson/advanced-spark-optimization-techniques-day-1-lecture-jan2025-p736-l812.md)
