Manage Data and Delta Tables with Apache Spark on Databricks

Manage Data and Delta Tables with Apache Spark on Databricks
.MP4, AVC, 1920x1080, 30 fps | English, AAC, 2 Ch | 1h 23m | 220 MB
Instructor: Dayo Bamikole
What you'll learn
Data engineers working on Databricks often struggle with the same set of problems: data that lands with no structure, schemas that break pipelines overnight, and tables that are painful to keep current without rebuilding from scratch. In this course, Manage Data and Delta Tables with Apache Spark on Databricks, you'll gain the ability to organize, manage, and operate Delta tables in a way that holds up in production.
First, you'll explore how Unity Catalog structures data access using catalogs, schemas, and tables, and how to read and write Delta tables using PySpark and Spark SQL. Next, you'll discover how Delta Lake enforces schemas, how to handle schema evolution safely using mergeSchema, and how to identify the risks that come with changing schemas on live tables. Finally, you'll learn how to use MERGE for upserts, work with semi-structured data, and use time travel to audit and recover from unintended changes.
When you're finished with this course, you'll have the skills and knowledge needed to build and maintain a reliable Delta table foundation that your team can depend on.
Homepage

