Discussion about this post

User's avatar
Kerby Lane's avatar

I'm not sure I see how example "4. Salting Every Key (Uniform Distribution Across N Buckets)" is different than than groupBy(youtuber_id, video_id), or hierarchical aggregation as you described later. There's no use of randomization to split up the video_id rows. The `salt` value calculation is deterministic. So all rows for a pair of youtuber_id and video_id will get the same value.

Jayasurya Pilli's avatar

Thank you for such a detailed explanation on solution options for handling data skew.

However, I have a question, probably a quick one...

Given that now we also have Liquid Clustering feature available in Databricks, should Liquid Clustering be considered as the first and the foremost recommended solution, even over the AQE please?

In other words, as a recommended approach, shouldn't Liquid Clustering be considered first, followed by AQE, then BroadcastHashJoin, then Salting.

Please correct me if I'm wrong in my understanding.

3 more comments...

No posts

Ready for more?