Introduction to Celeborn
What is Celeborn?
Apache Celeborn is a Remote Shuffle Service (RSS) designed to improve the efficiency, stability, and flexibility of shuffle operations in distributed compute engines. It supports Apache Spark, Apache Flink, Apache Tez, and MapReduce.
Why Celeborn?
Traditional shuffle frameworks have significant limitations that become critical at scale:
Problem | Traditional Shuffle | Celeborn Solution |
Network Efficiency | M × N connections between Mappers and Reducers | Consolidated M+N connections via Celeborn workers |
Disk I/O | Random I/O on compute nodes | Sequential I/O on dedicated shuffle nodes |
Dynamic Allocation | Limited by shuffle data locality | Full executor elasticity |
Node Failure | Shuffle data lost, job fails or retries | Data replicated — job continues without retry |
Storage | Large local disks required on compute nodes | Dedicated shuffle storage (local, HDFS, S3) |
Key Benefits
- Performance: 2–5× improvement in shuffle-heavy workloads
- Stability: Data replication prevents job failures from node loss
- Elasticity: Enables true dynamic resource allocation
- Disaggregation: Separates compute from shuffle storage
- Multi-Engine: Supports Spark, Flink, Tez, and MapReduce
