Apache Beam is an open source unified programming model for defining and executing both batch and streaming data-parallel processing pipelines. Beam provides a portable API layer for describing these pipelines independent of execution engines (or runners) such as Apache Spark, Apache Flink or Google Cloud Dataflow.
Why is Apache Beam used?
Apache Beam is an open source, unified model for defining both batch and streaming data-parallel processing pipelines. … These tasks are useful for moving data between different storage media and data sources, transforming data into a more desirable format, or loading data onto a new system.
Is Apache Beam good?
In terms of capabilities and generality, I think Apache Beam is currently the most advanced and most flexible framework for designing and implementing modern data-intensive applications. Perfectly able to specify both batch and streaming computations, plus in terms of streaming capabilities, it offers really a lot!