apache/ beam
Apache Beam is a unified programming model for Batch and Streaming data processing.
Java · Python 8.7k stars
replies in about a day
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
5 repos, most welcoming first
Only repos tagged big-data. Show all topics
Apache Beam is a unified programming model for Batch and Streaming data processing.
Java · Python 8.7k stars
replies in about a day
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
replies in 2 days
3 first-timers merged
Outside contributors get real replies here, and their work gets merged.
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.
Java 4.4k stars
replies in under a minute
4 first-timers merged
Outside contributors get real replies here, and their work gets merged.
The official home of the Presto distributed SQL query engine for big data
Java 17k stars
1 starter issuetaken
replies in 13 hours
10 first-timers merged
An issue to start with#26779 Add iceberg rewrite_table_path procedure to Presto (opens GitHub)1 open pull request
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
Java 13k stars
5 starter issuesall taken
replies in about a day
9 first-timers merged
An issue to start with#28176 dependabot-originating builds fail due to lack of secrets (opens GitHub)1 open pull request
Apache Beam is a unified programming model for Batch and Streaming data processing.
Java · Python 8.7k stars
replies in about a day
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
replies in 2 days
3 first-timers merged
Outside contributors get real replies here, and their work gets merged.
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.
Java 4.4k stars
replies in under a minute
4 first-timers merged
Outside contributors get real replies here, and their work gets merged.
The official home of the Presto distributed SQL query engine for big data
Java 17k stars
1 starter issuetaken
replies in 13 hours
10 first-timers merged
An issue to start with#26779 Add iceberg rewrite_table_path procedure to Presto (opens GitHub)1 open pull request
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
Java 13k stars
5 starter issuesall taken
replies in about a day
9 first-timers merged
An issue to start with#28176 dependabot-originating builds fail due to lack of secrets (opens GitHub)1 open pull request