reductstore/ reductstore
High-performance, time-indexed object storage for robotics and industrial IoT
Rust 373 stars
replies in 4 hours
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
12 repos, most welcoming first
Only repos tagged big-data. Show all topics
High-performance, time-indexed object storage for robotics and industrial IoT
Rust 373 stars
replies in 4 hours
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
replies in 3 hours
23 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)
Go 2.0k stars
replies in 16 hours
56 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Apache Beam is a unified programming model for Batch and Streaming data processing.
Java · Python 8.7k stars
replies in about a day
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
Scala · Python 5.2k stars
replies in 3 days
5 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Server for the ListenBrainz project, including the front-end (javascript/react) code that it serves and all of the data processing components that LB uses.
Python · TypeScript 1.0k stars
replies in 3 days
16 first-timers merged
Outside contributors get real replies here, and their work gets merged.
replies in 2 days
3 first-timers merged
Outside contributors get real replies here, and their work gets merged.
2 starter issuesall taken
replies in 2 days
10 first-timers merged
An issue to start with#23322 chore: Explore combining `SessionConfig` and `RuntimeConfig` into same framework (opens GitHub)1 open pull request
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.
Java 4.4k stars
replies in under a minute
4 first-timers merged
Outside contributors get real replies here, and their work gets merged.
The official home of the Presto distributed SQL query engine for big data
Java 17k stars
1 starter issuetaken
replies in 13 hours
10 first-timers merged
An issue to start with#26779 Add iceberg rewrite_table_path procedure to Presto (opens GitHub)1 open pull request
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
Java 13k stars
5 starter issuesall taken
replies in about a day
9 first-timers merged
An issue to start with#28176 dependabot-originating builds fail due to lack of secrets (opens GitHub)1 open pull request
3 starter issuesall taken
replies in 5 days
22 first-timers merged
An issue to start with#1437 Allow Kafka-UI container to use Kafka certificates directly (.key, .cert, .ca) without manual Java keystore conversion (opens GitHub)1 open pull request
High-performance, time-indexed object storage for robotics and industrial IoT
Rust 373 stars
replies in 4 hours
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
replies in 3 hours
23 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Fluid, elastic data abstraction and acceleration for BigData/AI applications in cloud. (Project under CNCF)
Go 2.0k stars
replies in 16 hours
56 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Apache Beam is a unified programming model for Batch and Streaming data processing.
Java · Python 8.7k stars
replies in about a day
10 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Simple and Distributed Machine Learning Python Library porting ML algorithms for Spark
Scala · Python 5.2k stars
replies in 3 days
5 first-timers merged
Outside contributors get real replies here, and their work gets merged.
Server for the ListenBrainz project, including the front-end (javascript/react) code that it serves and all of the data processing components that LB uses.
Python · TypeScript 1.0k stars
replies in 3 days
16 first-timers merged
Outside contributors get real replies here, and their work gets merged.
replies in 2 days
3 first-timers merged
Outside contributors get real replies here, and their work gets merged.
2 starter issuesall taken
replies in 2 days
10 first-timers merged
An issue to start with#23322 chore: Explore combining `SessionConfig` and `RuntimeConfig` into same framework (opens GitHub)1 open pull request
CrateDB is a distributed and scalable SQL database for storing and analyzing massive amounts of data in near real-time, even with complex queries. It is PostgreSQL-compatible, and based on Lucene.
Java 4.4k stars
replies in under a minute
4 first-timers merged
Outside contributors get real replies here, and their work gets merged.
The official home of the Presto distributed SQL query engine for big data
Java 17k stars
1 starter issuetaken
replies in 13 hours
10 first-timers merged
An issue to start with#26779 Add iceberg rewrite_table_path procedure to Presto (opens GitHub)1 open pull request
Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)
Java 13k stars
5 starter issuesall taken
replies in about a day
9 first-timers merged
An issue to start with#28176 dependabot-originating builds fail due to lack of secrets (opens GitHub)1 open pull request
3 starter issuesall taken
replies in 5 days
22 first-timers merged
An issue to start with#1437 Allow Kafka-UI container to use Kafka certificates directly (.key, .cert, .ca) without manual Java keystore conversion (opens GitHub)1 open pull request