Overview
FlareDB is an Apache Beam native streaming database for running Beam pipelines. It’s built in Rust, a modern systems programming language, and it uses a streams-tables architecture inspired by the ideas described in the Streaming Systems book (Chp 6).
The idea is that streams are data in motion, produced by computations (transforms), and a table is the same data at rest, during a windowing or grouping operation. FlareDB persists the PCollections as durable streams/tables on an append-only Apache Paimon table and runs the computations/transforms over the table using Apache DataFusion, a high-performance query engine.
As a result, FlareDB is lightweight, takes fewer resources to run, and makes the computational results queryable without the need for an external database.
The Beam Capability Matrix documents the supported capabilities of the FlareDB Runner.
Note: FlareDB is an independent and open-source runner for Apache Beam. It is not part of, or maintained by the Apache Beam project. The source code is available on GitHub.
How to use FlareDB Runner
Install FlareDB CLI to spawn up and manage FlareDB instance and run Beam pipelines.
1. Install the FlareDB CLI
If you are on Linux or macOS, please run the following command to install the CLI:
curl --proto '=https' --tlsv1.2 -LsSf https://github.com/flare-db/flare-db/releases/download/flare-cli-v0.3.2/flare-cli-installer.sh | sh
If you are on Windows use WSL.
2. Initialize FlareDB
After installing the CLI, run:
flare init
This command performs the initial setup by creating the required local directories and downloading the FlareDB binary and Apache Beam worker JAR.
The initialization only needs to be completed once. After that, you can use the flare up and flare down commands to manage the instance.
3. Start a FlareDB Instance
Start a local FlareDB instance with:
flare up
Once the instance is running, FlareDB is ready to accept pipeline jobs.
4. Configure Your Beam Pipeline
To run an Apache Beam pipeline on FlareDB, add the FlareDB Runner SDK as a dependency to your Beam project. The runner SDK submits the pipeline to the FlareDB instance as a Job.
- Java SDK
- Python SDK
Add the FlareDB Runner SDK to your pom.xml:
Set FlareRunner as the runner and configure the FlareDB instance and application JAR in your pipeline options:
Alternatively, you can pass these as CLI arguments while running the pipeline:
Check out the full Java WordCount example pipeline.
Create and activate a virtual environment, then install the flaredb-runner package. It includes the apache-beam dependency.
Set FlareRunner as the runner:
Run the pipeline:
Check out the full Python wordcount example pipeline.
5. Stop FlareDB instance
After executing pipelines, run this command to stop FlareDB instance
flare down
Pipeline options
The FlareDB Runner is configured through the following pipeline options:
| Option | Usage | Description |
|---|---|---|
| Runner | setRunner(FlareRunner.class) | Pipeline runner. |
| Job endpoint | setJobEndpoint("host:port") | URL of the FlareDB job service. Defaults to 127.0.0.1:8099. |
| Uber JAR | setUberJar("/path/to/app.jar") | Path to the fat JAR staged to workers. |
| Job name | setJobName("my-job") | Name of the submitted job. |
Next steps
- Browse the FlareDB documentation.
- See the Beam Capability Matrix to learn about features supported by FlareDB.
- Try more examples in the FlareDB repository.
- Report bugs or request features in the FlareDB issue tracker.
- Contributions are welcome. See the Contributing Guide to get started.
FlareDB is released under the Apache License 2.0.
Last updated on 2026/10/08
Have you found everything you were looking for?
Was it all useful and clear? Is there anything that you would like to change? Let us know!

