Overview

FlareDB is an Apache Beam native streaming database for running Beam pipelines. It’s built in Rust, a modern systems programming language, and it uses a streams-tables architecture inspired by the ideas described in the Streaming Systems book (Chp 6).


The idea is that streams are data in motion, produced by computations (transforms), and a table is the same data at rest, during a windowing or grouping operation. FlareDB persists the PCollections as durable streams/tables on an append-only Apache Paimon table and runs the computations/transforms over the table using Apache DataFusion, a high-performance query engine.


As a result, FlareDB is lightweight, takes fewer resources to run, and makes the computational results queryable without the need for an external database.

The Beam Capability Matrix documents the supported capabilities of the FlareDB Runner.

Note: FlareDB is an independent and open-source runner for Apache Beam. It is not part of, or maintained by the Apache Beam project. The source code is available on GitHub.

How to use FlareDB Runner

Install FlareDB CLI to spawn up and manage FlareDB instance and run Beam pipelines.

1. Install the FlareDB CLI

If you are on Linux or macOS, please run the following command to install the CLI:

curl --proto '=https' --tlsv1.2 -LsSf https://github.com/flare-db/flare-db/releases/download/flare-cli-v0.3.2/flare-cli-installer.sh | sh

If you are on Windows use WSL.

2. Initialize FlareDB

After installing the CLI, run:

flare init

This command performs the initial setup by creating the required local directories and downloading the FlareDB binary and Apache Beam worker JAR.

The initialization only needs to be completed once. After that, you can use the flare up and flare down commands to manage the instance.

3. Start a FlareDB Instance

Start a local FlareDB instance with:

flare up

Once the instance is running, FlareDB is ready to accept pipeline jobs.

4. Configure Your Beam Pipeline

To run an Apache Beam pipeline on FlareDB, add the FlareDB Runner SDK as a dependency to your Beam project. The runner SDK submits the pipeline to the FlareDB instance as a Job.

Add the FlareDB Runner SDK to your pom.xml:

<dependency>
  <groupId>com.flare-db</groupId>
  <artifactId>flaredb-runner</artifactId>
  <version>0.3.2</version>
</dependency>

Set FlareRunner as the runner and configure the FlareDB instance and application JAR in your pipeline options:

WordCountPipelineOptions options =
    PipelineOptionsFactory.fromArgs(args).as(WordCountPipelineOptions.class);

options.setRunner(FlareRunner.class);
options.setJobEndpoint("127.0.0.1:8099");
options.setUberJar("build/libs/wordcount-0.2.0-all.jar");

Pipeline pipeline = Pipeline.create(options);

Alternatively, you can pass these as CLI arguments while running the pipeline:

./gradlew :wordcount:run --args="\
  --runner=FlareRunner \
  --jobEndpoint=127.0.0.1:8099 \
  --uberJar=build/libs/wordcount-0.2.0-all.jar"

Check out the full Java WordCount example pipeline.

Create and activate a virtual environment, then install the flaredb-runner package. It includes the apache-beam dependency.

python3 -m venv .venv
source .venv/bin/activate
pip install flaredb-runner

Set FlareRunner as the runner:

import apache_beam as beam
from apache_beam.options.pipeline_options import PipelineOptions

from flaredb_runner.flare_runner import FlareRunner

pipeline_options = PipelineOptions(
    job_endpoint="127.0.0.1:8099",
)

with beam.Pipeline(runner=FlareRunner(), options=pipeline_options) as p:

Run the pipeline:

python3 wordcount.py

Check out the full Python wordcount example pipeline.

5. Stop FlareDB instance

After executing pipelines, run this command to stop FlareDB instance

flare down

Pipeline options

The FlareDB Runner is configured through the following pipeline options:

OptionUsageDescription
RunnersetRunner(FlareRunner.class)Pipeline runner.
Job endpointsetJobEndpoint("host:port")URL of the FlareDB job service. Defaults to 127.0.0.1:8099.
Uber JARsetUberJar("/path/to/app.jar")Path to the fat JAR staged to workers.
Job namesetJobName("my-job")Name of the submitted job.

Next steps

FlareDB is released under the Apache License 2.0.