Overview
Felix is a distributed data backend that serves three things most systems run three products for: event streams, work queues, and a key-value cache. All three are readings of the same replicated, append-only log, reached over one QUIC connection.
This page is the short version of how it fits together. The status table is the page to trust for what is and is not built.
The core idea: one log, three readings
Section titled “The core idea: one log, three readings”Internally there is a single primitive — an append-only log per shard. Each external API is a different way of reading it:
- Streams read by offset. Every subscription keeps its own cursor, so a slow reader never holds anyone else back.
- Queues read through a cursor a consumer group shares. A record goes to one consumer, is redelivered if unacknowledged, and dead-lettered when it keeps failing.
- Cache reads through a key index over the same log: latest value per key, with TTL.
Because it is one log, there is one durability path, one recovery path, one placement rule and one replication path. There is no cache-vs-stream consistency bug to have, because there is no second system to disagree with the first.
What Felix optimizes for
Section titled “What Felix optimizes for”Predictable latency under load, more than peak batch throughput. The choices that follow from that:
- QUIC as the only transport. Streams multiplex over one connection without head-of-line blocking, and every connection is TLS 1.3.
- Backpressure everywhere. Queues are bounded and overflow policy is explicit, so one slow path degrades locally instead of cascading.
- Slow-consumer isolation. A stalled subscriber loses its own events
(under the default
drop_newpolicy) rather than stalling the publisher or other subscribers. The slow-consumer demo runs both policies side by side. - Explicit tuning knobs. Batching bounds, flow-control windows, pool sizes, and fsync policy are configuration, not magic.
Measured numbers live on the Benchmarks page — one place, so they can’t drift page to page.
The pieces
Section titled “The pieces”Five crates, and the boundaries between them are the design. Everything below the dotted line is transport-independent: the broker core has no idea QUIC exists, which is what makes it testable in-process.
flowchart TB
subgraph app["Your application"]
direction LR
A1(["publisher"])
A2(["subscriber"])
A3(["cache client"])
end
SDK["felix-client<br/><small>publisher / subscription / cache APIs,<br/>connection + stream pools</small>"]
WIRE["felix-wire<br/><small>frame header, JSON control messages,<br/>binary data-plane frames</small>"]
TRANS["felix-transport<br/><small>QUIC endpoints, streams, flow control,<br/>dedicated I/O runtimes</small>"]
BRK["felix-broker<br/><small>stream registry, log, subscriber<br/>registry, fanout</small>"]
SVC["services/broker<br/><small>network service: handlers, auth,<br/>metrics, control-plane sync</small>"]
CP["services/controlplane<br/><small>tenants, namespaces, streams,<br/>tokens, RBAC</small>"]
A1 e1@--> SDK
A2 e2@--> SDK
A3 e3@--> SDK
SDK e4@--> WIRE
WIRE e5@--> TRANS
TRANS e6@-->|"QUIC + TLS 1.3"| SVC
SVC e7@--> BRK
SVC e8@-.->|"seeds metadata at startup"| CP
e1@{ animate: true }
e2@{ animate: true }
e3@{ animate: true }
e4@{ animate: true }
e5@{ animate: true }
e6@{ animate: true }
e7@{ animate: true }
classDef step fill:#e8f0fe,stroke:#4a6fa5,color:#1a2b40
classDef core fill:#fdf0e3,stroke:#b07d3a,color:#3d2a12
classDef endpoint fill:#e9f5ec,stroke:#4a8a5e,color:#16301f
class SDK,WIRE,TRANS,SVC step
class BRK,CP core
class A1,A2,A3 endpoint
felix-wireis the protocol: a fixed frame header, JSON control messages, binary frames on the data plane, and capability negotiation so old and new peers interoperate. Full details in Wire Protocol.felix-transportwraps QUIC: endpoints, streams, flow-control windows, and dedicated I/O runtimes so the transport driver is never starved by application tasks.felix-brokeris the data plane with no networking in it: the log, subscriber registry, fanout, cache index, consumer groups.felix-clientis the Rust SDK: publisher, subscription, and cache APIs over pooled connections, with reconnection and redirect-following in the cluster client. Python (crates/felix-python) and TypeScript (crates/felix-typescript) ship too, both as bindings over this client rather than reimplementations; Go and C# are not started.- The control plane (
services/controlplane) holds metadata — tenants, namespaces, streams, caches, nodes — behind a REST API, and assigns every shard to a broker. Brokers watch its assignment feed. It is not on the data path: a publish never waits on it.
Delivery guarantees, plainly
Section titled “Delivery guarantees, plainly”Configured per stream:
- Ephemeral streams are at-most-once. The broker acknowledges receipt, fans out from memory, and a subscriber that falls behind misses records. This is the low-latency mode, and the right one when stale data is worthless anyway.
- Durable streams are at-least-once: a record is persisted before it is
acknowledged, and can be replayed from any retained offset. With
consistency: quorum, the acknowledgement additionally waits until a majority of the replica set holds the record — so a failover cannot lose an acknowledged write. - Consumer groups redeliver anything unacknowledged, count attempts, and park repeat failures as dead letters you can list, discard, or redrive.
- Exactly-once is not implemented and is not close. If duplicates are unacceptable, deduplicate in the application.
A lost leader is replaced by a replica that provably holds the log — about a second on a local three-node cluster. This is tested against process kill, graceful stop, a leader frozen past its lease, and a partitioned broker that keeps heartbeating.
Security
Section titled “Security”Today: TLS 1.3 on every connection, OIDC token exchange at the control plane,
tenant-scoped tokens, RBAC enforced at the broker, and mutually authenticated
broker-to-broker QUIC once FELIX_INTERNAL_TLS_CERT, _KEY and _CA are set.
Not yet: end-to-end payload encryption, encryption at rest, audit logging.
Details in Security.
Running it
Section titled “Running it”A single broker is fine for development and for workloads that fit on one machine — durable storage and the cache work there; replication needs somewhere to replicate to.
A cluster is brokers plus a control plane:
The control plane keeps its metadata in one of three backends:
| Backend | What holds the metadata | When to pick it |
|---|---|---|
memory |
This process, nothing else. The default. | Development, and the harness |
postgres |
One external database the instances share | A platform that already runs an HA Postgres |
raft |
The control-plane instances themselves, replicated between them | No external database to operate — edge sites, appliances, or anywhere Postgres is a burden rather than a convenience |
Raft has shipped: the instances form a quorum and survive losing one without losing an acknowledged write. Postgres remains fully supported; the trade between the two, and the migration path, are in Metadata Raft and Control-plane HA.
Felix runs anywhere a process runs. For orchestrators it ships the pieces they expect: readiness and liveness endpoints that answer different questions, and a bounded graceful drain on SIGTERM — see Graceful shutdown.
Is Felix right for your workload?
Section titled “Is Felix right for your workload?”A good fit: real-time streaming with high fanout, low-latency caching, work distribution with retries and dead letters, and services that currently run a broker and a cache and a queue and would rather run one system.
A bad fit: petabyte-scale batch pipelines, complex stream processing (joins, windowing — use Flink or Kafka Streams), or anything that needs a mature connector ecosystem today. Felix is young and its ecosystem is one language deep.
The honest version of this list, kept current per capability, is What Felix Is For.
Where to next
Section titled “Where to next”- Quickstart — run a broker and the demos
- Installation — build from source
- System design — the architecture in depth
- Broker API — the wire-level API
