Observability
Three windows into a running Felix: structured logs, Prometheus metrics, and optional per-stage telemetry for performance work. This page lists what actually exists and which signals answer which questions — every metric named here is one the code emits.
Logging
Section titled “Logging”Felix logs through tracing, filtered by the standard RUST_LOG variable:
RUST_LOG=info # defaultRUST_LOG=felix_broker=debug # one module, louderRUST_LOG=felix_broker=trace,felix_wire=debug # several modulesLog lines are structured key-value events (tenant, stream, subscription id, error), so they grep and parse cleanly. There is no JSON output mode today; if your aggregation pipeline needs JSON, wrap the process output.
The lines worth knowing on sight:
- Startup: the QUIC listen address, and whether control-plane sync is on.
subscriber falling behind/events dropped for subscriber— a subscription hit its bounded queue. This is the log-side view offelix_sub_queue_dropped_total.- Drain lines during shutdown, saying which subsystems finished in time.
Metrics
Section titled “Metrics”The broker and the control plane each serve Prometheus text on their own
metrics endpoint (metrics_bind on the broker; /metrics, plus /live and
/ready for probes).
scrape_configs: - job_name: 'felix-broker' static_configs: - targets: ['broker-1:8080', 'broker-2:8080', 'broker-3:8080']Rather than an exhaustive list, here are the questions that come up and the metrics that answer them.
Is the publish path healthy?
felix_publish_requests_totalfelix_publish_bytes_totalfelix_publish_latency_ms # histogramfelix_broker_ingress_queue_depth # publish jobs waitingfelix_broker_ingress_dropped_total # overflow, by policyfelix_broker_ingress_rejected_totalfelix_client_publish_forwarded_total # client: publishes the broker had to relayfelix_broker_json_publishes_total # by frame: publishes still on JSONfelix_client_publish_cancelled_after_enqueue_total # client: publishes whose caller went awayfelix_client_publish_cancelled_after_enqueue_total is a client metric, and
non-zero is not an error. Cancelling a publish after it reaches the worker does
not cancel the publish — the record is sent and very likely lands, and only the
caller learning so is lost. It is here because those records have to be
explicable: a timeout around a publish means do not know, not did not
happen, and this is the number that says how often that happened.
A rising ingress depth means publishers are outrunning the broker; drops and rejections say the overflow policy fired, which is deliberate and visible.
felix_client_publish_forwarded_total is a client metric, labelled by the
owner the batch went to. Non-zero means this client is publishing to a broker
that does not own the shard, and each of those records is decrypted,
re-encrypted and decrypted again on the way — roughly half the throughput per
core. It is the client-side half of felix_broker_forwards_total, and the one
that says which client is mis-aimed.
felix_broker_json_publishes_total should be flat at zero. The data path is
binary; a client only falls back to JSON against a broker that did not advertise
the frame it wanted, so a non-zero rate means something in the deployment is
older than it looks — and it is paying for it, at roughly 70% of the binary
path’s throughput.
Are subscribers keeping up?
felix_subscribe_requests_totalfelix_sub_queue_enqueued_totalfelix_sub_queue_dropped_total # records lost to slow consumersfelix_sub_queue_drop_old_emulated_total # DropOld configured, DropNew behaviorfelix_sub_queue_lenfelix_subscriber_disconnect_totalfelix_sub_queue_dropped_total increasing is the signal that a subscriber is
missing records. On a durable stream the subscriber can detect this itself
from offset gaps and resume; on an ephemeral stream this counter is the only
witness.
Is durability the bottleneck? The storage layer’s metrics are designed around exactly this question — compare append time against sync time, and watch the group-commit fan-in:
felix_storage_append_duration_secondsfelix_storage_sync_duration_secondsfelix_storage_sync_batch_appends # appends served per device flushfelix_storage_unsynced_bytes # what a crash would lose right nowfelix_storage_sync_failures_total # non-zero: acknowledged durability in doubtIf sync dominates append, the fsync policy is the cost. A
sync_batch_appends near 1 under concurrent load means appends are
serializing on the device instead of sharing a flush.
Is the cluster healthy? Membership from both sides, replication, and leases:
felix_node_count # control plane: fleet size by lifecyclefelix_broker_membership_live # broker: does the cluster still count mefelix_broker_heartbeat_age_seconds # alert when this nears the expiry timeoutfelix_broker_replication_lag_recordsfelix_broker_replication_halted # a count; GET /replication/halted says whichfelix_broker_replication_rebuilding # halted followers the leader is rebuilding right nowfelix_broker_replication_rebuilds_total # by outcome: started, completed, refusedfelix_broker_replica_reports_per_request # shards per control-plane report; 1 on a busy broker means batching found nothingfelix_broker_lease_heldfelix_broker_lease_refusals_total # writes refused after a lease lapsedfelix_broker_credential_expires_in_seconds # counts down; -1 when the token carries no expfelix_broker_credential_refreshes_total # by outcome: ok, unavailablefelix_broker_credential_rotations_total # by outcome: ok, rejected — a token file rewritten from outsideAlert on felix_broker_credential_expires_in_seconds crossing a threshold,
not only on the refresh and rotation counters. The counters say renewal is
failing; the gauge says how long that has left to matter. It matters a lot: the
heartbeat carries this credential and the heartbeat is the lease renewal, so a
token that expires is not a degraded broker — it is one that stops serving the
shards it leads once the lease lapses. That is the safe outcome, and still an
outage.
A broker joining a cluster refuses to start when the credential expires and
neither FELIX_NODE_REFRESH_TOKEN_FILE nor FELIX_NODE_TOKEN_FILE is set, so
the commonest way to reach that state is caught at rollout rather than an hour
in.
Which replica stopped
Section titled “Which replica stopped”felix_broker_replication_halted is a count, and stays one: a label per shard
would be a label per stream per tenant. A leader rebuilds a halted follower on
its own, one at a time by default (FELIX_REPLICATION_REBUILD_MAX_CONCURRENT),
so a brief non-zero is a rebuild queue. When it stays above zero, ask the broker
which:
curl -s http://broker:9090/replication/halted | jq[ { "tenant_id": "t1", "namespace": "ns", "stream": "orders", "shard": 3, "kind": "stream", "node_id": "broker-b", "generation": 7, "next_offset": 120, "reason": "diverged", "remedy": "the follower holds different bytes at an offset this leader also holds, …" }]A halt does not clear on its own — that replica is out of every quorum until
someone acts — so any non-zero count is worth waking for. reason is stable
and safe to key a runbook off; remedy is prose and says whether the
follower’s data is wrong or merely incomplete, which is what decides whether a
rebuild is the right move.
A healthy broker answers [], not 404.
Bootstrap attempts
Section titled “Bootstrap attempts”felix_bootstrap_attempts_total{outcome,reason} covers the day-0 credential.
Bootstrap is presented once per tenant, by an operator, and never again — so a
rejected attempt is either a misconfigured deploy or someone guessing, and a
burst of already_initialized refusals against live tenants is what a leaked
token looks like.
# Anything but the two normal outcomes is worth waking for.rate(felix_bootstrap_attempts_total{outcome="rejected"}[5m])
# A leaked token being tried against tenants that already exist.rate(felix_bootstrap_attempts_total{reason="already_initialized"}[5m])reason is a small closed set — missing_token, malformed_token,
invalid_token, no_token_configured, already_initialized, ok — so it is
safe to group by. The token itself is never logged, including a near miss.
Example queries:
# Publish raterate(felix_publish_requests_total[1m])
# p99 publish latencyhistogram_quantile(0.99, rate(felix_publish_latency_ms_bucket[5m]))
# Records lost to slow consumersrate(felix_sub_queue_dropped_total[5m])
# Group-commit effectivenessrate(felix_storage_sync_batch_appends_sum[5m]) / rate(felix_storage_sync_batch_appends_count[5m])Shutdown and drain
Section titled “Shutdown and drain”Four signals, and the reason each one exists.
| Metric | Type | What it tells you |
|---|---|---|
felix_ready_state |
gauge | 1 while serving, 0 once draining. Distinguishes an instance that left rotation deliberately from one that vanished. |
felix_inflight_requests |
gauge | Requests being served right now. Watch it fall to zero during a drain. |
felix_drain_duration_ms |
gauge | How long the last drain took. |
felix_drain_forced_total |
counter, by subsystem |
Subsystems cancelled because the deadline expired. Non-zero means work was dropped. |
The last one is the point. A drain that finished in time and a drain that was cut off both take roughly the deadline to report, so duration alone cannot tell them apart — and the log line that says which does not survive the pod.
Alert on felix_drain_forced_total increasing. Everything else here is for
watching a rolling restart happen.
Probes
Section titled “Probes”The metrics listener serves /live and /ready, and they answer different
questions on purpose. /live says “this process can respond at all” and
touches nothing outside the process — a liveness probe drives restarts, and
restarting every instance because a dependency is down turns one outage into
a restart loop. /ready says “send this instance traffic,” and goes false
first thing during shutdown so load balancers steer away before anything
stops working. See Graceful shutdown.
Distributed tracing
Section titled “Distributed tracing”The broker builds an OTLP tracer provider on startup and installs a
tracing-opentelemetry layer when one is available. It is best-effort: if the
collector cannot be reached the broker starts anyway and logs without traces,
because losing telemetry must not stop a broker from serving.
Configuration is by environment variable. Felix does take a YAML config file
(FELIX_BROKER_CONFIG, see Configuration),
but it has no tracing keys — the exporter speaks OTLP over gRPC (tonic) and is
configured entirely through the standard OTel variables:
export OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317export RUST_LOG=info # the tracing subscriber's filterResource attributes are attached from the environment, so a span carries where it came from without the broker being told twice:
| Variable | Becomes |
|---|---|
FELIX_SERVICE_INSTANCE_ID, falling back to HOSTNAME |
service.instance.id |
K8S_CLUSTER_NAME |
k8s.cluster.name |
K8S_NAMESPACE_NAME |
k8s.namespace.name |
K8S_POD_NAME |
k8s.pod.name |
CLOUD_REGION |
cloud.region |
DEPLOYMENT_ENVIRONMENT |
deployment.environment |
Per-stage telemetry
Section titled “Per-stage telemetry”For performance investigations, both the broker and client can record per-stage timing samples — decode, fanout, write, and so on. It is off by default and behind a feature flag, because it is a profiling tool, not a production metrics system:
[dependencies]felix-client = { version = "0.1", features = ["telemetry"] }On the client, felix_client::frame_counters_snapshot() returns frame-level
counters, and felix_client::timings::take_samples() drains the recorded
per-stage samples. The benchmarks and the latency-demo binary are the
worked examples of reading them.
On the broker, FELIX_CONN_STATS_MS logs QUIC path statistics (MTU, cwnd,
RTT, loss, flow-control blocking) for healthy connections on an interval —
the data that says whether a throughput problem is transport-side or above
it. Off unless set.
Debugging quick answers
Section titled “Debugging quick answers”- Subscribers receive nothing: check the subscription was created (log
line), the stream exists, and the application is actually awaiting
next_event(). Then checkfelix_sub_queue_dropped_total. - Publish latency spiked: check
felix_broker_ingress_queue_depth(broker backed up), thenfelix_storage_sync_duration_seconds(durability is the cost), then the client-side telemetry to see which stage grew. - Cache misses you didn’t expect: TTL expiry, a broker restart on an ephemeral cache, or a key/scope mismatch — in that order of likelihood.
