OpenTelemetry Usage Metering Engine
go services that turn opentelemetry traces, logs and metrics into billable usage, published exactly once and reconciled against the invoice.

- role
- sole author
- stack
- Go · PostgreSQL 17 · ClickHouse · Kubernetes · Helm · React 19 · TanStack · OpenTelemetry
The problem
Usage-based billing sounds simple: count what customers use, multiply by a price. In practice the counting has to stay correct through retries, late data and partial failures, because every bug is either lost revenue or an angry customer. This engine is a personal build that works through those problems properly, end to end.
How data flows
- An ingest gateway receives OTLP data. It forwards to storage first, then records usage, then acknowledges. The worst failure is "stored but not billed", never "billed but not stored".
- Every batch is deduplicated over a 6-minute window, which covers the OpenTelemetry Collector's retry span.
- A rollup worker closes each hour and writes the hourly totals and their outgoing billing events in one transaction.
- A publisher sends those events to a Stripe-compatible billing provider.
- A reconciler compares the ledger, the outbox, the provider and the invoices, and can hold an invoice from finalising until the numbers agree.
Correctness, not just counting
- Append-only ledger. The app's database role can only read and insert usage rows. A correction is a new, reconciler-approved event, never an edit.
- Tenant isolation in the database. Row-level security scopes every transaction to one tenant and fails closed, and a lint rule bans opening transactions any other way.
- Exactly-once publication. Each billing event has a deterministic ID built from tenant, meter, signal, hour and sequence, plus a lease so only one writer publishes per tenant.
- Late data never reopens the past. It's counted in the hour it arrived, so a closed hour never changes. A written proof covers the clock-skew and timeout bounds, and a sentinel checks them in production.
- Drift is classified, not just detected. Differences are sorted into timing, structural or systemic. Three or more tenants drifting at once freezes publishing everywhere.
- Testable time. Reading the clock directly is banned outside one package, so a simulated clock can backfill three months of history in a test.
Running it
It runs on Kubernetes from day one, via a Helm chart with CloudNativePG, a ClickHouse operator and Keycloak. Observability is OpenTelemetry, Prometheus and Grafana, with an SLO dashboard, 20 unit-tested alert rules and 23 runbooks. A React 19 dashboard (TanStack Router and Query, types generated from OpenAPI) covers usage, invoices, budgets and alerts, and ingest keys.
By the numbers
- 11 services and tools, 81 Go packages
- 584 tests and fuzz targets across 185 test files
- 24 Postgres migrations, 14 design docs and 74 recorded decisions
- load test: 2,250 events per second sustained on a 10-CPU local cluster, limited by ClickHouse inserts