Gigflow docs
Engineering

ADR-0009: Engine read replica routing

Select primary or replica reads through the API request context.

Status: Proposed. Implementation available for review.

Date: 2026-09-23

Context

Engine uses asynchronous PostgreSQL replication. A standby can return data that predates a completed mutation. Engine uses service identities and shared market data. It has no team scope for Midday's shared mutation cache.

Decision

Keep Midday's database wrapper and request-context routing pattern. Bind each Drizzle method to its database instance. Writes, raw execute calls, transactions, and CTE statements use primary. CTE statements stay on primary because the builder also supports writes.

API queries use replicas by default. A shared primary-read-after-write middleware switches ctx.db to primary for mutations, recent mutations from that caller, subscriptions, and requests with x-force-primary: true. Endpoints do not need metadata or a separate replica procedure.

Use the authenticated internalIdentity.keyId as the cache key. Before a mutation starts, store an expiry timestamp in Redis with a 10-second TTL. Queries from that key use primary until the window expires. All API instances share this marker. The cache uses Engine's existing Redis/Sentinel settings and its configured key prefix: <prefix>:replication:<keyId>.

Authentication and role checks run before routing. The middleware passes a new context to the procedure. It does not change the shared database instance. If a cache read fails, use primary. If storing a mutation marker fails, reject the request before the mutation handler runs. Redis commands have a one-second timeout. The API closes its cache connection during shutdown.

Workers and one-off jobs use one primary pool. They expose the same database methods as the API client. Their executeOnReplica method uses primary.

ENGINE_DATABASE_URL selects the primary pooler. The optional ENGINE_DATABASE_REPLICA_URL selects the CNPG read-only pooler. If the replica URL is absent or equals the primary URL, the API shares the primary pool. The infrastructure must provide the configured read-only pooler.

Consequences

There is no async-local write tracking or application health probe. A standby's received and replayed WAL positions do not prove that it has received all WAL from primary. The application does not use these positions as a freshness guarantee.

The cache window starts before the mutation, as in Midday. It reduces stale reads but does not prove that a replica has caught up. A mutation that takes more than 10 seconds can outlast the window. Use an explicit primary read when freshness is required beyond this window.

One API key represents a service, not a user or team. If Admin shares a key, one Admin mutation sends every query through that key to primary for the window. Mutations through another key and direct worker writes do not update that caller's marker. Consumers that need to observe those writes must force primary reads.

Replica connection failures remain errors. The wrapper does not retry arbitrary SQL on primary.

Verification

Install the locked Bun dependencies and start Docker. From the repository root, run:

bun run --cwd apps/engine/apps/api test:e2e:replicas

The runner creates two disposable PostgreSQL containers with streaming replication and one Redis container. It uses a dedicated engine_test_replicas database and binds ports to localhost. It does not connect to the development or production database. It removes its containers and volumes after the run.

The HTTP test uses the real authentication middleware and configuration procedures in two separate API processes. It pauses WAL replay to verify primary reads after a mutation across those processes, then waits for the cache window to expire. It also checks caller isolation, cache failure, primary overrides, raw reads, transactions, worker reads, and access denial. The fixture creates only the table needed by these procedures. It does not validate CNPG failover or PgBouncer configuration.

Results, the tested revision, source fingerprints, and run logs are written to .context/read-replicas/. A failed setup produces a failed report.

On this page