AuctionEngine.
A real-time auction engine in Go, built around its invariants — the rules that must never break, enforced by the database rather than by hopeful application logic.
the problem
An auction is a contention problem wearing a costume. A thousand people race for one row, money changes hands at the end, and the interesting failures only appear when things are fast, duplicated, or killed halfway through.
the approach
I wrote down seven invariants first — statements that must hold no matter what — then made the schema enforce them. A bug in my Go code can't corrupt state. It can only get caught.
lock
SELECT … FOR UPDATEcompeting bids queue here
validate
window, minimum, idempotency key
write
bid + auction head + outbox event
publish
outbox → Kafka, keyed by auction
settle
charge exactly once
the numbers
1,126accepted bids per second4,145 req/s on one auction
3.3 msp99 latency under contention10 concurrent bidders
0lost updates at 1,000 biddersverified, not assumed
0double chargesat 50% simulated payment failure
three problems worth the trouble
- The stale clock. "No bid after the auction closes" sounds trivial until you learn
now()is fixed at transaction start — so a bid that waited on the row lock could be validated against a timestamp from before it waited. The fix is readingclock_timestamp()after the lock is acquired, in both Go and the trigger. A test caught this; reasoning alone would not have. - Exactly-once, across a boundary I don't control. Kafka delivers at least once, and a worker can die mid-charge. Settlement uses a transactional outbox — the event is written in the same transaction as the bid, so neither can exist without the other — plus idempotency keys at the payment provider. At 50% simulated payment failure, still zero double charges.
- Choosing the boring lock. I built optimistic concurrency too, then benchmarked both. Pessimistic row locking won at moderate contention (3.3 ms vs 7.0 ms p99) because optimistic wasted about three attempts per accepted bid. Optimistic won at extreme contention — and I wrote down that pool queueing confounds that result rather than claiming a clean win.
proving it, instead of asserting it
- An invariant checker that audits the database after every test and load run — and every check is itself proven to catch a violation planted for it.
- Mutation testing on the safety layers. Remove the app lock, the guard trigger, or the fork-proof constraint individually and the system still holds; remove all three and a real lost update appears — which the test catches.
- A self-verifying load generator. A run fails loudly unless the server counted every response the client saw, the client's accepted count matches the database, and every invariant held. A stub returning
201without storing anything is proven to fail it. - Fault injection kills API instances and workers mid-run, then re-checks all seven invariants.
shipping it
- Infrastructure as code. ECS, RDS, S3, and an ALB provisioned entirely in Terraform — no console clicking to reproduce.
- Every commit deploys. GitHub Actions runs lint,
go test -race, a page contract test, and an image build, then rolls out with health checks and no downtime. - Redis is an accelerator, never a source of truth. When it's slow or down, caching, rate limiting, and live fan-out degrade instead of failing.
- Every significant decision is an ADR — context, alternatives, consequences — with known risks tracked in the open questions file rather than quietly dropped.
7 invariants, schema-enforced
1,126 bids/sec · 3.3 ms p99
exactly-once settlement
Terraform + CI to AWS
See the code ↗