Examples/ML serving
Model inference API (example)
Customers call a REST API with API keys. Requests are batched onto GPU workers that load model weights by version hash, and usage is aggregated and billed monthly. Redis holds rate limits, Postgres holds the registry, and Kafka carries usage events.
Scale: 400 customers, 300 req/s, p95 under 800 ms for small inputs
Select a component to see what it is responsible for and which state it owns.
- 1Customer SDK → API gateway: Run inference
- 2API gateway → Rate limit store: Take a token
- 3API gateway → Inference API: Forward request
- 4Inference API → Model registry: Resolve key and model version
- 5Inference API → Request queues: Enqueue
- 6GPU workers → Request queues: Pull a batch
- 7GPU workers → Weights store: Load weights
- 8GPU workers → Results cache: Write result
- 9Inference API → Results cache: Wait for result
- 10GPU workers → Usage stream: Emit usage
- 11Billing worker → Usage stream: Consume usage
- 12Billing worker → Metered billing: Report usage
- Request / response
- Asynchronous
- Bulk data
Select a part to trace its flows.
- Request / response
- Asynchronous
- Bulk data
Parts 12
Flows 12
- 1Customer SDK → API gatewayRun inferenceRequest / response
- 2API gateway → Rate limit storeTake a tokenRequest / response
- 3API gateway → Inference APIForward requestRequest / response
- 4Inference API → Model registryResolve key and model versionRequest / response
- 5Inference API → Request queuesEnqueueAsynchronous
- 6GPU workers → Request queuesPull a batchRequest / response
- 7GPU workers → Weights storeLoad weightsBulk data
- 8GPU workers → Results cacheWrite resultRequest / response
- 9Inference API → Results cacheWait for resultRequest / response
- 10GPU workers → Usage streamEmit usageAsynchronous
- 11Billing worker → Usage streamConsume usageRequest / response
- 12Billing worker → Metered billingReport usageRequest / response
Invariants 3
Each request is billed once
Usage events carry the request id; the (customer, hour) aggregate update and an insert into processed_requests(request_id) run in one transaction; the primary key rejects a duplicate and aborts the update.
Kept by
A key never exceeds its rate
Check and decrement run as one Redis Lua script, so two concurrent requests cannot take the same last token.
Kept by
A request is served by the model version it was admitted with
Version hash recorded at admission; weights stored under immutable keys; never overwritten in place.
Kept by