Skip to content

Examples/ML serving

Model inference API (example)

Customers call a REST API with API keys. Requests are batched onto GPU workers that load model weights by version hash, and usage is aggregated and billed monthly. Redis holds rate limits, Postgres holds the registry, and Kafka carries usage events.

Scale: 400 customers, 300 req/s, p95 under 800 ms for small inputs

123456789101112CLIENTCustomer SDKEDGEAPI gatewayCACHERate limit storeSERVICEInference APIDATABASEModel registryQUEUERequest queuesWORKERGPU workersOBJECT STOREWeights storeCACHEResults cacheLOG / STREAMUsage streamWORKERBilling workerEXTERNALMetered billing

Select a component to see what it is responsible for and which state it owns.

  1. 1Customer SDK → API gateway: Run inference
  2. 2API gateway → Rate limit store: Take a token
  3. 3API gateway → Inference API: Forward request
  4. 4Inference API → Model registry: Resolve key and model version
  5. 5Inference API → Request queues: Enqueue
  6. 6GPU workers → Request queues: Pull a batch
  7. 7GPU workers → Weights store: Load weights
  8. 8GPU workers → Results cache: Write result
  9. 9Inference API → Results cache: Wait for result
  10. 10GPU workers → Usage stream: Emit usage
  11. 11Billing worker → Usage stream: Consume usage
  12. 12Billing worker → Metered billing: Report usage
  • Request / response
  • Asynchronous
  • Bulk data

Select a part to trace its flows.

  • Request / response
  • Asynchronous
  • Bulk data

Parts 12

Flows 12

  1. 1Customer SDK → API gatewayRun inferenceRequest / response
  2. 2API gateway → Rate limit storeTake a tokenRequest / response
  3. 3API gateway → Inference APIForward requestRequest / response
  4. 4Inference API → Model registryResolve key and model versionRequest / response
  5. 5Inference API → Request queuesEnqueueAsynchronous
  6. 6GPU workers → Request queuesPull a batchRequest / response
  7. 7GPU workers → Weights storeLoad weightsBulk data
  8. 8GPU workers → Results cacheWrite resultRequest / response
  9. 9Inference API → Results cacheWait for resultRequest / response
  10. 10GPU workers → Usage streamEmit usageAsynchronous
  11. 11Billing worker → Usage streamConsume usageRequest / response
  12. 12Billing worker → Metered billingReport usageRequest / response

Invariants 3

  • Each request is billed once

    Usage events carry the request id; the (customer, hour) aggregate update and an insert into processed_requests(request_id) run in one transaction; the primary key rejects a duplicate and aborts the update.

    Kept by

  • A key never exceeds its rate

    Check and decrement run as one Redis Lua script, so two concurrent requests cannot take the same last token.

    Kept by

  • A request is served by the model version it was admitted with

    Version hash recorded at admission; weights stored under immutable keys; never overwritten in place.

    Kept by