Skip to content
RY

Designing a Production-Ready REST API

Ram Laxman Yadav3 min read

Most REST API design advice stops at nouns-not-verbs and HTTP status codes. That gets you an API that looks correct in a slide deck. It says nothing about what happens when a client retries a POST after a timeout, or when you need to add a field without breaking every consumer that shipped last quarter.

This is the checklist I actually use when designing APIs for systems that move money, which is a domain with very little tolerance for “it works on my machine.”

Start with the resource model, not the routes

Before writing a single route, write down the resources and their relationships. A payments API, for example, usually needs at minimum:

  • PaymentIntent — the request to move money
  • Charge — an attempt to fulfill that intent
  • Refund — a reversal against a charge

Routes fall out of this model almost mechanically: POST /payment_intents, GET /payment_intents/{id}, POST /payment_intents/{id}/refunds. If you find yourself inventing a route that doesn’t map to a resource (POST /charge_and_notify_and_log), that’s a signal the resource model is missing something.

Idempotency is not optional

Any endpoint that causes a side effect — charging a card, sending an email, provisioning a resource — needs an idempotency story. Clients retry. Load balancers retry. Mobile networks drop packets mid-request. If a retried POST /charges can create two charges, you will eventually create two charges.

The pattern that has worked well for me:

POST /charges
Idempotency-Key: 7c2b9e2a-6e2b-4b7b-9b1a-2f7b6e2b9e2a

The server stores the result of the first request against that key for a bounded window (24–72 hours is typical) and replays the same response for duplicate keys, instead of re-executing the side effect.

Idempotency keys are not free

Storing and looking up idempotency keys adds a write to your hot path. Use a fast key-value store (Redis, or a dedicated indexed table) and set a TTL — don’t let the idempotency store grow unbounded.

Pagination that survives concurrent writes

Offset-based pagination (?page=2&per_page=20) breaks the moment rows are inserted or deleted between requests — a client can see the same row twice or skip one entirely. For anything with meaningful write volume, use cursor-based pagination instead:

GET /transactions?cursor=eyJpZCI6MTIzNDV9&limit=50

The cursor encodes a stable sort key (usually the primary key or a composite of created_at + id), so new rows inserted during pagination don’t shift the window.

Error responses are part of the contract

A 500 with an HTML stack trace is not an error response, it’s a bug. Every error should be structured, and the structure should be stable enough that clients can branch on it:

{
"error": {
"type": "invalid_request",
"code": "insufficient_funds",
"message": "The payment method has insufficient funds.",
"request_id": "req_8f2a1c9b"
}
}

The request_id matters more than it looks — it’s the thing that lets you join a client-reported error against your own logs in seconds instead of hours.

Version from day one

You don’t need a clever versioning scheme. You need a versioning scheme, decided before the first external consumer integrates. A path prefix (/v1/charges) is boring and that’s exactly why it works: it’s visible in every log line, every curl command, and every support ticket.

Checklist

  • Resources and relationships are modeled before routes are written
  • Every mutating endpoint accepts an idempotency key
  • Pagination is cursor-based for high-write collections
  • Errors are structured with a stable type/code and a request_id
  • API version is visible in the URL or a required header
  • Rate limits are documented and returned via Retry-After

None of this is novel. It’s also the first thing that gets skipped under deadline pressure, and the first thing that turns into an incident six months later.