Scale is usually described in request volume. In my experience, the harder scaling problem is organizational: keeping a system reliable as more people, more features, and more edge cases accumulate around it. Here are the lessons that have held up across a decade of building and operating payments and hospitality platforms.
Reconciliation is a feature, not an afterthought
Every system that moves money or inventory eventually drifts from its own ledger — a webhook arrives twice, a job retries after a partial write, a third-party API times out after already completing the action. The question isn’t whether drift happens, it’s whether you find out from your own reconciliation job or from a customer support ticket.
We built reconciliation as a first-class scheduled process from the start of the payments platform:
class ReconcileLedgerJob def perform(date) internal_total = Ledger.where(date: date).sum(:amount_cents) provider_total = PaymentProvider.settlement_report(date).total_cents
if internal_total != provider_total Alert.fire(:ledger_mismatch, date:, internal_total:, provider_total:) end endendIt’s deliberately simple. The value isn’t in the cleverness of the comparison, it’s in running it every single day without fail and alerting loudly the moment numbers disagree.
Observability has to answer “why,” not just “what”
Dashboards that show error rate and latency tell you that something is wrong. They rarely tell you why quickly enough during an incident. The investment that paid off repeatedly was structured, high-cardinality logging tied to a request ID that survives across service boundaries:
{ "request_id": "req_8f2a1c9b", "service": "payments-api", "event": "charge.declined", "decline_code": "insufficient_funds", "provider_latency_ms": 420}Tip
If an on-call engineer has to SSH into a box and grep a log file during an incident, that’s a gap in observability tooling, not a gap in the engineer’s skills.
Monoliths aren’t the enemy — premature services are
We migrated a hospitality booking monolith toward service-oriented architecture over roughly eighteen months, and the biggest lesson was about sequencing: we split off the inventory and pricing services after they had stable, well-understood contracts inside the monolith, not before.
Splitting a service out before its boundaries are stable just moves the chaos across a network call — you trade a messy function call for a messy, slower, harder-to-debug RPC call.
On-call load is a leading indicator, not a lagging one
A team whose on-call load is climbing is telling you about an architecture or process problem weeks before it becomes an outage. The habit that worked best as a manager was treating every page as a two-part question in the postmortem:
- What broke?
- What made this page necessary instead of self-healing?
The second question is the one that actually reduces on-call load over time — automated remediation, better defaults, or removing the failure mode entirely.
Mentorship scales systems more than architecture does
Twenty engineers who understand why a system is built the way it is will keep it reliable far longer than a brilliant architecture maintained by people who don’t. The highest-leverage thing I did at scale wasn’t a redesign — it was pairing consistently enough that on-call engineers could reason about failure modes without escalating.
Takeaways
- Build reconciliation before you need it, not after the first discrepancy
- Structured, request-scoped logging beats clever dashboards during incidents
- Split services along boundaries that have already proven stable
- Treat rising on-call load as an architecture signal
- Mentorship is infrastructure investment, not overhead