API performance under real load
A production API had severe latency spikes during peak traffic, expensive database paths and limited visibility into regressions.
- Established an APM and log baseline before changing the critical path
- Optimised queries, cache strategy and background jobs
- Added service-level dashboards and actionable alerts
- Outcome: core endpoint P95 moved from approximately 1.8 seconds to below 450ms, with stable errors and visible regressions
