Performance

Slow API response time: a practical diagnosis guide

Follow the time through your system, test the likely bottleneck, and verify that the fix helped real requests.

7 min read

The practical answer

Start with one affected endpoint and a defined time window. Compare latency, errors, traffic, and saturation; trace where requests spend time; then change the measured bottleneck and compare equivalent workloads.

A customer says the dashboard is unusable, but the infrastructure dashboard looks mostly green. You add capacity. The bill rises, and the customer still waits. Slow API response time is difficult to fix when the measurement does not describe the request the customer actually made.

This guide is for investigating persistent or recurring latency. If customers are currently unable to use the product, establish an incident owner and a recovery plan first. Then use the evidence below to distinguish a database problem, a waiting queue, application work, and an upstream dependency.

1. Define which requests are slow

Write down the endpoint, user action, first observed time, region, and affected customer group. Separate successful responses from failures. An endpoint that immediately returns an error can make an average response time look better while the product gets worse.

Compare the median with a high percentile such as p95 or p99. A p95 latency describes the response time at or below which roughly 95% of the measured requests fall; it is not a speed guarantee for every user. Always record the window, request count, and measurement location. Google’s SRE guidance explains why averages can conceal a slow tail.

Reference: Google SRE: Monitoring Distributed Systems

Client timing and server timing answer different questions. A browser may wait on connection setup, the network, server processing, and response transfer. A server-side timer may cover only part of that path. Compare measurements with clear boundaries before concluding that one of the dashboards is wrong.

  • Which route and operation are affected?
  • Does it affect all tenants, one region, large accounts, or a specific payload?
  • Did it begin after a deploy, configuration change, traffic increase, or scheduled job?
  • What did the same workflow look like during a healthy period?

2. Follow the request through the system

A trace lets you inspect operations within an individual request instead of guessing from an aggregate graph. OpenTelemetry represents those operations as spans, with timing and relationships that can connect work across services. Compare slow and healthy requests for the same operation.

Reference: OpenTelemetry: Traces and spans

Ask where the elapsed time accumulates: before application work starts, while obtaining a database connection, during a query, inside application processing, or while awaiting another service. Instrument a missing boundary when it would change the decision. A blank space in a trace is a question, not proof that the database is slow.

Use sanitized request IDs to connect evidence. Do not copy authentication tokens or customer payloads into an investigation document. The useful record is the operation, timing, relevant dimensions, and where the evidence can be inspected by the team.

3. Match the symptom to the next investigation

API latency hypotheses to test
Observed patternWhat to inspectA possible fix, if confirmed
One endpoint is slow; others are healthy.Its query count, payload size, and downstream calls.Remove unnecessary repeated work or reduce the response scope.
Requests wait before queries begin.Connection-pool wait time, active connections, and concurrency.Shorten connection holding time or adjust bounded concurrency.
Query time dominates the trace.Query plan, row counts, locks, and data distribution.Change the query or index the actual access pattern.
Latency rises with a scheduled job.Shared database, disk, CPU, and queue pressure.Isolate, reschedule, or limit the competing workload.
An upstream call dominates.Dependency timing, timeout settings, and retry attempts.Bound waiting and handle the dependency’s failure mode.
New capacity does not improve latency.Whether the bottleneck sits in a shared dependency.Address that shared limit before adding more callers.

Treat each row as a hypothesis. Low average CPU does not rule out a bottleneck elsewhere, and a busy database does not prove that every slow request is caused by a query. Choose the next check that can disprove your leading explanation.

4. Inspect database work with care

For PostgreSQL, begin with the query shape and an execution plan. Look for the operations doing the most work and large differences between estimated and observed row counts. An index is useful only when it supports the workload; adding indexes indiscriminately also changes the cost of writes and maintenance.

PostgreSQL’s EXPLAIN estimates a plan. EXPLAIN ANALYZE executes the statement to collect observations, and statements with side effects still perform them. Start in a representative test environment, use appropriate execution limits, and review the statement before running diagnostic work on production.

Reference: PostgreSQL: Using EXPLAIN

Query duration is only one part of the request. Waiting for a connection or a lock may happen outside the SQL timing shown by your application. A query that is fast in isolation can behave differently when many requests compete for the same resources. Reproduce the relevant concurrency and data shape before declaring the investigation finished.

A practical first experiment might compare one sanitized request with the current query and a proposed alternative in staging. Record returned rows and business results as well as timing. An empty or incomplete response is fast, but it is not an optimization.

5. Check whether retries are making the problem larger

A slow dependency can cause clients to retry while earlier attempts are still consuming resources. If several layers independently retry the same operation, one customer action can produce more downstream work than you expect. Record attempts per request, the total time budget, and whether repeating the operation is safe.

Bound retry attempts and use a delay policy appropriate to the dependency. AWS SDK documentation describes bounded retries and exponential backoff with jitter, including retry quotas. Check your actual client’s configuration rather than assuming defaults are identical across libraries or versions.

Reference: AWS SDKs: Retry behavior

Increasing a timeout can help a legitimately long operation finish, but it can also leave more callers waiting on the same bottleneck. Decide the acceptable customer experience, cancellation behavior, and error path before changing it. For writes, establish how duplicate attempts are identified and handled.

6. Verify the diagnosis with a controlled comparison

The first hypothesis is an expensive reporting query. Comparing traces shows that most of the extra time occurs before the query starts. Connection-pool observations show requests waiting while the import holds connections. The team tests a bounded import concurrency setting in staging, then plans a limited rollout.

That result would support investigating resource contention. It would not justify blindly raising the pool limit: the database may have its own capacity constraint. The useful fix is the one that improves the reporting workflow while preserving import correctness and completion requirements.

  • Compare the same route, response status, data shape, and similar load.
  • Record p50, p95, error rate, throughput, and the suspected waiting time.
  • Check that the result is complete and correct.
  • Watch a representative traffic or scheduled-job cycle after rollout.
  • Define what observation would trigger recovery or another investigation.

Keep one investigation note

API performance investigation
Affected workflow and measurement boundary:
Healthy comparison window:
Slow window and request count:
Latency distribution and error rate:
Where the time accumulates:
Leading hypothesis and contrary evidence:
Smallest experiment:
Correctness checks:
Rollout, recovery, and review owner:

This note should make it possible for another engineer to continue the investigation without repeating every dashboard search. If the data points toward a broader infrastructure incident, switch to the debugging playbook. If the same constraint keeps returning, capture it in the technical debt backlog with the evidence you now have.

References and further reading

All guides