Reliability

Infrastructure debugging: a production incident playbook

A practical path through failed deploys, intermittent errors, unhealthy containers, and capacity pressure.

7 min read

The practical answer

Define the customer impact, establish who is coordinating the response, and locate the failing boundary. Compare the last healthy state with recent changes, test one hypothesis at a time, and verify recovery through the affected user journey.

The service is returning errors, several dashboards are flashing, and three people are proposing three different fixes. Infrastructure debugging gets harder when every action changes the evidence. A shared sequence can keep a small team moving even when the cause is unclear.

This playbook covers recurring production symptoms across cloud and container environments. It gives you a way to narrow the investigation, plus a short Kubernetes inspection sequence. Commands are examples for an environment you operate; the surrounding decision process also applies to virtual machines and managed services.

1. Establish impact and keep one working record

Name the affected workflow, when it last worked, which customers or regions are affected, and the current workaround. Separate observations from hypotheses. “Requests to checkout began failing after 14:10 UTC” is an observation; “the release broke the database” is an explanation that still needs testing.

Choose a person to coordinate changes and a person to keep stakeholders informed, combining roles if the team is small. Keep a timestamped record of checks, interventions, and results. Google’s incident-response guidance emphasizes clear roles and a working record so responders can coordinate while restoring service.

Reference: Google SRE Workbook: Incident Response

Pause unrelated changes where practical. Preserve relevant logs and deployment information before replacing processes that may hold the only evidence. If you already have a tested, reversible mitigation, restoring the customer workflow can take priority over completing the diagnosis.

2. Find the boundary that is failing

Follow a representative request from the edge toward the dependency it needs. Can the name resolve? Can a connection be established? Does the load balancer have healthy targets? Does the application accept the request? Does the application reach its database or upstream service? The first failing boundary narrows the search.

A gateway error describes what the gateway observed. It does not, by itself, tell you whether the application crashed, a connection was refused, or a dependency took too long. Compare gateway logs with application observations for the same time window.

Symptoms, evidence, and conditional fixes
SymptomInspect nextFix only after confirming
502 or 504 responsesGateway logs, upstream health, connection failures, and request timing.Correct the failing route or upstream condition; adjust timeouts only with a clear time budget.
Failure begins during rolloutVersion distribution, configuration changes, health checks, and migration compatibility.Stop exposure to the suspect change or use the documented recovery path.
Containers keep restartingExit reason, termination state, events, and previous-container logs.Address the crash, invalid configuration, or resource constraint shown by the evidence.
Some instances work; others failVersion, node, zone, and configuration differences between instances.Repair or remove the affected subset using the service’s operating procedure.
Requests stall under loadQueues, connection waits, CPU or memory pressure, and dependency limits.Reduce contention or bound demand at the constrained layer.
Storage or writes failFree capacity, storage events, connection errors, and recent workload changes.Restore capacity or the failing dependency without discarding data to silence an alert.
DNS or TLS failuresResolution from the affected environment, hostname, certificate chain, and expiry.Correct the actual DNS or certificate configuration through its managed process.

These are investigation prompts, not a list of commands to run indiscriminately. The same visible symptom can come from several causes. Choose the check that best separates your current explanations.

3. Compare with the last healthy state

Review code deployments, environment configuration, secret or certificate rotation, network policy, scaling events, schema changes, and scheduled jobs. A change close to the start of the incident is a useful lead, but timing alone is not proof.

Write a testable prediction: “If this version is responsible, requests handled by the older version should still succeed under comparable conditions.” Then check whether the prediction holds. Compare equivalent traffic and customer cohorts; a quiet old instance is not a valid control for a busy new one.

4. Inspect a Kubernetes workload without changing it

First confirm the cluster context and namespace. Replace the uppercase placeholders below with your workload’s names. These commands inspect existing state and logs; they do not restart or delete the workload. Restrict log output to the affected interval and handle sensitive output within your team’s approved tools.

A bounded Kubernetes inspection sequence
kubectl config current-context
kubectl -n YOUR_NAMESPACE get pods -o wide
kubectl -n YOUR_NAMESPACE describe pod YOUR_POD
kubectl -n YOUR_NAMESPACE logs YOUR_POD -c YOUR_CONTAINER --since=15m --tail=200
kubectl -n YOUR_NAMESPACE logs YOUR_POD -c YOUR_CONTAINER --previous --tail=200

Reference: Kubernetes: Debug Running Pods

Compare restarts, termination reasons, events, and readiness with a healthy instance. A previous-container log can explain a crash even when the replacement is running. An OOMKilled termination points toward memory exhaustion; it still leaves you to investigate workload demand, configured limits, and application behavior.

If Pods appear healthy but the Service does not reach them, inspect the Service configuration and its EndpointSlices. Check that the intended Pods are selected, ready for traffic, and reached on the intended port. The Kubernetes service-debugging guide walks through these boundaries.

Reference: Kubernetes: Debug Services

Inspect the Service’s selected endpoints
kubectl -n YOUR_NAMESPACE get service YOUR_SERVICE -o yaml
kubectl -n YOUR_NAMESPACE get endpointslices -l kubernetes.io/service-name=YOUR_SERVICE

Avoid treating every restart as a reason to raise a memory limit, or every routing failure as a reason to recreate a Service. The observed reason should determine the experiment. For a managed database or virtual machine, use the equivalent resource events, process logs, health status, and connection checks.

5. Test one useful change and define the recovery signal

For each proposed intervention, record the hypothesis, expected result, affected scope, and a way to recover if it fails. Prefer a change small enough to evaluate. When several changes are applied together, a temporary recovery may leave you unsure which action mattered.

For example, suppose only instances running a new configuration reject requests, while the older configuration succeeds for the same operation. A limited reversal of that configuration may be a reasonable mitigation if its compatibility is understood. Record what happened afterward. This is an illustrative scenario, not evidence that every post-deploy failure should be rolled back.

  • Does the customer’s actual workflow succeed from the affected environment?
  • Have error rate and latency returned to an acceptable level?
  • Are retries or queued jobs hiding unfinished work?
  • Do records and side effects remain correct?
  • Does recovery hold through representative demand rather than one successful probe?

Make observations at the boundary that failed. A green process-health check may not exercise login, billing, data access, or the dependency the customer needs. Keep a watch owner and an agreed observation period appropriate to the workload.

6. Leave a useful handoff after recovery

Production incident worksheet
Impact and affected workflow:
First observed failure / last known healthy state:
Coordinator and communications owner:
Current observations:
Recent changes worth testing:
Hypothesis and next discriminating check:
Intervention, timestamp, and result:
Customer recovery signal:
Remaining backlog or data checks:
Follow-up owner and review date:

After recovery, separate the immediate mitigation from the underlying cause and the conditions that made the incident difficult to detect or resolve. Turn the most valuable follow-up into a bounded task with an owner. A recurring incident deserves a technical debt decision with evidence, rather than another promise to “improve monitoring” someday.

References and further reading

All guides