The practical answer
Define the customer impact, establish who is coordinating the response, and locate the failing boundary. Compare the last healthy state with recent changes, test one hypothesis at a time, and verify recovery through the affected user journey.
The service is returning errors, several dashboards are flashing, and three people are proposing three different fixes. Infrastructure debugging gets harder when every action changes the evidence. A shared sequence can keep a small team moving even when the cause is unclear.
This playbook covers recurring production symptoms across cloud and container environments. It gives you a way to narrow the investigation, plus a short Kubernetes inspection sequence. Commands are examples for an environment you operate; the surrounding decision process also applies to virtual machines and managed services.
1. Establish impact and keep one working record
Name the affected workflow, when it last worked, which customers or regions are affected, and the current workaround. Separate observations from hypotheses. “Requests to checkout began failing after 14:10 UTC” is an observation; “the release broke the database” is an explanation that still needs testing.
Choose a person to coordinate changes and a person to keep stakeholders informed, combining roles if the team is small. Keep a timestamped record of checks, interventions, and results. Google’s incident-response guidance emphasizes clear roles and a working record so responders can coordinate while restoring service.
Reference: Google SRE Workbook: Incident Response
Pause unrelated changes where practical. Preserve relevant logs and deployment information before replacing processes that may hold the only evidence. If you already have a tested, reversible mitigation, restoring the customer workflow can take priority over completing the diagnosis.
2. Find the boundary that is failing
Follow a representative request from the edge toward the dependency it needs. Can the name resolve? Can a connection be established? Does the load balancer have healthy targets? Does the application accept the request? Does the application reach its database or upstream service? The first failing boundary narrows the search.
A gateway error describes what the gateway observed. It does not, by itself, tell you whether the application crashed, a connection was refused, or a dependency took too long. Compare gateway logs with application observations for the same time window.
| Symptom | Inspect next | Fix only after confirming |
|---|---|---|
| 502 or 504 responses | Gateway logs, upstream health, connection failures, and request timing. | Correct the failing route or upstream condition; adjust timeouts only with a clear time budget. |
| Failure begins during rollout | Version distribution, configuration changes, health checks, and migration compatibility. | Stop exposure to the suspect change or use the documented recovery path. |
| Containers keep restarting | Exit reason, termination state, events, and previous-container logs. | Address the crash, invalid configuration, or resource constraint shown by the evidence. |
| Some instances work; others fail | Version, node, zone, and configuration differences between instances. | Repair or remove the affected subset using the service’s operating procedure. |
| Requests stall under load | Queues, connection waits, CPU or memory pressure, and dependency limits. | Reduce contention or bound demand at the constrained layer. |
| Storage or writes fail | Free capacity, storage events, connection errors, and recent workload changes. | Restore capacity or the failing dependency without discarding data to silence an alert. |
| DNS or TLS failures | Resolution from the affected environment, hostname, certificate chain, and expiry. | Correct the actual DNS or certificate configuration through its managed process. |
These are investigation prompts, not a list of commands to run indiscriminately. The same visible symptom can come from several causes. Choose the check that best separates your current explanations.
3. Compare with the last healthy state
Review code deployments, environment configuration, secret or certificate rotation, network policy, scaling events, schema changes, and scheduled jobs. A change close to the start of the incident is a useful lead, but timing alone is not proof.
Write a testable prediction: “If this version is responsible, requests handled by the older version should still succeed under comparable conditions.” Then check whether the prediction holds. Compare equivalent traffic and customer cohorts; a quiet old instance is not a valid control for a busy new one.
4. Inspect a Kubernetes workload without changing it
First confirm the cluster context and namespace. Replace the uppercase placeholders below with your workload’s names. These commands inspect existing state and logs; they do not restart or delete the workload. Restrict log output to the affected interval and handle sensitive output within your team’s approved tools.
kubectl config current-context
kubectl -n YOUR_NAMESPACE get pods -o wide
kubectl -n YOUR_NAMESPACE describe pod YOUR_POD
kubectl -n YOUR_NAMESPACE logs YOUR_POD -c YOUR_CONTAINER --since=15m --tail=200
kubectl -n YOUR_NAMESPACE logs YOUR_POD -c YOUR_CONTAINER --previous --tail=200Reference: Kubernetes: Debug Running Pods
Compare restarts, termination reasons, events, and readiness with a healthy instance. A previous-container log can explain a crash even when the replacement is running. An OOMKilled termination points toward memory exhaustion; it still leaves you to investigate workload demand, configured limits, and application behavior.
If Pods appear healthy but the Service does not reach them, inspect the Service configuration and its EndpointSlices. Check that the intended Pods are selected, ready for traffic, and reached on the intended port. The Kubernetes service-debugging guide walks through these boundaries.
Reference: Kubernetes: Debug Services
kubectl -n YOUR_NAMESPACE get service YOUR_SERVICE -o yaml
kubectl -n YOUR_NAMESPACE get endpointslices -l kubernetes.io/service-name=YOUR_SERVICEAvoid treating every restart as a reason to raise a memory limit, or every routing failure as a reason to recreate a Service. The observed reason should determine the experiment. For a managed database or virtual machine, use the equivalent resource events, process logs, health status, and connection checks.
5. Test one useful change and define the recovery signal
For each proposed intervention, record the hypothesis, expected result, affected scope, and a way to recover if it fails. Prefer a change small enough to evaluate. When several changes are applied together, a temporary recovery may leave you unsure which action mattered.
For example, suppose only instances running a new configuration reject requests, while the older configuration succeeds for the same operation. A limited reversal of that configuration may be a reasonable mitigation if its compatibility is understood. Record what happened afterward. This is an illustrative scenario, not evidence that every post-deploy failure should be rolled back.
- Does the customer’s actual workflow succeed from the affected environment?
- Have error rate and latency returned to an acceptable level?
- Are retries or queued jobs hiding unfinished work?
- Do records and side effects remain correct?
- Does recovery hold through representative demand rather than one successful probe?
Make observations at the boundary that failed. A green process-health check may not exercise login, billing, data access, or the dependency the customer needs. Keep a watch owner and an agreed observation period appropriate to the workload.
6. Leave a useful handoff after recovery
Impact and affected workflow:
First observed failure / last known healthy state:
Coordinator and communications owner:
Current observations:
Recent changes worth testing:
Hypothesis and next discriminating check:
Intervention, timestamp, and result:
Customer recovery signal:
Remaining backlog or data checks:
Follow-up owner and review date:After recovery, separate the immediate mitigation from the underlying cause and the conditions that made the incident difficult to detect or resolve. Turn the most valuable follow-up into a bounded task with an owner. A recurring incident deserves a technical debt decision with evidence, rather than another promise to “improve monitoring” someday.
