The practical answer
Investigate recurring backend bugs by identifying the condition that survived the previous fix. Reproduce that condition in a controlled environment, test the behavior that must hold, and verify the change through a representative operating cycle. Give unresolved follow-up work an owner.
I spent years working on startups that failed. I’ve also had plenty of projects go nowhere. Those experiences taught me a lot, but the learning came from going back through what happened and changing how I worked. Repeating the mistake by itself didn’t teach me anything.
I bring that same expectation to recurring backend bugs. If a problem returns after a fix, there is still something to understand. Maybe the fix covered one input. Maybe the restart cleared the evidence. Maybe two failures look identical from the outside and have different causes. I want to know which one we’re dealing with before the team spends another afternoon patching it.
1. Give recurring backend bugs a precise description
If customers are currently blocked, recovery comes first. Capture the evidence you can without delaying the response, then stabilize the service. The work below starts once there is room to investigate properly. Nobody needs to preserve a broken checkout just to make a debugging session more interesting.
After recovery, I want a sentence that describes the failed behavior. ‘The database broke again’ gives us very little. ‘Two requests created separate jobs for the same import, and both jobs processed it’ gives us an operation, an unexpected result, and a place to start. Use sanitized identifiers to connect the report to the relevant logs.
Compare that sentence with the previous incident. Did the same rule fail? Did it happen under the same deployment version, input shape, or timing? A matching error message is a useful clue, but I would still check the path that produced it. Grouping different failures together can hide the fact that an earlier fix worked.
Google’s postmortem guidance emphasizes recording impact, contributing causes, recovery, and preventive actions without blaming individuals. I use that as a starting point for the follow-up: explain what the system allowed, what information was available, and what needs to change.
Reference: Google SRE: Postmortem Culture
2. Reproduce the condition the first fix missed
Here’s a hypothetical example. An import endpoint checks whether a job already exists before creating one. Most requests work. Occasionally, the same import runs twice. A patch adds another existence check, and the issue appears to go away until two requests arrive together again.
| Step | Request A | Request B |
|---|---|---|
| 1 | Checks for the job; none exists. | Has not checked yet. |
| 2 | Has not inserted yet. | Checks for the job; none exists. |
| 3 | Inserts a new job. | Still holds its earlier result. |
| 4 | Returns success. | Inserts another job. |
The useful reproduction forces those two checks to finish before either insert proceeds. In a test environment, a synchronization barrier can make that ordering deliberate. Waiting an arbitrary number of milliseconds and hoping the requests overlap makes the test depend on machine speed and luck. Give the test a deadline so a broken synchronization step fails clearly.
Then simplify everything that doesn’t affect the failure. You may need a real database to reproduce this race, but you probably don’t need a real email provider, a browser, and a full customer account. I want the smallest setup that preserves the behavior we’re trying to explain. That makes it easier to see whether the proposed fix actually addresses it.
3. Put the test at the boundary that matters
For this example, write down the rule: each tenant can have only one job for a particular import identifier. A test should check the stored result after concurrent requests. Checking that a helper function ran twice tells us very little about whether duplicate work can still enter the system.
A PostgreSQL unique constraint on the tenant and import identifier can enforce that particular rule across rows. Make the identifiers non-null if missing values are invalid; ordinary unique constraints allow multiple null values by default. Existing duplicates need a deliberate cleanup plan before adding the constraint.
Reference: PostgreSQL: Unique and not-null constraints
The application also needs a defined response when another request wins. In this hypothetical API, returning the existing job could be appropriate. Sending a generic server error and letting callers retry indefinitely would create a separate problem. Test both the database outcome and the response promised to the caller. Job execution still needs its own protection against duplicate delivery; one stored job does not settle that question.
Google’s reliability testing guidance describes regression tests as a way to preserve knowledge of previous failures. For this change, I’d first verify that the test fails against the broken behavior, then that it passes with the fix. If it passes both versions, investigate whether the test actually exercises the failure.
Reference: Google SRE: Testing for Reliability
4. Follow the assumption into nearby code
Once I understand an assumption that failed, I want to know where else we rely on it. In the import example, search for other ways to create the same job: a scheduled task, an administrator tool, a retry handler, or an older API route. Fixing only the public endpoint leaves those paths worth investigating.
Keep this search bounded. Start with callers that share the same data and rule. Finding one concurrency bug is not enough evidence to rewrite the entire backend. It is enough evidence to inspect the other writers and make sure the rule holds where the data is stored. Record unrelated findings separately so they don’t swallow the repair.
- Which entry points can create or change this state?
- Can a background worker replay the operation after a crash?
- Does a timeout leave the caller unsure whether the first request completed?
- Can an older deployed version bypass the new behavior?
- What existing data needs repair, and how will that repair be checked?
This is where I want curiosity to earn its keep. Every extra investigation should answer a question that changes the fix, the rollout, or our confidence in it. When the answer has no bearing on this failure, write it down and keep moving.
5. Decide what will make the fix complete
A passing test gives us evidence about the conditions it covers. Production adds a separate check. For the import example, inspect duplicate job creation, caller errors, and completed imports through a representative busy period. Include the scheduled workload that previously triggered the issue if one was involved.
Write the review condition before rollout. ‘We’ll keep an eye on it’ leaves too much to memory. ‘The owner will compare the next import cycle with the previous failure window and inspect every uniqueness conflict’ gives someone a concrete job. A conflict may mean the protection is working, while a sharp rise could reveal a caller that still needs attention.
Failed behavior and customer impact:
Condition the previous fix left possible:
Reproduction and original failing version:
Rule the new test protects:
Other writers or callers inspected:
Existing data repair, if needed:
Rollout and recovery conditions:
Production evidence to review:
Owner and review date:I expect to keep making mistakes. There’s always more to learn. What I want from each failure is something the next piece of work can benefit from: a clearer rule, a useful test, a better signal, or a smaller set of ways to break the same thing. That’s how the time spent debugging starts paying us back.
