←Back to Blogs

What Systems Believe When Things Go Wrong

January 12, 2026

Data EngineeringSystem DesignDistributed SystemsIdempotencyFailure Handling

This blog explores how data systems behave when things don’t go as planned. Instead of focusing on tools or ideal flows, it looks at partial failures, retries, and ambiguous states—and how the assumptions systems make during those moments quietly shape correctness, trust, and long-term reliability.

Most Systems Are Designed for the Happy Path

Happy path diagram showing the ideal flow of data through a system

When we build data systems, our thinking almost always starts with the happy path. Data arrives when it’s supposed to. Processing finishes within expected time. Writes succeed. Dashboards refresh. As long as these steps happen in sequence, everything feels predictable and controlled.

Architecturally, this creates a comforting mental model. Data comes in, gets transformed, is stored somewhere durable, and is later queried. Each stage feels discrete, well-defined, and cleanly separated from the next.

The problem is that real systems rarely break in dramatic, obvious ways. They don’t usually crash outright or stop functioning entirely. Instead, they fail partially — quietly and inconsistently.

A subset of records fails while the rest succeed. A batch job runs long and times out halfway through. A retry kicks in while a previous attempt is still in progress. A downstream consumer sees the same input twice and processes it twice because nothing explicitly told it not to.

In those moments, the question subtly changes. It’s no longer “does the system work?”
The real question becomes: what does the system believe happened?


Partial Failure Is Where Systems Become Ambiguous

Partial completion is one of the hardest things to reason about in distributed systems, largely because it forces you to confront what “done” actually means.

Imagine a pipeline where 10,000 records enter for processing and 9,700 of them are written successfully. No fatal error is thrown. The job reaches the end of its execution. On the surface, everything looks fine.

But what is the correct interpretation of that outcome?

Is the job successful because most of the data made it through?
Is it failed because some records didn’t?
Or is it incomplete because the system never reconciled expectation with reality?

Most systems don’t answer these questions explicitly. They infer meaning from execution behavior. If no exception bubbles up, the job is marked successful. If an exception appears, the job is marked failed. Completion becomes an assumption, not a fact.

Different systems make different assumptions. Some retry the entire batch, risking duplicate writes for the records that already succeeded. Some retry only the failures, assuming the successful records are immutable and safely persisted. Others do nothing at all, quietly accepting partial output as “good enough.”

Each of these choices embeds a belief about state, correctness, and trust. Most of the time, those beliefs are never written down.


A Concrete Failure Walkthrough

Partial failure in parallel batch processing with retries and duplicates

Consider a simple ingestion job that runs once a day.

It reads a file containing transaction records, processes them in parallel chunks, and writes the results to a database. Each chunk contains 1,000 records. Ten chunks run concurrently.

On a normal day, all ten chunks finish and the job marks itself complete.

One day, chunk number seven fails halfway through due to a transient database timeout. The other nine chunks succeed and commit their writes. The job framework retries chunk seven automatically.

This is where things start to get interesting.

During the retry, chunk seven reprocesses all 1,000 records. But 400 of those records were already written successfully before the failure occurred. If the system doesn’t have a strong idempotency guarantee at the database layer, those 400 records are now duplicated.

If the database does enforce uniqueness, the retry fails again — but for a different reason. The system now sees constraint violations instead of timeouts. Depending on how retries are configured, the job may exhaust its retry limit and land in a dead-letter state.

At this point, what is the truth?

Nine thousand four hundred records are safely written.
Six hundred records are missing.
The job is marked as failed.
Re-running the job risks duplicating everything unless carefully scoped.

If someone asks whether the data for that day is “ready”, the answer depends entirely on how well the system tracked state. Did it record which chunks completed? Did it persist partial progress? Did it distinguish between not processed yet and processed but not acknowledged?

If it didn’t, the system itself doesn’t actually know what happened. Humans are left to reconstruct reality from logs.


Why This Is Harder Than It Looks

What makes these scenarios tricky isn’t the failure itself. It’s the ambiguity that follows.

Retries blur causality. Parallelism breaks linear reasoning. Success signals don’t necessarily mean correctness. Without explicit state tracking, systems end up guessing whether they are complete.

This is why two systems built with the same tools can behave radically differently under stress. The difference isn’t the framework or the infrastructure. It’s whether the design makes state, expectation, and completion explicit.


The Shift That Changed How I Think

Idempotency and state transitions: running twice should give the same result

At some point, I stopped asking whether a system could scale or handle load. I started asking a simpler but more uncomfortable question:

If this runs twice, what exactly happens?

That question forces clarity. It exposes hidden assumptions about uniqueness, ordering, retries, and time. It reveals whether correctness is enforced or merely hoped for.

If I can explain a system’s behavior when things partially fail — without hand-waving or relying on tool names — I usually understand it. If I can’t, the system probably doesn’t either.