Eduard Tamsa

Eduard Tamsa

Software thinkerer

All articles

Reliable Systems Make Their Assumptions Visible

Many systems appear simple because their assumptions are invisible.

A deployment works because a particular repository exists, a secret has the expected name, the network can reach a dependency, and somebody remembered to create a namespace months ago. None of those conditions may be written down. They are simply part of the environment until one of them changes.

Then the system fails in a way that feels surprising.

The failure is not always caused by a bad change. Often it is caused by an old assumption finally becoming false.

Reliable systems make those assumptions visible before production has to reveal them.

Defaults Are Still Decisions

Every default carries an opinion.

A timeout assumes how long an operation should take. A retry policy assumes which failures are temporary. A naming convention assumes who will read the name. A default region assumes where users, data, and dependencies live.

Defaults reduce repeated work, which is useful. The danger begins when teams stop treating them as decisions.

If a default affects reliability, security, cost, or recovery, document why it exists and where it applies. A value without context becomes difficult to review. People either preserve it forever because they are afraid to change it, or replace it casually because it looks arbitrary.

Good defaults should be easy to use and easy to question.

Preconditions Belong Beside the Automation

Automation often starts in the middle of a process.

A pipeline deploys an application, but assumes the cloud account, permissions, DNS zone, certificates, and observability already exist. A script rotates a credential, but assumes every consumer can reload it. A job deletes old resources, but assumes labels are accurate.

These preconditions are part of the automation even if the code does not create them.

Put them close to the workflow. Validate the ones a machine can check. Link to the owner and setup procedure for the ones it cannot. Fail early with a message that explains the missing condition instead of continuing until a later command produces a confusing error.

The best failure is often the one that says, clearly, what the system expected.

Make Environmental Knowledge Queryable

Teams regularly store important facts in memory.

Which cluster handles production? Which account owns the DNS zone? Which queue can be replayed? Which database is authoritative? Which service must be restarted after a certificate change?

If the answer depends on finding the right person, the information is not operationally available.

Move stable facts into versioned configuration, service catalogs, runbooks, or machine-readable metadata. Prefer sources that can be queried by both people and automation. Avoid copying the same fact into five documents that will drift in five different directions.

The goal is not to build a perfect inventory. It is to make the next important decision depend less on memory.

Test the Conditions Around the Happy Path

Most tests confirm that the code behaves correctly when its environment matches expectations.

Reliability improves when tests also challenge those expectations.

What happens when a required variable is missing? When a dependency responds slowly? When the identity has read access but not write access? When the input uses an old schema? When the resource already exists in a partially configured state?

These are not exotic edge cases. They are normal ways that real environments differ from a developer’s laptop.

Test the boundaries that matter most. You do not need to simulate every possible failure. Start with assumptions that could produce data loss, prolonged downtime, security exposure, or difficult recovery.

Observability Should Explain Context

A metric that says an operation failed is useful. A metric that also shows which assumption failed is much better.

Record the environment, version, dependency, policy decision, and correlation identifier needed to understand the event. Log expected and actual states without exposing secrets. Make dashboards reflect the customer-facing outcome, not only the health of individual processes.

This context reduces the amount of reconstruction required during an incident.

Observability should not force responders to rediscover the system’s operating model while the system is already broken.

Review Assumptions Like Interfaces

An assumption shared by multiple teams behaves like an interface.

If a platform team changes a label, directory structure, authentication flow, or deployment convention, every consumer that relies on it may be affected. The fact that no formal API exists does not make the dependency less real.

Identify these contracts and review changes accordingly. Give consumers a migration path. Detect old usage where possible. Set a removal date only after the replacement is usable and visible.

This is especially important for internal platforms. Their most important interfaces are often conventions rather than endpoints.

Remove Assumptions That No Longer Help

Not every assumption deserves better documentation. Some should be eliminated.

If every repository must contain the same copied configuration, generate it or provide a shared component. If every team must request the same permission manually, consider a safe default role. If a deployment requires steps in a specific undocumented order, encode the dependency.

Visibility is the first step because it shows where complexity lives. Simplification is the next step.

The strongest platform is not the one with the longest list of documented rules. It is the one that needs fewer rules to produce a safe result.

Final Thought

Systems are built on assumptions about people, infrastructure, timing, data, and failure.

Those assumptions will change. Teams will reorganize. Dependencies will become slower. Names will evolve. Permissions will tighten. Traffic will grow beyond the original design.

Make important assumptions visible. Validate them early. Include them in tests and observability. Treat shared conventions as interfaces, and remove unnecessary conditions whenever you can.

Reliability is not only about handling known failures. It is also about reducing the number of invisible beliefs that production is expected to keep true.