Cloud database reliability engineering

Database reliability decisions your team can defend.

I help engineering teams running PostgreSQL and AWS Aurora/RDS diagnose performance and resilience risks, test candidate fixes, and deliver an evidence-backed change plan.

Alex Ivanov · 35+ years across databases and data platforms

PostgreSQL · AWS Aurora/RDS · Snowflake · Redshift · SQL Server · Oracle

When to call

A consequential database decision needs independent depth.

Before change

A migration, upgrade, failover design, or architecture choice carries material risk.

During uncertainty

Performance symptoms persist and several remedies appear equally plausible.

After repetition

Incidents recur because evidence, ownership, runbooks, and change controls do not connect.

What you receive

Enough evidence to act—and clear boundaries where evidence stops.

01

Findings

What is known, what is risky, and what still needs evidence.

02

Options

Candidate remedies compared against explicit tradeoffs.

03

Change plan

The recommended action, prerequisites, rollback, and authority boundaries.

04

Verification

Acceptance criteria, evidence record, and operational follow-through.

Engagements

Bounded work. Concrete outputs.

Expert-led consulting delivered with product discipline—not a SaaS product and not open-ended platform transformation.

Selected case study

The plan improved. The outcome did not.

A PostgreSQL index changed the execution plan but failed the latency test. The investigation continued until a different design produced a material local result.

See the decision, evidence, and limitations →
99.09%median reporting improvement
97.05%median mixed-workload improvement
randomized local repetitions
0production approvals claimed

Field Notes

What difficult database incidents actually teach.

Short accounts from the field: the problem, the decision, the outcome, and the operating practice that followed.

Recovery · SQL Server

A backup is not real until you can restore it.

A Friday release damaged a live e-commerce database. The immediate recovery mattered; the durable result was a restore discipline the company could rely on.

Read the field note →
Diagnosis · PostgreSQL

The PostgreSQL database that was never lost.

PostgreSQL was running and the expected data was gone. The breakthrough came from questioning which instance was actually answering on port 5432.

Read the field note →
Fleet reliability · AWS Aurora

A fleet report is not an operating model.

A reusable assessment surfaced likely database pressure and likely excess capacity. It also revealed why visibility without ownership and follow-through does not become reliability.

Read the field note →

How I work

From symptom to verified decision.

Automation can accelerate evidence and drafts. Production authority remains explicit, reviewable, and human.

01Observe
02Diagnose
03Compare
04Decide
05Control
06Verify
Alex Ivanov, founder of Cloud DBRE Tech

Alex Ivanov

Tools can draft a fix. Judgment decides whether it is safe to trust.

I started in software and database work in 1990 and moved through DBA, database engineering, data engineering, and cloud platforms. I bring the context AI does not own: operational tradeoffs, contradictory evidence, failure history, change authority, and accountability for a recommendation your team can review.

Start with the decision

What database risk or performance question does your team need to resolve?

Discuss the problem