FEATURED PROJECT

Recovering 56 Healthcare Clinics

A large-scale production recovery involving dozens of electronic medical record databases, coordinated technical teams, and restoring critical healthcare systems safely and consistently.

The phone call that changed my vacation.

I was on vacation when I received a phone call from work. During a software deployment, a database restore had been performed against the wrong production database, resulting in a production overwrite that affected dozens of healthcare clinics. Electronic Medical Record systems became unavailable, and a large, coordinated recovery effort was immediately launched.

I immediately shifted from vacation mode into incident mode, coordinating the database recovery effort and helping guide the work across infrastructure, application, and support teams. The priority wasn't simply restoring databases—it was safely restoring the Electronic Medical Record systems that physicians, nurses, and clinic staff relied on every day to care for patients.

As more information became available, it was clear this would be one of the largest recovery efforts of my career.

The recovery was more complex than simply restoring a single database. Approximately 56 clinic schemas existed within the affected environment, and not every clinic had been impacted by the overwrite. Some remained fully operational, which meant we had to identify exactly which schemas required recovery while ensuring that healthy clinic data was left untouched.

Every decision required careful validation. Restoring the wrong schema or recovering data unnecessarily could have introduced additional risk. Accuracy was just as important as speed.

The database team quickly came together to develop a structured recovery plan. Working closely with infrastructure, application, and technical support teams, we prioritized affected clinics, coordinated recovery activities, and maintained clear communication throughout the incident so that every team understood the current status and next steps.

Recover safely. Recover consistently.

Recovering a single production database requires careful validation. In this case, approximately 56 clinic schemas existed within the affected database, and each represented a live healthcare environment. Some schemas required recovery while others remained healthy, so we could not simply restore the entire database and overwrite everything.

We developed a parallel recovery strategy using multiple sources. One recovery stream used the disaster recovery environment, placing the standby database into snapshot standby so that data could be recovered without disrupting the production recovery process. Another stream restored a Veeam backup to a separate server. From those environments, we exported the affected clinic schemas using Oracle Data Pump and imported them individually back into production.

Our five-person DBA team divided the work so recovery could happen in parallel. Two DBAs focused on Data Pump exports from the Veeam and DR recovery environments, while the remaining DBAs handled imports into production, validation, and returning recovered clinics to service.

Every schema had to be tracked carefully. Before anything was restored, we needed to confirm that the clinic had actually been affected, identify the correct recovery source, and ensure that healthy production data was never overwritten. The challenge was not simply moving data quickly—it was performing dozens of precise recoveries without creating a second incident.

Multiple recovery paths. One coordinated effort.

Because not every clinic schema had been affected, we could not simply restore the entire production database. We needed a controlled process that identified the affected schemas, recovered them from safe sources, validated the data, and returned each clinic to service without disturbing healthy production data.

01

Identify Affected Schemas

Confirm which clinic schemas had been overwritten and which remained healthy in production.

02

Establish Recovery Sources

Use multiple recovery paths, including the disaster recovery environment and a Veeam backup restored to a separate server.

03

Prepare DR for Recovery

Place the standby database into snapshot standby so data could be recovered safely without disrupting the ongoing production recovery effort.

04

Export Clinic Data

Use Oracle Data Pump to export affected clinic schemas from the available recovery environments.

05

Restore in Parallel

Split the five-person DBA team into parallel work streams, with DBAs handling exports while others imported recovered schemas back into production.

06

Validate Before Release

Verify each recovered schema before returning the clinic to service and confirm that unaffected production data remained untouched.

Bringing structure to a complex recovery.

I helped coordinate the database recovery effort, working across the DBA, infrastructure, application, and technical support teams to keep the recovery organized and moving forward. Within the DBA team, we divided the workload across the available recovery sources, tracked which clinic schemas required action, and coordinated exports, imports, validation, and return-to-service activities.

My focus was on bringing structure to a situation that could easily have become chaotic. We needed to move quickly, but every recovery had to be deliberate and validated. I helped keep the team aligned, identify priorities, communicate progress, and make sure that speed never came at the expense of data integrity.

Recovering service while improving future processes.

The affected clinic schemas were successfully recovered and returned to service without overwriting the clinics that had remained healthy. What began as a large production incident became a coordinated, parallel recovery effort involving multiple backup sources, five DBAs, and close collaboration across several technical teams.

The incident reinforced something I had learned throughout my career: a backup is only one part of recoverability. Successful recovery also depends on knowing where your data is, having multiple recovery options, documenting the process, validating every step, and having people who can work together calmly when the pressure is high.

Every major incident leaves behind something valuable.

Preparation Matters

Recovery begins long before an outage occurs.

Communication Builds Confidence

Keeping teams informed reduces confusion and allows technical experts to focus on solving problems.

Stay Calm

The best decisions during production incidents come from clear, methodical thinking—not panic.