Skip to content

IT disaster recovery at the U: Planning for the ‘what ifs’

Kim Tanner, director, Platform Services, Chief Technology Officer organization

You may have heard “disaster recovery,” or DR, mentioned in a meeting without explanation — but what does it mean?

In IT, a disaster is any event that seriously disrupts the systems the university depends on, whether that’s a ransomware attack, a fire or flood, or network failure. Recovery is the planning and practice that gets those systems back up quickly, so things like payroll and student registration aren’t disrupted for long. It’s more than restoring a backup: it’s deciding which services come back first, knowing who does what, and making sure teams coordinate when it matters.

Where the program started

At the University of Utah Health, Phil Kimball, associate director for Service Management in the Chief Technology Officer (CTO) organization and Mike Madsen, manager for the ITSM Process Support and Disaster Recovery team, have been “tag-teaming DR responsibilities for years,” according to Emily Rushton, senior IT disaster recovery program coordinator for Madsen’s team. Kimball and Madsen, she said, have been running tabletop exercises — discussion-based drills based on a realistic scenario — as far back as fiscal year 2018.

Emily Rushton, senior IT disaster recovery program coordinator, ITSM Process Support and Disaster Recovery, CTO organization

On campus, a similar story has played out. For close to a decade, UIT has led a series of projects to expand DR coverage for applications and services, scoped and budgeted a year at a time.

What changed a year and a half ago is that U of U Health funded a dedicated DR program team. Rushton was hired in February 2025 as the first person at the U fully dedicated to IT disaster recovery, and Brian Stephens soon joined as a DR cloud architect for the IT Enterprise Architecture team.

On campus, PeopleSoft has been the most recent focus for Jason Moeller, senior director for University Support Services (USS), Kim Tanner, director for CTO Platform Services, and John Nielson, manager for Application Platform Support Services. PeopleSoft is the platform behind many IT systems, software, and web applications used by students, staff, and faculty at the U, like finance, HR, and student records.

Tanner said the CTO DR teams meet regularly to review progress, align priorities, and coordinate recovery planning efforts. The regular collaboration, she said, “strengthens partnerships across teams and helps maintain a consistent approach to disaster recovery planning.”

DR and business continuity: What’s the difference?

“Disaster recovery” and “business continuity” are often used interchangeably, but they mean different things. Disaster recovery is about IT: how data and systems are recovered after something goes sideways. Business continuity is about the business: how critical operations keep running even while IT is down.

Rushton’s go-to example is Epic, U of U Health’s core clinical system. If Epic goes down, DR is what brings it back, including a cloud-based environment built for that purpose. Business continuity is hospital staff switching to paper charting so patient care doesn’t stop.

For the full policy framework behind all of this, access Information Security Policy 4-004, which governs how the University of Utah approaches business continuity and disaster recovery planning.

The biggest tabletop exercise yet

On June 12, 2026, the U’s DR program ran its largest tabletop exercise to date — 166 participants joined a Microsoft Teams call for a six-hour exercise. A tabletop exercise is a structured, hypothetical scenario that tests how teams communicate and make decisions under pressure, without touching a single production system.

A tabletop, Tanner said, “enables our organization to validate our current capabilities, uncover operational gaps, and refine response frameworks, including communication protocols and recovery plans.”

What Rushton wants people to take away isn’t the scenario itself but the mindset behind it. A tabletop is inexpensive compared to a real outage. It’s meant to be the place where decisions stall, where access is missing, where dependencies hide, and where communication frays — on purpose, in a low-stakes setting, so those things get fixed before they matter.

“The whole point is to surface the ‘we don’t know’ moments now, not during an actual disaster,” she said.

Looking ahead, the DR program is already in early scenario definition for next year’s exercise, considering a mix of approaches, including smaller, more frequent exercises focused on just one or two teams alongside the large annual event.

DR support for a growing list of applications

Elements of a DR plan

Communication

The team responsible for creating, implementing, and managing the disaster recovery plan must designate and communicate roles and responsibilities.

Recovery timeline

Time frames for systems returning to normal operations must be identified and should address:

  • Recovery time objective (RTO), a metric that determines the maximum amount of time that passes before disaster recovery is completed.
  • Recovery point objective (RPO), the maximum amount of time acceptable for data loss after a disaster.

Data backups

Options include cloud storage, vendor-supported backups, and internal offsite data backups. To account for natural disasters, backups should not be onsite.

Testing

DR plans must be regularly tested. Similarly, security and data protection strategies should be updated frequently to prevent unauthorized access.


Source: Amazon Web Services

The PeopleSoft DR effort brings together teams from across UIT, in coordination with Moeller, whose team works closely with business leaders to identify their recovery needs and determine system priorities. Building out and delivering these priorities is a joint effort across multiple areas, from CTO Platform Services, Core Infrastructure Services, Identity & Access Management in the Information Security Office, IT Enterprise Architecture team, Project Management Office, and several others.

Teams work together to define the scope of the DR requirement, assess level of effort, identify resources required, and develop a funding proposal. Tanner said the work is organized as fiscal-year projects, though it sometimes stretches into the next year depending on volume and testing involved.

UIT recently documented a tiered list of services, from core infrastructure and immediate safety at the top to services USS doesn’t manage — mirroring a tiering system the hospitals and clinics already use, since applications can’t come back up until the IT systems underneath them are running.

Achieving this is harder than it sounds, largely because of how interconnected university systems have become, Tanner said, comparing it to building a house where intricacies multiply well beyond the original blueprint.

Upcoming DR training

A new self-paced DR training module, currently in the pilot phase, will be available through U of U Health’s Learning Management System (LMS) and automatically assigned to all Information Technology Services (ITS) employees, with completion tracked. Rushton said the program will help satisfy audit requirements and accreditation reviews that ask whether the university has a formal DR training program. The goal is to have it live before year’s end, after which UIT teams will adapt the content for UIT employees and roll it out through the LearningHub LMS on the same annual basis.

What you can do

Asked what she’d want every IT professional at the U to take away, Rushton said if you work in IT, “disaster recovery isn’t someone else’s job, it’s yours, too.”

Her challenge to colleagues?

“Think about what keeps you up at night. What in your area worries you? Is your recovery process documented? Have you tested your service recently? And then — do something about it.”

Share this article:

 

Node 4

Our monthly newsletter includes news from UIT and other campus/ University of Utah Health IT organizations, features about UIT employees, IT governance news, and various announcements and updates.

Subscribe

Categories

Featured Posts

Last Updated: 9/30/26