The PACS Outage Runbook: A Step-by-Step Checklist for Imaging Teams

A practical PACS outage runbook: confirm scope, set severity, protect clinical workflow, escalate, communicate, restore, and learn from every incident.

The PACS Outage Runbook: A Step-by-Step Checklist for Imaging Teams

By Trisha Seal — September 24, 2026. Trisha documents incident runbooks, escalation paths, and recovery procedures within RAD365's vendor-agnostic PACS support practice. RAD365 does not read or interpret studies.

Why Every Imaging Department Needs a PACS Outage Runbook

A PACS outage runbook is the difference between a calm, ten-minute response and an hour of phone calls at 2 a.m. When PACS goes down, the people on shift are rarely the people who built the system. A written checklist tells them exactly what to check, who to call, and how to keep patient care moving while the problem is fixed.

This guide gives you a practical runbook structure you can adapt. It pairs well with our quick-reference posts on what to check first when PACS is down and what to check when PACS is slow.

The Runbook at a Glance

PhaseGoalKey output
1. Detect and confirmKnow what is actually brokenScope statement (who and what is affected)
2. ClassifySet urgencySeverity level and response clock
3. Protect workflowKeep patient care movingDowntime procedures activated
4. InvestigateFind the failing layerLikely cause and evidence
5. EscalateGet the right people involvedL2, networking, and manufacturer engaged
6. CommunicateKeep stakeholders informedStatus updates on a set cadence
7. Restore and verifyReturn to normal safelyQueued studies delivered, users confirmed
8. ReviewPrevent recurrenceRoot cause and runbook updates

Phase 1: Detect and Confirm

Most outages are first reported by a clinician, not an alert. The first job is to confirm scope:

Write the answers into the ticket. A clear scope statement stops the team from chasing the wrong layer.

Phase 2: Classify Severity

Use written definitions so nobody debates urgency in the moment. RAD365's PACS support framework uses four levels:

SeverityDefinitionResponse commitment
Severity 1System down with clinical impact15 minutes; continuous work until resolved
Severity 2Degraded service1 hour
Severity 3Single-user or non-urgent issue4 business hours
Severity 4Standard request or changeNext business day

Phase 3: Protect Clinical Workflow

While engineers investigate, clinical teams need a way to keep working. Your runbook should spell out, in plain language agreed with radiology leadership:

Phase 4: Investigate Layer by Layer

Work from the outside in so you do not restart a healthy server for a network problem:

  1. Network: Can the workstation reach the PACS servers? Are switches, firewalls, and VPN links up?
  2. Authentication: Is the directory or single sign-on service working? Has a service account or certificate expired?
  3. Application services: Are the PACS application and web services running on every node?
  4. Database: Is the database online and responsive, with enough disk space for logs?
  5. Storage and archive: Are storage volumes mounted, healthy, and not full?
  6. Interfaces: Are DICOM and HL7 connections to modalities, RIS, and the EHR flowing, or are queues backing up?

Phase 5: Escalate on a Clock

A good runbook sets escalation triggers in advance rather than relying on judgment at 3 a.m. In a two-tier model, L1 handles application, access, and clinical-user support, while L2 handles infrastructure, DICOM and HL7 engineering, interface workflows, backup, and disaster recovery. List when each tier is engaged, when hospital networking is paged, and when the PACS manufacturer is called, along with the support number, contract ID, and the logs they will ask for.

Phase 6: Communicate on a Cadence

Silence creates more calls than the outage itself. Set an update rhythm by severity and use a simple template: what is affected, what is being done, what users should do meanwhile, and when the next update arrives.

Phase 7: Restore and Verify

Service is not restored just because the application starts. Before closing, confirm that:

Phase 8: Review and Improve

Every outage should end with a short root cause review and a runbook update. RAD365 follows an ITIL-compliant lifecycle of logging, root cause analysis, categorization and priority, escalation, resolution, and closure, so the same incident is less likely to return.

A Copy-Ready Outage Checklist

Keeping the Runbook Alive

A runbook written once and filed away is almost as risky as none. Many hospitals, especially those relying on one PACS administrator, hand ownership to a managed PACS support partner that keeps runbooks, contacts, and monitoring current across every shift. If you are between providers or your administrator has left, interim PACS support can put a runbook in place quickly.

RAD365 is an operations and workflow partner providing PACS Support and the Radiology Workflow Manager only. It does not read or interpret studies and provides no preliminary-read services.

Put a Tested Runbook Behind Your PACS

RAD365 builds runbooks, escalation paths, and monitoring into onboarding so every shift knows exactly what to do.

See the PACS Support Framework →

Frequently Asked Questions

Runbook Basics

What is a PACS outage runbook?

A PACS outage runbook is a written, step-by-step procedure that tells whoever is on shift how to confirm an outage, classify its severity, protect clinical workflow, investigate likely causes, escalate to the right people, communicate status, restore service, and close the incident. It turns individual memory into a repeatable process.

How is a runbook different from a disaster recovery plan?

A disaster recovery plan covers large-scale loss, such as a data center or archive failure, and how systems are rebuilt or failed over. A runbook covers the far more common day-to-day outages and degradations, and it usually links to the disaster recovery plan as one escalation path.

Who should own the PACS outage runbook?

One named owner should maintain it, typically the PACS administrator or the support partner's L2 team, with input from radiology leadership, IT networking, and the PACS manufacturer. Ownership matters because an unowned runbook quickly drifts out of date.

How often should a PACS runbook be reviewed?

Review it after every significant incident, after any upgrade, migration, or interface change, and on a regular calendar at least a few times a year. Each review should confirm contacts, access paths, and system dependencies are still accurate.

During an Outage

What is the very first step when PACS goes down?

Confirm scope before acting: is the problem one workstation, one site, one modality, or everyone? That single answer decides the severity level and which part of the runbook to follow, and it prevents the team from restarting servers for a problem that is really one user's network cable.

How do you classify the severity of a PACS outage?

Use written definitions. In RAD365's model, Severity 1 is a system down with clinical impact, Severity 2 is degraded service, Severity 3 is a single-user or non-urgent issue, and Severity 4 is a standard request or change. The severity sets the response clock and escalation path.

What downtime procedures should the runbook include for clinical staff?

It should tell technologists and radiologists how to keep working while PACS is unavailable, such as where modalities hold studies, how to view images at the modality or an approved backup viewer, and how urgent cases are communicated. These steps should be agreed with radiology leadership in advance.

When should the PACS manufacturer be contacted?

Contact the manufacturer once local checks point to an application, database, or product defect, or as soon as a Severity 1 incident is not quickly resolved. The runbook should list the support number, contract or entitlement ID, and the evidence to send, so no time is lost searching.

How often should status updates go out during an outage?

Set a cadence in the runbook based on severity, for example frequent updates during a Severity 1 incident and less frequent updates for degraded service. Each update should state what is affected, what is being done, and when the next update will come, even if nothing has changed.

After Recovery

What should happen after PACS service is restored?

Confirm that modalities have sent any queued studies, that the worklist and interfaces are current, and that users can open recent and prior exams. Then close the incident only after a clear verification step, not as soon as the application starts.

Why is root cause analysis part of the runbook?

Without root cause analysis, the same outage tends to return. An ITIL-style lifecycle includes logging, root cause analysis, categorization and priority, escalation, resolution, and closure, so the fix addresses the cause and the runbook is updated with what was learned.

How do you test a PACS outage runbook without a real outage?

Run tabletop exercises where the team walks through a scenario step by step, and schedule controlled failover or restore tests during agreed maintenance windows. Testing exposes missing contacts, expired credentials, and steps nobody actually knows how to perform.

Can a small hospital with one PACS administrator maintain a runbook?

Yes, and it matters even more there, because the runbook is what lets someone else respond when that one person is unavailable. Many smaller hospitals pair their runbook with an outside support partner so there is always a staffed team able to follow it.

Does RAD365 help build PACS runbooks?

Yes. Runbooks, escalation mapping, and monitoring are part of RAD365's onboarding, which typically takes two to four weeks and can be compressed to under a week when a hospital is already exposed. RAD365 supports PACS operations only and does not read or interpret studies.

Related Services