The PACS Outage Runbook: A Step-by-Step Checklist for Imaging Teams
A practical PACS outage runbook: confirm scope, set severity, protect clinical workflow, escalate, communicate, restore, and learn from every incident.
By Trisha Seal — September 24, 2026. Trisha documents incident runbooks, escalation paths, and recovery procedures within RAD365's vendor-agnostic PACS support practice. RAD365 does not read or interpret studies.
Why Every Imaging Department Needs a PACS Outage Runbook
A PACS outage runbook is the difference between a calm, ten-minute response and an hour of phone calls at 2 a.m. When PACS goes down, the people on shift are rarely the people who built the system. A written checklist tells them exactly what to check, who to call, and how to keep patient care moving while the problem is fixed.
This guide gives you a practical runbook structure you can adapt. It pairs well with our quick-reference posts on what to check first when PACS is down and what to check when PACS is slow.
The Runbook at a Glance
| Phase | Goal | Key output |
|---|---|---|
| 1. Detect and confirm | Know what is actually broken | Scope statement (who and what is affected) |
| 2. Classify | Set urgency | Severity level and response clock |
| 3. Protect workflow | Keep patient care moving | Downtime procedures activated |
| 4. Investigate | Find the failing layer | Likely cause and evidence |
| 5. Escalate | Get the right people involved | L2, networking, and manufacturer engaged |
| 6. Communicate | Keep stakeholders informed | Status updates on a set cadence |
| 7. Restore and verify | Return to normal safely | Queued studies delivered, users confirmed |
| 8. Review | Prevent recurrence | Root cause and runbook updates |
Phase 1: Detect and Confirm
Most outages are first reported by a clinician, not an alert. The first job is to confirm scope:
- Is it one workstation, one department, one site, or every user?
- Is it one modality failing to send, or all modalities?
- Can users log in at all, or do they log in but see no images?
- Did anything change recently: an upgrade, a patch, a network change, a certificate renewal?
Write the answers into the ticket. A clear scope statement stops the team from chasing the wrong layer.
Phase 2: Classify Severity
Use written definitions so nobody debates urgency in the moment. RAD365's PACS support framework uses four levels:
| Severity | Definition | Response commitment |
|---|---|---|
| Severity 1 | System down with clinical impact | 15 minutes; continuous work until resolved |
| Severity 2 | Degraded service | 1 hour |
| Severity 3 | Single-user or non-urgent issue | 4 business hours |
| Severity 4 | Standard request or change | Next business day |
Phase 3: Protect Clinical Workflow
While engineers investigate, clinical teams need a way to keep working. Your runbook should spell out, in plain language agreed with radiology leadership:
- Where modalities hold studies while PACS is unavailable, and how long they can hold them.
- How technologists and radiologists can view images in the meantime, such as at the modality console or an approved backup viewer.
- How urgent and emergency cases are communicated to the care team.
- Who tells the emergency department and inpatient units that imaging is in downtime.
Phase 4: Investigate Layer by Layer
Work from the outside in so you do not restart a healthy server for a network problem:
- Network: Can the workstation reach the PACS servers? Are switches, firewalls, and VPN links up?
- Authentication: Is the directory or single sign-on service working? Has a service account or certificate expired?
- Application services: Are the PACS application and web services running on every node?
- Database: Is the database online and responsive, with enough disk space for logs?
- Storage and archive: Are storage volumes mounted, healthy, and not full?
- Interfaces: Are DICOM and HL7 connections to modalities, RIS, and the EHR flowing, or are queues backing up?
Phase 5: Escalate on a Clock
A good runbook sets escalation triggers in advance rather than relying on judgment at 3 a.m. In a two-tier model, L1 handles application, access, and clinical-user support, while L2 handles infrastructure, DICOM and HL7 engineering, interface workflows, backup, and disaster recovery. List when each tier is engaged, when hospital networking is paged, and when the PACS manufacturer is called, along with the support number, contract ID, and the logs they will ask for.
Phase 6: Communicate on a Cadence
Silence creates more calls than the outage itself. Set an update rhythm by severity and use a simple template: what is affected, what is being done, what users should do meanwhile, and when the next update arrives.
Phase 7: Restore and Verify
Service is not restored just because the application starts. Before closing, confirm that:
- Modalities have sent every queued study and nothing is stuck in a send queue.
- The worklist, RIS orders, and EHR links are current.
- Users can open new studies and their priors at normal speed.
- Any studies viewed during downtime are reconciled in PACS.
Phase 8: Review and Improve
Every outage should end with a short root cause review and a runbook update. RAD365 follows an ITIL-compliant lifecycle of logging, root cause analysis, categorization and priority, escalation, resolution, and closure, so the same incident is less likely to return.
A Copy-Ready Outage Checklist
- ☐ Scope confirmed and written in the ticket
- ☐ Severity assigned and response clock started
- ☐ Downtime procedures activated and clinical teams notified
- ☐ Network, authentication, services, database, storage, and interfaces checked
- ☐ L2, networking, and manufacturer engaged per escalation triggers
- ☐ Status updates sent on schedule
- ☐ Queued studies delivered and users verified
- ☐ Root cause documented and runbook updated
Keeping the Runbook Alive
A runbook written once and filed away is almost as risky as none. Many hospitals, especially those relying on one PACS administrator, hand ownership to a managed PACS support partner that keeps runbooks, contacts, and monitoring current across every shift. If you are between providers or your administrator has left, interim PACS support can put a runbook in place quickly.
RAD365 is an operations and workflow partner providing PACS Support and the Radiology Workflow Manager only. It does not read or interpret studies and provides no preliminary-read services.
Put a Tested Runbook Behind Your PACS
RAD365 builds runbooks, escalation paths, and monitoring into onboarding so every shift knows exactly what to do.
See the PACS Support Framework →Frequently Asked Questions
Runbook Basics
What is a PACS outage runbook?
A PACS outage runbook is a written, step-by-step procedure that tells whoever is on shift how to confirm an outage, classify its severity, protect clinical workflow, investigate likely causes, escalate to the right people, communicate status, restore service, and close the incident. It turns individual memory into a repeatable process.
How is a runbook different from a disaster recovery plan?
A disaster recovery plan covers large-scale loss, such as a data center or archive failure, and how systems are rebuilt or failed over. A runbook covers the far more common day-to-day outages and degradations, and it usually links to the disaster recovery plan as one escalation path.
Who should own the PACS outage runbook?
One named owner should maintain it, typically the PACS administrator or the support partner's L2 team, with input from radiology leadership, IT networking, and the PACS manufacturer. Ownership matters because an unowned runbook quickly drifts out of date.
How often should a PACS runbook be reviewed?
Review it after every significant incident, after any upgrade, migration, or interface change, and on a regular calendar at least a few times a year. Each review should confirm contacts, access paths, and system dependencies are still accurate.
During an Outage
What is the very first step when PACS goes down?
Confirm scope before acting: is the problem one workstation, one site, one modality, or everyone? That single answer decides the severity level and which part of the runbook to follow, and it prevents the team from restarting servers for a problem that is really one user's network cable.
How do you classify the severity of a PACS outage?
Use written definitions. In RAD365's model, Severity 1 is a system down with clinical impact, Severity 2 is degraded service, Severity 3 is a single-user or non-urgent issue, and Severity 4 is a standard request or change. The severity sets the response clock and escalation path.
What downtime procedures should the runbook include for clinical staff?
It should tell technologists and radiologists how to keep working while PACS is unavailable, such as where modalities hold studies, how to view images at the modality or an approved backup viewer, and how urgent cases are communicated. These steps should be agreed with radiology leadership in advance.
When should the PACS manufacturer be contacted?
Contact the manufacturer once local checks point to an application, database, or product defect, or as soon as a Severity 1 incident is not quickly resolved. The runbook should list the support number, contract or entitlement ID, and the evidence to send, so no time is lost searching.
How often should status updates go out during an outage?
Set a cadence in the runbook based on severity, for example frequent updates during a Severity 1 incident and less frequent updates for degraded service. Each update should state what is affected, what is being done, and when the next update will come, even if nothing has changed.
After Recovery
What should happen after PACS service is restored?
Confirm that modalities have sent any queued studies, that the worklist and interfaces are current, and that users can open recent and prior exams. Then close the incident only after a clear verification step, not as soon as the application starts.
Why is root cause analysis part of the runbook?
Without root cause analysis, the same outage tends to return. An ITIL-style lifecycle includes logging, root cause analysis, categorization and priority, escalation, resolution, and closure, so the fix addresses the cause and the runbook is updated with what was learned.
How do you test a PACS outage runbook without a real outage?
Run tabletop exercises where the team walks through a scenario step by step, and schedule controlled failover or restore tests during agreed maintenance windows. Testing exposes missing contacts, expired credentials, and steps nobody actually knows how to perform.
Can a small hospital with one PACS administrator maintain a runbook?
Yes, and it matters even more there, because the runbook is what lets someone else respond when that one person is unavailable. Many smaller hospitals pair their runbook with an outside support partner so there is always a staffed team able to follow it.
Does RAD365 help build PACS runbooks?
Yes. Runbooks, escalation mapping, and monitoring are part of RAD365's onboarding, which typically takes two to four weeks and can be compressed to under a week when a hospital is already exposed. RAD365 supports PACS operations only and does not read or interpret studies.