DOCS / EN

Backup restore test — safe procedure

A detailed safe backup restore test procedure covering authorisation, isolation, restore-point selection, validation, evidence, cleanup, and retesting.

This procedure checks whether a selected backup can be restored and used safely. It does not replace a continuity plan, vendor instructions, or decisions by the system owner. One successful test confirms only that backup, scope, and test conditions.

Use the backup restore test worksheet to record the exercise. Ransomware-resistant backups: boundaries, retention, and restore tests explains boundaries, retention, and separate credentials. If this test forms part of a ransomware response, structure the decisions first with the first-hour decision card.

Critical rule: never attach the only backup to a suspect, compromised, or unverified environment. Work from a controlled working copy or through a safe read-only access mechanism, leaving the original intact.

1. Define scope and authorisation

  1. Name the test owner, restore operator, system owner, and person who will accept the result. At least one named person must have authority to stop the exercise.
  2. Obtain written approval to use the backup, restore data into the test environment, and remove that copy later. Include privacy and evidence-retention requirements.
  3. Define the exact scope: system, database, machine, directory, mailbox, or object set; data and period; exclusions; and the expected level of operation after recovery.
  4. Set pass and stop criteria, including an unexpected production connection, malicious code, risk of overwriting the backup, or data disclosure.
  5. Give the exercise a change, ticket, or incident number. Record dates and times with their timezone.

Do not begin without an owner, approval, and safe location. Scope and responsibilities can be developed through system administration, IT administration, or, during an event, incident response.

2. Record time and data-loss objectives

Record the RTO: the target time from starting recovery until the agreed function is usable. In plain language, how long the system may remain unavailable. Define the measurement endpoints; extracted files alone do not make an operational service.

Record the RPO: maximum acceptable data loss expressed as time. In plain language, how old the restored data may be. With a four-hour RPO, a twelve-hour-old point fails even if usable. Record missing objectives as risks; do not invent them during the exercise.

3. Select and protect the restore point

  1. Identify the unwanted change and record its earliest and latest known time plus the information source.
  2. Choose a point that evidence shows predates the change, not simply the newest backup with a green status. Use trusted job logs, version history, an incident timeline, file metadata, application logs, or monitoring.
  3. Add a safety margin when the start of the change is uncertain. A point from before discovery may already contain damage or a persistence mechanism.
  4. Record the job identifier, location, version, creation time, retention, and known verification state. Do not alter the original’s metadata.
  5. If an incident is suspected, preserve relevant logs and metadata before operating. A restore test is not malware analysis; escalate through the response plan.

4. Prepare an isolated environment

Build a separate machine, network, cloud account, tenant, or directory. Never test over production. Deny production and internet traffic by default, allowing only justified connections. Use non-production DNS; disable messaging, scheduled jobs, webhooks, replication, management agents, and automatic share mounting.

Provide adequate capacity, compatible software, and the correct timezone. Record network and isolation rules. Agree malware-scanning tools and timing with the response team so required evidence is not modified before preservation.

5. Check credentials and dependencies

Confirm encryption keys, certificates, service accounts, licences, software versions, database schemas, DNS, identity, queues, storage, and startup order. Decide which integrations need a stub.

Retrieve credentials from an approved vault and grant them for the test under least privilege. In the record, capture the vault name or entry identifier, technical account, access scope, and who approved use and when. Never record passwords, tokens, private keys, recovery codes, or complete connection strings in the worksheet, tickets, command logs, or screenshots. Remove temporary access after the exercise and rotate any secret that may have been exposed.

6. Perform the restore

  1. Record the procedure version, backup-tool version, and test-environment component versions. Enable activity logging without capturing secrets.
  2. Start timing. Record retrieval, target preparation, restoration, startup, and validation separately.
  3. Create a working copy where the technology supports it. Mount the source read-only; if the tool must write metadata, document the control protecting the original.
  4. Run the tool’s built-in verification. Save its job number, result, and complete log.
  5. Restore in dependency order: required configuration and keys, data, application state, then service. Follow vendor consistency instructions.
  6. Do not connect the restored application to the production directory, DNS, payment system, email platform, or clients. Start it in a restricted mode first.
  7. Record warnings, skipped objects, repairs, and version differences. Record the first error before retrying.

7. Validate integrity

Do not treat a green job status as proof of usability. Compare restored data with trusted checks recorded before the test: checksums, a signed manifest, expected record counts, control queries, an application report, or a known sample. If no such values exist, use a documented verify or check mechanism, a database consistency test, or vendor validation, and state the limitation clearly.

Check completeness; a representative sample; permissions, ownership, ACLs, timestamps, and encoding; database consistency; encryption; and log errors. Sample important, varied data. A post-restore hash supports later comparison but alone does not prove a match with the original.

8. Run functional and user checks

The system owner or a nominated user performs the pre-agreed scenarios: sign in with a test account, read a record, search for a document, open an attachment, generate a report, or perform a controlled write. Use test data and test recipients. Check important dependencies through stubs or safe connections, and confirm that blocked integrations sent nothing.

Record the expected and actual result, tester, and time for every scenario. Mark a condition “not tested” when it could not be verified; do not silently turn it into a pass.

9. Assess timing, outcome, and failures

Stop the timer only when the agreed usable state has been reached. Compare total time with the RTO and restored-data age with the RPO. Separate operator effort from waiting for transfer, hardware, approval, and dependencies. Record bottlenecks; a small-scale test may understate the duration of a production recovery.

Assign an outcome: pass, pass with limitations, fail, or stopped. When an error occurs, preserve its message, log, stage, time, and the action taken. Stop if continuing could threaten production, the original backup, confidentiality, or evidence. Do not delete a failed environment until the evidence owner confirms what must be retained. Escalate signs of compromise, failed integrity, and unauthorised traffic.

10. Preserve evidence and clean up

Evidence should include approved scope, restore-point ID, isolation, tool versions, timeline, verification logs, control results, secret-free screenshots, deviations, acceptance, and action owners. Store it with controlled access and retention.

After acceptance, stop services, unmount the source, and securely remove restored data, snapshots, disks, temporary files, caches, and exports according to classification. Confirm cloud deletion. Remove temporary access and network rules, rotate exposed secrets, and check for leftover DNS, schedules, or integrations.

11. Remediate, retest, and set cadence

Turn every deviation into an action with an owner, deadline, and closure criterion: correct retention, instructions, monitoring, key access, capacity, versions, or automation. After remediation, retest the same failed criterion. If the restore point, tool, or scope changes, identify the work as a new test. Do not close the problem merely because documentation was updated.

Set cadence by criticality and change rate. Test critical systems more often, combining full restores with representative-data exercises. Retest after backup, retention, encryption, key, application, or infrastructure changes, and after failures or incidents. Review RTO, RPO, roles, and scope at least annually; a date without an owner is not a control.

Want this in a lab or on production?

Docs stay free. The form is for scope, not a paywall.

The inquiry is stored on the server. You will get a short confirmation from hello@tomek.st. The operator is also notified on Telegram and email.