Call us — 0113 322 3083
Mon–Fri · 9am–5:30pm · No fix, no fee
Start a free diagnostic →

Data Recovery Case File · NAS & Network Storage · The Array Is Probably Fine

RAID 5 Server: the OS Failed, Not the Array

His enquiry was three lines of specification and one problem. An HP ProLiant server: "set up as a RAID 5 with eight disks. OS has failed, containing work files that need recovering. Can you help?" Yes — and the first thing worth separating is the two halves of that sentence, because they are far less connected than they sound. An operating system failing is a volume problem: files that the server boots from have become damaged or unbootable. The array underneath is a different layer entirely, and in the great majority of these cases it is completely healthy — eight disks quietly holding intact data while the machine above them refuses to start. Which means the danger here is not the fault. It is what gets done next in an attempt to fix it.

MediaEight-disk RAID 5 array in a rack server — operating system failed to start; work files held on the array; member disks not reported as failed
Reported situationServer operating system non-functional · array configuration known and documented by the owner · work files required · no individual disk failure reported
Fault classOperating system or volume failure over an intact array — data layer typically unaffected; risk concentrated in reconstructive attempts rather than in the original fault
Equipment usedArray state established before any action · all members imaged individually write-blocked (Atola TaskForce 2) · parameters and parity order derived from the images · array assembled offline from the copies · files extracted, verified by opening and delivered

The decode: two layers, and the three things that must not happen

Why the layers are separate: a RAID controller presents eight physical disks to the server as one logical volume, and it does not care what is written on that volume. The operating system, the filesystem and the files all sit above that layer. So an OS that will not boot — a corrupted system volume, a failed update, a damaged boot configuration — is a problem with what was written, not with the disks that hold it. The array can be in perfect health while the server appears completely dead, and usually is.

The three things that must not happen: and this is where these cases are lost, almost always by someone trying to help. Do not rebuild. A rebuild reads every remaining member intensively to reconstruct one, and on an eight-disk array that is the most punishing operation it will ever perform — it is precisely when a second marginal disk fails, and on RAID 5 a second failure means the volume is gone. Do not initialise or re-create the array. Controllers offer to create a new array from the same disks, and accepting writes fresh metadata over the configuration a reconstruction depends on. Do not reinstall the operating system. Reinstalling writes to the volume, which is where the work files live. Each of those is a reasonable-looking step in a server room and each is worse than the original fault.

What RAID 5 actually tolerates: worth being precise, since it governs urgency. A RAID 5 array survives the loss of one member and no more. If a disk has already failed and he has not been told — which happens, because a degraded array keeps serving data and the alert sits in a management console nobody is watching — then the array is running without protection and any further failure is terminal. So the array's real state should be established before anything, and that is the first question rather than an afterthought.

How the data actually comes back: not by repairing the server. Every member is imaged individually and write-blocked, the array's parameters — member order, block size, parity rotation and offsets — are derived from those images by analysis, and the volume is assembled offline from the copies. Nothing is written to any original disk, no controller is asked to interpret anything, and a wrong hypothesis about the layout costs nothing because it is tested against a copy. The filesystem is then read from the reassembled volume and the work files extracted. The failed operating system never has to be fixed at all.

On the bench

The array's true state was established first — whether all eight members were healthy or the array was already running degraded — because that determines the urgency and whether anything must be done before imaging. Every member was then imaged individually and write-blocked on the Atola TaskForce 2, with bay positions recorded, so that the originals were secured before any interpretation was attempted. The parameters were derived from the images by testing candidate layouts against the filesystem's own structures until the volume resolved coherently, and the array was assembled offline from the copies. The work files were extracted, verified by opening, and delivered on fresh media.

The outcome

All members imaged, the array reconstructed offline from the copies, and the work files verified and delivered without the server ever being repaired. Free assessment, one fixed written figure including VAT; where a drive has to be opened, 50% of parts and labour is payable upfront with the balance only on success — otherwise no recovery, no fee. The decode, for anyone with a server that will not start: the operating system and the array are separate layers, and an OS failure usually says nothing about the disks beneath it — the array is very often entirely healthy; the danger is in the remedies, so do not rebuild, do not initialise or re-create the array, and do not reinstall the operating system, because each writes over what a recovery depends on and a rebuild is when second disks fail; establish whether the array is already running degraded, since RAID 5 tolerates exactly one failure; and note that the data comes back by reassembling images offline, so the server never needs fixing first.

Server that will not boot with work data on the array

Power it down and change nothing. The operating system and the array are separate layers, and an OS that won't start usually means a damaged system volume rather than failed disks — your array is probably intact. What loses these cases is the remedies. Don't let anything rebuild: a rebuild reads every remaining member hard, and that is exactly when a second marginal disk fails, which on RAID 5 ends the volume. Don't let a controller initialise or re-create the array from the same disks, because that writes new configuration over the one a recovery reads. And don't reinstall the operating system onto the volume holding your files. Check your management console for whether a member has already failed, since a degraded array keeps serving data quietly and you may be running without protection. Then label every disk with its bay position before removing anything.

Server down with the array probably fine?
Don't rebuild anything — call Leeds Data Recovery on 0113 322 3083; array state established first, every member imaged write-blocked, volume reassembled offline from the copies.
Request a quote online →

Our case files are drawn from genuine enquiries received by our laboratory over the past ten years, anonymised to protect client confidentiality. Each one describes the diagnostic and recovery procedure our engineers apply to that fault, using the equipment listed.

0113 322 3083