Why a rebuild kills more often than the initial failure
A RAID array can cope with the failure of one drive. What it often cannot cope with is the attempt to recover from it. In practice the rebuild is one of the most common causes of permanent data loss.
What redundancy really means
A RAID with redundancy distributes the data across several drives in such a way that what was on a failed drive can be reconstructed from the remaining ones. With RAID 5 the array copes with one missing drive, with RAID 6 with two.
The important word is “copes”. After a failure the system keeps running, but without any reserve. At that moment it is no longer an array with fault tolerance but a stack of drives where the next error gets through. And it is precisely in this state that a rebuild is usually started.
Why something goes wrong at exactly that moment
The rebuild is the heaviest load in a drive’s life
To reconstruct a missing member, every remaining one has to be read in full — at today’s capacities that means hours or days in one go, without a pause. A drive that was inconspicuous in everyday use because it was only addressed occasionally now shows its weaknesses under this sustained load.
The drives are the same age and from the same source
An array is as a rule bought and populated all at once. The drives therefore come from the same production series, have the same operating hours and the same environmental conditions behind them. If one dies of old age, the others are not far behind.
A single unreadable sector is enough
During a rebuild truly every sector has to be read — including those nobody has touched for years. If an unreadable spot is found on one of the remaining drives, the process aborts. With a RAID 5 without reserve that means: the array is offline.
The rebuild writes
And it writes over the existing redundancy. If it runs halfway through and then fails, the state afterwards is worse than before — the old parity information has been partly replaced by new information based on incomplete data.
The core of it in one sentence
What to do instead
- 01
Shut down
Switch off cleanly. A degraded system that keeps running keeps writing — and every write operation reduces the room for manoeuvre.
- 02
Label
Before removing them, mark each drive with its bay number. The order is part of the information from which the array can be reconstructed.
- 03
Confirm nothing
Do not accept any prompt to initialise, format or adopt a foreign configuration. Do not run a file system check.
- 04
Have a diagnosis made
First every member is imaged individually and write-protected. The array is then rebuilt from the copies — offline, without the originals being written to.
How a reconstruction proceeds
An array is more than a collection of drives. For a coherent file system to emerge from them again, four things have to be right: the size of the blocks written in alternation; the order of the drives; the pattern by which the parity information moves from drive to drive; and the offset at which all of it begins.
On some systems these values are held in administrative data on the drives themselves — on others they are not, or they are damaged. They are then derived from the data: combinations are tried against the copies and checked to see which produces a valid file system. That is computation, not guesswork, and it takes place exclusively on the images.
And ransomware?
Network storage is a popular target. What helps in such cases is the design of some file systems: when something changes they overwrite nothing but write a new version and keep the old one as long as there is space. From such earlier states it is sometimes possible to extract a condition from before the infection.
Whether that succeeds depends on how much has been written since the incident. So here too: switch the system off, delete nothing, create nothing new.
Answered briefly.
My NAS reports that a drive has failed and offers a rebuild. Should I?
If the data is important and there is no current backup: no. First secure whatever is still reachable — or have the array examined. The rebuild is the measure with the highest risk at exactly the moment when you can afford none.
I have already started the rebuild and it aborted. Is everything lost?
Not necessarily. Even a partly overwritten array can often still be reconstructed. What matters is not to attempt anything further now and to leave the system switched off.
Does RAID count as a backup?
No. An array protects against the failure of one drive — not against accidental deletion, ransomware, theft, fire or an error in the system itself. All of those hit every drive at the same time.
We will take a look at your storage device.
The Economy diagnosis is free of charge. You find out what is possible before anything costs money — and what it would cost is put in writing beforehand.
Wasserweg 8–10 · 60594 Frankfurt am Main

