๐ง RAID for Linux Admins (The Ultimate Intro Guide)

๐ Article Type: Comprehensive Overview
โฑ๏ธ Read Time: ~18-22 minutes
๐ฏ Audience: Linux admins (beginner to intermediate)
๐
Last Updated: October 2025
๐ What Is RAID?
Whether you manage Linux servers, NAS systems, or production databases โ understanding RAID isnโt optional.
Itโs what separates a hobbyist from a systems admin who never panics when a disk dies.
RAID stands for Redundant Array of Independent Disks.
It means combining multiple disks to act as one logical storage system.
The goal?
โ
Speed โ read/write faster using multiple disks together
โ
Reliability โ keep data safe if one disk fails
โ
Scalability โ manage big storage more easily
๐ก What Is Redundancy?
In general:
Redundancy means a backup inside the system itself โ an extra copy or layer that takes over if something fails.
In storage:
Redundancy means data is protected from disk failure โ either by mirroring (copying) or by parity (rebuild math).
Example:
If you save your file on two disks:
Disk1 โ has file
Disk2 โ has same file
If Disk1 dies, Disk2 still has it.
โ That's redundancy.
๐งฑ Types of RAID
1. Software RAID
Managed by the Linux kernel using the mdadm utility
No special hardware needed
Cheaper, more flexible, great for servers and learning
2. Hardware RAID
Managed by a dedicated RAID controller card
The OS only sees one logical disk
Better performance and battery-backed cache
Used in enterprise setups โ more expensive, less flexible
โ๏ธ Core RAID Concepts
Before we dive into RAID levels, understand these three core techniques:
| Concept | Meaning | Example |
| Striping | Split data across disks to increase speed | Word "USAMA" โ Disk1: U S A; Disk2: A M A |
| Mirroring | Same data copied to all disks | Disk1 = [U S A M A], Disk2 = [U S A M A] |
| Parity | A mathematical backup. It doesnโt store your files twice โ instead, it stores โcalculated hintsโ that help rebuild missing data if a disk fails. | A1 โ A2 = P1 โ lose A2 โ rebuild using A1 โ P1 |
| Redundancy | Data protection using mirror or parity | RAID 1, 5, 6, 10 have redundancy |
๐งฎ Quick RAID Math Cheat Sheet
RAID 0: n ร disk_size
RAID 1: (n/2) ร disk_size
RAID 5: (nโ1) ร disk_size
RAID 6: (nโ2) ร disk_size
RAID 10: (n/2) ร disk_size
โก RAID 0 โ The Speed Demon (No Safety)
๐ What Is RAID 0?
RAID 0 uses striping โ it splits your data into small chunks and writes them across multiple disks simultaneously.
Think of it like this: Instead of writing a 100MB file to one disk, RAID 0 writes 50MB to Disk1 and 50MB to Disk2 at the same time. This makes it twice as fast.
๐ How It Actually Works
Example: Saving the word "COMPUTER" on 2 disks
Without RAID 0 (normal):
Disk1: C O M P U T E R
Disk2: (empty)
With RAID 0:
Disk1: C M U E
Disk2: O P T R
The data is split (striped) across both disks. When you read it back, both disks work together to deliver it faster.
๐ฏ Real-World Scenario
You're a video editor working with 4K footage:
Normal disk: Plays 4K video at 30fps, sometimes stutters
RAID 0 (2 disks): Plays 4K video smoothly at 60fps, no stutters
Why? Because two disks reading together = double the speed.
โ Common Questions
Q: Why is it called RAID 0 if there's no redundancy?
A: The "0" means zero fault tolerance โ if any disk fails, ALL data is lost.
Q: What happens if one disk fails?
A: Imagine tearing a book in half and giving each half to two people. If one person loses their half, you can't read the book anymore. Same with RAID 0 โ lose one disk = lose everything.
Q: Can I use different size disks?
A: Yes, but RAID 0 will only use the smallest disk's size for both. Example: 2TB + 4TB = RAID 0 will use only 2TB from each = 4TB total.
โ Pros
Very fast read/write (2ร with 2 disks, 3ร with 3 disks)
Full storage capacity usable (no space wasted)
Simple to set up
โ Cons
If ANY disk fails โ 100% total data loss (catastrophic)
No fault tolerance whatsoever
Risk increases with more disks (more disks = more failure points)
โ When NOT to Use
Never for important data (documents, photos, databases)
Never for production systems
Never for operating system drives
๐ฐ Cost Example (4ร4TB disks)
Usable space: 16TB (100%)
Cost efficiency: Best (no wasted space)
๐ฏ Best Use Cases
Video editing scratch disk (temporary render files)
Gaming storage (games can be re-downloaded)
Cache storage
Non-critical temporary data
Min disks: 2
Fault tolerance: 0 disks
๐ง Remember
RAID 0 = Speed without safety. Only use for data you can afford to lose.
๐ก๏ธ RAID 1 โ The Mirror (Safety First)
๐ What Is RAID 1?
RAID 1 uses mirroring โ it writes identical copies of your data to two or more disks. Every file exists on every disk.
Think of it like writing the same letter on two pieces of paper. If you lose one, you still have the other.
๐ How It Actually Works
Example: Saving the word "COMPUTER" on 2 disks
RAID 1 (mirror):
Disk1: C O M P U T E R
Disk2: C O M P U T E R
Both disks have the exact same data. They're perfect mirrors of each other.
๐ฏ Real-World Scenario
You're running a company website on a server:
Normal disk: Disk fails at 3 AM โ website goes down โ you lose business
RAID 1: Disk fails at 3 AM โ website keeps running on the second disk โ you replace the failed disk during business hours
โ Common Questions
Q: If I have RAID 1, do I still need backups?
A: YES! RAID 1 protects against disk failure, NOT against:
Accidentally deleting files (deleted from both disks)
Virus infection (both disks get infected)
File corruption (corrupted on both disks)
Ransomware (both disks get encrypted)
Q: Does RAID 1 make writes slower?
A: Slightly, because data must be written to both disks. But the difference is minimal with modern hardware.
Q: Does RAID 1 make reads faster?
A: Yes! The system can read from both disks simultaneously, roughly doubling read speed.
Q: Can I use 3 disks in RAID 1?
A: Yes! You can mirror across 3 or more disks. More mirrors = more safety (but more expensive).
Q: What happens if both disks fail?
A: You lose all data. But the probability of both failing simultaneously is very low.
โ Pros
Excellent data protection
Simple to understand and manage
Fast read performance (can read from multiple disks)
Easy recovery (just use the surviving disk)
Can lose all but one disk and still function
โ Cons
Only 50% of total space is usable (expensive)
Write performance is the same as single disk
Not ideal for large storage needs
โ When NOT to Use
When you need lots of storage space (very expensive)
When budget is tight (wastes 50% capacity)
For large media libraries (consider RAID 5/6 instead)
๐ฐ Cost Example (4ร4TB disks)
Usable space: 8TB (50%)
Cost efficiency: Poor (half the space wasted)
๐ฏ Best Use Cases
Operating system drives
Boot partitions
Small but critical databases
Application servers
Configuration files
Any data where safety > space
Min disks: 2
Max disks: Typically 2-4 (more is possible but uncommon)
Fault tolerance: nโ1 disks (if you have 2 disks, 1 can fail; if 3 disks, 2 can fail)
๐ง Remember
RAID 1 = Maximum safety, minimum complexity. Perfect for critical small data.
๐งฎ RAID 5 โ Balanced Performer (1-Disk Parity)
๐ What Is RAID 5?
RAID 5 uses striping + distributed parity. It's like RAID 0 (fast) but with a mathematical safety net.
Instead of copying entire files (like RAID 1), RAID 5 stores "parity" โ a mathematical calculation that can rebuild lost data.
๐ How It Actually Works
Let's use simple math to understand parity:
Example with numbers:
Disk1: 5
Disk2: 3
Parity: 5 โ 3 = 6 (using XOR operation)
If Disk2 fails:
We know: Disk1 = 5, Parity = 6
Calculate: 5 โ 6 = 3 (we recovered Disk2!)
Example: Saving "COMPUTER" on 3 disks
Disk1: C O M P1 (P1 = parity of row 1)
Disk2: P U T P2 (P2 = parity of row 2)
Disk3: E R P3 P4 (P3, P4 = parity of rows 3,4)
Notice: Parity is distributed across all disks (not stuck on one disk)
If Disk2 fails, we can rebuild it using data from Disk1 and Disk3.
๐ฏ Real-World Scenario
You're setting up a file server with 4ร4TB disks:
Option 1: RAID 1
Usable: 8TB (50%)
Cost: $400 in disks for 8TB usable
Option 2: RAID 5
Usable: 12TB (75%)
Cost: $400 in disks for 12TB usable
Still protected against 1 disk failure
RAID 5 gives you 50% more space while maintaining safety.
โ Common Questions
Q: Why is parity distributed across all disks?
A: If parity was on one disk, that disk would wear out faster (bottleneck). Distributing it balances the wear.
Q: What is the "write penalty" everyone talks about?
A: Every write in RAID 5 requires 4 operations:
Read old data block
Read old parity block
Calculate new parity (using old data + old parity + new data)
Write new data + new parity
This makes writes 2-3ร slower than RAID 10.
Q: What is URE and why is it dangerous?
A: URE = Unrecoverable Read Error. Modern disks have ~1 error per 10^14 bits read.
When rebuilding a 4TB disk:
You read ~32 trillion bits
Statistically, you'll likely hit 1+ URE
If URE happens during rebuild โ rebuild fails โ data loss
This is why RAID 5 is declining for disks >2TB.
Q: How long does a rebuild take?
A: Depends on disk size:
1TB disk: 3-6 hours
4TB disk: 12-24 hours
8TB disk: 24-48 hours
modern disks such nvmes take less time
During rebuild, if another disk fails โ total data loss.
Q: Can I add more disks later?
A: Very difficult with traditional RAID 5. You'd need to backup, recreate array, and restore data.
โ Pros
Good balance of speed, space, and safety
More usable space than RAID 1 (only 1 disk worth is "wasted")
One disk can fail safely
Good read performance
More economical than RAID 1 for large storage
โ Cons
Write penalty: Writes are significantly slower (4 operations per write)
URE risk: Dangerous with disks >2TB during rebuild
Slow rebuild times (hours to days)
During rebuild, array is vulnerable (no redundancy)
Complex recovery compared to RAID 1
โ When NOT to Use
Disks larger than 2TB (high URE risk)
Write-heavy workloads (databases with lots of updates)
Critical production systems (RAID 10 is better)
SSDs in write-intensive scenarios
๐ฐ Cost Example (4ร4TB disks)
Usable space: 12TB (75%)
Formula: (n-1) ร disk_size = (4-1) ร 4TB = 12TB
๐ก Pro Tip: Hot Spare Disk
Add an extra disk that sits idle. When a disk fails, it automatically becomes part of the array and rebuild starts immediately. This reduces the vulnerable window.
๐ฏ Best Use Cases
File servers (read-heavy)
Media storage (videos, photos)
Backup targets
Home NAS systems
Archive storage
Min disks: 3
Recommended: 4-6 disks
Fault tolerance: 1 disk
๐ง Remember
RAID 5 = Good middle ground, but risky with modern large disks. Consider RAID 6 or 10 instead.
๐งฎ RAID 6 โ Double Parity Protection
๐ What Is RAID 6?
RAID 6 is like RAID 5 but with two parity blocks instead of one. This means it can survive two simultaneous disk failures.
Think of it as RAID 5 with an extra safety cushion.
๐ How It Actually Works
Example: Saving data on 4 disks
Disk1: A1 A2 P3 Q4
Disk2: A3 P2 Q3 A5
Disk3: P1 Q2 A4 A6
Disk4: Q1 A7 A8 A9
P = First parity (like RAID 5)
Q = Second parity (different calculation)
If Disk1 AND Disk2 both fail, you can still rebuild using P and Q parity from Disk3 and Disk4.
๐ฏ Real-World Scenario
You have 8ร8TB disks (64TB raw) in a NAS:
RAID 5 scenario:
Disk3 fails โ rebuild starts (takes 30 hours)
During rebuild, Disk5 also fails โ TOTAL DATA LOSS
RAID 6 scenario:
Disk3 fails โ rebuild starts
During rebuild, Disk5 also fails โ Array still works! Data is safe
You can replace both failed disks and rebuild
โ Common Questions
Q: Why is RAID 6 safer than RAID 5?
A: During RAID 5 rebuild (which can take 24+ hours), if another disk fails, you lose everything. RAID 6 can handle this second failure.
Q: Is RAID 6 slower than RAID 5?
A: Yes, writes are slower because it must calculate TWO parity blocks (P and Q) instead of one. Reads are similar speed.
Q: When should I use RAID 6 instead of RAID 5?
A: When:
Using large disks (>2TB)
Managing large arrays (6+ disks)
Data is critical and rebuild time is long
You can afford the extra disk for second parity
Q: How much space do I lose?
A: You lose 2 disks worth of space (for the two parity blocks).
4 disks ร 4TB = 16TB raw โ 8TB usable (50%)
6 disks ร 4TB = 24TB raw โ 16TB usable (66%)
8 disks ร 4TB = 32TB raw โ 24TB usable (75%)
Q: Can it survive 3 disk failures?
A: No. RAID 6 can only survive 2 disk failures. For 3+ failures, you'd need RAID 60 or other advanced setups.
โ Pros
Can survive 2 simultaneous disk failures
Much safer than RAID 5 for large disks
Good read performance
Better for arrays with 6+ disks
Safer during long rebuild times
โ Cons
Heavier write penalty (must calculate P and Q parity)
Slower writes than RAID 5
Rebuilds still take long (but safer during rebuild)
Requires minimum 4 disks
More CPU intensive (complex parity calculations)
โ When NOT to Use
Small arrays (< 4 disks) โ wasteful
Write-intensive databases โ too slow
When write performance is critical
๐ฐ Cost Example (4ร4TB disks)
Usable space: 8TB (50%)
Formula: (n-2) ร disk_size = (4-2) ร 4TB = 8TB
Note: With only 4 disks, RAID 6 = same usable space as RAID 10, but RAID 10 is faster.
๐ฏ Best Use Cases
Large NAS systems (8+ disks)
Long-term archival storage
Enterprise backup systems
Any large array where rebuild time is measured in days
Media production houses with huge storage needs
Min disks: 4
Recommended: 6+ disks
Fault tolerance: 2 disks
๐ง Remember
RAID 6 = RAID 5 with double safety. Essential for large modern disks and big arrays.
โก๐ก๏ธ RAID 10 โ Speed + Safety Hybrid (The Production Favorite)
๐ What Is RAID 10?
RAID 10 (also called RAID 1+0) combines RAID 1 (mirroring) and RAID 0 (striping).
How it works:
First, create RAID 1 mirror pairs
Then, stripe across those mirror pairs
You get the speed of RAID 0 AND the safety of RAID 1.
๐ How It Actually Works
Setup: 4 disks
Step 1: Create two mirror pairs
Mirror Pair 1: Disk1 โ Disk2
Mirror Pair 2: Disk3 โ Disk4
Step 2: Stripe data across the mirror pairs
Write "COMPUTER":
Mirror Pair 1 (Disk1 & Disk2): C O M P
Mirror Pair 2 (Disk3 & Disk4): U T E R
Both Disk1 and Disk2 have: C O M P
Both Disk3 and Disk4 have: U T E R
๐ฏ Real-World Scenario
You're running a production database that needs:
Fast reads/writes (many transactions per second)
Zero downtime tolerance
Quick recovery if disk fails
RAID 5:
Write penalty slows down transactions
Rebuild takes 20+ hours (risky)
RAID 10:
Fast writes (no parity calculation)
Rebuild takes 2-4 hours (just copy from mirror)
Database stays fast during rebuild
This is why 70-80% of production databases use RAID 10.
โ Common Questions
Q: What's the difference between RAID 10 and RAID 01?
A: Different order:
RAID 10: Mirror first, then stripe (better)
RAID 01: Stripe first, then mirror (worse fault tolerance)
Always use RAID 10, not RAID 01.
Q: How many disks can fail?
A: Best case: One from each mirror pair
4 disks setup:
โ
Can survive: Disk1 + Disk3 failure (different pairs)
โ Cannot survive: Disk1 + Disk2 failure (same pair)
Probability: With 4 disks and 1 failure, there's a 66% chance the second failure will be in a different pair (safe) and 33% chance it'll be in the same pair (data loss).
Q: Why is RAID 10 faster than RAID 5/6?
A: No parity calculation needed!
Write to RAID 5: Read old data, read old parity, calculate, write data, write parity (4-5 operations)
Write to RAID 10: Write to Disk1, write to Disk2 (2 operations)
Q: Is RAID 10 expensive?
A: Yes, because you only get 50% usable space. But for critical systems, the speed and reliability are worth it.
Q: Can I use odd number of disks?
A: No, RAID 10 requires even numbers (2, 4, 6, 8, 10, etc.) because you need pairs for mirroring.
Q: What happens during rebuild?
A: Super fast! Just copy from the surviving mirror disk. A 4TB disk can rebuild in 2-4 hours vs 20+ hours for RAID 5.
โ Pros
Excellent read performance (stripe + multiple mirrors)
Excellent write performance (no parity calculation)
Very fast rebuild (just copy from mirror)
High fault tolerance (can lose multiple disks if in different pairs)
Simple to understand and manage
Best performance + redundancy balance
โ Cons
Only 50% space usable (expensive)
Requires minimum 4 disks (even numbers only)
Most expensive RAID option per usable TB
โ When NOT to Use
When budget is very tight
When you need maximum storage capacity
Home users with limited disks
Cold storage / archival (use RAID 6 instead)
๐ฐ Cost Example (4ร4TB disks)
Usable space: 8TB (50%)
Cost: $400 for 8TB usable
Formula: (n/2) ร disk_size = (4/2) ร 4TB = 8TB
๐ฏ Best Use Cases
Production databases (MySQL, PostgreSQL, Oracle)
Virtual machine storage (VMware, Hyper-V)
Email servers (Exchange, Postfix)
Application servers (web apps, APIs)
Any critical system where speed + safety matter
E-commerce platforms
Min disks: 4 (even numbers only)
Recommended: 4-8 disks
Fault tolerance: Up to n/2 disks (if in different mirror pairs)
๐ Advanced: RAID 10 Fault Tolerance Deep Dive
Scenario: 4 disks (2 mirror pairs)
Mirror Pair 1: Disk1 โ Disk2
Mirror Pair 2: Disk3 โ Disk4
Failure scenarios:
โ
Disk1 fails โ Safe (use Disk2)
โ
Disk1 + Disk3 fail โ Safe (use Disk2 + Disk4)
โ
Disk1 + Disk4 fail โ Safe (use Disk2 + Disk3)
โ Disk1 + Disk2 fail โ Data loss (entire mirror pair gone)
Probability after 1st disk fails:
- 2 remaining disks: 1 is safe, 1 is dangerous
- 50% chance second failure is safe
๐ง Remember
RAID 10 = The gold standard for production systems. Fast, safe, simple. Worth the 50% space cost.
๐งฑ RAID 50 / RAID 60 โ Nested RAID (Advanced / Optional)
๐ What Are They?
RAID 50 = Multiple RAID 5 groups + RAID 0 striping across them
RAID 60 = Multiple RAID 6 groups + RAID 0 striping across them
These are "nested" or "hybrid" RAID levels for very large storage systems.
๐ How RAID 50 Works
Example: 6 disks
Step 1: Create two RAID 5 arrays
RAID 5 Array 1: Disk1, Disk2, Disk3
RAID 5 Array 2: Disk4, Disk5, Disk6
Step 2: Stripe across them (RAID 0)
Data is striped between Array 1 and Array 2
Result: Faster than pure RAID 5, can survive 1 disk failure per RAID 5 group.
โ When to Consider These?
Only consider RAID 50/60 if:
You have 8+ disks
You're building enterprise storage (50TB+)
You need high performance + redundancy
You have experienced storage admins
For most Linux admins: Stick with RAID 0/1/5/6/10. These cover 95% of real-world needs.
๐ง Remember
RAID 50/60 = Enterprise-level complexity. Master the basics first.
โ๏ธ Comparison Tables
Quick Decision Guide
| RAID | Best For | Risk Level | Usable Space (4ร4TB) | Rebuild Time |
| 0 | Speed / temp data | โ ๏ธ High | 16TB (100%) | N/A |
| 1 | Critical data / OS | โ Low | 8TB (50%) | 2-4 hours |
| 5 | File servers | โ ๏ธ Medium | 12TB (75%) | 12-24 hours |
| 6 | Large archives | โ Low | 8TB (50%) | 12-24 hours |
| 10 | Databases / VMs | โ Low | 8TB (50%) | 2-4 hours |
Technical Specifications
| RAID | Min Disks | Write Penalty | Fault Tolerance | Read Speed | Write Speed |
| 0 | 2 | None | 0 disks | ๐ฅ๐ฅ๐ฅ Excellent | ๐ฅ๐ฅ๐ฅ Excellent |
| 1 | 2 | Low | nโ1 disks | ๐ฅ๐ฅ Good | ๐ฅ Normal |
| 5 | 3 | High | 1 disk | ๐ฅ๐ฅ Good | โก Slow |
| 6 | 4 | Very High | 2 disks | ๐ฅ๐ฅ Good | โก Slower |
| 10 | 4 | Low | 1 per pair | ๐ฅ๐ฅ๐ฅ Excellent | ๐ฅ๐ฅ Very Good |
๐งฉ Visual Overview
RAID 0: (speed, no safety)
Disk1: A1 A3 A5
Disk2: A2 A4 A6
โ Fast but risky
RAID 1: (mirror)
Disk1: A1 A2 A3 A4
Disk2: A1 A2 A3 A4
โ Safe but expensive
RAID 5: (1 parity)
Disk1: A1 A2 P3
Disk2: A3 P2 A4
Disk3: P1 A5 A6
โ Balanced
RAID 6: (2 parity)
Disk1: A1 A2 P3 Q4
Disk2: A3 P2 Q3 A5
Disk3: P1 Q2 A4 A6
Disk4: Q1 A7 A8 A9
โ Extra safe
RAID 10: (mirror + stripe)
Disk1/Disk2: A1 A3 A5 (mirrored)
Disk3/Disk4: A2 A4 A6 (mirrored)
โ Fast and safe
โ ๏ธ CRITICAL: RAID โ Backup
What RAID Protects Against:
โ
Single disk hardware failure
โ
Disk bad sectors (with redundancy)
โ
System stays running during disk replacement
What RAID DOES NOT Protect Against:
โ Accidental file deletion (deletes from all disks)
โ Virus/malware infection (infects all disks)
โ Data corruption (corrupts all disks)
โ Ransomware encryption (encrypts all disks)
โ Accidental formatting (formats all disks)
โ Building fire/flood/theft (destroys all disks)
โ Controller failure (makes all disks inaccessible)
The 3-2-1 Backup Rule
Always maintain:
3 copies of your data (original + 2 backups)
2 different storage types (e.g., local + cloud)
1 copy offsite (different physical location)
Real Horror Story
Scenario: Company had RAID 5 server with critical data
Monday 3 PM: Employee accidentally runs "rm -rf /" command
Result: All data deleted from all RAID disks instantly
RAID couldn't help: It dutifully deleted from all disks
Recovery: Impossible (no backups existed)
Outcome: Company went out of business
RAID = High availability, NOT data protection
๐ฏ Key Takeaways
Choose Your RAID:
Need maximum speed, don't care about data loss? โ RAID 0
Small critical data (OS, boot drive)? โ RAID 1
File server with <2TB disks? โ RAID 5
Large storage with 2TB+ disks? โ RAID 6
Production database/VM/critical apps? โ RAID 10 (always!)
Golden Rules:
RAID is NOT backup โ Always maintain separate backups
Monitor your RAID โ A failed disk you don't know about = no redundancy
Use hot spares โ Automatic replacement reduces risk
Test your recovery โ Practice rebuilding before disaster strikes
RAID 5 caution โ Avoid with >2TB disks (URE risk)
RAID 10 for production โ Worth the 50% space cost
In short: RAID keeps your system running โ backups keep your job safe.
Master both, and youโll never fear a failed disk again.




