Why your backup is far bigger than your data
The most common backup-sizing mistake is buying storage to match your live data. But a backup repository doesn’t hold one copy — it holds a retention chain: every daily incremental and every weekly, monthly and yearly full you choose to keep. Retention, not raw data size, is what fills a repository. Get the retention wrong and you either run out of space in month three or massively over-buy.
Two things claw the capacity back. Deduplication and compression remove redundant blocks — typically 3:1 for files and VMs, more for highly repetitive estates, but never for already-compressed media, which stays 1:1. And XFS/ReFS fast-clone lets a synthetic full reference blocks that already exist, so it costs a fraction of a real full. This calculator applies both — using conservative, sourced ratios you can see and edit — so the number it gives you is defensible, not optimistic.
The 3-2-1-1-0 rule — and why 3-2-1 is no longer enough
The classic 3-2-1 rule (3 copies, 2 media, 1 offsite) predates modern ransomware, which now deliberately hunts down and encrypts backups first. The current standard adds two digits: 1 immutable or air-gapped copy that an attacker cannot alter or delete, and 0 errors — proven by automated recovery verification. The calculator sizes a dedicated tier for each part of the rule, so what you buy actually survives an attack.
The fast, on-site repository your day-to-day restores come from — sized to your operational (daily) retention and how often you take restore points. This is where an all-flash or hybrid array with XFS/ReFS fast-clone pays off, keeping synthetic fulls near-spaceless.
Find the right primary array →The "1 immutable" copy — locked with object-lock or a hardened Linux repository so ransomware cannot alter or delete it within the lock window. Increasingly a compliance and cyber-insurance requirement, not an optional extra.
Veeam immutable backup →A second-media, offsite copy of the full retention chain for a site-loss or disaster-recovery scenario — a second data centre, a colocation, or the cloud. This is what you fail over to when the primary site is gone.
Offsite & DR storage →Monthly and yearly retention on air-gapped media — LTO tape or offline object storage — for compliance and worst-case recovery. Physically offline media is the ultimate ransomware backstop and the cheapest £/TB for long retention.
LTO tape storage →Backup change-rate & data-reduction reference
The planning assumptions this tool uses per workload — the daily change rate that drives incremental size, and the typical dedup+compression ratio. Every value is a sourced range, shown so you can sanity-check the maths. Media is hard-locked at 1:1 (pre-compressed data cannot be reduced).
| Workload | Daily change (default) | Change range | Reduction (default) | Reduction range | Source |
|---|---|---|---|---|---|
| Virtual machines / VDI | 5%/day | 3–8% | 5:1 | 5:1–10:1 | StorageMath |
| Databases (SQL / Oracle) | 10%/day | 5–15% | 4:1 | 4:1–10:1 | Veeam B&R Best Practice |
| File servers / unstructured | 3%/day | 2–5% | 3:1 | 3:1–8:1 | StorageMath |
| Email / Exchange (on-prem) | 3%/day | 2–5% | 3:1 | 3:1–5:1 | StorageMath |
| Microsoft 365 — Exchange Online | 1%/day | 0.2–1% | 3:1 | 3:1–5:1 | Veeam Backup for M365 |
| Microsoft 365 — OneDrive / SharePoint | 1%/day | 0.5–1% | 3:1 | 3:1–5:1 | Veeam Backup for M365 |
| Media / pre-compressed | 2%/day | 1–4% | 1:1 (locked) | 1:1–1:1 | StorageMath |
| Mixed / general estate | 3%/day | 2–5% | 3:1 | 2:1–5:1 | StorageMath |
Overheads applied: +15% metadata/catalogue, 10% free working space, 1.25× workspace on the largest full, long-term archive reduction up to 40:1. Sources: Veeam B&R Best Practice · Veeam B&R Best Practice · Veeam Backup for M365 · Veeam Restore Point Simulator · StorageMath · StorageMath · SNIA · Commvault.
Backup & DR sizing — FAQs
How much backup storage do I need?
Far more than your live data — because you keep many restore points, not one copy. As a rule of thumb, a full GFS retention (e.g. 14 daily + 4 weekly + 12 monthly + 1 yearly) with data reduction lands a single repository at roughly 1.5–3× your source data, and a full 3-2-1-1-0 architecture (production + immutable + offsite + air-gap) multiplies that across copies. This calculator does the exact maths from your workloads, change rate, retention and RPO/RTO rather than a rule of thumb.
What is the 3-2-1-1-0 backup rule?
The modern, ransomware-resilient backup standard: 3 copies of your data, on 2 different media, with 1 copy offsite, 1 copy immutable or air-gapped, and 0 errors (verified by regular test-restores). It extends the classic 3-2-1 rule with an immutable/offline copy because attackers now target backups first. The calculator sizes a dedicated tier for each part of the rule.
Why is my backup repository bigger than my data?
Because a repository stores a retention chain, not a single copy. Every daily incremental and every weekly/monthly/yearly full you keep adds capacity. Retention — not raw data size — is what fills a repository. Data reduction (deduplication + compression) claws some of that back, and XFS/ReFS fast-clone makes synthetic fulls near-spaceless, both of which the calculator accounts for.
How does dedup and compression change the numbers?
Deduplication and compression typically reduce backup data by around 3:1 for file/VM data up to 10:1 or more for highly redundant estates — but never for already-compressed media (video, images, encrypted data), which stays 1:1. This tool applies a conservative, cited reduction ratio per workload type and hard-locks media at 1:1 so it never over-promises savings.
What do RPO and RTO have to do with capacity?
They drive the architecture, which drives capacity. A near-zero RPO means continuous replication or many restore points per day (more capacity); a 24-hour RPO needs only a daily backup. A minutes-level RTO needs a hot on-site copy for instant recovery; a days-level RTO can restore from cheaper tape. The calculator turns your RPO/RTO into the right tier design and sizes each tier accordingly.
How accurate is the calculator?
The GFS/retention maths is deterministic and matches the Veeam Restore Point Simulator model (validated against a canonical worked example). The change-rate and data-reduction ratios are conservative defaults drawn from Veeam best-practice, StorageMath and SNIA — every one is shown on screen, sourced, and fully editable. It is an accurate planning estimate; a Servnet engineer validates the design and confirms exact capacity and drives before any quotation.
Do you show prices?
No — we never publish distributor pricing, and drive/array pricing moves too often to quote reliably online. The calculator builds the right capacity and architecture; request a quote and we return a firm bill of materials with finance options for the whole solution.
Explore backup, storage & DR at Servnet
Talk to a UK specialist
Get expert advice or a no-obligation quote — servers, storage, networking, maintenance, finance and cloud. We reply the same working day.