“Tier III” and “Tier IV” get used loosely in data center marketing, often as if they were just two grades of “very reliable.” They are not. They describe two genuinely different things a facility can promise — and the gap between them is exactly the gap between planned resilience and unplanned resilience.
当您为切不可承受丢失风险的基础设施选址时,这种差异在实践中体现在以下方面:
Uptime Institute 的等级划分体系
该等级分类归属于 Uptime Institute — 自 20 世纪 90 年代起专门对数据中心进行权威认证的独立机构。它共划分为四个等级:
- Tier I — 基础容量,无冗余能力。
- Tier II — 具备冗余容量组件,但仅为单路分配路径。
- Tier III — concurrently maintainable: any component can be taken offline for maintenance without stopping IT operations.
- Tier IV — fault tolerant: the facility keeps running through an unplanned failure of any single component, automatically.
这些等级具有累加属性 — Tier IV 涵盖了其下方等级的所有要求 — 且仅在通过官方认证后方可生效,而非凭自宣自评。
等级背后真正的确凿保障
The decisive line is between Tier III and Tier IV, and it comes down to one word: fault.
Tier III guarantees concurrent maintainability. You can service any part of the power or cooling system — replace a UPS, test a generator, work on a chiller — without taking IT down. That removes planned maintenance as a source of downtime. But Tier III does not promise to survive an unplanned failure: if a component fails unexpectedly while the system is in a normal state, an outage is still possible.
Tier IV adds fault tolerance. Every capacity system and every distribution path is independently duplicated (commonly described as 2N), so any single unplanned failure — a breaker, a pump, a cable — is absorbed automatically, with no interruption to IT. Tier IV is concurrently maintainable and fault tolerant at the same time.
可用性(Aptime)的精算逻辑
这两个等级在预期可用性上的表现截然不同:
- Tier III ≈ 99.982% 的可用性 — 相当于每年约 1.6 小时的停机时间。
- Tier IV ≈ 99.995% 的可用性 — 相当于每年约 26 分钟的停机时间。
纸面上零点零几的百分比差异看似微不足道。但若换算成实际时间,这就是“整整缺失大半个工作日”与“仅相当于一次喝咖啡间歇”之间的本质区别。
“容错能力”并不等同于“硬件冗余”
最常见的误区是认为“具备冗余”就等于“具备容错能力”。事实并非如此。
Redundancy means there is a spare — an N+1 design has one more unit than it strictly needs. Fault tolerance means the spare is already wired, energised and carrying its share, on an independent path, so a failure is survived without anyone having to act. Tier III is typically redundant. Tier IV is fault tolerant. When a component fails at 3 a.m., that distinction is the whole game.
停机成本
等级之间的差异是否具有实质意义,取决于一个小时的停机对您意味着多少损失。对于展示型企业官网而言,损失几乎为零。但对于支付清算、交易平台、医院核心系统,或是巅峰旺季下的电商平台,一个小时的停机损失可能远超在整个服务周期内为更高可靠性等级所付出的溢价。
一个实用的做法是用具体数字将其量化:每小时损失的营业额,加上恢复成本,加上合同或监管罚款,再加上难以估量的声誉损失。将这个数字与不同等级之间的差价进行对比,决策通常就会显而易见。
何时选择 Tier IV 是合理的
Tier IV is the right choice when an unplanned outage is genuinely unacceptable — financial services, government and public infrastructure, healthcare, large-scale AI/ML training, and any business whose product is its uptime. If your workload can absorb a short, rare interruption, Tier III may be sufficient. If it cannot, the only honest place for it is a fault-tolerant facility.
Akashi 严格按照 Tier IV 标准建设,目标可用性达 99.995%。这是因为中亚地区开始运行的业务负载 — AI 基础设施、金融系统、主权云 — 已经明确跨过了那条线,进入了绝不可接受非计划停机的范畴。
常见问题
Tier III 与 Tier IV 的区别是什么?
Tier III 提供并行维护能力 — 任何组件均可在不停机的情况下进行维护 — 但非计划故障仍可能引发停机。Tier IV 增加了容错能力:任何单一非计划故障均可被系统自动吸收,不会造成任何中断。
Tier IV 保证多少可用性?
Tier IV 对应的可用性约为 99.995% — 相当于每年约 26 分钟的停机时间。Tier III 对应的可用性约为 99.982%,即每年约 1.6 小时。
“硬件冗余”和“容错能力”是一回事吗?
不是。硬件冗余意味着存在备用组件。容错能力则意味着通过独立路径构成的双路备用系统已经在实时运行,因此发生故障时无需人工干预即可自动吸收。Tier IV 具备容错能力;而 Tier III 通常具备硬件冗余。
我需要 Tier IV 数据中心吗?
当非计划停机不可接受时 — 如金融服务、医疗卫生、政务部门、AI 训练以及收入直接取决于系统可用性的业务 — 您就需要 Tier IV。对于能够承受罕见短时间中断的负载,Tier III 可能就足够了。
Akashi Data Center 属于哪个等级?
Akashi 严格按照 Tier IV 标准(Uptime Institute 标准)建设,目标可用性达 99.995% — 作为中亚地区首个 Tier IV 设施。