What is a network incident?
A network incident is any unplanned event that degrades or interrupts the
connectivity between systems, services, or users. It sits at the intersection of
infrastructure and application reliability: a service can be perfectly healthy in
isolation and still be "down" because the packets can't get to it.
Unlike a pure application bug, network incidents are often **invisible at the
source**. The failing component blames its dependency, the dependency blames the
client, and the truth lives somewhere in the wires, DNS records, load balancers,
or certificate stores in between.
The anatomy of an incident
Most network incidents move through the same five phases. Naming them helps a
team know where they are while the adrenaline is flowing.
- Trigger — a change or external event (a deploy, a config push, an expired
cert, a BGP withdrawal, a traffic spike).
- Degradation — latency climbs, error rates rise, or connections time out.
- Detection — a human or a monitor notices. The gap between phases 2 and 3
is your *time to detect (TTD)* — often the single biggest lever on total
damage.
- Mitigation — restore service, even if the root cause isn't fully
understood yet. Rollback, failover, drain, or reroute.
- Resolution & learning — fix the root cause and change the system so it
can't recur the same way.
Common categories
1. DNS failures
DNS is the phone book of the internet, and it fails in spectacular, confusing
ways. Symptoms: intermittent "host not found," works-on-my-machine, or
region-specific outages. Causes: expired domains, misconfigured records, TTL
caching masking a bad change, or resolver overload.
Rule of thumb: "It's always DNS" is a meme because it's frequently true.
Check name resolution *first*, not last.
2. Certificate & TLS expiry
A silent time bomb. Everything works until a certificate expires at an exact
timestamp, then *every* client fails simultaneously. Automate renewal (ACME /
Let's Encrypt), monitor expiry dates as a first-class metric, and alert weeks
ahead — not hours.
3. Authentication & authorization edges
Not "network" in the cable sense, but they present identically: connections
that reach the server and get rejected. A rotated or expired token, a revoked
OAuth grant, or a stale hardcoded Authorization header will all surface as a
401/403 that looks like an outage to the caller. The tell is that the TCP
and TLS handshake *succeed* — the failure is at the application/auth layer.
4. Routing & connectivity
BGP misconfigurations, dropped routes, firewall rule changes, security-group
edits, and NAT exhaustion. These tend to be binary (fully reachable or fully
gone) and correlate tightly with a recent change.
5. Capacity & congestion
Connection-pool exhaustion, bandwidth saturation, or a thundering herd after a
cache expires. Latency degrades gracefully at first, then falls off a cliff.
6. Dependency & third-party outages
Your CDN, your cloud provider's control plane, a payment gateway, an upstream
API. You didn't cause it and can't fix it — but you own the customer
experience, so graceful degradation is your responsibility.
Detection: seeing the incident early
You cannot respond to what you cannot see. Effective detection rests on three
layers:
- Metrics (the golden signals): latency, traffic, errors, and saturation.
Alert on rates and ratios, not raw counts.
- Logs: structured, correlated by request/trace ID, retained long enough to
reconstruct a timeline.
- Synthetic & real-user monitoring: probes that exercise the full path from
the outside, plus telemetry from actual clients. Internal health checks lie
when the problem is *between* the user and you.
Good alerts are actionable, attributable, and rare. An alert that fires
constantly trains people to ignore it — alert fatigue is itself an incident
waiting to happen.
Response: the first 30 minutes
- Declare early. A lightweight declaration ("we have an incident") is
cheaper than a delayed one. Under-declaring costs more than over-declaring.
- Assign roles. At minimum: an *incident commander* (coordinates, decides),
an *operations lead* (hands on keyboard), and a *communications lead* (updates
stakeholders). One person should not wear all three hats.
- Establish a single source of truth. One channel, one document. Timestamp
everything.
- Mitigate before you diagnose. Stopping the bleeding — rollback, failover,
feature-flag off — buys time to think. Root cause can wait; users cannot.
- Change one thing at a time. In the fog of an incident, parallel changes
make it impossible to know what actually helped.
- Communicate on a cadence. Even "no update yet, next update in 15 minutes"
keeps stakeholders calm and off the responders' backs.
A quick triage checklist
When connectivity breaks, walk *up* the stack — it isolates the layer fast:
- [ ] DNS — does the name resolve, to the right address?
- [ ] TCP — does the connection open on the expected port?
- [ ] TLS — does the handshake complete? Is the cert valid and unexpired?
- [ ] HTTP status — reachable but returning
4xx/5xx? That's auth or
application, not the network.
- [ ] Recent changes — what deployed, rotated, or expired in the last hour?
- [ ] Blast radius — one user, one region, or everyone?
The status code is a map: a clean handshake followed by a 401 points at
credentials, not cables. A connection that never opens points at routing,
firewalls, or DNS.
Learning: the blameless postmortem
The incident isn't over when service is restored — it's over when the system is
stronger than before.
- Blameless by default. People act reasonably given the information and
incentives they have. If a human "caused" it, the real defect is a system that
*allowed* a single human to cause it.
- Build a factual timeline. When did it start, when detected, when mitigated,
when resolved. These give you TTD and TTR (time to recover) — track them.
- Ask "why" past the first answer. "The cert expired" is a symptom. "We had
no expiry monitoring and renewal was manual" is closer to the cause.
- Turn findings into owned, dated action items. A postmortem with no tracked
follow-ups is a diary entry, not an improvement.
Prevention: engineering for resilience
- Redundancy & failover — no single point of failure in the critical path.
- Timeouts, retries with backoff, and circuit breakers — so one slow
dependency doesn't cascade into total collapse.
- Graceful degradation — serve stale data or a reduced experience rather
than an error page.
- Automate the expiry-prone — certificates, tokens, and domains should renew
themselves and alert when they can't.
- Chaos & game days — inject failure on purpose, in controlled conditions,
so the first time you handle a failover isn't at 3 a.m.
- Reduce change risk — canary deploys, gradual rollouts, and fast rollback
paths. Most incidents trace back to a change; make changes safer.
Key metrics to track over time
| Metric | What it tells you |
|---|---|
| TTD (time to detect) | How good your monitoring is |
| TTR (time to recover) | How good your response is |
| MTBF (mean time between failures) | How reliable the system is trending |
| Incident count by category | Where to invest prevention effort |
| % incidents caused by change | Whether your release process is safe |
Closing thought
Network incidents are not a sign of failure — in any system of meaningful scale,
they are inevitable. Mature teams don't aim for zero incidents; they aim to
detect fast, recover faster, and never be surprised the same way twice. The
goal is a system, and a culture, that gets a little more resilient after every
outage.
---
*Written 2026-09-06.*
---
日本語まとめ:ネットワーク障害の実務ガイド
ネットワーク障害(network incident)とは、システム間・サービス間・ユーザーとの「つながり」が予期せず劣化・断絶する事象です。アプリ自体は正常でも、経路上の DNS・ロードバランサ・証明書・認証などが原因で「落ちている」ように見えることが多く、原因が発生源からは見えにくいのが特徴です。
障害の5つのフェーズ
- 引き金 — デプロイ、設定変更、証明書失効、経路の消失、急激な負荷。
- 劣化 — レイテンシ増加、エラー率上昇、タイムアウト。
- 検知 — TTD(検知までの時間)。ここが被害の大きさを最も左右します。
- 緩和 — 原因究明より先に、まず復旧。ロールバック・フェイルオーバー・切り替え。
- 解決と学び — 根本原因を直し、同じ形で再発しない仕組みへ。
よくある種類
DNS 障害、TLS 証明書の失効、認証・トークン切れ(TCP/TLS は成功するのに 401/403)、経路・ファイアウォール設定、容量・輻輳、サードパーティ障害。
最初の30分
早めの「インシデント宣言」→ 役割分担(指揮官・実務・広報)→ 単一の情報源 → 診断より先に緩和 → 変更は一度に一つ → 定期的な状況共有。
トリアージ
DNS → TCP → TLS → HTTP ステータス → 直近の変更 → 影響範囲、の順にスタックを上へ辿ると切り分けが速くなります。ハンドシェイク成功後の 401 は「配線」ではなく「認証」を指します。
振り返り(ポストモーテム)
担当者を責めない(blameless)姿勢が前提です。「証明書が切れた」は症状であり、「失効監視が無く更新が手作業だった」が真因に近い。TTD / TTR を計測し、期限付き・担当付きのアクションに落とし込みます。
予防
冗長化とフェイルオーバー、タイムアウト・指数バックオフ・サーキットブレーカー、グレースフルデグラデーション、証明書やトークンの自動更新、カオスエンジニアリング、カナリアリリース。
成熟したチームは「障害ゼロ」ではなく、速く検知し、より速く復旧し、同じ失敗を二度は繰り返さないことを目指します。



