Network diagram with one failing node highlighted in red

Network Incidents: A Practical Guide to Detection, Response, and Learning

A field guide for engineers: what network incidents are, how to detect them, the first-30-minutes response playbook, a layer-by-layer triage checklist, blameless postmortems, and engineering for resilience.

What is a network incident?

A network incident is any unplanned event that degrades or interrupts the

connectivity between systems, services, or users. It sits at the intersection of

infrastructure and application reliability: a service can be perfectly healthy in

isolation and still be "down" because the packets can't get to it.

Unlike a pure application bug, network incidents are often **invisible at the

source**. The failing component blames its dependency, the dependency blames the

client, and the truth lives somewhere in the wires, DNS records, load balancers,

or certificate stores in between.

The anatomy of an incident

Most network incidents move through the same five phases. Naming them helps a

team know where they are while the adrenaline is flowing.

  1. Trigger — a change or external event (a deploy, a config push, an expired

cert, a BGP withdrawal, a traffic spike).

  1. Degradation — latency climbs, error rates rise, or connections time out.
  2. Detection — a human or a monitor notices. The gap between phases 2 and 3

is your *time to detect (TTD)* — often the single biggest lever on total

damage.

  1. Mitigation — restore service, even if the root cause isn't fully

understood yet. Rollback, failover, drain, or reroute.

  1. Resolution & learning — fix the root cause and change the system so it

can't recur the same way.

Common categories

1. DNS failures

DNS is the phone book of the internet, and it fails in spectacular, confusing

ways. Symptoms: intermittent "host not found," works-on-my-machine, or

region-specific outages. Causes: expired domains, misconfigured records, TTL

caching masking a bad change, or resolver overload.

Rule of thumb: "It's always DNS" is a meme because it's frequently true.

Check name resolution *first*, not last.

2. Certificate & TLS expiry

A silent time bomb. Everything works until a certificate expires at an exact

timestamp, then *every* client fails simultaneously. Automate renewal (ACME /

Let's Encrypt), monitor expiry dates as a first-class metric, and alert weeks

ahead — not hours.

3. Authentication & authorization edges

Not "network" in the cable sense, but they present identically: connections

that reach the server and get rejected. A rotated or expired token, a revoked

OAuth grant, or a stale hardcoded Authorization header will all surface as a

401/403 that looks like an outage to the caller. The tell is that the TCP

and TLS handshake *succeed* — the failure is at the application/auth layer.

4. Routing & connectivity

BGP misconfigurations, dropped routes, firewall rule changes, security-group

edits, and NAT exhaustion. These tend to be binary (fully reachable or fully

gone) and correlate tightly with a recent change.

5. Capacity & congestion

Connection-pool exhaustion, bandwidth saturation, or a thundering herd after a

cache expires. Latency degrades gracefully at first, then falls off a cliff.

6. Dependency & third-party outages

Your CDN, your cloud provider's control plane, a payment gateway, an upstream

API. You didn't cause it and can't fix it — but you own the customer

experience, so graceful degradation is your responsibility.

Detection: seeing the incident early

You cannot respond to what you cannot see. Effective detection rests on three

layers:

  • Metrics (the golden signals): latency, traffic, errors, and saturation.

Alert on rates and ratios, not raw counts.

  • Logs: structured, correlated by request/trace ID, retained long enough to

reconstruct a timeline.

  • Synthetic & real-user monitoring: probes that exercise the full path from

the outside, plus telemetry from actual clients. Internal health checks lie

when the problem is *between* the user and you.

Good alerts are actionable, attributable, and rare. An alert that fires

constantly trains people to ignore it — alert fatigue is itself an incident

waiting to happen.

Response: the first 30 minutes

  1. Declare early. A lightweight declaration ("we have an incident") is

cheaper than a delayed one. Under-declaring costs more than over-declaring.

  1. Assign roles. At minimum: an *incident commander* (coordinates, decides),

an *operations lead* (hands on keyboard), and a *communications lead* (updates

stakeholders). One person should not wear all three hats.

  1. Establish a single source of truth. One channel, one document. Timestamp

everything.

  1. Mitigate before you diagnose. Stopping the bleeding — rollback, failover,

feature-flag off — buys time to think. Root cause can wait; users cannot.

  1. Change one thing at a time. In the fog of an incident, parallel changes

make it impossible to know what actually helped.

  1. Communicate on a cadence. Even "no update yet, next update in 15 minutes"

keeps stakeholders calm and off the responders' backs.

A quick triage checklist

When connectivity breaks, walk *up* the stack — it isolates the layer fast:

  • [ ] DNS — does the name resolve, to the right address?
  • [ ] TCP — does the connection open on the expected port?
  • [ ] TLS — does the handshake complete? Is the cert valid and unexpired?
  • [ ] HTTP status — reachable but returning 4xx/5xx? That's auth or

application, not the network.

  • [ ] Recent changes — what deployed, rotated, or expired in the last hour?
  • [ ] Blast radius — one user, one region, or everyone?

The status code is a map: a clean handshake followed by a 401 points at

credentials, not cables. A connection that never opens points at routing,

firewalls, or DNS.

Learning: the blameless postmortem

The incident isn't over when service is restored — it's over when the system is

stronger than before.

  • Blameless by default. People act reasonably given the information and

incentives they have. If a human "caused" it, the real defect is a system that

*allowed* a single human to cause it.

  • Build a factual timeline. When did it start, when detected, when mitigated,

when resolved. These give you TTD and TTR (time to recover) — track them.

  • Ask "why" past the first answer. "The cert expired" is a symptom. "We had

no expiry monitoring and renewal was manual" is closer to the cause.

  • Turn findings into owned, dated action items. A postmortem with no tracked

follow-ups is a diary entry, not an improvement.

Prevention: engineering for resilience

  • Redundancy & failover — no single point of failure in the critical path.
  • Timeouts, retries with backoff, and circuit breakers — so one slow

dependency doesn't cascade into total collapse.

  • Graceful degradation — serve stale data or a reduced experience rather

than an error page.

  • Automate the expiry-prone — certificates, tokens, and domains should renew

themselves and alert when they can't.

  • Chaos & game days — inject failure on purpose, in controlled conditions,

so the first time you handle a failover isn't at 3 a.m.

  • Reduce change risk — canary deploys, gradual rollouts, and fast rollback

paths. Most incidents trace back to a change; make changes safer.

Key metrics to track over time

| Metric | What it tells you |

|---|---|

| TTD (time to detect) | How good your monitoring is |

| TTR (time to recover) | How good your response is |

| MTBF (mean time between failures) | How reliable the system is trending |

| Incident count by category | Where to invest prevention effort |

| % incidents caused by change | Whether your release process is safe |

Closing thought

Network incidents are not a sign of failure — in any system of meaningful scale,

they are inevitable. Mature teams don't aim for zero incidents; they aim to

detect fast, recover faster, and never be surprised the same way twice. The

goal is a system, and a culture, that gets a little more resilient after every

outage.

---

*Written 2026-09-06.*

---

日本語まとめ:ネットワーク障害の実務ガイド

ネットワーク障害(network incident)とは、システム間・サービス間・ユーザーとの「つながり」が予期せず劣化・断絶する事象です。アプリ自体は正常でも、経路上の DNS・ロードバランサ・証明書・認証などが原因で「落ちている」ように見えることが多く、原因が発生源からは見えにくいのが特徴です。

障害の5つのフェーズ

  1. 引き金 — デプロイ、設定変更、証明書失効、経路の消失、急激な負荷。
  2. 劣化 — レイテンシ増加、エラー率上昇、タイムアウト。
  3. 検知 — TTD(検知までの時間)。ここが被害の大きさを最も左右します。
  4. 緩和 — 原因究明より先に、まず復旧。ロールバック・フェイルオーバー・切り替え。
  5. 解決と学び — 根本原因を直し、同じ形で再発しない仕組みへ。

よくある種類

DNS 障害、TLS 証明書の失効、認証・トークン切れ(TCP/TLS は成功するのに 401/403)、経路・ファイアウォール設定、容量・輻輳、サードパーティ障害。

最初の30分

早めの「インシデント宣言」→ 役割分担(指揮官・実務・広報)→ 単一の情報源 → 診断より先に緩和 → 変更は一度に一つ → 定期的な状況共有。

トリアージ

DNS → TCP → TLS → HTTP ステータス → 直近の変更 → 影響範囲、の順にスタックを上へ辿ると切り分けが速くなります。ハンドシェイク成功後の 401 は「配線」ではなく「認証」を指します。

振り返り(ポストモーテム)

担当者を責めない(blameless)姿勢が前提です。「証明書が切れた」は症状であり、「失効監視が無く更新が手作業だった」が真因に近い。TTD / TTR を計測し、期限付き・担当付きのアクションに落とし込みます。

予防

冗長化とフェイルオーバー、タイムアウト・指数バックオフ・サーキットブレーカー、グレースフルデグラデーション、証明書やトークンの自動更新、カオスエンジニアリング、カナリアリリース。

成熟したチームは「障害ゼロ」ではなく、速く検知し、より速く復旧し、同じ失敗を二度は繰り返さないことを目指します。