Skip to content
DnsLister Forum

Where domain hunters compare notes

Twenty minutes into the incident, three people are testing three different theories and nobody has written down what is actually broken.

Everyone is competent. That is not the failure. The failure is that "the API is slow" is not a symptom, it is a feeling, and four people will each optimise against a different version of it.

Structure is a hard thing to argue for during an incident. It feels like paperwork while the site is down. The trade is real: a few minutes defining the symptom, in exchange for not spending an hour eliminating hypotheses that were never consistent with it.

This template runs symptom, blast radius, layered hypotheses, elimination, confirmation, then fix. It is long, and it is meant to be kept open rather than read once.

This is the long one in the set. Keep it open beside the incident rather than reading it through:

Context Variables: - System/Component: [WHAT SYSTEM OR COMPONENT IS AFFECTED] - Symptom: [WHAT IS THE OBSERVABLE PROBLEM] - Environment: [PRODUCTION / STAGING / DEVELOPMENT] - When It Started: [WHEN WAS THE ISSUE FIRST OBSERVED] - What Changed Recently: [DEPLOYMENTS, CONFIG CHANGES, TRAFFIC PATTERNS] - Impact: [NUMBER OF USERS AFFECTED, REVENUE IMPACT, SLA STATUS] --- ## Thought Template: Technical Troubleshooting Reasoning ### Phase 1: Symptom Precision Before diagnosing, define the symptom precisely: 1) What EXACTLY is happening? (Not "it's slow" — "P95 latency increased from 200ms to 1400ms") 2) What is NOT happening that should be? (Or what IS happening that shouldn't be?) 3) When does it happen? (Always? Under load? At specific times? For specific users?) 4) Where does it happen? (All regions? Specific servers? Specific endpoints?) **Heuristic**: If you can't measure the symptom, you can't verify the fix. Convert every vague symptom into a measurable metric. **Checkpoint 1**: State the symptom as: "[Metric] is [current value] when it should be [expected value], affecting [scope], since [time]." ### Phase 2: Blast Radius Assessment Before deep-diving into root cause, understand the impact: 1) Is this getting worse, stable, or intermittent? 2) Is there a workaround? (If yes, implement it NOW, then continue diagnosis) 3) Is this a single-point failure or systemic? 4) What is the business cost per hour of this issue? **Heuristic - Severity Classification:** - SEV1 (Critical): Revenue impact or data loss, no workaround → All hands, fix first, investigate later - SEV2 (High): Degraded service, workaround exists → Dedicated responder, investigate within hours - SEV3 (Medium): Limited impact, workaround exists → Scheduled investigation - SEV4 (Low): Cosmetic or edge case → Backlog ### Phase 3: Hypothesis Generation (5 Whys × MECE) Generate hypotheses using the 5 Whys framework, ensuring they are MECE: **Layer 1 - Application:** - Code bug introduced in recent deployment? - Configuration change? - Memory leak or resource exhaustion? **Layer 2 - Infrastructure:** - Server/container health? CPU, memory, disk? - Network connectivity or DNS issues? - Load balancer misconfiguration? **Layer 3 - Dependencies:** - Database performance? Query plans changed? - External API latency or errors? - Cache hit rate dropped? **Layer 4 - Data:** - Data volume spike? - Corrupted or malformed data? - Missing or stale data? **Layer 5 - External:** - Traffic pattern change (DDoS, viral content, bot traffic)? - Third-party service degradation? - Cloud provider issue? **Heuristic**: Check the simplest explanations first. In order: Was anything deployed recently? → Is the infrastructure healthy? → Are dependencies responding normally? → Is the data correct? ### Phase 4: Systematic Elimination For each hypothesis, define a quick test: | Hypothesis | Test (< 5 min) | Expected Result if True | Result | Verdict | |-----------|----------------|------------------------|--------|--------| | Recent deployment bug | Rollback or check deploy logs | Issue resolves after rollback | | | | Resource exhaustion | Check CPU/mem/disk metrics | Resources > 90% utilized | | | | Database issue | Check slow query log, connection pool | Slow queries or pool exhaustion | | | | External dependency | Check dependency health dashboard | Dependency showing errors | | | | Traffic spike | Check request volume metrics | Unusual traffic pattern | | | **Checkpoint 2**: After testing, which hypotheses are eliminated? Which remain? Narrow to top 2 candidates. ### Phase 5: Root Cause Confirmation For the remaining candidates: 1) Can you reproduce the issue in a non-production environment? 2) Can you correlate the symptom timeline with the hypothesized cause timeline? 3) Does the hypothesis explain ALL observed symptoms (not just some)? **Heuristic**: A root cause must explain 100% of the symptoms. If your hypothesis only explains 80%, either it's a contributing factor (not root cause) or there are multiple issues. ### Phase 6: Fix and Verify 1) **Immediate fix**: [What to do right now] 2) **Verification**: [How to confirm the fix works — specific metrics to watch] 3) **Rollback plan**: [How to undo the fix if it makes things worse] 4) **Permanent fix**: [What structural change prevents recurrence] 5) **Post-incident review**: [What to document and share] --- ## Validation Gate - [ ] Symptom is precisely defined with measurable metrics - [ ] Impact has been assessed and communicated to stakeholders - [ ] At least 3 layers of hypotheses were tested - [ ] Root cause explains all observed symptoms - [ ] Fix has been verified with metrics (not just "it looks better") - [ ] A post-incident review is scheduled --- ## Output Format 1. **Incident Summary** (symptom, impact, duration, root cause — 3 sentences) 2. **Diagnosis Path** (hypotheses tested → eliminated → confirmed) 3. **Root Cause Statement** (single sentence) 4. **Fix Applied** (what was done + verification metrics) 5. **Prevention Plan** (structural changes to prevent recurrence) 6. **Timeline** (when detected → diagnosed → fixed → verified) 

Serving notes

Phase 1 pays for the rest. Converting "it is slow" into a metric with a before value, an after value and a scope gives you something you can test a fix against later. Without that, "it looks better now" is the only verification available.

The elimination table caps each test at five minutes. That constraint is what stops the exercise turning into an investigation. Cheap tests first, in order, and the point is to remove candidates rather than to be right early.

The rule that a root cause must account for everything observed is the one worth arguing over. An explanation that covers most of the symptoms is a contributing factor, and stopping there is how the same incident returns next month wearing a different shirt.

Scope note: this is for systems and services. Nothing here belongs anywhere near work that authorises physical activity. AI prepares, humans decide, and a permit is a human's signature.

Full prompt page: https://www.nerdychefs.ai/pack/reasoning-playbook-buffer-of-thoughts/technical-troubleshooting-reasoning-template?utm_source=reddit&utm_medium=social&utm_campaign=daily_special_set6&utm_content=d10

During an incident, who on your team is responsible for writing down the symptom before anyone starts fixing?

Source: r/NerdyChefs · by /u/Difficult-Sugar-4862

Leave a Reply

Your email address will not be published. Required fields are marked *