Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This adds a skill and guide for how to figure out how and why a particular flake is flaking, in order to get at the root cause, and fix it at the core instead of papering over the top of it.
LLM-generated technical summary
Adds
agentic-skills/incident-report.mdand a row for it in theagentic-skillsREADME table.The skill covers two things: investigating a failed or flaky run (test, deploy or environment) down to its mechanism, and writing that up as a blame-free report. It is written as a guide to pick from rather than a fixed procedure, and says so up front; the developer driving the investigation can narrow or reorder it.
The investigation part lists ways in: pin the exact failure signatures and when each wait started, check which commit the run actually tested, put CI output, deploy output, container logs, metrics, config and code on one timeline, diff against the last green run, test the explanations already on the table against the records, find the mechanism in config and code, reproduce locally, and measure how often with margins rather than pass/fail. It asks for the trigger to be explained and not only the gap, and for measured event counts to win over conclusions drawn from how the config ought to behave.
The report part separates certain claims (read directly from a record) from inferred ones (a conclusion, with what it rests on and a confidence in words, never an unmeasured percentage), allows one labelled best guess when the cause stays unknown, and gives an optional section list: TL;DR, timeline, mechanism, comparison with green runs, reproduction, frequency, considered and ruled out, fix, what does not fix it, open questions. The fix section is limited to what would have prevented the incident, and the checklist before publishing covers names, secrets copied from logs, and the impact outside test environments.
The content is drawn from a handful of recent incident write-ups and the investigations behind them.