A chaos-engineering and self-healing agent for Kubernetes — induces real and synthetic infrastructure failures, then detects, diagnoses, and remediates them with OpenAI and a human-in-the-loop approval gate.
Architecture · Roadmap · Sentinel Platform
Phoenix is the resilience tier of the Sentinel platform — the counterpart to Argus (security). Where Argus watches for attackers, Phoenix manufactures failure on purpose and proves the cluster (and the agent) can recover from it.
- Provisioning Simulator — a FastAPI service that mimics enterprise cloud infrastructure operations (volume create/attach, VLAN/subnet create, instance provision) and is intentionally faultable, so Phoenix can induce infrastructure-operation failures, not just pod failures
- Chaos injection engine — wraps Chaos Mesh (pod kill, network latency, packet loss, IO delay) and the simulator's fault hooks behind one control surface
- LangGraph agent — detect (Prometheus/Loki anomaly watch) → diagnose (OpenAI reads logs/metrics and reconstructs the causal chain — "X failed because Y") → heal (executes remediation via MCP tools) → human-approval gate (risky actions pause for sign-off with full rationale + predicted outcome) → verify (confirm recovery, log MTTR)
- Dashboard — live chaos-scenario grid, blast-radius graph, healing pipeline swim lanes, action ledger, and failure-mode catalog, in the platform's dark command-center style
- Causal incident chains reconstructed from the dependency graph + event timeline
- Blast-radius prediction before a chaos scenario runs
- Resilience score per datacenter/component from MTTR, recovery rate, cascade-prevention rate
- Failure-mode taxonomy auto-classifying every failure (transient, cascading, resource-exhaustion, network-partition, quota-limit)
- Predictive healing with confidence — "seen this 7×, scaling replicas fixed 6, failover needed 1. Confidence 85%. Suggested action: …"
- Synthetic user-journey simulation measuring failure propagation end to end
The journeys/ service provides a cluster-free, seeded resilience
proof loop. It generates realistic customer operations, load profiles, and
faults; enforces safety budgets and human approval for high-risk experiments;
then reports availability, recovery, MTTR, error-budget consumption, and whether
the original journey passed after healing. The same seed replays the same test.
| Module | Description | Status |
|---|---|---|
| M1.1 — Provisioning Simulator | Faultable volume/subnet/instance lifecycle APIs (/sim) |
Complete |
| M1.2 — Chaos Injection Engine | Chaos Mesh wrapper + simulator fault control surface (/chaos) |
Complete |
| M1.3 — Fault Library & Taxonomy Classifier | Failure-mode classification from observed events (/faultlib) |
Complete |
| M1.4 — Blast-Radius Graph Builder | Dependency graph from live k8s + Hubble topology (/graph) |
Complete |
| Build Week Phase 4 — Customer Journey Resilience Lab | Seeded load/fault generation, safety gates, recovery evidence (/journeys) |
Complete (cluster-free adapter) |
| M2 — Phoenix Agent | LangGraph detect → diagnose → heal → approve → verify state machine, predictive-healing memory store, event publisher | Complete |
| M3 — Phoenix Dashboard | React console, blast-radius graph, healing pipeline swim lanes, incident feed, fleet weakness map | Complete |
See the milestones for the full build sequence and issue backlog.
For the complete Build Week judge experience, use the cluster-free one-command platform path from the sibling Argus checkout. It starts a disposable local SOG plus the real local Argus, Phoenix, and Sentinel services; builds an explicitly synthetic service topology; publishes deterministic cross-agent evidence; and verifies the correlated Sentinel incident before reporting success.
Projects/
├── argus-k8s/ # run the command here
└── sentinel-stack/
├── phoenix/
├── sentinel/
└── sentinel-platform/
Diagnose the portable path without starting a container, process, or publishing evidence:
make -C ../../argus-k8s doctorLaunch the complete demo:
make -C ../../argus-k8s demo-platformOpen Argus at http://127.0.0.1:5173, Phoenix at
http://127.0.0.1:5174, and Sentinel at http://127.0.0.1:5175. The command
installs missing local dependencies and labels every topology fixture as
synthetic_fixture/demo-data=synthetic. Argus evidence is replayed; Phoenix outcomes
are simulator. A bounded feed updates the dashboards during the presentation. No
Kubernetes API, Hubble relay, Falco workload, or Chaos Mesh fault is used. Ctrl-C
stops the local services and removes the disposable Redis container.
For the guarded real k3s proof instead:
kubectl config use-context argus
make -C ../../argus-k8s doctor-live
make -C ../../argus-k8s demo-platform-liveThe doctor is read-only and ends with a READY/NOT READY verdict plus an exact next
action. The live command requires the exact context and the phrase
INJECT LIVE FAULT, creates only sentinel-live-demo, and continuously probes a
two-replica HTTP target. Phoenix must create a real Chaos Mesh PodChaos for one
disposable replica; the proof passes only after Kubernetes supplies a new Ready pod,
both replicas are Ready, measured availability is recorded, and Sentinel correlates the
observed Argus evidence with Phoenix's verified recovery. Ctrl-C deletes only the
isolated namespace.
Phoenix persists complete scenario and agent-run documents in SQLite: correlation IDs,
backend references, lifecycle status, diagnosis, approvals, actions, verification,
errors, MTTR, and timestamps. Kubernetes mounts separate persistent volume claims for
both histories, so completed evidence remains visible after pod restarts. An unfinished
action is retained and marked for manual review after restart; Phoenix never resumes a
risky action implicitly. Both health endpoints expose persistent_history; false
means the service is using a deliberately ephemeral test database.
For Phoenix by itself, choose one of the paths below. The local path proves Phoenix's safe simulator workflow without touching Kubernetes. The cluster path adds live topology, Hubble flows, Chaos Mesh, and the OpenAI-powered recovery agent.
| Path | Kubernetes | Dashboard | What you can inject |
|---|---|---|---|
| Safe local demo | No | http://127.0.0.1:5174 |
Synthetic provisioning faults only |
| Live k3s demo | Yes | http://127.0.0.1:3000 |
Synthetic faults plus bounded Chaos Mesh experiments |
Run the simulator, chaos API, taxonomy service, and dashboard in four terminals. The dashboard intentionally reports cluster topology and agent panels as unavailable; those are live-only signals, not mocked data.
# Terminal 1 — provisioning simulator
python3 -m venv sim/.venv
sim/.venv/bin/pip install -r sim/requirements.txt
sim/.venv/bin/python -m uvicorn main:app --app-dir sim/src --host 127.0.0.1 --port 8083# Terminal 2 — safe simulator-domain injection API
python3 -m venv chaos/.venv
chaos/.venv/bin/pip install -r chaos/requirements.txt
SIMULATOR_URL=http://127.0.0.1:8083 \
chaos/.venv/bin/python -m uvicorn main:app --app-dir chaos/src --host 127.0.0.1 --port 8082# Terminal 3 — fault catalog and scenario rankings
python3 -m venv faultlib/.venv
faultlib/.venv/bin/pip install -r faultlib/requirements.txt
CHAOS_URL=http://127.0.0.1:8082 \
faultlib/.venv/bin/python -m uvicorn main:app --app-dir faultlib/src --host 127.0.0.1 --port 8081# Terminal 4 — Phoenix console
npm --prefix dashboard install
VITE_ARGUS_URL=http://127.0.0.1:5173 \
VITE_SENTINEL_URL=http://127.0.0.1:5175 \
npm --prefix dashboard run devOpen http://127.0.0.1:5174, choose Safe Simulation, and inject a simulator fault. This path never creates a Kubernetes or Chaos Mesh resource.
Phoenix reuses Argus's real three-node k3s cluster. Verify the context before deploying:
kubectl config use-context argus
kubectl get nodes
kubectl get pods -n chaos-meshExpect k3s-master, k3s-worker1, and k3s-worker2 to be Ready. Then deploy all
Phoenix services and keep the supervised port-forwards running:
./deploy.shOpen http://127.0.0.1:3000. Safe Simulation remains non-disruptive. Live
k3s shows the exact namespace, selector, fault, duration, and blast radius and
requires the typed confirmation INJECT LIVE FAULT before creating a bounded Chaos
Mesh resource.
For prerequisites, existing-cluster checks, service-by-service development, and troubleshooting, follow setup.md.
FastAPI + LangGraph agent, OpenAI Responses API for reasoning, MCP tools for cluster actions
The Phoenix top bar switches directly between all three platform consoles. Destinations are never inferred from fixed ports: configure them explicitly for each environment. Missing destinations render as disabled instead of navigating to an assumed address:
For the local three-console demo, the reserved ports are:
| Console | URL |
|---|---|
| Argus | http://127.0.0.1:5173 |
| Phoenix | http://127.0.0.1:5174 |
| Sentinel | http://127.0.0.1:5175 |
Set the switcher destinations before starting Phoenix locally:
VITE_ARGUS_URL=http://127.0.0.1:5173 \
VITE_SENTINEL_URL=http://127.0.0.1:5175 \
npm --prefix dashboard run devFor deployed environments, provide their public URLs during the dashboard build:
VITE_ARGUS_URL=https://argus.example.com \
VITE_SENTINEL_URL=https://sentinel.example.com \
npm --prefix dashboard run build(kubectl, PromQL, Loki, Chaos Mesh, provisioning sim), React + TypeScript + Vite + Tailwind dashboard. Reuses the existing Prometheus/Grafana/Loki + Cilium stack and k3s cluster from argus-k8s — no duplicate infrastructure.