Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
53 commits
Select commit Hold shift + click to select a range
19faed1
feat(k8s): trace-instrument the workflow — every error clear, waits live
pablovilas Aug 7, 2026
8170c7a
feat(k8s): trace waits as sub-steps with live progress, cover every w…
pablovilas Aug 12, 2026
f6e51cf
feat(k8s): SWM-grade lineage — what each step wrote, read and offers
pablovilas Aug 12, 2026
5f8ed97
chore(tracing): re-vendor nptrace.sh with the lint fix from catalog-t…
pablovilas Aug 12, 2026
ddbbb53
chore(tracing): re-vendor nptrace.sh after the readability refactor i…
pablovilas Aug 12, 2026
9f71a9c
feat(k8s): speak the platform's step vocabulary — SWM-shaped plans fo…
pablovilas Aug 12, 2026
39d2429
fix(k8s): close the networking verification as skipped on flavors wit…
pablovilas Aug 12, 2026
9c3610e
feat(k8s): declare the cluster flavor separately from the DNS type
pablovilas Aug 12, 2026
e49592c
fix(k8s): a station azure can never verify is not declared there — no…
pablovilas Aug 12, 2026
54bb8ca
feat(k8s): declare job definitions so plans preview before any run
pablovilas Aug 12, 2026
755d051
feat(k8s): full SWM narrative — counted io, health meter, severity ex…
pablovilas Aug 12, 2026
a125e92
feat(k8s): complete the annotation coverage — every workflow, every f…
pablovilas Aug 12, 2026
314f41a
chore(tracing): re-vendor nptrace.sh with the complete surface (catal…
pablovilas Aug 12, 2026
529cbcc
fix(k8s): failed applies carry kubectl's actual reason; restarts carr…
pablovilas Aug 12, 2026
cc3882c
feat(k8s): show the REAL error, proactively — kubernetes' words and t…
pablovilas Aug 12, 2026
35291db
fix(k8s): surface the provider's real error in IAM and resource lookups
pablovilas Aug 13, 2026
e044c80
feat: consume catalog-tracing-sh as a git submodule
pablovilas Aug 13, 2026
2792c19
feat(k8s): state each error's cause to the workflow engine via np_ste…
pablovilas Aug 13, 2026
1f3aaba
feat(k8s): a traced run missing the bundled SDK says so, loudly
pablovilas Aug 25, 2026
6bdbcf0
fix(k8s): each workflow's story lands on its OWN station
pablovilas Aug 29, 2026
0865820
feat(k8s): user-facing step titles everywhere; plumbing hidden
pablovilas Aug 29, 2026
8884739
fix(tracing): vendor nptrace.sh at the repo root — agent clones carry…
pablovilas Aug 29, 2026
beeb855
fix(k8s): a step's diagnostics attach to the row the user clicks
pablovilas Aug 29, 2026
050af7c
fix(tracing): vendor SDK with adopted-step coordinate triple
pablovilas Aug 29, 2026
f591ef6
fix(tracing): adopted-step observations accumulate in one bag
pablovilas Aug 29, 2026
dc25cc3
feat(tracing): phase-titled checks, human narratives, curated diagnose
pablovilas Aug 29, 2026
98d6a7e
fix(tracing): the adopted-node cache survives subshells
pablovilas Aug 29, 2026
cb79dd3
fix(tracing): land the wait's final truth and the apply lineage now
pablovilas Aug 29, 2026
2e5d16f
fix(tracing): derive the wait's phase title from its step identity
pablovilas Aug 29, 2026
b6858b8
fix(tracing): progress speaks the wire's unit vocabulary
pablovilas Aug 29, 2026
994f135
fix(tracing): finalize milestones carry their station group
pablovilas Aug 29, 2026
895a4ba
fix(tracing): honest milestones, declared job identity, per-increment…
pablovilas Aug 29, 2026
21cd41e
fix(tracing): deployment workflows trace declared steps only
pablovilas Aug 30, 2026
78d6c9d
fix(tracing): vendored SDK sends wire-shaped edge bindings
pablovilas Aug 30, 2026
dd1772a
fix(tracing): waits declare their signal, not label chips
pablovilas Aug 30, 2026
14096fb
fix(tracing): applied/removed manifests named by their literal k8s kind
pablovilas Aug 30, 2026
537730f
feat(tracing): job definitions scoped per provider
pablovilas Aug 30, 2026
ed9dc05
feat(workflows): every traced workflow declares what a human calls it
pablovilas Aug 30, 2026
83791cc
test(tracing): pin the wait, binding and progress contracts as they n…
pablovilas Aug 30, 2026
0817dc5
chore(vendor): track the SDK at its merged commit on main
pablovilas Aug 30, 2026
01de7d4
feat(workflows): finalize and rollback say what each step is for, and…
pablovilas Aug 30, 2026
c071460
feat(tracing): a stuck wait says why, and what to do about it
pablovilas Aug 30, 2026
c419557
fix(tracing): a multi-word reason survives the reason list whole
pablovilas Aug 30, 2026
b2d612c
fix(k8s): name the CAUSE of a crash loop, not the loop
pablovilas Aug 31, 2026
4e26364
fix(k8s): every crash cause explains itself, not just the OOM
pablovilas Aug 31, 2026
8f45ad8
fix(k8s): suggest the setting the console shows, not the field key
pablovilas Aug 31, 2026
133403e
feat(k8s): every workflow says what a human calls it
pablovilas Aug 31, 2026
fa88bcd
fix(k8s): a diagnostic check says what it FOUND on its own step
pablovilas Aug 31, 2026
aa3c825
test(k8s): assert the full hint lines, not fragments
pablovilas Aug 31, 2026
c64249e
feat(k8s): lifecycle jobs declare their identity, so the page finds t…
pablovilas Sep 1, 2026
65b9ac4
feat(k8s): scope lifecycle jobs declare their identity, so consumers …
pablovilas Sep 1, 2026
7894642
refactor(tracing): drop explanatory comments from the traced scripts …
pablovilas Sep 2, 2026
856b8b3
merge(beta): bring the diagnose results payload fix into trace-instru…
pablovilas Sep 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/pr-checks.yml
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,8 @@ jobs:
runs-on: ubuntu-24.04
steps:
- uses: actions/checkout@v4
with:
submodules: true # catalog-tracing-sh — the tracing SDK the tests exercise

- name: Install dependencies
run: sudo apt-get update && sudo apt-get install -y bats jq
Expand Down
4 changes: 4 additions & 0 deletions .gitmodules
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
[submodule "vendor/catalog-tracing-sh"]
path = vendor/catalog-tracing-sh
url = https://github.com/nullplatform/catalog-tracing-sh.git
branch = main
1 change: 1 addition & 0 deletions azure-aro/values.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
configuration:
DNS_TYPE: azure
K8S_FLAVOR: aro
USE_ACCOUNT_SLUG: false
IMAGE_PULL_SECRETS:
ENABLED: false
Expand Down
1 change: 1 addition & 0 deletions azure/values.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,6 @@
configuration:
DNS_TYPE: azure
K8S_FLAVOR: aks
USE_ACCOUNT_SLUG: false
IMAGE_PULL_SECRETS:
ENABLED: false
Expand Down
53 changes: 51 additions & 2 deletions k8s/apply_templates
Original file line number Diff line number Diff line change
Expand Up @@ -31,9 +31,39 @@ while IFS= read -r TEMPLATE_FILE; do
IGNORE_NOT_FOUND="--ignore-not-found=true"
fi

if ! kubectl "$ACTION" -f "$TEMPLATE_FILE" $IGNORE_NOT_FOUND; then
log error " ❌ Failed to apply"
TRACE_STEP_KEY=""
if [[ "$ACTION" == "apply" ]] && command -v np_scope_step_begin >/dev/null 2>&1; then
case "$FILENAME" in
deployment-*) TRACE_STEP_KEY="create-deployment" ;;
secret-*) TRACE_STEP_KEY="create-secret" ;;
scaling-*) TRACE_STEP_KEY="create-hpa" ;;
service-*) TRACE_STEP_KEY="create-service" ;;
pdb-*) TRACE_STEP_KEY="create-pod-disruption-budget" ;;
ingress-*)
if [[ -n "${DEPLOYMENT_ID:-}" ]]; then
TRACE_STEP_KEY="configure-ingress"
else
TRACE_STEP_KEY="create-ingress"
fi ;;
esac
[[ -n "$TRACE_STEP_KEY" ]] && np_scope_step_begin "$TRACE_STEP_KEY"
fi

if KUBECTL_OUT=$(kubectl "$ACTION" -f "$TEMPLATE_FILE" $IGNORE_NOT_FOUND 2>&1); then
[[ -n "$KUBECTL_OUT" ]] && echo "$KUBECTL_OUT"
[[ -n "$TRACE_STEP_KEY" ]] && command -v np_scope_step_end >/dev/null 2>&1 && np_scope_step_end 0
if [[ "$ACTION" == "apply" ]] && command -v np_scope_k8s_applied >/dev/null 2>&1; then
np_scope_k8s_applied "${K8S_NAMESPACE:-}" "$KUBECTL_OUT"
fi
if [[ "$ACTION" == "delete" ]] && command -v np_scope_k8s_deleted >/dev/null 2>&1; then
np_scope_k8s_deleted "${K8S_NAMESPACE:-}" "$KUBECTL_OUT"
fi
else
[[ -n "$KUBECTL_OUT" ]] && echo "$KUBECTL_OUT" >&2
log error " ❌ Failed to apply $FILENAME${KUBECTL_OUT:+: $KUBECTL_OUT}"
[[ -n "$TRACE_STEP_KEY" ]] && command -v np_scope_step_end >/dev/null 2>&1 && np_scope_step_end 1
fi
TRACE_STEP_KEY=""
fi

DEST_DIR="${BASE_DIR}/$ACTION"
Expand All @@ -52,4 +82,23 @@ if [[ "$DRY_RUN" == "true" ]]; then
exit 1
fi

if [[ "${TRACE_TRAFFIC_SWITCH:-false}" == "true" ]] && [[ -n "${CONTEXT:-}" ]] \
&& command -v np_scope_explain >/dev/null 2>&1; then
TRAFFIC_TO=$(echo "$CONTEXT" | jq -r '.deployment.strategy_data.desired_switched_traffic // 100')
TRAFFIC_FROM=$(echo "$CONTEXT" | jq -r '.deployment.strategy_data.switched_traffic // 0')
if [[ "$TRAFFIC_TO" =~ ^[0-9]+$ ]] && [[ "$TRAFFIC_FROM" =~ ^[0-9]+$ ]]; then
np_scope_labels "deployment.id=${DEPLOYMENT_ID:-}" "scope.id=${SCOPE_ID:-}" "action=traffic-switch"
np_scope_explain --title "Switch traffic for deployment ${DEPLOYMENT_ID:-}" \
--what "Switching blue/green traffic for deployment ${DEPLOYMENT_ID:-} from ${TRAFFIC_FROM}% to ${TRAFFIC_TO}%"
np_scope_affordance "{\"kind\":\"traffic-switch\",\"deployment_id\":\"${DEPLOYMENT_ID:-}\",\"current_traffic\":$TRAFFIC_TO,\"new_traffic\":$TRAFFIC_TO,\"old_traffic\":$((100 - TRAFFIC_TO)),\"target_traffic\":100}"
np_scope_input traffic "{\"from\":$TRAFFIC_FROM,\"desired\":$TRAFFIC_TO}"
np_scope_output traffic "{\"switched\":$TRAFFIC_TO}"
np_scope_progress "$TRAFFIC_TO" 100 percent
fi
fi

if command -v np_trace_flush >/dev/null 2>&1; then
NP_TRACE_FLUSH_TIMEOUT=5 np_trace_flush
fi

source "$SERVICE_PATH/backup/backup_templates" --action="$ACTION" --files "${APPLIED_FILES[@]}"
28 changes: 25 additions & 3 deletions k8s/deployment/print_failed_deployment_hints
Original file line number Diff line number Diff line change
Expand Up @@ -124,7 +124,12 @@ diagnose_failure() {
if [[ -n "$pods_json" ]] && echo "$pods_json" | jq -e . >/dev/null 2>&1; then
FAILURE_REASON=$(echo "$pods_json" | jq -r '
[.items[].status.containerStatuses[]?
| (.state.waiting.reason // .lastState.terminated.reason // empty)
| (.state.waiting.reason // "") as $w
| (.lastState.terminated.reason // "") as $t0
| (if $t0 == "Completed" then "" else $t0 end) as $t
| (if ($w == "CrashLoopBackOff" or $w == "BackOff") and $t != "" then $t
elif $w != "" then $w
else $t end)
] | map(select(. != "" and . != "Completed")) |
group_by(.) | max_by(length) | .[0] // empty' 2>/dev/null)

Expand Down Expand Up @@ -181,13 +186,20 @@ diagnose_failure() {
CrashLoopBackOff|BackOff)
HUMAN_MESSAGE="The container started and crashed repeatedly."
SUGGESTED_FIX="Review application logs for startup errors (failed dependencies, bad config, panics)." ;;
Error|ContainerStatusUnknown)
if [[ -n "$FAILURE_EXIT_CODE" ]]; then
HUMAN_MESSAGE="The container started and exited with code ${FAILURE_EXIT_CODE}."
else
HUMAN_MESSAGE="The container started and crashed repeatedly."
fi
SUGGESTED_FIX="Review application logs for startup errors (failed dependencies, bad config, panics)." ;;
OOMKilled)
if [[ -n "$req_memory" ]]; then
HUMAN_MESSAGE="The container exceeded its memory limit (${req_memory}Mi) and was terminated."
else
HUMAN_MESSAGE="The container exceeded its memory limit and was terminated."
fi
SUGGESTED_FIX="Increase ram_memory for scope '$scope_name' or reduce application memory usage." ;;
SUGGESTED_FIX="Raise the RAM Memory of scope '$scope_name', or reduce how much memory the application uses." ;;
CreateContainerConfigError)
HUMAN_MESSAGE="The container configuration is invalid."
SUGGESTED_FIX="Check for missing secrets or configmaps referenced by the deployment." ;;
Expand Down Expand Up @@ -223,18 +235,28 @@ diagnose_failure() {
elif [[ "$UNHEALTHY_MESSAGE" =~ statuscode:[[:space:]]*([0-9]+) ]]; then
SUGGESTED_FIX="The app responded with HTTP ${BASH_REMATCH[1]} on $health_check_path — inspect application logs for startup errors; the process is running but $health_check_path is not returning 2xx."
elif [[ "$UNHEALTHY_MESSAGE" == *"context deadline exceeded"* || "$UNHEALTHY_MESSAGE" == *"Client.Timeout"* || "$UNHEALTHY_MESSAGE" == *"i/o timeout"* ]]; then
SUGGESTED_FIX="The probe timed out — the app may be slow to start or $health_check_path is blocking. Consider increasing startup probe initialDelaySeconds/timeoutSeconds, or making $health_check_path lighter."
SUGGESTED_FIX="The probe timed out — the app may be slow to start, or $health_check_path is blocking. Raise the health check's Initial Delay or Timeout on scope '$scope_name', or make $health_check_path lighter."
else
SUGGESTED_FIX="Ensure the app listens on port 8080 and returns 2xx on $health_check_path within the readiness window."
fi ;;
FailedCreate|FailedCreatePodSandBox)
HUMAN_MESSAGE="Kubernetes could not create the pod sandbox."
SUGGESTED_FIX="Check node health, CNI configuration, and pod security policies." ;;
Evicted)
HUMAN_MESSAGE="The pod was evicted from its node."
SUGGESTED_FIX="The node ran out of memory or disk. Lower the scope's resource requests, or free capacity on the cluster." ;;
Unschedulable)
HUMAN_MESSAGE="No node can accept the pod."
SUGGESTED_FIX="Reduce requested resources, free cluster capacity, or review nodeSelector/affinity rules." ;;
DeadlineExceeded)
HUMAN_MESSAGE="The pod ran past its active deadline and was stopped."
SUGGESTED_FIX="Find why the workload runs longer than the deadline it was given, or allow it more time." ;;
"")
HUMAN_MESSAGE=""
SUGGESTED_FIX="" ;;
*)
HUMAN_MESSAGE="Pods are failing with reason: $FAILURE_REASON"
# Empty on purpose: an empty fix routes the reader to the generic checklist below.
SUGGESTED_FIX="" ;;
esac
}
Expand Down
10 changes: 10 additions & 0 deletions k8s/deployment/scale_deployments
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,12 @@ GREEN_DEPLOYMENT_ID=$DEPLOYMENT_ID
BLUE_REPLICAS=$(echo "$CONTEXT" | jq -r .blue_replicas)
BLUE_DEPLOYMENT_ID=$(echo "$CONTEXT" | jq .scope.current_active_deployment -r)

if [ "$DEPLOY_STRATEGY" != "rolling" ]; then
if command -v np_step_skip >/dev/null 2>&1; then
np_step_skip "instance counts are set by the manifests for the $DEPLOY_STRATEGY strategy — nothing to scale"
fi
fi

if [ "$DEPLOY_STRATEGY" = "rolling" ]; then
GREEN_DEPLOYMENT_NAME="d-$SCOPE_ID-$GREEN_DEPLOYMENT_ID"
BLUE_DEPLOYMENT_NAME="d-$SCOPE_ID-$BLUE_DEPLOYMENT_ID"
Expand Down Expand Up @@ -41,6 +47,10 @@ if [ "$DEPLOY_STRATEGY" = "rolling" ]; then
unset TIMEOUT
unset SKIP_DEPLOYMENT_STATUS_CHECK

if command -v np_scope_output >/dev/null 2>&1; then
np_scope_output replicas "{\"green\": $GREEN_REPLICAS, \"previous\": $BLUE_REPLICAS}"
fi

log debug ""
log info "✨ Deployments scaled successfully"
fi
107 changes: 80 additions & 27 deletions k8s/deployment/tests/print_failed_deployment_hints.bats
Original file line number Diff line number Diff line change
Expand Up @@ -83,7 +83,8 @@ assert_not_contains() {
assert_contains "$output" "📋 Reason: The container exceeded its memory limit (512Mi)"
assert_contains "$output" "📋 Detected: OOMKilled on container app (exit 137)"
assert_contains "$output" "📋 Details: out of memory"
assert_contains "$output" "💡 Suggested fix: Increase ram_memory for scope 'my-app'"
assert_contains "$output" "💡 Suggested fix: Raise the RAM Memory of scope 'my-app'"
assert_not_contains "$output" "ram_memory"
assert_not_contains "$output" "⚠️ Application Startup Issue Detected"
}

Expand Down Expand Up @@ -229,8 +230,9 @@ assert_not_contains() {
run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
assert_contains "$output" "did not pass its health check at /health"
assert_contains "$output" "💡 Suggested fix: Ensure the app listens on port 8080 and returns 2xx on /health"
assert_contains "$output" "📋 Reason: The application did not pass its health check at /health."
assert_contains "$output" "📋 Detected: Unhealthy on container api"
assert_contains "$output" "💡 Suggested fix: Ensure the app listens on port 8080 and returns 2xx on /health within the readiness window."
assert_not_contains "$output" "⚠️ Application Startup Issue Detected"
}

Expand All @@ -250,13 +252,11 @@ assert_not_contains() {
run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
# HUMAN_MESSAGE retains the base sentence and appends the translated probe failure
assert_contains "$output" "did not pass its health check at /health"
assert_contains "$output" "Detected: Startup probe"
assert_contains "$output" "not yet listening"
# SUGGESTED_FIX is targeted: tells the user the app is not binding the port
assert_contains "$output" "not listening on port 8080"
# Generic fallback fix must NOT appear
assert_contains "$output" "📋 Reason: The application did not pass its health check at /health. Detected: Startup probe — app is not yet listening on /health."
assert_contains "$output" "📋 Detected: Unhealthy on container api"
assert_contains "$output" "📋 Recent warnings:"
assert_contains "$output" " • Unhealthy (×1)"
assert_contains "$output" "💡 Suggested fix: The container is not listening on port 8080 — verify the start command runs, the process binds to 0.0.0.0:8080, and nothing is crashing before it accepts connections."
assert_not_contains "$output" "returns 2xx on /health within the readiness window"
}

Expand All @@ -276,11 +276,9 @@ assert_not_contains() {
run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
assert_contains "$output" "Detected: Startup probe"
assert_contains "$output" "HTTP 502"
# SUGGESTED_FIX cites the status code and points to app logs
assert_contains "$output" "responded with HTTP 502"
assert_contains "$output" "inspect application logs"
assert_contains "$output" "📋 Reason: The application did not pass its health check at /health. Detected: Startup probe — app responded with HTTP 502 (expected 2xx)."
assert_contains "$output" "📋 Detected: Unhealthy on container api"
assert_contains "$output" "💡 Suggested fix: The app responded with HTTP 502 on /health — inspect application logs for startup errors; the process is running but /health is not returning 2xx."
}

@test "print_failed_deployment_hints: enriches Unhealthy with timeout detail and targeted fix" {
Expand All @@ -299,16 +297,14 @@ assert_not_contains() {
run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
assert_contains "$output" "Detected: Startup probe"
assert_contains "$output" "timed out"
# SUGGESTED_FIX mentions timing knobs
assert_contains "$output" "initialDelaySeconds"
assert_contains "$output" "📋 Reason: The application did not pass its health check at /health. Detected: Startup probe — request timed out on /health."
assert_contains "$output" "📋 Detected: Unhealthy on container api"
assert_contains "$output" "💡 Suggested fix: The probe timed out — the app may be slow to start, or /health is blocking. Raise the health check's Initial Delay or Timeout on scope 'my-app', or make /health lighter."
assert_not_contains "$output" "initialDelaySeconds"
}

@test "print_failed_deployment_hints: falls back to raw Unhealthy message when translation is impossible" {
export K8S_NAMESPACE="ns" DEPLOYMENT_ID="d1"
# Message does not match any known probe pattern → translate_probe_message returns non-zero.
# The raw text must still be surfaced in the hint instead of being silently dropped.
export ALL_EVENTS='{"items":[{"type":"Warning","reason":"Unhealthy","lastTimestamp":"2026-05-20T13:13:42Z","message":"completely unknown probe failure format from a future K8s"}]}'

kubectl() {
Expand All @@ -323,10 +319,8 @@ assert_not_contains() {
run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
# Raw message appears verbatim in the reason line
assert_contains "$output" "completely unknown probe failure format from a future K8s"
# Base sentence is still there
assert_contains "$output" "did not pass its health check at /health"
assert_contains "$output" "📋 Reason: The application did not pass its health check at /health. Detected: completely unknown probe failure format from a future K8s"
assert_contains "$output" "💡 Suggested fix: Ensure the app listens on port 8080 and returns 2xx on /health within the readiness window."
}

@test "print_failed_deployment_hints: Unhealthy picks the latest event when multiple are present" {
Expand Down Expand Up @@ -399,8 +393,6 @@ assert_not_contains() {
# health_check_path default "/" must apply when CONTEXT is unset.
assert_contains "$output" "health check at /."
assert_contains "$output" "returns 2xx on /"
# Guard against the previous escape bug: a literal backslash in the message
# would indicate jq received {\} instead of {} and silently failed.
assert_not_contains "$output" "{\\"
}

Expand Down Expand Up @@ -488,3 +480,64 @@ assert_not_contains() {
[ "$status" -eq 0 ]
assert_contains "$output" "📊 Progress at failure: 1/3 ready, 2/3 available"
}

@test "print_failed_deployment_hints: an OOM kill inside a crash loop is diagnosed as the OOM" {
export K8S_NAMESPACE="ns" DEPLOYMENT_ID="d1"
kubectl() {
case "$*" in
"get pods"*)
echo '{"items":[{"status":{"containerStatuses":[{"name":"app","state":{"waiting":{"reason":"CrashLoopBackOff","message":"back-off 20s restarting failed container=app pod=d-1_ns(abc)"}},"lastState":{"terminated":{"reason":"OOMKilled","exitCode":137}}}]}}]}'
;;
esac
}
export -f kubectl

run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
assert_contains "$output" "📋 Reason: The container exceeded its memory limit (512Mi) and was terminated."
assert_contains "$output" "📋 Detected: OOMKilled on container app (exit 137)"
assert_contains "$output" "📋 Details: back-off 20s restarting failed container=app pod=d-1_ns(abc)"
assert_contains "$output" "💡 Suggested fix: Raise the RAM Memory of scope 'my-app', or reduce how much memory the application uses."
assert_not_contains "$output" "The container started and crashed repeatedly."
}

@test "print_failed_deployment_hints: a plain non-zero exit in a crash loop keeps its startup-log advice" {
export K8S_NAMESPACE="ns" DEPLOYMENT_ID="d1"
kubectl() {
case "$*" in
"get pods"*)
echo '{"items":[{"status":{"containerStatuses":[{"name":"app","state":{"waiting":{"reason":"CrashLoopBackOff","message":"back-off 20s restarting failed container=app"}},"lastState":{"terminated":{"reason":"Error","exitCode":1}}}]}}]}'
;;
esac
}
export -f kubectl

run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
assert_contains "$output" "📋 Reason: The container started and exited with code 1."
assert_contains "$output" "📋 Detected: Error on container app (exit 1)"
assert_contains "$output" "📋 Details: back-off 20s restarting failed container=app"
assert_contains "$output" "💡 Suggested fix: Review application logs for startup errors (failed dependencies, bad config, panics)."
assert_not_contains "$output" "Pods are failing with reason: Error"
}

@test "print_failed_deployment_hints: an evicted pod says why and what to do" {
export K8S_NAMESPACE="ns" DEPLOYMENT_ID="d1"
kubectl() {
case "$*" in
"get pods"*)
echo '{"items":[{"status":{"containerStatuses":[{"name":"app","state":{"waiting":{"reason":"Evicted"}}}]}}]}'
;;
esac
}
export -f kubectl

run bash "$BATS_TEST_DIRNAME/../print_failed_deployment_hints"

[ "$status" -eq 0 ]
assert_contains "$output" "📋 Reason: The pod was evicted from its node."
assert_contains "$output" "📋 Detected: Evicted on container app"
assert_contains "$output" "💡 Suggested fix: The node ran out of memory or disk. Lower the scope's resource requests, or free capacity on the cluster."
}
Loading
Loading