Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 10 additions & 5 deletions docs/observability.md
Original file line number Diff line number Diff line change
Expand Up @@ -360,18 +360,23 @@ as loaded and healthy and could not fire for any input ([#63](https://github.com
`prometheus/tests/*.test.yaml` holds `promtool test rules` unit tests, which
feed a rule synthetic series and assert it fires — paired with a case asserting
it stays quiet, because a test that only ever expects silence would have passed
against the broken rule too. Coverage is twenty-seven rules of 56 so far — the five
against the broken rule too. Coverage is twenty-eight rules of 56 so far — the five
in `blackbox.rules.yaml`, both in `dns.rules.yaml`, `ContainerHighMemory`,
`ContainerNearMemoryLimit`, `ContainerRestartLoop` and
`ContainerNearMemoryLimit`, `ContainerRestartLoop`, `ContainerCpuThrottled` and
`PrometheusSizeRetentionActive`, `Watchdog`, the three iLO rules from
[#76](https://github.com/Gerrrt/HomeLab/issues/76), all five in
`backup.test.yaml`, all five in `deploy.test.yaml`, `RemoteWriteJobStale`,
and `SuricataStopped`.
The other 29 are still validated for syntax only, which is exactly the
The other 28 are still validated for syntax only, which is exactly the
standing #63 had. Both numbers are checked by `scripts/check_docs.py` — the
sentence they replaced claimed six and named two, and had been wrong for
weeks. Keep each count on one line: the checker reads prose line by line, so a
phrase wrapped mid-claim is a claim it cannot see.
weeks.

`ContainerCpuThrottled` is the odd one in that list: it is
inert in production and cannot fire against anything cAdvisor
currently reports, because no service sets a CPU quota. Its tests are what make
the rule's correctness checkable anyway, which is the #63 lesson applied before
rather than after the fact ([#185](https://github.com/Gerrrt/HomeLab/issues/185)).

Disk alerting is predictive rather than a fixed threshold — `predict_linear` over
a 6-hour window, firing when the extrapolation reaches zero within a day *and*
Expand Down
27 changes: 27 additions & 0 deletions stacks/observability/prometheus/rules/containers.rules.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,33 @@ groups:
Anonymous memory, so page cache is not inflating it. The container
is OOM-killed if this reaches 100%.

# INERT BY DESIGN, and deliberately so — see #185.
#
# CFS accounting only exists when a CPU quota exists. No service in any
# stack sets `cpus` or `deploy.resources`, so the cgroup reports
# nr_periods=0 and nr_throttled=0 forever, both rates are 0, and the
# numerator can never be non-zero. `promtool check rules` passes it —
# it parses PromQL and never asks whether an expression can be true —
# and it shows loaded and healthy on the Prometheus status page. That is
# #63 exactly, one rule up in this same file.
#
# The tempting fix is the wrong one. Adding `cpus:` would make this live
# and should not be done: the whole stack peaks at ~0.4 of 4 cores, and
# HostHighLoad already fires at load15 > 2 x cores. There is no saturation
# to prevent, and CPU degrades gracefully anyway — CFS fair-shares, so a
# runaway produces slowness rather than death. Meanwhile a 0.20-core-
# average service saturates a modest quota inside individual 100ms CFS
# windows during a scrape burst, so the limit would manufacture exactly
# the alert fatigue this repository avoids, to justify a control that
# prevents nothing.
#
# So it is kept, unfireable, against the day a quota is set for some other
# reason — a noisy-neighbour guest, a service that genuinely needs
# bounding. What has to change for it to become live: any service gains
# `cpus:` or `deploy.resources.limits.cpus`. Nothing else. The unit tests
# in containers.test.yaml assert it CAN fire against synthetic non-zero
# CFS series, so its correctness is checked even while production cannot
# produce the input — which is the whole lesson of #63.
- alert: ContainerCpuThrottled
expr: |
rate(container_cpu_cfs_throttled_periods_total{name!=""}[15m])
Expand Down
64 changes: 64 additions & 0 deletions stacks/observability/prometheus/tests/containers.test.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -347,3 +347,67 @@ tests:
- eval_time: 20m
alertname: ContainerRestartLoop
exp_alerts: []

# --- ContainerCpuThrottled: fires ----------------------------------------
# The rule is inert in production and always has been: CFS accounting only
# exists when a CPU quota exists, no service sets one, and both counters stay
# at 0 forever (#185). That is precisely why it needs a firing case — the
# input it would fire on is one production cannot currently produce, so
# nothing else in this repository can tell a correct expression from a broken
# one. #63 is what that looks like when nobody checks.
#
# Synthetic, and realistic: the default CFS period is 100ms, so a container
# pinned by a quota accrues 600 periods a minute. 180 of them throttled is
# 30%, comfortably over the 0.25 threshold and a plausible figure for a
# service being squeezed rather than a contrived 100%.
#
# `rate()` returns a value from the second sample onward — it does not wait
# for the 15m window to fill — so the ratio is over threshold from 1m, the
# alert goes pending there and fires at 31m. Evaluated at 50m, with 25m as
# the paired "not yet" so the `for:` is asserted rather than assumed.
- interval: 1m
input_series:
- series: 'container_cpu_cfs_throttled_periods_total{name="loki",instance="prometheus",host="prometheus",job="integrations/cadvisor"}'
values: "0+180x60"
- series: 'container_cpu_cfs_periods_total{name="loki",instance="prometheus",host="prometheus",job="integrations/cadvisor"}'
values: "0+600x60"
alert_rule_test:
- eval_time: 50m
alertname: ContainerCpuThrottled
exp_alerts:
- exp_labels:
alertname: ContainerCpuThrottled
name: loki
instance: prometheus
host: prometheus
job: integrations/cadvisor
component: containers
severity: info
category: saturation
exp_annotations:
summary: "Container loki throttled 30% of CPU periods"

# Pending, not firing — for: 30m has not elapsed at 25m.
- eval_time: 25m
alertname: ContainerCpuThrottled
exp_alerts: []

# --- ContainerCpuThrottled: quiet on what production actually reports -----
# Both counters flat at 0, which is what cAdvisor reports for every container
# in this estate and will keep reporting until something sets a CPU quota.
# clamp_min keeps the denominator at 1 rather than dividing by zero, so the
# expression evaluates to 0 and stays silent instead of producing NaN.
#
# This is the case that documents the rule's real state. If a quota is ever
# set, this fixture stops resembling production — and the firing case above
# is what says the rule was ready for it.
- interval: 1m
input_series:
- series: 'container_cpu_cfs_throttled_periods_total{name="loki",instance="prometheus",host="prometheus",job="integrations/cadvisor"}'
values: "0x60"
- series: 'container_cpu_cfs_periods_total{name="loki",instance="prometheus",host="prometheus",job="integrations/cadvisor"}'
values: "0x60"
alert_rule_test:
- eval_time: 50m
alertname: ContainerCpuThrottled
exp_alerts: []