Conversation
|
Hi @elkaix — apologies, the schema version has moved under you twice while this sat, and that is on me for not sequencing it sooner. Where things stand: This PR should claim 1.5.0. Nothing else is queued behind it, so that number is yours. Rebasing onto For the two schema files: take One thing that will save you a confusing CI failure — do not hand-edit the schema JSON. It is generated, and That writes One more worth a careful eye: Everything else in the 32 files still applies. Thanks for your patience on this one — the pg_stat_io work is the part of the backlog I most want to land. |
|
@elkaix pls resolve conflicts |
…1.6.0) Double-sample pg_stat_io (PG16+) per backend_type × object × context into physical read/write/fsync rates and, with track_io_timing on, mean per-op latency. Latency — not the miss count — is what separates a page-cache hit from a wait on the device. Also whitelist the PG18 AIO knobs (io_method, io_workers, io_max_concurrency), effective_io_concurrency, hash_mem_multiplier, plan_cache_mode and max_slot_wal_keep_size so rules can read them. Additive: Context gains io_stats, contract bumps to 1.3.0.
…retention rules - io_read_latency_high: mean physical read latency ≥ 5 ms over ≥ 500 reads (critical at 20 ms) — the storage-bound verdict, with the heaviest reader named. Remediation orders the fixes: read fewer blocks, fit the working set, then the IO knobs, then the volume. - io_concurrency_low: effective_io_concurrency ≤ 1 on an SSD-backed provider. - plan_cache_mode_forced: a cluster-wide generic/custom pin removes the planner's defence against parameter-sensitive plans. - slot_wal_keep_unbounded: max_slot_wal_keep_size = -1 with slots present. - work_mem_overcommit now computes the real envelope: shared_buffers + work_mem × hash_mem_multiplier × max_connections. - PoolSizing (Little's law from pg_stat_statements totals) feeds connections_overprovisioned and a pool line in pgbot tune. - inspect --full prints the pg_stat_io line; checked-line and catalog wired.
…ract Evidence → hypothesis → refused reflex, in the layer order query/schema → concurrency → memory/IO/WAL/vacuum → replicas → partitioning → sharding. Names the three inputs to ask for (workload class, p99 target, topology) and the shape of the answer (attribution, ranked hypotheses, cost/verify/rollback).
…io verdict in why (schema 1.6.0) - partition_skew: the hottest leaf takes ≥ 4× the per-partition average scans (or holds ≥ 4× the rows). What pgbot can see of the hot-shard problem — a partition key with a hot value. Rollup SQL now returns the hot/big leaf. - autovacuum_table_tuning: a ≥ 1M-row write-active table still on the global scale factor; reports the current trigger and the per-table override (0.02 / 1000) that would replace it. Surfaced by pgbot tune. - wait_io_bound and why --duration now bracket the window with pg_stat_io: ms-per-read says the device served the misses, µs says the page cache did, and the next step differs (fewer blocks vs cache/IO knobs). WaitStudy gains io; PartitionRollup gains hot/big partition fields. Contract 1.3.0 → 1.4.0.
Image templates take the ghcr namespace from IMAGE_REPO (set to the releasing repository), the GitHub release targets the repository the workflow runs in, and brew-smoke only runs upstream where the tap credential exists. A fork's tag now produces binaries, SBOMs and an image instead of failing on pgrundev-owned targets.
The image-runs and cosign-verify jobs pulled from and verified against pgrundev/pgbot by name, so on a fork they 404'd after a successful goreleaser. Derive both from GITHUB_REPOSITORY (lowercased for ghcr).
pgrundev#113 landed after pgrundev#114 removed pgx from go.mod, so the test's pgx import left the collect package unbuildable under go vet / go test ./...
33aae27 to
8098779
Compare
…he latency verdict - pg_stat_io wal rows (PG18) are timed by track_wal_io_timing, not track_io_timing: counted but zero-timed by default. They now count toward the rates and stay out of the latency means and reads_in_window. - Latency only when track_io_timing was on at both ends of the window. - autovacuum_table_tuning applies PG18's autovacuum_vacuum_max_threshold cap (default 100M) to the trigger it reports. - idle_replication_slot_timeout is PG18+, not PG17+. - model.IOStats.JudgedReadLatency owns the "enough reads, timing on" rule and the 1 ms device boundary; findings and why both use it instead of their own copies of 500 / 1.0. - why distinguishes "too few reads" from a timing reset. - The wait study's closing pg_stat_io sample survives Ctrl+C under its own budget, so an interrupted window still gets its IO verdict. - partitions.sql reads pg_stat_user_tables once. - work_mem_overcommit no longer prints "shared_buffers 0 B" when the setting fails to parse.
|
@alexshapalov rebased onto main and conflicts resolved.
CI is waiting on "Approve and run" (fork PR). |
Summary
Adds the "attribute before you tune" layer: a
pg_stat_iocollector with per-read latency, six new deterministic findings, a memory-envelope fix forwork_mem_overcommit, a Little's-law pool-size line inpgbot tune, and the storage verdict inpgbot why --duration. Everything stays read-only; the JSON contract goes 1.5.0 → 1.6.0 additively (one bump on top of main's 1.5.0; schema generated bytools/schemagen,TestSchema_matchesModelgreen).Collectors / model
io_stats(pg_stat_io, PG16+): double-sampled read/write/fsync rates per backend_type × object × context and, withtrack_io_timingon for the whole window, mean ms per op. PG18walrows count toward the rates but stay out of the latency means (they're timed bytrack_wal_io_timing). Latency, not the miss count, is what separates a page-cache hit from a wait on the device.settings.paramsgains the PG18 AIO knobs (io_method,io_workers,io_max_concurrency),effective_io_concurrency,hash_mem_multiplier,plan_cache_mode,max_slot_wal_keep_size,autovacuum_vacuum_max_threshold(PG18).PartitionRollupgains the hottest/largest leaf;WaitStudygainsio(pg_stat_io bracketed around the sampling window).Findings (each with a docs page + unit tests)
io_read_latency_highio_concurrency_loweffective_io_concurrency≤ 1 on an SSD-backed managed providerplan_cache_mode_forcedplan_cache_modepinned away fromautocluster-wideslot_wal_keep_unboundedmax_slot_wal_keep_size = -1with replication slots presentpartition_skewautovacuum_table_tuningautovacuum_vacuum_max_threshold)work_mem_overcommitnow computesshared_buffers + work_mem × hash_mem_multiplier × max_connectionsvseffective_cache_size.wait_io_boundand the livewhystorage branch carry the pg_stat_io verdict: "device served them → cut blocks read first" vs "page cache served them → volume, not latency".connections_overprovisionedandpgbot tuneprint the average-busy-backends estimate and a pool size.Docs / skill
skills/postgres-diagnostics/SKILL.md: evidence → next hypothesis → refused reflex table, the layer order (query/schema → concurrency → memory/IO/WAL/vacuum → replicas → partitioning → sharding), and the one-change-with-cost/verify/rollback hand-off shape.Also on this branch
test(collect): port the replica-identity fixture to pggo: feat(findings): replica_identity_missing — a published table that can't be updated #113 landed after Switch the PostgreSQL driver from pgx to pgGo #114 dropped pgx from go.mod, sogo vet/go test ./...on main fail at package setup (every CI job on main is red for this). Happy to split it into its own PR.Release workflow
IMAGE_REPO, GitHub release targets the running repository,brew-smokegated to upstream. No behavior change forpgrundev/pgbot.Test plan
Rebased on main @
acd485d(2026-10-01).go test ./...,go vet,gofmt,golangci-lint(0 issues),scripts/gate.sh(4 arches)PGBOT_TEST_SUPERUSER_DSN) againstpostgres:18withtrack_io_timing=on, a hash-partitioned table with a hot tenant, and a 1.2M-row updated table:partition_skewrollup fields populated,autovacuum_table_tuningsurfaced bytune,why --durationJSON carriesstudy.iogo test -race ./..., vet, golangci-lint (0 issues); integration on postgres:16/17/18 (doc-verify on 18: 114 statements); new tests verified to fail with each fix revertedgoreleaser check(config valid; thebrewsdeprecation predates this PR)