Skip to content

inspect --all-instances: every Aurora writer and reader behind one endpoint (experimental) - #38

Merged
alexshapalov merged 4 commits into
mainfrom
feat/aurora-all-instances
Sep 29, 2026
Merged

alexshapalov merged 4 commits into
mainfrom
feat/aurora-all-instances

Conversation

@alexshapalov

@alexshapalov alexshapalov commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Implements #23 within its own constraints: SQL and DNS only, no AWS credentials, CLI, SDK, or API.

How it works

  • aurora_replica_status() lists the members; exactly one writer is expected, anything else is refused rather than mislabeled.
  • Each instance endpoint is derived from the entry endpoint's DNS name (<instance>.<cluster-id>.<region>.rds.amazonaws.com). Cluster, reader, custom, and instance endpoints all work as entry points; a custom domain is followed through its CNAME; RDS Proxy and non-RDS names are refused.
  • aurora_db_instance_identifier() must confirm each derived endpoint reached the member it names before anything is collected.
  • The --all-databases fan-out is generalized to (member, database) targets, so the two flags compose. Writer first, then readers, one member at a time. Cluster-wide findings dedupe per member.
  • Output: text banners each target; --json carries server.instance and server.instance_role (contract 1.3.0, additive); SARIF/JUnit objects are prefixed instance:<id>/; Prometheus series gain instance and role labels.
  • Any member that can't be reached is reported as partial coverage and the run exits 3.

Needs validation on a real cluster. I have no Aurora to run this against, so it is marked experimental and should not merge until someone with a cluster confirms it. @paul-enz, since you asked for this and validated the PgDog work so carefully, would you try it?

git fetch origin pull/38/head:aurora && git checkout aurora && go build -o /tmp/pgbot ./cmd/pgbot

# 1. discovery + derivation + identity check, no collection
PGBOT_AURORA_TEST_DSN="$AURORA_CLUSTER_URL" go test ./internal/conn/ -run TestIntegration_auroraInstances -v

# 2. the real thing
/tmp/pgbot inspect "$AURORA_CLUSTER_URL" --all-instances
/tmp/pgbot inspect "$AURORA_CLUSTER_URL" --all-instances --all-databases --json | jq '.[] | .server | {instance, instance_role, database}'

What I most want to know: whether every member was found and reached, whether the writer/reader roles are right, and, if derivation fails, the exact shape of your cluster endpoint (hostname with the ids redacted is fine). Also worth trying: the reader endpoint and an instance endpoint as the entry point, and the run through RDS Proxy, which should refuse cleanly.

…d one endpoint (experimental)

An Aurora cluster endpoint stands for several instances, and pg_stat_* on one
says nothing about the others (#23). Discovery uses SQL and DNS only:
aurora_replica_status() lists the members (exactly one writer expected;
anything else is refused), each instance endpoint is derived from the entry
endpoint's DNS name (<instance>.<cluster-id>.<region>.rds.amazonaws.com; a
custom domain is followed through its CNAME; RDS Proxy and non-RDS names are
refused), and aurora_db_instance_identifier() must confirm a derived endpoint
reached the member it names before anything is collected. No AWS credentials,
CLI, SDK, or API.

The --all-databases fan-out is generalized to (member, database) targets, so
the two flags compose; output is writer first, then readers, one member at a
time. Cluster-wide findings dedupe per member (parameter groups differ per
instance). Text banners each target; JSON carries server.instance and
server.instance_role (SchemaVersion 1.3.0, additive); SARIF/JUnit objects are
prefixed instance:<id>/; Prometheus series gain instance and role labels so two
members' samples for one database are distinct. Missing members are reported
as partial coverage and fail the run with exit 3.

conn.ConnectDBAt overrides the host (and the TLS server name, so verify-full
validates the member's own certificate). Unit tests cover endpoint derivation
across cluster/reader/custom/instance endpoints and the GovCloud and China
partitions, role parsing, target composition, dedupe, merge tagging, and the
Prometheus labels; TestIntegration_auroraInstances (PGBOT_AURORA_TEST_DSN) is
the opt-in end-to-end check against a real cluster.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013qGZKWgfGTBCoHsDjy1SuB
@paul-enz

paul-enz commented Sep 9, 2026

Copy link
Copy Markdown

@alexshapalov sorry I haven't got around to this yet. I do intend to test it out tomorrow and I'll get back to you.

@paul-enz

Copy link
Copy Markdown

Hey @alexshapalov Sorry for the wait. I gave this a run against an Aurora PostgreSQL 18.3 cluster, connecting straight to the cluster endpoint, and --all-instances refuses to start:

pgbot: discover instances: not an Aurora cluster (aurora_version() is missing) — --all-instances needs a native Aurora endpoint

It is a real Aurora cluster. The catalog probe in internal/conn/connect.go:249-253 is what misses it:

-- Aurora exposes aurora_version(). Look it up in the catalog rather
-- than CALLING it: on every other server the call fails, which writes
-- an ERROR to the server log and books a rollback in pg_stat_database
-- on each pgbot run — the very counter pgbot reports.
(SELECT count(*) FROM pg_proc WHERE proname = 'aurora_version') > 0

On 18.3 that returns 0. There are no aurora% rows in pg_proc at all, even though the functions are there and return exactly what discovery wants:

select aurora_version();            -> 18.3.4
select server_id, session_id from aurora_replica_status();
  -> one writer row (MASTER_SESSION_ID) plus one reader row

With IsAurora false, detectProvider falls through to rds, so AuroraInstances bails out before it ever runs the query.

Swap the probe to to_regprocedure. It finds them where pg_proc does not, and resolves the signature instead of calling it, so nothing lands in the server log or in pg_stat_database:

--- a/internal/conn/connect.go
+++ b/internal/conn/connect.go
@@ -246,7 +246,11 @@
 		       pg_is_in_recovery(),
-		       -- Aurora exposes aurora_version(). Look it up in the catalog rather
-		       -- than CALLING it: on every other server the call fails, which writes
-		       -- an ERROR to the server log and books a rollback in pg_stat_database
-		       -- on each pgbot run — the very counter pgbot reports.
-		       (SELECT count(*) FROM pg_proc WHERE proname = 'aurora_version') > 0`
+		       -- Aurora exposes aurora_version() and aurora_replica_status(), but
+		       -- does not catalogue them in pg_proc on 18.3. to_regprocedure finds
+		       -- them by signature without CALLING them, so no ERROR is written to
+		       -- the server log and no rollback is booked in pg_stat_database.
+		       to_regprocedure('aurora_version()') IS NOT NULL
+		       OR to_regprocedure('aurora_replica_status()') IS NOT NULL`

and the marker comment goes stale too:

--- a/internal/conn/provider.go
+++ b/internal/conn/provider.go
@@ -30,1 +30,1 @@
-	IsAurora    bool // aurora_version() exists in pg_proc
+	IsAurora    bool // aurora_version() or aurora_replica_status() is resolvable

Both predicates are true on 18.3. I built that locally and discovery gets past detection and on to host derivation.

--all-instances refused to start against a real Aurora PostgreSQL 18.3
cluster connected through its writer endpoint:

    pgbot: discover instances: not an Aurora cluster (aurora_version()
    is missing) — --all-instances needs a native Aurora endpoint

Aurora 18.3 exposes aurora_version() and aurora_replica_status() but
catalogues neither in pg_proc, so the count(*) probe returned 0, IsAurora
stayed false, detectProvider fell through to rds, and AuroraInstances bailed
out before its query ever ran.

to_regprocedure resolves a function by signature without calling it, so it
finds them where pg_proc does not while keeping the property the original
probe was written for: no ERROR in the server log and no rollback booked in
pg_stat_database — the very counter pgbot reports. Verified on PostgreSQL
15.12, where the predicate returns false and xact_rollback is unchanged
across repeated runs, while SELECT aurora_version() raises an ERROR and
increments it.

Reported-by: paul-enz, who diagnosed this against Aurora PostgreSQL 18.3
and confirmed both predicates resolve there.
#55 landed the collation_version_mismatch finding and took SchemaVersion
1.3.0, which this branch was also claiming. Three files conflicted:

- internal/model/schema_version.go — both bumped the const to 1.3.0. Keep
  main's 1.3.0 collation note and add a 1.4.0 note for instance/instance_role,
  which are omitted on a single-instance run, so a 1.3.0 consumer still parses
  1.4.0 output.
- schema/pgbot-context-1.3.0.json — add/add. Restored byte-identical to main's
  (the collation schema); this branch's additions now live in a new
  pgbot-context-1.4.0.json regenerated by tools/schemagen, so
  TestSchema_matchesModel stays green.
- CHANGELOG.md — both added a first entry under Unreleased. Kept both, and
  noted the 1.4.0 bump on the --all-instances entry.

The conflict also explained why CI never ran on 7e8a373: GitHub cannot build
the merge ref for a conflicted PR, so no pull_request workflow was dispatched.
@alexshapalov
alexshapalov merged commit a4be025 into main Sep 29, 2026
20 checks passed
@alexshapalov
alexshapalov deleted the feat/aurora-all-instances branch September 29, 2026 22:11
@alexshapalov

Copy link
Copy Markdown
Contributor Author

Merged, with your to_regprocedure fix — thank you for chasing this down. Resolving both functions by signature rather than looking them up in pg_proc is exactly right, and it keeps the property the original probe was written for. I confirmed that half on plain PostgreSQL 15.12: the predicate returns false without raising, and pg_stat_database.xact_rollback is unchanged across repeated runs, while SELECT aurora_version() does raise and does increment it.

The branch also had to move its schema bump to 1.4.0 — 1.3.0 went to the collation finding in #55 while this was open.

Could I ask you for one more run against your cluster, now on main? Your earlier report confirmed detection gets past and on to host derivation, which is the part that was broken — but not that a full --all-instances run completes. That is the half nobody has seen yet, and it is the part most likely to have a second problem hiding in it.

Most useful to know:

  1. Does the run finish, and does it reach every member — one writer plus each reader?
  2. Does endpoint derivation produce reachable hostnames for your cluster endpoint shape? That is the piece I would most expect to break on a cluster that is not shaped like the common case, and it is why the flag ships marked experimental.
  3. If a member is unreachable, does it fail loudly and exit 3 rather than silently reporting partial coverage?
  4. Anything odd in --json — server.instance and server.instance_role should name each member.

If derivation does fail, the cluster endpoint shape is the thing to paste; that is what I would need to fix it.

No rush, and thanks again — this would have shipped broken for every Aurora 18.3 user without your report.

@alexshapalov

Copy link
Copy Markdown
Contributor Author

@paul-enz sorry for late response!

lofoneh added a commit to lofoneh/pgbot that referenced this pull request Sep 30, 2026
…'t be updated

A table published for UPDATE or DELETE needs a replica identity so the
subscriber can find the row. Without one Postgres rejects the write itself:
"cannot update table … because it does not have a replica identity and
publishes updates". Reads and INSERTs keep working, so the failure lands on
the first UPDATE after a migration, not at deploy time.

A catalog-only gauge collector reads pg_publication and pg_class: DEFAULT with
no primary key, NOTHING, or USING INDEX whose index is gone or invalid.
Insert-only publications need no identity and are not reported. Scope is
schema, so it runs under --profile=schema and pgbot lint on an empty CI
database — the migration-PR path where this is worth catching.

The integration fixture proves the UPDATE fails before asserting the finding,
then that a primary key clears both.

New replica_identity section in --json; SchemaVersion 1.5.0 (additive) — pgrundev#38
took 1.4.0 first, so this is the next minor.
alexshapalov pushed a commit that referenced this pull request Sep 30, 2026
…'t be updated (#113)

A table published for UPDATE or DELETE needs a replica identity so the
subscriber can find the row. Without one Postgres rejects the write itself:
"cannot update table … because it does not have a replica identity and
publishes updates". Reads and INSERTs keep working, so the failure lands on
the first UPDATE after a migration, not at deploy time.

A catalog-only gauge collector reads pg_publication and pg_class: DEFAULT with
no primary key, NOTHING, or USING INDEX whose index is gone or invalid.
Insert-only publications need no identity and are not reported. Scope is
schema, so it runs under --profile=schema and pgbot lint on an empty CI
database — the migration-PR path where this is worth catching.

The integration fixture proves the UPDATE fails before asserting the finding,
then that a primary key clears both.

New replica_identity section in --json; SchemaVersion 1.5.0 (additive) — #38
took 1.4.0 first, so this is the next minor.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants