Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 11 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -349,6 +349,13 @@ nullable or has a DEFAULT. `ALTER TYPE … ADD VALUE` needs a `.sql.conf` sideca
PENDING_AI → PENDING_REVIEW (bytes-scanned cap, no estimate, bytes_cap_missing_estimate=
REQUIRE_REVIEW — #941; suppresses the same auto-approve paths as a SQL
review BLOCK, never softens AUTO_REJECT)
PENDING_AI → REJECTED (data budget — #942; a SELECT whose submitter has used up an applying
per-user data budget with breach_action=REJECT. Decided right after the
bytes-scanned cap, no routing_decision row, QueryAutoRejectedEvent with a
null policy id; audited as QUERY_DATA_BUDGET_ENFORCED)
PENDING_AI → PENDING_REVIEW (data budget used up, breach_action=REQUIRE_REVIEW — #942; suppresses
the same auto-approve paths as a SQL review BLOCK, never softens
AUTO_REJECT)
PENDING_REVIEW → APPROVED or REJECTED (external ticket resolution — AF-453; a channel with
bidirectional_sync=true maps a ServiceNow/Jira ticket resolution onto a
decision via workflow.api.ExternalDecisionService. System-attributed:
Expand All @@ -373,7 +380,10 @@ nullable or has a DEFAULT. `ALTER TYPE … ADD VALUE` needs a `.sql.conf` sideca
APPROVED → EXECUTED (break-glass run — audit action QUERY_BREAK_GLASS_EXECUTED — AF-385)
APPROVED → FAILED (execution error; also the bytes-scanned cap re-checked just before
execution refusing the run — #941, scheduled / recurring /
break-glass included)
break-glass included; an exhausted data budget re-checked there —
#942, REJECT always, REQUIRE_REVIEW unless the exhausted budget
itself forced the review (data_budget_review_forced); break-glass
is counted but never capped or refused by a budget)
```

Illegal transitions must throw a domain exception, not silently succeed. **Break-glass /
Expand Down
3 changes: 2 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,7 @@ A glance at the day-to-day flows engineers and approvers actually use.

- **Proxy-first execution** — no user ever holds production credentials; the proxy holds them encrypted and opens connections only after approval. Single SQL statements run with autocommit; multi-statement INSERT/UPDATE/DELETE batches wrapped in `BEGIN; … COMMIT;` execute atomically inside one JDBC transaction (mixed SELECT/DML batches are rejected at parse time), with homogeneous INSERT runs collapsed into JDBC `executeBatch()` for bulk-load throughput. Optional **multi-replica read load balancing**: attach any number of replica endpoints to a datasource and SELECT traffic round-robins across the healthy ones — per-node health checks with circuit-breaker failover skip downed replicas, and only full replica-set exhaustion falls back to the primary (with an audit row). Optional **SELECT result caching**: opt a datasource into a Redis-backed result cache (per-datasource TTL) keyed over the security-rewritten query — masking and row-level security still apply — and invalidated on any proxied write to a referenced table.
- **Configurable review workflows** — per-datasource review plans, multi-stage sequential approval chains, optional auto-approve for reads, approval timeouts with auto-reject. Reviewers can be scoped **per-datasource** (directly or via groups) so different teams see only the queues that belong to them. Reviewers going away can set an **out-of-office delegation** naming a colleague to cover their review duty for a window — across queries, governed API calls, and grouped requests — with every decision recording both identities. A delegate can never act on the delegator's own requests, delegation never grants a permission they lack, and it does not chain. Plans can also **escalate** a request that nobody has decided on to the reviewers at its current stage plus your admins before the approval timeout auto-rejects it, and **nudge** those same reviewers on a cadence — both optional, both notify-only, so waiting never changes who may approve.
- **Policy-as-code routing** — ordered, attribute-based routing policies decide a query's path after AI analysis and before reviewers see it: **auto-approve**, **auto-reject**, **require N approvals**, or **escalate**. Conditions match on query type, referenced tables (glob), AI risk level / score, requester role or group, time-of-day / day-of-week, WHERE / LIMIT presence, **query shape** (joins, set operations, subqueries, CTEs, GROUP BY / HAVING, aggregates, window functions), the transactional flag, the pre-flight estimate (**estimated rows, estimated bytes scanned**, scan type), and the submission **client context** — source IP / CIDR, user-agent, time-since-last-approval, and CI/CD origin (API key or `X-AccessFlow-CI` header) — combined with AND / OR / NOT. Client-context conditions **fail closed** (missing context never auto-approves), so an off-network or stale-approval query escalates to stricter review instead. First match by priority wins; on no match the query falls through to the datasource's review plan. Every automated decision is recorded in the audit log.
- **Policy-as-code routing** — ordered, attribute-based routing policies decide a query's path after AI analysis and before reviewers see it: **auto-approve**, **auto-reject**, **require N approvals**, or **escalate**. Conditions match on query type, referenced tables (glob), AI risk level / score, requester role or group, time-of-day / day-of-week, WHERE / LIMIT presence, **query shape** (joins, set operations, subqueries, CTEs, GROUP BY / HAVING, aggregates, window functions), the transactional flag, the pre-flight estimate (**estimated rows, estimated bytes scanned**, scan type), the submitter's **data-budget usage**, and the submission **client context** — source IP / CIDR, user-agent, time-since-last-approval, and CI/CD origin (API key or `X-AccessFlow-CI` header) — combined with AND / OR / NOT. Client-context conditions **fail closed** (missing context never auto-approves), so an off-network or stale-approval query escalates to stricter review instead. First match by priority wins; on no match the query falls through to the datasource's review plan. Every automated decision is recorded in the audit log.
- **Policy simulator** — dry-run a draft routing, row-security, or masking policy against your own historical query traffic before you save it. AccessFlow replays the window **twice** — once against your current policies, once with the draft applied — and reports the difference between those two runs: how many past queries would change, which users would be affected, and which queries a row predicate would newly filter, deny, or reject outright. It is strictly read-only — no connection to your database, nothing executed, nothing persisted — and it names its own approximations (memberships are read as they are now; an engine that cannot classify a query shape offline is reported as *unclassifiable*, never as safe) instead of implying a precision the data does not have.
- **Deterministic SQL review rules (#860)** — a named rule catalog that judges every SQL query alongside the AI verdict and routing policies, and shows its findings **in the editor as you type**. Fourteen built-in rules derived from the parsed statement alone (missing `WHERE` on `UPDATE` / `DELETE`, an always-true `WHERE`, `SELECT *`, unbounded reads, cross joins, leading-wildcard `LIKE`, `DROP` / `TRUNCATE` / DDL, banned functions, protected tables by glob, DML outside a transaction), each at an admin-set **`OFF` / `WARN` / `BLOCK`** severity configured per **environment** (`DEVELOPMENT` / `TEST` / `STAGING` / `PRODUCTION`, plus an organisation default). **`BLOCK` escalates, it never rejects**: a blocking finding suppresses every auto-approve path and sends the query to a human reviewer; a routing auto-reject still rejects. Evaluated synchronously at submission so findings exist even when AI analysis is off or fails, rendered in each reader's language on the query detail, review queue, break-glass retro-review and request-group detail; break-glass runs record findings but are never gated by them, and non-relational engines report *not applicable*. See [`docs/19-sql-review.md`](https://github.com/bablsoft/accessflow/blob/main/docs/19-sql-review.md).
- **Just-in-time (JIT) access requests** — users self-request temporary, scoped access to a datasource (read/write/DDL, optional schema/table scope) or an API connection (read/write, optional operation allow-list) for an ISO-8601 duration. Requests flow through the same approval engine, a time-boxed permission is granted on approval, and a clustered scheduler auto-revokes it on expiry (admins can also revoke early). A grant can opt into **query pre-approval**: while it is active, queries it covers skip human review and are auto-approved with the grant recorded as the approval provenance — routing policies, high-risk AI verdicts, and behavioural anomalies still override.
Expand All @@ -85,6 +85,7 @@ A glance at the day-to-day flows engineers and approvers actually use.
- **Query playground / dry-run sandbox** — preview a query's impact before submitting it for review. A **Dry run** action returns the engine's execution plan (node type, estimated rows, cost, filters) and a best-effort estimated row impact **without executing or mutating data**, shown in the editor next to the AI panel. It runs through the same governance as a real execution (datasource access, table allow-list, row-level security) but creates no review. Dialect-aware across every engine with a plan concept — relational `EXPLAIN` (PostgreSQL / MySQL / MariaDB / Oracle / SQL Server), MongoDB `explain`, Couchbase / Neo4j `EXPLAIN`, Elasticsearch / OpenSearch query validation; engines without one degrade gracefully. Cloud warehouses additionally report the estimated bytes the query would scan (BigQuery's native dry-run job) — the direct cost signal for bytes-billed engines. SELECT dry-runs prefer the read replica when configured.
- **Pre-flight cost & blast-radius estimation (AF-624)** — every submitted query automatically gets a persisted cost estimate before review: the engine's own `EXPLAIN` plan (estimated rows, scan type, cost) plus, for UPDATE/DELETE, an **exact affected-row count** computed with a governed, non-mutating count (relational `SELECT COUNT(*)` rewrite; MongoDB `countDocuments`, SQL++ `COUNT(*)`, Cypher `count(*)`, Elasticsearch `_count`). Reviewers see it on the query detail page ("this DELETE touches ~2.4M rows via Seq Scan"), routing policies can match on `estimated_rows` / `scan_type` (e.g. auto-escalate full-scan DELETEs over 100k rows), and the estimate is folded into the AI analyzer's prompt so risk scoring reflects the actual blast radius. Engines without a plan concept degrade gracefully to an "estimate unavailable" state that never blocks submission.
- **Bytes-scanned cost caps** — on BigQuery, Snowflake and Databricks, where a ten-row result can scan a terabyte, set a maximum bytes-scanned per query on the datasource and, tighter, per grant. A query whose pre-flight scan estimate exceeds the cap is refused before it runs, with a message naming both numbers; the cap is checked again right before execution, so scheduled, recurring, grouped and break-glass runs honour it too. When no estimate exists, the datasource decides explicitly: send the query to a person, or refuse it. For an advisory signal instead, route on `estimated_bytes_scanned` to escalate expensive queries.
- **Per-user data-volume budgets** — bound how much data each person can read from a datasource over a rolling window (an hour up to 31 days): a row limit, a result-size limit, or both, for everyone or for chosen roles, groups or users — each person gets their own allowance, never a shared pool. While allowance remains, a query's result is capped to what is left; once it is used up, new reads are refused or sent to human review, per budget. Every delivered read counts — interactive, scheduled, recurring, grouped, table previews and break-glass (which is counted but never capped or blocked). Users see their remaining allowance in the query editor, get a warning past a threshold, and admins are told when someone runs out — the slow-exfiltration control that per-query caps cannot be. Route on `data_budget_used_percent` to escalate heavy readers before they hit the limit.
- **Approval-likelihood prediction (AF-645)** — a triage signal for busy review queues. AccessFlow trains a small per-organization statistical model on your own historical review decisions and shows reviewers the probability that a query gets approved — a bare percentage badge on the review queue, and a labelled "Historical approval likelihood: 78%" card on the query detail page. No LLM call and no external service — it's plain logistic regression over data you already collect (query type, AI risk verdict, cost estimate, time of day, per-submitter and per-datasource approval history), retrained daily. **Advisory only**: it never approves or rejects anything, and it is never an input to routing, grant coverage, or any other decision path. Auto-approved, break-glass, and grant-covered queries are excluded from training because they carry no reviewer judgment, and until an organization has enough decided history for the model to clear its quality gate the UI says "not enough review history yet" rather than showing a number.
- **Text-to-query** — opt-in per datasource. Users describe what they want in plain language ("order numbers for the last 5 days") and the AI drafts a schema-grounded query into the editor — in the datasource engine's **native query language** (SQL for relational engines, plus MongoDB shell/JSON, Cypher, CQL, Elasticsearch Query DSL, redis-cli, SQL++, and PartiQL for the NoSQL engines). Reuses the datasource's AI configuration and is grounded in its introspected schema (restricted columns are never referenced). The generated query is only a draft — it's still submitted through the normal pipeline, so AI risk analysis and human review always apply.
- **RAG knowledge base** — opt-in per AI configuration. Attach knowledge documents (data-governance policies, naming conventions, schema notes); AccessFlow embeds them with a dedicated embedding model and, at analysis / text-to-SQL time, retrieves the most relevant chunks and injects them into the prompt — so analysis reflects your house rules. Pluggable vector store: in-app **pgvector** (self-contained, auto-provisioned where the DB role permits — and optional: AccessFlow still starts if the extension is absent, with the in-app store disabled) or external **Qdrant**. Retrieval is best-effort and never blocks analysis.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,15 @@ public enum AuditAction {
* {@code estimated_bytes} and {@code outcome}.
*/
QUERY_BYTES_SCANNED_CAP_ENFORCED,
/**
* An exhausted data-volume budget (#942) refused a query or turned an automatic approval into
* human review ({@code stage=decision}), or refused an execution ({@code stage=execution}).
*/
QUERY_DATA_BUDGET_ENFORCED,
/** A data-volume budget (#942) was created, updated or deleted. */
DATA_BUDGET_CREATED,
DATA_BUDGET_UPDATED,
DATA_BUDGET_DELETED,
ROW_SECURITY_POLICY_CREATED,
ROW_SECURITY_POLICY_UPDATED,
ROW_SECURITY_POLICY_DELETED,
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -37,6 +37,7 @@ public enum AuditResourceType {
SERVICE_ACCOUNT("service_account"),
ROW_SECURITY_POLICY("row_security_policy"),
ROW_LIMIT_POLICY("row_limit_policy"),
DATA_BUDGET("data_budget"),
CONNECTOR("connector"),
QUERY_COMMENT("query_comment"),
DATA_CLASSIFICATION_TAG("data_classification_tag"),
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
package com.bablsoft.accessflow.core.api;

import java.util.List;
import java.util.UUID;

/**
* Admin CRUD for per-user data-volume budgets on a datasource (#942). All methods are
* organization-scoped: a datasource outside {@code organizationId} raises
* {@link DatasourceNotFoundException}; {@code applies_to} targets must belong to the organization.
*/
public interface DataBudgetAdminService {

List<DataBudgetView> listForDatasource(UUID datasourceId, UUID organizationId);

DataBudgetView create(UUID datasourceId, UUID organizationId, DataBudgetCommand command);

DataBudgetView update(UUID budgetId, UUID datasourceId, UUID organizationId,
DataBudgetCommand command);

void delete(UUID budgetId, UUID datasourceId, UUID organizationId);
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
package com.bablsoft.accessflow.core.api;

/**
* What happens to a SELECT once a data budget (#942) is exhausted. {@link #REJECT} beats
* {@link #REQUIRE_REVIEW} when several exhausted budgets disagree.
*/
public enum DataBudgetBreachAction {
REJECT,
REQUIRE_REVIEW;

public static DataBudgetBreachAction strictest(DataBudgetBreachAction a, DataBudgetBreachAction b) {
if (a == null) {
return b;
}
if (b == null) {
return a;
}
return a == REJECT || b == REJECT ? REJECT : REQUIRE_REVIEW;
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
package com.bablsoft.accessflow.core.api;

import java.util.List;
import java.util.UUID;

/**
* Create/update payload for a data budget (#942). Updates are total: every field is re-applied.
* At least one of {@code maxRows} / {@code maxBytes} must be set; {@code windowMinutes} and
* {@code breachAction} fall back to 1440 and {@link DataBudgetBreachAction#REQUIRE_REVIEW}.
*/
public record DataBudgetCommand(
String name,
Long maxRows,
Long maxBytes,
Integer windowMinutes,
DataBudgetBreachAction breachAction,
Integer warnThresholdPercent,
List<String> appliesToRoles,
List<UUID> appliesToGroupIds,
List<UUID> appliesToUserIds,
Boolean enabled) {
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
package com.bablsoft.accessflow.core.api;

import java.util.UUID;

/**
* One budget's consumption for one user over its trailing window (#942). A null limit means the
* budget does not bound that metric; the matching {@code remaining*} is then null too.
*/
public record DataBudgetConsumption(
UUID budgetId,
String name,
Long maxRows,
Long maxBytes,
int windowMinutes,
DataBudgetBreachAction breachAction,
Integer warnThresholdPercent,
long usedRows,
long usedBytes) {

public Long remainingRows() {
return maxRows == null ? null : Math.max(0, maxRows - usedRows);
}

public Long remainingBytes() {
return maxBytes == null ? null : Math.max(0, maxBytes - usedBytes);
}

public boolean exhausted() {
return (maxRows != null && usedRows >= maxRows) || (maxBytes != null && usedBytes >= maxBytes);
}

/** The larger of the rows and bytes percentages used, unclamped (may exceed 100). */
public double usedPercent() {
return Math.max(percent(usedRows, maxRows), percent(usedBytes, maxBytes));
}

/** This consumption after {@code rows} / {@code bytes} more were read. */
public DataBudgetConsumption plus(long rows, long bytes) {
return new DataBudgetConsumption(budgetId, name, maxRows, maxBytes, windowMinutes,
breachAction, warnThresholdPercent, usedRows + rows, usedBytes + bytes);
}

private static double percent(long used, Long limit) {
return limit == null ? 0d : used * 100d / limit;
}
}
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
package com.bablsoft.accessflow.core.api;

public sealed class DataBudgetException extends RuntimeException
permits DataBudgetNotFoundException, IllegalDataBudgetException {

protected DataBudgetException(String message) {
super(message);
}
}
Loading
Loading