diff --git a/CHANGELOG.md b/CHANGELOG.md index 619ae9c..d2a46f5 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -4,6 +4,58 @@ Notable changes to JevGate. Versions follow [Semantic Versioning](https://semver ## [Unreleased] +A pull request check now judges only what the change touches and fails only on what has been measured right, and OpenRouter and Vercel AI Gateway keys work as TypeSafe keys do. A check with the defaults (`jevgate check --base`, as the action runs it), replayed from the answer cache on the last commit of 118 corpus projects (90 open-source, 28 of the maintainer's own private repositories), failed 18 of them with 0.25.0 and left one more incomplete (its replay lacked a cached answer), on 61 reviews, 32 of the 46 labeled right (9 wrong, 5 debatable). With 0.26 it fails 9, on 11 function-simplification reviews, 9 of them right (none wrong, 2 debatable). 8 of those 9 projects, and 10 of the 11 findings, are the maintainer's own; the other is ky's `Ky` constructor, labeled right. It reports 29 reviews and 70 considers where 0.25.0 reported 61 and 140. On the 25 of those projects never used for tuning, it fails 2 instead of 4. + +- **The default gate fails only on what is measured right.** By default, only the rules and levels measured right on projects JevGate was never tuned on fail the check. The default gate level is now `mature`: a rule's reviews or considers fail the check when at least 80% of them were right on the 25 projects never used for tuning (11 held out, 14 fresh), over at least 20 findings labeled by hand from the code. Every other finding is still reported, marked as still being measured, and the check passes. Two levels are mature: function-simplification reviews (20 of 23 right, 87%) and agent-context considers (22 of 24, 92%). Agent context is a documentation rule, which runs only when selected: with `--rule documentation` or `--rule all`, its considers fail the check, the first consider level to do so. Its 22 right findings on unseen projects come from 4 of the maintainer's own repositories. On the unseen projects' full runs, the default gate failed on 122 findings, 57% of the labeled ones right (64% leaving the 13 debatable ones out), and on 17 of 22 projects; it now fails on 23, 87% right, and on 9 projects. On the 72 projects used for tuning: 468 findings at 73% to 95 at 83%, and 61 projects to 28. The 49 right reviews on unseen projects that no longer fail the check are still reported. The concern probability could not do this: unseen reviews were right 55%, 46%, 56% and 61% of the time with a probability below 0.90, below 0.95, below 0.98 and above. The labels were made on 0.24.1's findings and joined to 0.25.0's, replayed from the answer cache; 0.25.0 had turned 17 labeled security findings into notes, one on an unseen project, none in a mature level. Bend 2's labels are kept apart, so `tests/laws`, which judges only Bend 2 code, has no row: its findings are reported without failing the check, and say they were labeled only on Bend 2 projects. Right / labeled on unseen and tuned projects, a debatable label counting as not right: + + | Rule | Reviews, unseen | Reviews, tuned | Considers, unseen | Considers, tuned | + |---|---|---|---|---| + | File organization | 2/5 | 15/30 | 17/29 | 19/43 | + | Function simplification | **20/23 (87%)** | 57/69 | 85/126 (67%) | 147/197 | + | Shared logic | 46/85 (54%) | 181/240 (75%) | 85/158 (54%) | 146/244 | + | Hardcoded values | 1/8 | 15/28 | 5/29 | 32/57 | + | Injection | 3/4 | 81/96 | 5/13 | 27/47 | + | Sensitive data | 10/24 | 41/64 | 0/5 | 12/15 | + | Unsafe settings | 2/4 | 53/72 | 0/4 | 16/21 | + | Access control | - | 2/5 | - | 5/14 | + | Workflows | 1/1 | 0/1 | - | - | + | Test value | 3/5 | 5/11 | 1/2 | 20/27 | + | Test redundancy | 1/1 | 4/4 | 27/45 (60%) | 27/34 | + | Agent context | - | - | **22/24 (92%)** | 64/68 | + | Large docs | - | - | 1/1 | 1/5 | + | Staleness | - | - | 2/2 | 14/15 | + | Duplication | - | - | 3/20 | 5/22 | + | Code comments | - | - | 39/72 (54%) | 97/150 | + + - Any level you set replaces the default exactly as it says: `fail_on = ["review"]` or `--fail-on review` fails on every review, as before; `--fail-on security=consider` sets one group and leaves the others at `mature`, which can also be set by name (`security = "mature"` in `[rules]`). Undecided answers never fail the check under `mature`. + - A `jevgate.toml` written by `jevgate init` before 0.26 sets `maintainability = "review"` and `tests = "review"`, which keeps every review of those groups failing the check and judges hardcoded values; delete the two lines for the new default. While they are there with `init`'s comments, every command that reads `jevgate.toml` says so on stderr; deleting the comments keeps the levels and stops the notice. `jevgate init` now writes each group as a commented example. + - The output says what fails. The agent text marks each finding that fails the gate `(fails the gate)`, and says when reviews did not fail it because their rules are still being measured and how to make them fail it. A GitHub warning, a SARIF result and a GitLab issue for such a finding say so, with how often its rule and level were right. The JSON report records how the gate counted each new finding as `gate` (`fails`, `measuring` or `advisory`), and `fail_on_mature` says what `mature` stands for among the selected rules; the HTML report and the MCP server's `jevgate_findings` show both. Where a list is capped (the agent text's 10 considers, GitHub's 10 warning annotations a step and 50 summary rows, the MCP server's 50 findings), the findings that fail the gate come first, then reviews, each by rank: ranked with considers, 267 of 424 reviews still being measured fell past GitHub's tenth warning in whole-repository runs of 94 corpus projects. + - `jevgate rules` shows the levels that fail by default and how often each rule's reviews and considers were right on unseen projects, with the number labeled; `--format json` adds each level's labels on unseen and tuned projects as `maturity`. +- Hardcoded values no longer runs by default: 6 of its 37 labeled reviews and considers were right on projects JevGate was never tuned on (16%), against 47 of 85 on the projects it was tuned on (55%). Without it, a default run asks 44% fewer first-pass requests (22,370 to 12,623 on 94 corpus projects) and uploads 39% fewer bytes. It stays in the `maintainability` group and in `all`; `--rule default --rule hardcoded-values` adds it to the default rules. +- **`--base` judges only what the change touches.** It selected whole changed files, so a pull request check asked about and reported every unit of a touched file: on each corpus project's last commit, 54% of the review and consider findings in the changed files sat on lines the commit did not touch. A check with `--base` now asks about and reports the functions, tests, comments, values and security units whose lines the change added or modified, or removed lines between; copies where either copy changed; a file's outline, and a large document's, only when the change adds a member or heading its base version lacks; and a document the change left alone only in a section that names a path it deleted or renamed. A new or untracked file is judged whole, and `--whole-files` keeps the whole-file check, asking exactly what 0.25.0 asked. On the last commits of 118 corpus projects (all rules, tests included): + + | | 0.25.0 | 0.26 | + |---|---|---| + | Review and consider findings off the changed lines | 167 of 311 (54%) | 9 of 149 (6%) | + | Findings on changed lines | 144 | 140: 137 the same, 3 new; 7 of 0.25.0's not reported | + | Labeled right, on changed lines | 73% | 75% | + | First-pass requests, nothing cached | 4,828 | 2,259 | + | First-pass input tokens, nothing cached | 11.05M ($0.46) | 4.55M ($0.19) | + + The 9 off the changed lines are there on purpose: 5 outlines the change added members to, whose finding names an existing group, 3 copy pairs reported at the copy that did not change, and a section with lines removed inside it. Of the 7 not reported, 3 are outlines of files the commit added no member to, which are not asked (one labeled wrong, one debatable), one is a consider over six comments, labeled right, of which the commit touched one, now a note, and 3 are function-simplification considers asked beside fewer functions, now notes (one labeled wrong): 1 right, 2 wrong, 1 debatable and 3 unlabeled. The 3 new ones include comments past the 80 per file that a whole-file check never asks: the cap now counts only the comments the change touched. About $0.12 on the corpus, both versions' runs included. + - After upgrading, a `--base` check asks its touched units once more, in changed-lines packs its cache has not seen: with a cache holding the whole-file answers of the same commits, the corpus's first pass needed 716 new requests and 1.34M input tokens, where 0.25.0 needed 765 and 1.06M (dry runs). `--whole-files` asks nothing again. + - Functions, values, comments, security units, instruction sections and laws are still packed within runs whose ends are computed over the whole file, so the changed functions of one run share a pack and no other pack is sent. A later push that changes another function of the run asks that pack again whole: over the last two commits of 16 corpus projects that edit one file twice, the second push re-asked 93 units the first had asked, 1% of the bytes it sent (3% judging whole files). A unit asked beside other functions can answer differently: of 11,693 first-pass answers about the same units, 88% were the same as with whole-file packs, and 36 crossed 0.50 or 0.80 (13 up, 23 down). + - The report says what a check judged: `scope` is `changed-lines` or `whole-files` in the JSON report, and the headline and the HTML report say `changed lines since 1a2b3c4`. After a check of changed lines, `baseline --merge` keeps the other accepted findings of the files it checked, and an earlier finding it did not judge is `non-comparable` in the report's changes, not `resolved`. The MCP tool `jevgate_check` takes `whole_files`. +- `--base` in a repository whose `jevgate.toml` sits in a subdirectory, such as one package of a monorepo, found no tracked change and passed as `no-changed-source`: Git printed the changed paths from its top level. They are now read from the directory holding `jevgate.toml`. +- **Gateway keys.** An OpenRouter or Vercel AI Gateway key works as a TypeSafe key does; both gateways serve TypeSafe's API and the same model at the same price. `jevgate auth login` asks which kind of key it is (`--provider typesafe|openrouter|vercel` answers it for scripts) and saves the key with its provider. A check reads `TYPESAFE_API_KEY` from the environment, then `--env-file` (`TYPESAFE_API_KEY`, `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY`; the repository's `.env` is read only for `TYPESAFE_API_KEY`, since a gateway's key there is usually the application's own), then the saved key, and only then `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY` from the environment: other tools read those two, and one exported for them must not move a check to another account's credits, another data processor and a model name no cached answer was asked with. A key saved before 0.26 has no provider recorded beside it, so while a gateway's variable is set a check reads the credential store to find it, which an interactive run may ask permission for. With the JevGate action, leave `api-key` out and set the gateway's variable in the step's `env`; jevgate-action 1.2 adds `api-key-kind: openrouter` or `vercel` for it. `jevgate auth status` shows the provider, where requests go and the keys set but not used, and the headline says `via OpenRouter` when a gateway answers. The default model follows the key: `typesafe/jev-1.13` on OpenRouter (the 1.13 line; `~typesafe/jev-latest` would move to new major versions) and `typesafe-ai/jev` on Vercel, both aliases whose answers expire after `cache_ttl_secs`; TypeSafe's stays `jev-1.13.0`, so a TypeSafe key's cached answers still count. A key goes only to its own provider: nothing in `jevgate.toml` or `.env` chooses the host, a key that starts as another provider's keys do (`sk-or-`, `vck_`) is refused, and `JEVGATE_BASE_URL`, read only from the environment, points requests at a self-hosted proxy over https (or http on this machine). A gateway's key is saved as its provider's name and the key, which versions before 0.26 refuse rather than send to TypeSafe. Tested against a local server in each gateway's shape, and through OpenRouter on 2026-09-28 (key check, answers, usage, price and request ids as expected); Vercel AI Gateway has not been tried with a key. +- Models: a model name without an `x.y.z` version is an alias, whose cached answers expire after `cache_ttl_secs`, and gateways' names are accepted (`typesafe/jev-1.13`, `~typesafe/jev-latest`, `typesafe-ai/jev`). Only `jev-latest` and `jev-preview` were aliases: any other name was taken for a pinned version, cached forever and held to answering with exactly that name, and a `/` or `~` in the answering model failed every answer. A pinned name still accepts only its own version, with or without a gateway's namespace. The default `jev-1.13.0` keeps its cached answers. +- Cost: a run is priced by the model that answered each request, not by the name it asked for. A `jev-latest` run showed no cost, although TypeSafe answers it with `jev-1.13.0`. The 1.13 line is priced under a gateway's namespace and as a dated snapshot too: OpenRouter lists `typesafe/jev-1.13` as the endpoint `typesafe/jev-1.13-20260917` at the same $0.042 per million input tokens. Vercel AI Gateway's `typesafe-ai/jev` names no version, so a run answered under that name shows its cost as unknown. A response without `usage`, which a gateway need not send, is accepted, and the run's cost is shown as unknown rather than $0, in the headline and in the HTML report, which says why; its tokens stay out of the bytes-per-token calibration. The JSON report adds `paid_models` (input tokens by the model that answered them), `unmetered_requests` and `estimated_usd` (null when unknown). +- Provider errors: an error names the provider's request id when it sent one (`x-typesafe-request-id`, else the response's own `id`), and each judgment in the report and each cached answer keeps the id of the request that answered it, to quote to the provider's support. A 402 says the credits are exhausted and where to add them. A 422 names each invalid field and its error type (`body.questions.q1.criteria missing`), never the provider's message or input, which can echo the source. A model the provider does not know says so: TypeSafe answers `jev-1.13`, a name its docs use, with an HTTP 400 "unknown model" (checked on 2026-09-28), and a gateway may answer 404. A 413, a gateway's refusal of an oversized request, counts as beyond the model's context like TypeSafe's own `max_tokens_exceeded`. A retry waits as long as `retry-after-ms` asks, else `Retry-After` in seconds or as an HTTP date; JevGate read whole seconds only. The pause is still at most 30 seconds. +- An incomplete run says why in the agent text: after the findings, each reason files failed, with how many gave it (`Failed 1: TypeSafe HTTP 402 (credits exhausted; …)`), as skipped files already were. The reasons were only in the JSON report and with `--verbose`, so the MCP server's `jevgate_check`, which returns the agent text, could not tell exhausted credits from a missing key. +- Pacing: requests start at least 50 ms apart, TypeSafe's documented limit of 1,200 a minute, and at most 6 are sent at once with a TypeSafe key, the default, and 3 with an OpenRouter or Vercel AI Gateway key: TypeSafe's limit is for one account, and through a gateway that account is the gateway's, shared with its other customers. That is a precaution rather than a measured fix: on 2026-09-28 TypeSafe's own endpoint answered 503 as often as OpenRouter did (65% of attempts, against 63%), and OpenRouter's rounds of a few requests at once fared only a little better (50% of attempts answered 503, against 65% in rounds of up to six). `--concurrency` and `concurrency` in `jevgate.toml` set it for any key, the file's value capping the flag's, and the report's `concurrency` says what ran. A higher `--concurrency`, which 0.25.0 accepted up to 8, is lowered to 6 with a notice on stderr, and a `concurrency` of 7 or 8 in `jevgate.toml`, which 0.25.0 accepted too, means 6. Six workers made 18 to 20 requests a second on the corpus's largest runs (0.3 s a request), so pacing leaves them as fast; eight would make about 27. Each attempt times out after 20 seconds instead of 60 (TypeSafe's SDKs wait 10), and a timed-out request is still sent once more. +- Retries: an answer worth retrying (a rate limit, overload, or a server or gateway error) is sent up to 6 times instead of 4, pausing 1, 2, 4, 8 and 8 seconds, each up to a quarter longer (the first three as before), or longer when the provider asks, up to 30 seconds as before. On 2026-09-28 TypeSafe answered 503 to about two attempts in three for at least ten minutes, directly (22 of 34 attempts) and through OpenRouter (141 of 224), each failed attempt taking about 10 seconds: with 4 attempts, a self-check with a TypeSafe key and all four canaries through OpenRouter ended incomplete, 1 of 13 and 16 of 99 requests having given up. Simulating the queue, 6 attempts leave 6% of requests unanswered in that brownout instead of 16%; when 1 attempt in 5 fails, a 1,000-request run completes 95% of the time instead of 21%; and a provider failing every attempt is outlasted for 26 seconds instead of 8. They cost time only while attempts fail: a hard outage takes up to three times as long to end a run incomplete, 8 minutes instead of 3 for 100 requests. A timed-out request is still sent twice at most, and one whose connection failed before sending 4 times. +- The test suite leaves alone the repository it runs in. Git exports `GIT_DIR` to a hook, `git rebase --exec` or `git bisect run` in a linked worktree, and `GIT_INDEX_FILE` to a pre-commit hook; `cargo test` run there made the tests' `git init`, `add` and `commit` act on that repository instead of their temporary projects, which set `core.bare = true` in a clone's shared configuration and committed a test's files into a worktree. The tests' Git, and a check's in the unit tests, now runs without the variables that point Git at a repository, and `tests/lint_policy.rs` rejects starting Git anywhere else. A check still honors them: `jevgate check --base HEAD` in a pre-commit hook reads the index being committed, and a repository kept apart from its work tree is found through `GIT_DIR`. + ## [0.25.0] - 2026-09-27 Fixes from running JevGate on widely used projects under daily development (rtk, headroom, paperclip, hermes-agent, cc-switch, freellmapi, herdr, multica, OmniRoute, dify, openclaw, n8n), each checked against the code. diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index 84fa444..3d81508 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -24,6 +24,7 @@ The tests run offline and need no API key. `jevgate check --dry-run --show-reque ## Code conventions - Unused code is deleted, not silenced, and long parameter lists are grouped into a type. `tests/lint_policy.rs` rejects `allow` or `expect` for `dead_code`, `unused`, `too_many_arguments` and `complexity`. Any other exception uses `#[expect(lint, reason = "…")]`. +- Tests run Git through `tests/support/git.rs`, which drops the variables that point Git at a repository (`GIT_DIR`, `GIT_INDEX_FILE` and the others Git exports to hooks and `git rebase --exec`), so `cargo test` run there leaves that repository alone. `tests/lint_policy.rs` rejects starting Git anywhere but there and `src/revision.rs`. - Commit subjects say what changed, in the imperative ("Report overlapping tests by groups"). - Messages and findings are plain sentences that name the code and say what to do next. diff --git a/README.md b/README.md index 202b125..b1bc24e 100644 --- a/README.md +++ b/README.md @@ -7,11 +7,11 @@ **JevGate is a code-review gate. It asks small, precise questions about your code and turns the answers into findings you can act on.** -JevGate parses your repository locally and builds small units of evidence: a function, a file outline, a pair of copies, a test, a documentation section. It asks [TypeSafe Jev](https://docs.typesafe.ai) short, typed questions about each one. Code, not a chat model, combines the answers into a verdict. Each finding has a location, a probability and a concrete next step, so an agent or CI job can act on it and a person can check it quickly. +JevGate parses your repository locally and builds small units of evidence: a function, a file outline, a pair of copies, a test, a documentation section. It asks [TypeSafe Jev](https://docs.typesafe.ai) short, typed questions about each one. Code, not a chat model, combines the answers into a verdict. Each finding has a location, a probability and a concrete next step, so an agent or CI job can act on it and a person can check it quickly. By default the gate fails only on the rules and levels measured right at least 80% of the time on projects JevGate was never tuned on: among the default rules, function-simplification reviews, right 20 of the 23 times they were labeled there (87%). -![JevGate's terminal output on zoxide: the gate fails on 3 review findings (a function mixing separate jobs, and two sets of importers repeating the same steps), 3 consider findings (a group of file helpers that could be a module, branching that hides a main path, an unexplained 7) and optional notes on unnamed values](site/src/images/terminal.svg) +![JevGate's terminal output on zoxide: the gate fails on 1 function-simplification review (a function mixing separate jobs); 3 more reviews (a file holding several features, two sets of importers repeating the same steps) and a consider (branching that hides a main path) are reported without failing it, since their rules and levels are still being measured](site/src/images/terminal.svg) -JevGate 0.22.0 on [zoxide](https://github.com/ajeetdsouza/zoxide/tree/09a18b4424b3f1033094ffd97da6d47585e38259), rerun from its answer cache, so it cost nothing; 9 of the 11 notes are left out. +JevGate 0.26.0 on [zoxide](https://github.com/ajeetdsouza/zoxide/tree/09a18b4424b3f1033094ffd97da6d47585e38259), rerun from its answer cache, so it cost nothing. **[Documentation](https://tech-byte-frontier.github.io/jevgate/)** · [Rules](https://tech-byte-frontier.github.io/jevgate/reference/rules.html) · [Configuration](https://tech-byte-frontier.github.io/jevgate/configuration.html) · [CI](https://tech-byte-frontier.github.io/jevgate/ci.html) · [Troubleshooting](https://tech-byte-frontier.github.io/jevgate/troubleshooting.html) · [Changelog](CHANGELOG.md) @@ -19,7 +19,7 @@ JevGate 0.22.0 on [zoxide](https://github.com/ajeetdsouza/zoxide/tree/09a18b4424 | Group | Rules | On | |---|---|---| -| Maintainability | File organization, function simplification, shared logic, hardcoded values | by default | +| Maintainability | File organization, function simplification, shared logic; hardcoded values (opt-in) | by default | | Tests | Test value (mock-only checks, expected values recomputed with the code's own logic), test redundancy | with `--include-tests` | | Security | Injection, sensitive data, unsafe settings, SQL access control, GitHub workflows; each finding names a CWE | `--rule security` | | Documentation | Agent instruction files, large and stale docs, duplicated sections, code comments | `--rule documentation` | @@ -35,13 +35,13 @@ cargo binstall jevgate # any platform, with cargo-binstall cargo install jevgate --locked # build from source; needs Rust 1.90 or later ``` -Releases have binaries for Linux, macOS and Windows with checksums and build provenance. Reviewing needs a [TypeSafe API key](https://console.typesafe.ai/settings/keys). [Install](https://tech-byte-frontier.github.io/jevgate/install.html) covers verifying a download, shell completions and man pages. +Releases have binaries for Linux, macOS and Windows with checksums and build provenance. Reviewing needs an API key from [TypeSafe](https://console.typesafe.ai/settings/keys) or [OpenRouter](https://openrouter.ai/settings/keys), which serve the same model at the same price; a [Vercel AI Gateway](https://vercel.com/docs/ai-gateway/authentication-and-byok/api-keys) key is accepted too, but has not been tried with a real key yet. [Install](https://tech-byte-frontier.github.io/jevgate/install.html) covers verifying a download, shell completions and man pages. ## Quick start ```sh jevgate init # write a commented jevgate.toml for this repository -jevgate auth login # validate and save your TypeSafe API key +jevgate auth login # validate and save your API key: TypeSafe, OpenRouter or Vercel jevgate check --dry-run --show-requests # see exactly what would be uploaded; free and offline jevgate check --report # review, then open a local HTML dashboard jevgate baseline # accept today's findings; later checks fail only on new ones @@ -49,7 +49,7 @@ jevgate baseline # accept today's findings; later check `jevgate check --report` writes the same findings to a local dashboard you can filter by path and classification, with each file's findings, undecided units and the answers behind them: -![JevGate's HTML report on zoxide: totals for files, review and consider findings, notes and cost, then a list of files by classification, with src/util.rs open to show its review and consider findings, their next steps and one undecided unit](site/src/images/report.png) +![JevGate's HTML report on zoxide: the gate's result and what fails it by default, totals for files, findings, notes and cost, then a list of files by classification, with src/util.rs open to show its two review findings, the one that fails the gate marked, and how each rule classified the file](site/src/images/report.png) `jevgate --help` gives the workflow, exit codes, files and environment, and `jevgate check --help` explains each flag and the JSON report. Coding agents can also call JevGate as a tool through its MCP server, `jevgate mcp` ([coding agents](https://tech-byte-frontier.github.io/jevgate/coding-agents.html)). @@ -75,7 +75,7 @@ jobs: version: 0.25.0 ``` -It reviews only the changed files, annotates each finding on its line and writes a job summary; unchanged code is answered from the cache for free. [Continuous integration](https://tech-byte-frontier.github.io/jevgate/ci.html) covers pre-commit, other CI systems, pull requests from forks, budgets and a gate policy the change cannot edit. +It reviews only what the pull request changed (the functions, tests and comments on changed lines, and copies where either copy changed), annotates each finding on its line and writes a job summary; unchanged code is answered from the cache for free. [Continuous integration](https://tech-byte-frontier.github.io/jevgate/ci.html) covers pre-commit, other CI systems, pull requests from forks, budgets and a gate policy the change cannot edit. ## Output and exit codes @@ -87,13 +87,13 @@ It reviews only the changed files, annotates each finding on its line and writes | 1 | Gate failed | | 2 | Run incomplete, invalid configuration or invalid usage | -Findings are `review` (act on it), `consider` (worth a look) or `note` (optional). A file whose answers stay undecided is `uncertain`, never hidden or counted as clear. `--fail-on` and `jevgate.toml` set what fails the gate, per rule and per path. `jevgate baseline` accepts today's findings, and a `jevgate: allow(RULE) reason` comment accepts one where it is. +Findings are `review` (act on it), `consider` (worth a look) or `note` (optional). A file whose answers stay undecided is `uncertain`, never hidden or counted as clear. By default only the rules and levels measured right at least 80% of the time on projects JevGate was never tuned on fail the gate (`jevgate rules` shows them); the other findings are reported without failing it. `--fail-on` and `jevgate.toml` set what fails the gate, per rule and per path. `jevgate baseline` accepts today's findings, and a `jevgate: allow(RULE) reason` comment accepts one where it is. ## Privacy and cost - **What is uploaded:** only the selected units of source, bounded by `upload_allow` and `upload_deny`; `--dry-run --show-requests` prints every request body offline. - **Secrets:** out of scope on purpose, because judging secrets would mean uploading them. Use a local secret scanner. -- **Cost:** every run prints its input tokens and an estimated cost, and cached answers cost nothing. +- **Cost:** every run prints its input tokens and an estimated cost, and cached answers cost nothing, so a rerun of unchanged code sends no request. Jev 1.13 costs $0.042 per million input tokens: checked as pull requests with every rule and nothing cached, the last commits of 118 corpus projects sent 4.55M first-pass input tokens, $0.19 for all 118. [Privacy and cost](https://tech-byte-frontier.github.io/jevgate/privacy-and-cost.html) and [limits](https://tech-byte-frontier.github.io/jevgate/limits.html) say more, and [how it works](https://tech-byte-frontier.github.io/jevgate/how-it-works.html) explains the evidence units and how code turns answers into findings. diff --git a/docs/classification-cascade.md b/docs/classification-cascade.md index abd310e..82381dc 100644 --- a/docs/classification-cascade.md +++ b/docs/classification-cascade.md @@ -132,6 +132,22 @@ signatures, or one candidate pair. a full pass but saved a half to two thirds as much per edit, and left whole files in one run (`clones.rs`, `literals.rs`); one in four overtakes it after 31 to 40 edits. + With `--base`, only what the change touched is asked: units whose lines + it added or modified, or removed lines inside; copies where either copy + changed; a file's outline, or a large document's, only when the change + adds a member or heading its base version lacks; and a document it left + alone only in a section that names a path it deleted or renamed. The + touched functions of one run share a pack, in runs that end where they + end for the whole file, so no other pack is sent. A later push that + changes another function of the run adds it to that pack, which is asked + again whole: on 16 corpus projects whose last two commits edit the same + file, the second push re-asked 93 units the first had asked, in 45 of + its 575 new packs and 1% of the bytes it sent (judging whole files, 363 + units in 129 of 759 packs, 3%). A unit asked beside other functions can + answer differently: of 11,693 first-pass answers about the same units on + the corpus's last commits, 88% were the same as with whole-file packs, + the others moved 0.03 on average, and 36 crossed 0.50 or 0.80 (13 up, 23 + down), which made three function-simplification considers notes. Tests are sent one per request, because unrelated tests in the same state left more answers undecided. State uses literal paths such as `functions[2].source`; group IDs are Choice options. Stage and freshness @@ -692,8 +708,13 @@ signatures, or one candidate pair. rather than splitting it, so a consider left naming no group is a note and a review says to split the whole file. 6. **Gate.** `--fail-on`, `[[scope]]` levels per path and the baseline act on - composed findings only. Baseline entries can carry a reason (`intended`, - `later`, `wrong`) that survives rewrites; `baseline stats` counts them. + composed findings only. The default level, `mature`, fails only on the + rules and levels whose findings were right at least 80% of the time on + projects never used for tuning, over at least 20 hand labels + (`maturity::TABLE`); a probability says how sure an answer is, not how + often such findings are right. Baseline entries can carry a reason + (`intended`, `later`, `wrong`) that survives rewrites; `baseline stats` + counts them. ## Constraints diff --git a/jevgate-baseline.json b/jevgate-baseline.json index 54dc997..d20369b 100644 --- a/jevgate-baseline.json +++ b/jevgate-baseline.json @@ -1,5 +1,15 @@ { "version": 1, - "created_at": 1790514454, - "findings": [] + "created_at": 1790587702, + "findings": [ + { + "fingerprint": "62ce38c353079b7b9773080b905d226e8d8e473ef040d8b1b828e25cca8fc4f1", + "rule": "maintainability/hardcoded-values", + "path": "src/units/tests/mod.rs", + "line": 312, + "strength": "review", + "message": "One of this file's constants fixes a value that differs between deployments (0.98). The constant is `HARDCODED`.", + "reason": "wrong" + } + ] } diff --git a/jevgate.schema.json b/jevgate.schema.json index 9042d77..3ab6041 100644 --- a/jevgate.schema.json +++ b/jevgate.schema.json @@ -6,6 +6,7 @@ "enum": [ "review", "consider", + "mature", "uncertain", "report", "none", @@ -18,6 +19,7 @@ "enum": [ "review", "consider", + "mature", "uncertain", "report", "none", @@ -168,6 +170,7 @@ "enum": [ "review", "consider", + "mature", "uncertain", "report", "none" @@ -260,13 +263,13 @@ "description": "JevGate configuration. The command line wins over the file, except that upload patterns and budgets in the file are ceilings that flags can only narrow. Unknown keys are errors.", "properties": { "cache_ttl_secs": { - "description": "Cache lifetime in seconds for the `jev-latest` and `jev-preview` aliases; pinned versions never expire. Default: 3600.", + "description": "Cache lifetime in seconds for an alias, a model name without an x.y.z version such as `jev-latest`; pinned versions never expire. Default: 3600.", "minimum": 0, "type": "integer" }, "concurrency": { - "description": "Ceiling on simultaneous requests (1-8). Default: 6.", - "maximum": 8, + "description": "Most simultaneous requests; flags can only lower it. JevGate sends at most 6 at once, so a higher value means 6. Default: 6 with a TypeSafe key, 3 with an OpenRouter or Vercel AI Gateway key.", + "maximum": 6, "minimum": 1, "type": "integer" }, @@ -278,11 +281,12 @@ "type": "array" }, "fail_on": { - "description": "The level for rules without their own, like `--fail-on`. Default: [\"review\"].", + "description": "The level for rules without their own, like `--fail-on`. Default: [\"mature\"], which fails only on the levels of a rule measured right at least 80% of the time on projects JevGate was never tuned on; `jevgate rules` shows them.", "items": { "enum": [ "review", "consider", + "mature", "uncertain", "report", "none" @@ -318,7 +322,7 @@ "type": "integer" }, "model": { - "description": "TypeSafe model; a pinned version keeps results repeatable. `--model` overrides it.", + "description": "Model, as the key's provider names it; a pinned version keeps results repeatable. `--model` overrides it. Default: `jev-1.13.0` with a TypeSafe key, `typesafe/jev-1.13` with an OpenRouter key, `typesafe-ai/jev` with a Vercel AI Gateway key.", "type": "string" }, "rules": { diff --git a/site/generate.py b/site/generate.py index 5d47ade..2aa7515 100644 --- a/site/generate.py +++ b/site/generate.py @@ -15,7 +15,7 @@ COMMANDS = ["auth", "check", "baseline", "rules", "init", "completions", "man", "serve", "mcp"] GROUPS = { - "maintainability": "On by default.", + "maintainability": "On by default, except hardcoded values: add it with `--rule default --rule hardcoded-values`, or a level for it in `[rules]`.", "tests": "On by default. Test value and test redundancy are judged with `--include-tests` or `include_tests = true`; the laws of Bend 2 code are judged without it.", "security": "Opt-in: `--rule security`, or a level in `[rules]`.", "documentation": "Opt-in: `--rule documentation`, or a level in `[rules]`.", @@ -36,14 +36,24 @@ def rules_page(binary): "its key or its group anywhere a rule is accepted: `--rule`, `--skip-rule`,", "`--fail-on TARGET=LEVEL`, `[rules]`, `[[scope]]` and `jevgate: allow(…)` comments.", "", - "| Rule | Key | Default | Question |", - "|---|---|---|---|", + "*Fails by default* names the levels that fail the check under the default gate level,", + "`mature`: those right at least 80% of the time over at least 20 findings labeled by hand on", + "projects JevGate was never tuned on; an opt-in rule's levels fail it once the rule is selected.", + "The other findings are reported without failing it.", + "*Reviews right* and *considers right* give the share of labeled findings that were right on", + "those projects, and how many were labeled; a debatable one counts as not right.", + "`tests/laws` is labeled only on Bend 2 projects, which these numbers leave out.", + "", + "| Rule | Key | Default | Fails by default | Reviews right | Considers right | Question |", + "|---|---|---|---|---|---|---|", ] for rule in rules: anchor = rule["id"].replace("/", "-") default = "yes" if rule["default_enabled"] else "opt-in" + maturity = rule["maturity"] lines.append( - f"| [`{rule['id']}`](#{anchor}) | `{rule['key']}` | {default} | {cell(rule['inspection'])} |" + f"| [`{rule['id']}`](#{anchor}) | `{rule['key']}` | {default} | {blocks(maturity)}" + f" | {right(maturity, 'review')} | {right(maturity, 'consider')} | {cell(rule['inspection'])} |" ) group = None for rule in rules: @@ -78,6 +88,20 @@ def rules_page(binary): return "\n".join(lines) + "\n" +def blocks(maturity): + """The levels that fail the default gate, or "no".""" + return ", ".join(level for level in ("review", "consider") if maturity.get(level, {}).get("mature")) or "no" + + +def right(maturity, level): + """"87% of 23": labeled findings of a level right on unseen projects, or "-".""" + unseen = maturity.get(level, {}).get("unseen") + if not unseen or not unseen["labeled"]: + return "-" + percent = (200 * unseen["right"] + unseen["labeled"]) // (2 * unseen["labeled"]) # half up, as JevGate rounds + return f"{percent}% of {unseen['labeled']}" + + def configuration_page(schema_path): schema = json.loads(Path(schema_path).read_text()) lines = [ diff --git a/site/src/ci.md b/site/src/ci.md index 90e59b7..82f43b2 100644 --- a/site/src/ci.md +++ b/site/src/ci.md @@ -20,13 +20,13 @@ jobs: version: 0.25.0 ``` -The action installs a checked release binary, keeps `.jevgate/cache` in the Actions cache and runs `jevgate check --base --format github`; `args` passes more flags, such as `--rule security`. It runs on Linux, macOS and Windows runners. +The action installs a checked release binary, keeps `.jevgate/cache` in the Actions cache and runs `jevgate check --base --format github`; `args` passes more flags, such as `--rule default --rule security` to add the security rules to the default ones. Naming a rule replaces the selection, and a selection without a mature rule level never fails the default gate: `--rule security` alone reports security findings without ever failing the gate. It runs on Linux, macOS and Windows runners. For an OpenRouter or Vercel AI Gateway key, leave `api-key` out and set the key's variable in the step's `env`, such as `OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}`, with `version` 0.26.0 or later; jevgate-action 1.2 adds `api-key-kind: openrouter` or `api-key-kind: vercel` for the same. -`--format github` annotates the changed lines with each finding. A finding that fails the gate is an error; the others are warnings. A Markdown table goes to the job summary, and the usual text goes to the log. The full JSON report is always at `.jevgate/latest.json` if you want to keep it as an artifact. +`--format github` annotates the changed lines with each finding. A finding that fails the gate is an error; the others are warnings, and a warning whose rule and level are still being measured says so, with how often such findings were right. A Markdown table goes to the job summary, and the usual text goes to the log. The full JSON report is always at `.jevgate/latest.json` if you want to keep it as an artifact. -- **Changed files only:** `--base` reviews what changed since the fork point with that revision, the same files a pull request diff shows, plus uncommitted and untracked files. It needs the history, so check out with `fetch-depth: 0`. When no supported file changed, the run passes without any request. -- **Cache:** answers are stored under a hash of the exact request: source, questions and model. Restoring an older cache is always safe, and unchanged code costs nothing on the next run. -- **Advisory or blocking:** `fail_on = ["none"]` in `jevgate.toml` or `--fail-on none` reports findings without failing. A run that could not finish (missing key, provider rejection, request budget reached) still exits 2, so an outage never passes as a clean review. +- **Only what changed:** `--base` reviews what changed since the fork point with that revision, as a pull request diff shows it, plus uncommitted and untracked files. It asks about and reports only what the change touches: functions, tests, comments and values on changed lines, copies where either copy changed, a file's outline when the change adds members to it, and a document when a section names a path the change deleted or renamed. A new file is judged whole. On the last commits of 118 corpus projects (90 open-source, 28 of the maintainer's own private repositories), this took the findings off the changed lines from 54% to 6% and halved the first-pass requests. `--whole-files` judges every unit of each changed file instead, as `--base` did before 0.26. It needs the history, so check out with `fetch-depth: 0`. When no supported file changed, the run passes without any request. +- **Cache:** answers are stored under a hash of the exact request: source, questions and model. Restoring an older cache is always safe, and unchanged code costs nothing on the next run. A gateway's model names are aliases, so with an OpenRouter or Vercel AI Gateway key answers expire after `cache_ttl_secs` (an hour by default); raise it to reuse answers across runs further apart, at the price of noticing a new model version later. +- **Advisory or blocking:** by default only the rules and levels measured right at least 80% of the time on projects JevGate was never tuned on fail the check ([what fails by default](configuration.md#what-fails-the-check-by-default)); the other findings are warnings. On the last commits of those 118 projects, a pull request check with the defaults fails 9 of them, all on function-simplification reviews, where 0.25.0 failed 18; 8 of the 9 are the maintainer's own repositories. `fail_on = ["review"]` in `jevgate.toml` or `--fail-on review` fails on every review, and `fail_on = ["none"]` or `--fail-on none` reports findings without failing. A run that could not finish (missing key, provider rejection, request budget reached) still exits 2, so an outage never passes as a clean review. - **A policy the change cannot edit:** a pull request can edit `jevgate.toml`. To apply the reviewed policy of the base branch instead, read it with `--config`: ```sh @@ -36,7 +36,7 @@ The action installs a checked release binary, keeps `.jevgate/cache` in the Acti - **Forks:** GitHub withholds secrets from pull requests opened from forks, so there the run exits 2 with "No API key configured". Skip the job for forks, or run it only on branches of the repository. - **Budgets:** `max_requests` caps the API attempts of one run. Reaching it leaves the run incomplete instead of passing on partial evidence. `--dry-run` counts the planned requests the cache already answers, so its estimate covers only what the cache lacks; follow-ups depend on answers and are not counted. -- **Transient failures:** rate limits, overload and server or edge errors (HTTP 408, 429, 500, 502–504, 520–524, 529) are retried up to four attempts; a timeout or dropped connection is retried once, since the first send may have run. +- **Transient failures:** rate limits, overload and server or edge errors (HTTP 408, 429, 500, 502–504, 520–524, 529) are retried up to six attempts, with pauses of 1 to 8 seconds; an attempt that has not answered in 20 seconds, or whose connection drops, is retried once, since the first send may have run. A provider that fails every attempt ends a run of 100 requests incomplete after 8 minutes or more (16 with a gateway's key, which sends 3 requests at once). - **Report-only paths:** give tooling its own level with `[[scope]]` (below), so scripts are reported while product code gates. Before each commit, with [pre-commit](https://pre-commit.com), review what is staged: @@ -49,7 +49,7 @@ repos: - id: jevgate-system # the jevgate on PATH; `jevgate` builds it with Rust instead ``` -On GitLab, a merge request pipeline can show the findings in the merge request with a Code Quality report. Set `TYPESAFE_API_KEY` as a masked CI/CD variable: +On GitLab, a merge request pipeline can show the findings in the merge request with a Code Quality report. Set `TYPESAFE_API_KEY` as a masked CI/CD variable (or `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY` for a gateway's key): ```yaml jevgate: @@ -70,4 +70,4 @@ jevgate: - if: $CI_PIPELINE_SOURCE == "merge_request_event" ``` -Other CI systems work the same way: install with `install.sh` or `cargo binstall`, set `TYPESAFE_API_KEY`, keep `.jevgate/cache` between runs, and read the exit code or the JSON report. +Other CI systems work the same way: install with `install.sh` or `cargo binstall`, set `TYPESAFE_API_KEY` (or a gateway's variable), keep `.jevgate/cache` between runs, and read the exit code or the JSON report. diff --git a/site/src/coding-agents.md b/site/src/coding-agents.md index fefcddd..ddd5bb4 100644 --- a/site/src/coding-agents.md +++ b/site/src/coding-agents.md @@ -7,15 +7,16 @@ JevGate's default output is written for coding agents as much as for people: ran Ask the agent to review its own change before it reports back, for example in `AGENTS.md` or `CLAUDE.md`: ```markdown -Before finishing, run `jevgate check --base origin/main`. Fix each `review` finding; -for a `consider`, fix it or say why the code should stay as it is. +Before finishing, run `jevgate check --base origin/main`. Fix each finding marked +"fails the gate". Weigh the other `review` and `consider` findings: fix one when it +is right, or say why the code should stay as it is. ``` -`--base` limits the review to the files changed since that revision, plus uncommitted and untracked files, so a check costs only what the change touches, and cached answers make reruns free. The exit code says what to do next: +`--base` limits the review to what changed since that revision, uncommitted and untracked changes included: the functions, tests and comments on changed lines, and copies where either copy changed. A check asks about and reports only what the change touches, and cached answers make reruns free. The exit code says what to do next: | Exit code | Meaning for the agent | |---|---| -| 0 | The gate passed; `consider` findings may still be worth a look | +| 0 | The gate passed; `consider` findings, and reviews from rules still being measured, may still be worth fixing | | 1 | The gate failed: act on the findings listed | | 2 | The run could not finish (no key, provider rejection, request budget); report it, don't treat it as a pass | @@ -35,7 +36,7 @@ The second form is for clients configured with JSON, such as Cursor. The server | Tool | What it does | |---|---| -| `jevgate_check` | Runs `jevgate check` in the repository with `base`, `paths`, `rules`, `include_tests`, `dry_run` or `verbose`, and returns the ranked findings. An incomplete run (exit 2) is a tool error, never a pass | +| `jevgate_check` | Runs `jevgate check` in the repository with `base`, `whole_files`, `paths`, `rules`, `include_tests`, `dry_run` or `verbose`, and returns the ranked findings. An incomplete run (exit 2) is a tool error, never a pass | | `jevgate_findings` | Reads the last report's findings, optionally under one path, without running anything | | `jevgate_rules` | Lists every rule with the question it asks | @@ -60,4 +61,4 @@ A check runs as a child process with the repository's `jevgate.toml` and key, so ## Documentation for agents -The opt-in documentation rules judge the instruction files agents load at the start of every session (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, and Cursor, Copilot, Windsurf, Cline, Kiro, Junie and Roo Code rules): sections that only restate the manifest or generic advice, and text loaded in every session that applies to one directory. `jevgate check --rule documentation` also estimates the tokens each harness loads. +The opt-in documentation rules judge the instruction files agents load at the start of every session (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, and Cursor, Copilot, Windsurf, Cline, Kiro, Junie and Roo Code rules): sections that only restate the manifest or generic advice, and text loaded in every session that applies to one directory. Their considers on instruction files fail the check by default, the one consider level that does (22 of 24 were right on projects JevGate was never tuned on). `jevgate check --rule documentation` also estimates the tokens each harness loads. diff --git a/site/src/configuration.md b/site/src/configuration.md index e9358c3..2640f46 100644 --- a/site/src/configuration.md +++ b/site/src/configuration.md @@ -9,9 +9,9 @@ include_tests = true max_requests = 300 [rules] # a level per group or rule -maintainability = "review" # judge, and fail the gate on review findings +maintainability = "review" # judge every rule of the group, and fail on its reviews tests = "consider" -security = "consider" # opt-in group, enabled by naming it +security = "mature" # opt-in group, enabled by naming it; fails only on levels measured mature "maintainability/hardcoded-values" = "report" # judge but never fail; "off" skips it [[scope]] # levels for the files these paths match @@ -27,17 +27,29 @@ rules = { security = "consider" } # except these | `generated` | built-in names | Globs of generated files, which are skipped | | `tests` | built-in conventions | Globs of additional test files | | `context` | none | Files always sent as related evidence, like `--context` | -| `rules` | the `default` group | A list selects rules. A table gives each group or rule a level: `review`, `consider`, `uncertain`, `report` (judge, never fail) or `off` | +| `rules` | the `default` group | A list selects rules. A table gives each group or rule a level: `review`, `consider`, `mature`, `uncertain`, `report` (judge, never fail) or `off`; a level for a group judges every rule of it, opt-in ones included | | `[[scope]]` | none | `paths` (globs), with `fail_on` for every rule and `rules` for rules or groups, as above; `off` is not accepted (use `upload_deny`). The last scope that matches a file and addresses a rule wins; flags win over scopes | -| `fail_on` | `["review"]` | The level for rules without their own, like `--fail-on` | +| `fail_on` | `["mature"]` | The level for rules without their own, like `--fail-on` | | `include_tests` | `false` | Judge tests, like `--include-tests` | -| `model` | `jev-1.13.0` | TypeSafe model; a pinned version keeps results repeatable | -| `cache_ttl_secs` | `3600` | Cache lifetime for the `jev-latest` and `jev-preview` aliases; pinned versions never expire | +| `model` | the key's provider's | The model, as the key's provider names it: `jev-1.13.0` for TypeSafe, `typesafe/jev-1.13` for OpenRouter, `typesafe-ai/jev` for Vercel AI Gateway. A pinned version keeps results repeatable; a repository that sets it for one provider needs `--model` with another provider's key | +| `cache_ttl_secs` | `3600` | Cache lifetime for an alias: a model name without an `x.y.z` version, such as `jev-latest` or `jev-1.13`. Pinned versions such as `jev-1.13.0` never expire | | `max_requests` | unlimited | Ceiling on API attempts per invocation | -| `concurrency` | `6` | Ceiling on simultaneous requests (1–8) | +| `concurrency` | `6`, or `3` with a gateway's key | Most simultaneous requests; `--concurrency` can only lower it. JevGate sends at most 6 at once, so a higher value, which releases before 0.26 accepted up to 8, means 6, and `--concurrency` above 6 is lowered to 6 with a notice. Requests also start at least 50 ms apart, TypeSafe's limit of 1,200 a minute for an account. With an OpenRouter or Vercel AI Gateway key the default is 3: the gateway's account on TypeSafe is shared by its other customers | | `max_file_bytes` | `262144` | Files larger than this are reported as needs-context, never truncated; generated and vendored files are skipped instead | | `max_context_bytes` | `32768` | Ceiling on context bytes per request | Rules are named by ID (`maintainability/shared-logic`), key (`shared_logic`) or group (`maintainability`, `tests`, `security`, `documentation`, `default`, `all`). The same names work in `--rule`, `--skip-rule` and `--fail-on TARGET=LEVEL`, and the most specific entry wins. +## What fails the check by default + +The default level, `mature`, fails the check only on the rules and levels measured *mature*: their findings were right at least 80% of the time on projects JevGate was never tuned on, over at least 20 findings labeled by hand from the code. Today those are function-simplification reviews (20 of 23 right) and, when the documentation rules run, agent-context considers (22 of 24). Every other finding is reported and marked as still being measured, without failing the check. `jevgate rules` shows each rule's levels that fail by default and how often its reviews and considers were right. + +Any level you set replaces the default exactly as it says, for the rules and paths it addresses: `fail_on = ["review"]` (or `--fail-on review`) fails on every review, as releases before 0.26 did; `--fail-on security=consider` sets one group and leaves the others at `mature`; `mature` itself can be set, such as for one group after a stricter `fail_on`. A later release can mark more levels mature as labels accumulate, or fewer; set `fail_on` to keep a fixed policy. Undecided answers never fail the check under `mature`. + +A `jevgate.toml` written by `jevgate init` before 0.26 sets `maintainability = "review"` and `tests = "review"`: those lines keep every review of the two groups failing the check, and judge hardcoded values. Delete them for the default rules and gate. While they are there with the comments `init` wrote after them, every command that reads `jevgate.toml` says so on stderr; to keep the levels, delete the comments. + +## Keys and where requests go + +`jevgate.toml` has no key for the provider or its address: the change under review can edit that file, so it must not be able to send your key elsewhere. The provider follows the key ([Install](install.md) lists the three kinds), and a key goes only to its own provider. `JEVGATE_BASE_URL`, read only from the environment, replaces the provider's API root for a self-hosted proxy or a test server: `https://` to any host, or `http://` only to `localhost`, `127.0.0.1` or `[::1]`. Each check says on stderr when it is set, and `jevgate auth status` checks the key against `/v1/models` there. + The [configuration reference](reference/configuration.md) lists every key with its type, and the rule names and levels it accepts. diff --git a/site/src/how-it-works.md b/site/src/how-it-works.md index 5d85194..30bf0aa 100644 --- a/site/src/how-it-works.md +++ b/site/src/how-it-works.md @@ -150,6 +150,22 @@ signatures, or one candidate pair. a full pass but saved a half to two thirds as much per edit, and left whole files in one run (`clones.rs`, `literals.rs`); one in four overtakes it after 31 to 40 edits. + With `--base`, only what the change touched is asked: units whose lines + it added or modified, or removed lines inside; copies where either copy + changed; a file's outline, or a large document's, only when the change + adds a member or heading its base version lacks; and a document it left + alone only in a section that names a path it deleted or renamed. The + touched functions of one run share a pack, in runs that end where they + end for the whole file, so no other pack is sent. A later push that + changes another function of the run adds it to that pack, which is asked + again whole: on 16 corpus projects whose last two commits edit the same + file, the second push re-asked 93 units the first had asked, in 45 of + its 575 new packs and 1% of the bytes it sent (judging whole files, 363 + units in 129 of 759 packs, 3%). A unit asked beside other functions can + answer differently: of 11,693 first-pass answers about the same units on + the corpus's last commits, 88% were the same as with whole-file packs, + the others moved 0.03 on average, and 36 crossed 0.50 or 0.80 (13 up, 23 + down), which made three function-simplification considers notes. Tests are sent one per request, because unrelated tests in the same state left more answers undecided. State uses literal paths such as `functions[2].source`; group IDs are Choice options. Stage and freshness @@ -759,8 +775,13 @@ signatures, or one candidate pair. rather than splitting it, so a consider left naming no group is a note and a review says to split the whole file. 6. **Gate.** `--fail-on`, `[[scope]]` levels per path and the baseline act on - composed findings only. Baseline entries can carry a reason (`intended`, - `later`, `wrong`) that survives rewrites; `baseline stats` counts them. + composed findings only. The default level, `mature`, fails only on the + rules and levels whose findings were right at least 80% of the time on + projects never used for tuning, over at least 20 hand labels + (`maturity::TABLE`); a probability says how sure an answer is, not how + often such findings are right. Baseline entries can carry a reason + (`intended`, `later`, `wrong`) that survives rewrites; `baseline stats` + counts them. ### Constraints diff --git a/site/src/images/report.png b/site/src/images/report.png index c499e7c..5b4e40c 100644 Binary files a/site/src/images/report.png and b/site/src/images/report.png differ diff --git a/site/src/images/terminal.svg b/site/src/images/terminal.svg index a4ea8f4..6bce55a 100644 --- a/site/src/images/terminal.svg +++ b/site/src/images/terminal.svg @@ -1,40 +1,33 @@ - + - - + + jevgate check · zoxide -~/zoxide $ jevgate check --verbose -JevGate: review · gate failed: 3 new review findings · 27 files · 0 API requests · 0 input tokens · ~$0.0000 -Review (3): - src/util.rs:269 [maintainability/function-simplification] `resolve_path` mixes separate jobs in long blocks; - splitting it would make it easier to understand (0.86). +~/zoxide $ jevgate check +JevGate: review · gate failed: 1 new review finding · 27 files · 0 API requests · 0 input tokens · ~$0.0000 +Review (4): + src/util.rs:269 [maintainability/function-simplification] (fails the gate) `resolve_path` mixes separate + jobs in long blocks; splitting it would make it easier to understand (0.86). → Extract each separate job into its own named function - src/import/autojump.rs:64 [maintainability/shared-logic] `Iter::next` (src/import/autojump.rs:64) and - `Iter::next` (src/import/z.rs:66) perform the same steps for the same purpose (1.00). - → Move the shared steps into one implementation - src/import/fasd.rs:16 [maintainability/shared-logic] `Fasd::dirs` (src/import/fasd.rs:16), `ZshZ::dirs` - (src/import/zsh_z.rs:16) and 2 more copies perform the same steps for the same purpose (0.99). - → Move the shared steps into one implementation -Consider (3): - src/util.rs:154 [maintainability/file-organization] Some members of this file could move to a separate - module (0.97). G3 (`write`, `tmpfile`, `rename`, `canonicalize`, `current_dir`, `path_to_str` and 1 - more) or G1 (`Fzf`, `Fzf::new`, `Fzf::enable_preview`, `Fzf::args`, `Fzf::env`, `Fzf::envs`) would be - most useful as its own module. - → Consider moving that set of members into its own module + src/util.rs:154 [maintainability/file-organization] This file holds several features that would be easier + to find apart (0.81). G3 (`write`, `tmpfile`, `rename`, `canonicalize`, `current_dir`, `path_to_str` + and 1 more) or G1 (`Fzf`, `Fzf::new`, `Fzf::enable_preview`, `Fzf::args`, `Fzf::env`, `Fzf::envs`) + would be most useful as its own module. + → Move that group into its own module + src/import/autojump.rs:64 [maintainability/shared-logic] `Iter::next` (src/import/autojump.rs:64) and + `Iter::next` (src/import/z.rs:66) perform the same steps for the same purpose (1.00). + → Move the shared steps into one implementation + src/import/fasd.rs:16 [maintainability/shared-logic] `Fasd::dirs` (src/import/fasd.rs:16), `ZshZ::dirs` + (src/import/zsh_z.rs:16) and 2 more copies perform the same steps for the same purpose (0.99). + → Move the shared steps into one implementation +Consider (1): src/import/atuin.rs:75 [maintainability/function-simplification] `Iter::next` has branching that likely hides its main path (0.94). → Consider guard clauses, early returns or a lookup table - src/cmd/query.rs:32 [maintainability/hardcoded-values] `Query::query_interactive` uses a value whose meaning - a reader must guess (0.82). The value is 7. - → Give the value a descriptive constant name -Notes (11, optional): - src/db/mod.rs:121 [maintainability/hardcoded-values] `Database::age` likely uses a value whose meaning a - reader must guess. The value is 0.9. It is written once in its file, so it is a note. - → Optional: name or configure the value if it changes - src/db/dir.rs:59 [maintainability/hardcoded-values] `DirDisplay::fmt` likely uses a value whose meaning a - reader must guess. The value is 9999.0. It is written once in its file, so it is a note. - → Optional: name or configure the value if it changes - … 9 more notes -2 files with uncertain units. +2 optional notes on code that reads well as it is; --verbose shows them. +3 reviews and 1 consider did not fail the gate: by default only rules and levels right at least 80% of the + time on projects JevGate was never tuned on fail it, and theirs are still being measured. `jevgate + rules` shows each one's precision; `--fail-on review` makes every review fail the gate. +Skipped 1: Operational script. Maintainability gates apply to application and library source. diff --git a/site/src/install.md b/site/src/install.md index e6d01b9..375c84b 100644 --- a/site/src/install.md +++ b/site/src/install.md @@ -11,4 +11,14 @@ Each [release](https://github.com/Tech-Byte-Frontier/jevgate/releases) has binar `jevgate completions bash|zsh|fish|powershell` prints a shell completion script and `jevgate man` a man page; Homebrew installs both. -Reviewing needs a [TypeSafe API key](https://console.typesafe.ai/settings/keys). Git is needed only for `--base` and the staleness rule. +Reviewing needs an API key. Jev, the model JevGate asks, is served by TypeSafe and by two gateways, at the same price: + +| Key | Create one at | Environment variable | Default model | +|---|---|---|---| +| TypeSafe | [console.typesafe.ai](https://console.typesafe.ai/settings/keys) | `TYPESAFE_API_KEY` | `jev-1.13.0` | +| OpenRouter | [openrouter.ai](https://openrouter.ai/settings/keys) | `OPENROUTER_API_KEY` | `typesafe/jev-1.13` | +| Vercel AI Gateway | [the Vercel dashboard](https://vercel.com/docs/ai-gateway/authentication-and-byok/api-keys) | `AI_GATEWAY_API_KEY` | `typesafe-ai/jev` | + +`jevgate auth login` asks which kind of key it is and saves it with its provider; in CI, set the variable from a secret. A check uses the first key it finds: `TYPESAFE_API_KEY` in the environment, then `--env-file` or `TYPESAFE_API_KEY` in the repository's `.env`, then the saved key, then `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY` in the environment (an empty variable counts as unset). The gateways' variables come last because other tools read them too: one exported for another tool does not move JevGate off the key you gave it. `jevgate auth status` shows which key a check uses and the keys it leaves unused. Only TypeSafe offers a pinned version (`jev-1.13.0`): the gateways' names are aliases, whose cached answers expire after `cache_ttl_secs` (an hour by default). + +Git is needed only for `--base` and the staleness rule. diff --git a/site/src/introduction.md b/site/src/introduction.md index f4c58bf..bf15326 100644 --- a/site/src/introduction.md +++ b/site/src/introduction.md @@ -2,7 +2,7 @@ **JevGate is a code-review gate. It asks small, precise questions about your code and turns the answers into findings you can act on.** -JevGate parses your repository locally and builds small units of evidence: a function, a file outline, a pair of copies, a test, a documentation section. It asks [TypeSafe Jev](https://docs.typesafe.ai) short, typed questions about each one. Code, not a chat model, combines the answers into a verdict. Each finding has a location, a probability and a concrete next step, so an agent or CI job can act on it and a person can check it quickly. +JevGate parses your repository locally and builds small units of evidence: a function, a file outline, a pair of copies, a test, a documentation section. It asks [TypeSafe Jev](https://docs.typesafe.ai) short, typed questions about each one. Code, not a chat model, combines the answers into a verdict. Each finding has a location, a probability and a concrete next step, so an agent or CI job can act on it and a person can check it quickly. By default the gate fails only on the rules and levels measured right at least 80% of the time on projects JevGate was never tuned on: among the default rules, function-simplification reviews, right 20 of the 23 times they were labeled there (87%). ```text JevGate: consider · gate passed · 42 files · 118 API requests · 263410 input tokens · ~$0.0111 diff --git a/site/src/limits.md b/site/src/limits.md index 099b274..84bc76a 100644 --- a/site/src/limits.md +++ b/site/src/limits.md @@ -3,4 +3,5 @@ - **Languages and frameworks:** see [the support table](languages.md). Astro, Vue and Svelte markup is not read, only their scripts. PHP's inline HTML is read only for its `` echoes, and variables a page gets from the files it includes are not followed there: they are judged where those files set them. - **Security scope:** one function plus at most one hop of callers. This is not whole-program data-flow analysis. Access control reads the final state of policies, SECURITY DEFINER functions and grants across a project's SQL files in path order, leaving out uninstall, teardown, rollback and down scripts; with `--base`, unchanged migrations are read for that state but not judged. It does not judge application-level authorization or dynamic SQL inside database functions. - **Documentation scope:** staleness works only from the paths, scripts, tags and deletions that Git and the manifests show; it does not compare prose with code behavior. Paraphrases that share little wording are not found as duplicates, nor are code examples that share only code; a translation is not a duplicate. Sphinx and AsciiDoc includes are not followed, and MDX expressions are not evaluated. A code comment is judged with the code next to it, not against what the whole program does, so a comment that no longer matches its code is not found. Token counts are estimates at four bytes per token. +- **What a change touches:** with `--base`, a unit is judged when the change added or modified one of its lines, or removed lines between two of them. Lines removed just above or below a unit, such as a decorator over a function, do not count, since without the old file they belong as much to the code beside it. A change elsewhere that alters a unit's evidence, such as a function it calls, a caller or an error class, does not judge the unit again, and copies are looked for among the changed files, so code a change copies from a file it left alone is not found. `--whole-files` judges every unit of each changed file. - **Probabilities:** these are model judgments, not measured accuracy. JevGate complements linters, type checkers, tests and dedicated security scanners; it does not replace them. diff --git a/site/src/output.md b/site/src/output.md index 11e2382..67368d8 100644 --- a/site/src/output.md +++ b/site/src/output.md @@ -6,11 +6,13 @@ | `json` | The full report: every file, finding, raw answer and probability, gate and usage | | `jsonl` | One compact report per line; one per evaluation with `--watch` | | `github` | GitHub Actions annotations and job summary, then the agent text | -| `sarif` | A SARIF 2.1.0 log for [GitHub code scanning](https://docs.github.com/en/code-security/code-scanning/integrating-with-code-scanning/uploading-a-sarif-file-to-github) and other SARIF readers: the findings the annotations show, `error` when they fail the gate | +| `sarif` | A SARIF 2.1.0 log for [GitHub code scanning](https://docs.github.com/en/code-security/code-scanning/integrating-with-code-scanning/uploading-a-sarif-file-to-github) and other SARIF readers: the findings the annotations show, `error` when they fail the gate, with how the gate counted each as its `gate` property | | `gitlab` | A [GitLab Code Quality](https://docs.gitlab.com/ci/testing/code_quality/) report for merge requests: the same findings, `major` when they fail the gate and `minor` otherwise | Agent output is colored on a terminal; `--color never`, or `NO_COLOR` set to any value, turns it off, and `--color always` or `CLICOLOR_FORCE` turns it on for pipes and logs. +The first line, also the title of the GitHub job summary, sums up the run: its status, the gate and why it failed, the files, what a `--base` check judged (`changed lines since 1a2b3c4`, or `whole files changed since 1a2b3c4` with `--whole-files`; the report's `scope` records it), the API requests (`via OpenRouter` or `via Vercel AI Gateway` when a gateway answered), the input tokens and the estimated cost, which is `cost unknown` when a response reported no token usage or the model that answered has no known price. + Findings are `review` (act on it), `consider` (worth a look) or `note` (optional, shown with `--verbose`, never failing the gate). A file whose answers stay undecided is `uncertain`, and one that cannot be judged without more evidence is `needs-context`; neither is hidden or counted as clear. A finding's message shows the probability that set its level; a note shows none, and the JSON report keeps every raw value. Finished plans that share a directory are one finding. A hardcoded-value finding that cannot name its value is one level lower. | Exit code | Meaning | @@ -19,7 +21,11 @@ Findings are `review` (act on it), `consider` (worth a look) or `note` (optional | 1 | Gate failed | | 2 | Run incomplete, invalid configuration or invalid usage | -`--fail-on review|consider|uncertain|none` sets what fails the gate; `--fail-on security=consider` sets it for one group or rule. Baselined findings, findings allowed by a comment, and notes never fail it. +After the findings, the agent text gives each reason files failed or were skipped, with how many files give it, such as `Failed 3: TypeSafe HTTP 402 (credits exhausted; …)`, so an incomplete run says why without `--verbose`, and so does the MCP server's `jevgate_check`, which returns this text. + +`--fail-on review|consider|mature|uncertain|none` sets what fails the gate; `--fail-on security=consider` sets it for one group or rule. The default, `mature`, fails only on the rules and levels measured right at least 80% of the time on projects JevGate was never tuned on; `jevgate rules` lists them, and [configuration](configuration.md#what-fails-the-check-by-default) explains it. Baselined findings, findings allowed by a comment, and notes never fail the gate. + +The agent text marks each finding that fails the gate with `(fails the gate)`, and says when reviews did not fail it because their rules are still being measured. The JSON report records how the gate counted each new finding in its `gate` field: `fails`, `measuring` (reported without failing: the level is `mature` and its rule and level are still being measured) or `advisory` (the level in force does not count it, as `review` does not count a consider). `fail_on_mature` says what `mature` stands for among the selected rules. A single finding can also be accepted where it is, with a comment on its line or directly above it (doc comments and attributes may sit in between). The comment names a rule ID (`security/injection`), its name (`injection`), its key or a group, and needs a reason; without one it is ignored and the finding says so: diff --git a/site/src/privacy-and-cost.md b/site/src/privacy-and-cost.md index 8dcbcb9..7781547 100644 --- a/site/src/privacy-and-cost.md +++ b/site/src/privacy-and-cost.md @@ -3,6 +3,7 @@ - **What is uploaded:** only the selected units of source, bounded by `upload_allow` and `upload_deny`. `--dry-run --show-requests` prints every initial request body without credentials or network access. - **Instruction files:** uploaded only when a documentation rule is selected, and still bounded by the upload patterns. - **README opening:** a sensitive-data finding about error details is asked who reads the error text, with the first 1,200 characters of the root README's prose (images, badges and HTML left out), unless `upload_deny` covers the README or `upload_allow` leaves it out. -- **Credentials:** a check reads `TYPESAFE_API_KEY` from the environment, then `--env-file` or the repository's `.env`, then the key saved by `jevgate auth login` (OS credential store, or an owner-only file). The key is never printed or written to reports. -- **Cost:** every run prints its input tokens and an estimated cost. Cached answers cost nothing. +- **Credentials:** a check uses the first key it finds: `TYPESAFE_API_KEY` in the environment, then the file `--env-file` names (`TYPESAFE_API_KEY`, `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY`) or `TYPESAFE_API_KEY` in the repository's `.env`, then the key saved by `jevgate auth login` (OS credential store, or an owner-only file), then `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY` in the environment. A gateway's variable exported for another tool therefore never sends your source through that gateway while JevGate has a key of its own. The key is never printed or written to reports, and it goes only to its own provider: a key that starts as another provider's keys do (`sk-or-`, `vck_`) is refused, and nothing in the repository can choose the host ([Configuration](configuration.md#keys-and-where-requests-go)). +- **Where requests go:** to TypeSafe, or through OpenRouter or Vercel AI Gateway, which pass TypeSafe's API through. TypeSafe commits not to train models on user data and offers zero data retention to enterprise customers ([Privacy Policy](https://typesafe.ai/legal/privacy-policy)). A gateway also handles each request under its own terms: Vercel AI Gateway's model list marks Jev `zdr: none`, no zero data retention, and `no_training: all` (checked 2026-09-28). +- **Cost:** every run prints its input tokens and an estimated cost, priced by the model that answered: Jev 1.13 costs $0.042 per million input tokens, under any of its names (`jev-1.13.0`, OpenRouter's `typesafe/jev-1.13` and its dated `typesafe/jev-1.13-20260917`), and output is free. When a response reports no token usage, or names a model without a version, such as Vercel AI Gateway's `typesafe-ai/jev`, whose price follows whatever version it points to, the cost is shown as unknown, never as $0. Cached answers cost nothing. - **Secrets:** out of scope on purpose, because judging secrets would mean uploading them. Use a local secret scanner. diff --git a/site/src/quick-start.md b/site/src/quick-start.md index d9717e1..35ad41f 100644 --- a/site/src/quick-start.md +++ b/site/src/quick-start.md @@ -2,7 +2,7 @@ ```sh jevgate init # write a commented jevgate.toml for this repository -jevgate auth login # validate and save your TypeSafe API key +jevgate auth login # validate and save your API key: TypeSafe, OpenRouter or Vercel jevgate check --dry-run --show-requests # see exactly what would be uploaded; free and offline jevgate check --report # review, then open a local HTML dashboard jevgate baseline # accept today's findings; later checks fail only on new ones @@ -16,7 +16,7 @@ jevgate check --rule default --rule security # add the security group jevgate check --rule documentation # agent instruction files, project docs and code comments jevgate check --rule comments # only code comments jevgate check --include-tests # also judge tests -jevgate check --base origin/main --format json # changed files only, for agents and scripts +jevgate check --base origin/main --format json # only what changed, for agents and scripts jevgate check --watch # re-check on save jevgate baseline --merge # after a --base or path check: accept its findings, keep the rest jevgate baseline mark wrong src/api/search.ts:41 # record why an accepted finding was accepted diff --git a/site/src/stability.md b/site/src/stability.md index 3bbd637..d51e8ea 100644 --- a/site/src/stability.md +++ b/site/src/stability.md @@ -24,6 +24,8 @@ Deprecated flags and keys keep working for at least one minor release, with a wa These are judgments or presentation, and any release can change them; the changelog says how: - **Which findings a rule reports**, their levels, wording, probabilities and next steps. Findings are model judgments composed by code, and improving them is most of what releases do. A rule's `version` changes when its questions or composition change, and its cached answers are asked again. +- **Which rules and levels fail the check by default.** The default level, `mature`, follows the labeled findings: a release marks a rule's reviews or considers mature once they are right at least 80% of the time on projects JevGate was never tuned on, over at least 20 labels, and can drop one that stops measuring up. The changelog gives the numbers. Set `fail_on`, or a level per rule, to keep a fixed policy. +- **Which rules run by default.** A rule can leave the default group, as hardcoded values did in 0.26, or join it; `--rule` and `[rules]` keep an explicit selection. - **The agent text** (`--format agent`): it is written for people and coding agents to read. Scripts should read JSON. - **Request bodies and the answer cache**: the cache is safe to delete or restore at any version; unmatched entries are simply not used. - **The default model**: a release can pin a newer model version, which re-asks every unit once. Set `model` in `jevgate.toml` to keep one. diff --git a/site/src/troubleshooting.md b/site/src/troubleshooting.md index c0b1a62..ddcf669 100644 --- a/site/src/troubleshooting.md +++ b/site/src/troubleshooting.md @@ -4,8 +4,14 @@ Exit code 2 means the run could not finish, or the configuration or command line is invalid. The message says which; an outage never passes as a clean review. -**`No API key configured. Run jevgate auth login, set TYPESAFE_API_KEY, or provide --env-file PATH`** -: A check reads `TYPESAFE_API_KEY` from the environment, then `--env-file` or the repository's `.env`, then the key saved by `jevgate auth login`. `jevgate auth status` shows which one a check would use and verifies it. On GitHub Actions, pull requests from forks don't receive secrets: skip the job for them (`if: github.event.pull_request.head.repo.full_name == github.repository`). +**`No API key configured. Run jevgate auth login, set TYPESAFE_API_KEY (or OPENROUTER_API_KEY, AI_GATEWAY_API_KEY), or provide --env-file PATH`** +: A check reads `TYPESAFE_API_KEY` from the environment, then `--env-file` (or `TYPESAFE_API_KEY` in the repository's `.env`), then the key saved by `jevgate auth login`, then `OPENROUTER_API_KEY` or `AI_GATEWAY_API_KEY` from the environment. A gateway's key in the repository's `.env` is read only with `--env-file .env`, since it is usually the application's own; the message says so when one is there. `jevgate auth status` shows which key a check would use and verifies it. On GitHub Actions, pull requests from forks don't receive secrets: skip the job for them (`if: github.event.pull_request.head.repo.full_name == github.repository`). + +**`TYPESAFE_API_KEY environment variable: the key was issued by OpenRouter (it starts with sk-or-), not by TypeSafe`** +: A key goes only to the provider that issued it. Set it in the variable the message names, or save it with `jevgate auth login --provider openrouter`. + +**`TypeSafe HTTP 400 (unknown model; check the model name)`**, or **`OpenRouter HTTP 404 (not found; check the model name)`** +: A model name is sent as written, and each provider has its own: `jev-1.13.0`, `jev-latest` or `jev-preview` on TypeSafe (which refuses `jev-1.13`, though its docs use it), `typesafe/jev-1.13` on OpenRouter, `typesafe-ai/jev` on Vercel AI Gateway. A `model` in `jevgate.toml` written for one provider needs `--model` with another provider's key. **`Cannot find revision …; in CI, fetch it (for example fetch-depth: 0)`**, or **`… and HEAD share no history`** : `--base` needs the history back to the fork point. Check out with `fetch-depth: 0`. @@ -16,22 +22,41 @@ Exit code 2 means the run could not finish, or the configuration or command line **`Cannot connect to TypeSafe; request was not sent`** : A network problem before anything was sent. Rerun; cached answers are kept. +**`OpenRouter HTTP 503; gave up after 6 attempts`**, or **`TypeSafe HTTP 503; gave up after 6 attempts`** +: The provider stayed overloaded through six attempts and 23 seconds or more of pauses. On 2026-09-28 TypeSafe answered 503 to about two attempts in three for at least ten minutes, directly and through OpenRouter alike. The run exits 2 and keeps the answers it received: rerun later, and only the unanswered units are asked. + +**`TypeSafe HTTP 402 (credits exhausted; add credits or turn on auto-refill at https://console.typesafe.ai)`** +: The account's prepaid credits ran out. The run stops sending requests and exits 2; the answers it received are kept in the cache. Add credits and rerun: only the unanswered units are asked. + +**`TypeSafe HTTP 422 (invalid request: body.questions.q1.criteria missing)`** +: TypeSafe refused a request as malformed. The message names each invalid field and its error type, never the text TypeSafe sends with it, which can quote your source. It is a JevGate bug: please [report it](https://github.com/Tech-Byte-Frontier/jevgate/issues/new) with the request id. + +A provider error ends with the provider's request id when it sent one (`; request id req_…`); quote it to the provider's support. The report also keeps the id of the request behind each answer (`files[].judgments[].request_id`). + **`Session API request budget exhausted; restart with an explicit larger --max-requests`** : `max_requests` or `--max-requests` capped the run. Raise it, or check fewer files with `--base` or paths; `--dry-run` estimates what a run will ask. **`Another JevGate session owns latest.json`** : Another `check` or `--watch` is running in the same repository. Stop it first. -Rate limits, overload and server errors (HTTP 408, 429, 500, 502–504, 520–524, 529) are retried up to four attempts before the run gives up, and a timeout or dropped connection is retried once. +Rate limits, overload and server errors (HTTP 408, 429, 500, 502–504, 520–524, 529) are retried up to six attempts before the run gives up, pausing 1, 2, 4, 8 and 8 seconds (each up to a quarter longer, so requests spread out), or as long as the provider asks when that is longer (`retry-after-ms`, or `Retry-After` in seconds or as a date, at most 30 seconds). A pause holds every request of the run. An attempt that has not answered within 20 seconds, or whose connection drops, is retried once, and a connection that fails before anything is sent is tried four times. Requests start at least 50 ms apart, within TypeSafe's limit of 1,200 a minute, and at most 6 are sent at once with a TypeSafe key, 3 with an OpenRouter or Vercel AI Gateway key (`--concurrency` or `concurrency` sets it). ## Many files are uncertain A file is `uncertain` when some of its answers stayed undecided after the follow-up questions. JevGate reports this instead of hiding it or counting the file as clear. `--verbose` lists each undecided unit and the question it stayed undecided on. It never fails the gate unless you ask for that with `--fail-on uncertain`. +## A review did not fail the check + +By default only the rules and levels measured right at least 80% of the time on projects JevGate was never tuned on fail the check; `jevgate rules` shows which, and how often each rule's reviews and considers were right. The other findings are reported, and the output says their rules are still being measured. To fail on them, set a level: `--fail-on review` or `fail_on = ["review"]` for every rule, or `--fail-on maintainability/shared-logic=review` for one. + ## A finding is wrong Accept it with `jevgate baseline`, and record why with `jevgate baseline mark wrong PATH:LINE`; `jevgate baseline stats` counts each rule's mistaken findings. Reporting it with the [wrong finding template](https://github.com/Tech-Byte-Frontier/jevgate/issues/new?template=wrong_finding.yml), with the finding from `.jevgate/latest.json` and a small piece of the code, is how the rules improve. +## A `--base` check leaves out a finding + +With `--base`, only what the change touches is asked about and reported: units on changed lines, copies where either copy changed, a file's outline when the change adds members, and documents naming a path it removed. A finding elsewhere in a changed file comes back with `--whole-files`, or in a check without `--base`. The report's `scope` says which a check used. + ## A file is skipped Skipped files are listed with the reason: generated, vendored or minified code, migrations, an unsupported language, syntax errors, a parser that did not finish within 10 seconds, Bend 1 code (JevGate reads Bend 2), or a path outside the upload patterns. `generated`, `tests` and the upload patterns in `jevgate.toml` change what is selected. A file larger than `max_file_bytes` is not skipped but reported as `needs-context`, never truncated, and so is a unit whose request the provider refuses as beyond the model's context. diff --git a/site/src/what-it-finds.md b/site/src/what-it-finds.md index d73a317..5124bf2 100644 --- a/site/src/what-it-finds.md +++ b/site/src/what-it-finds.md @@ -1,13 +1,13 @@ # What it finds -**Maintainability** (on by default) +**Maintainability** (on by default, except hardcoded values) | Rule | Example finding | |---|---| | File organization | This file holds several features that would be easier to find apart; the upload helpers would be most useful as their own module. Test files are judged too, at most as a consider. | | Function simplification | `sync_accounts` mixes separate jobs in long blocks; lines 40–71 would be most useful as their own function. | | Shared logic | `createInvoice` and `createReceipt` perform the same steps; one shared implementation would serve both. | -| Hardcoded values | Module constants fix a value that differs between deployments; `apply_discount` special-cases one specific customer. | +| Hardcoded values (opt-in: `--rule default --rule hardcoded-values`) | Module constants fix a value that differs between deployments; `apply_discount` special-cases one specific customer. On projects JevGate was never tuned on, 6 of its 37 labeled findings were right, against 47 of 85 on the projects it was tuned on, so it no longer runs by default. | **Tests** (with `--include-tests`; file organization judges test files without it, and the laws of Bend 2 code are judged where they are) @@ -37,6 +37,6 @@ | Duplication | Section `Release Workflow` of `CLAUDE.md` states everything section `Release` of `README.md` states. | | Code comments | `save_skill` has 3 comments to clean up: at lines 214, 218 and 222 they repeat the code (`# Create skill directory` above `skill_dir.mkdir(…)`). A module docstring saying it was "split out of `portfolio.py` to stay under the 500-line budget" narrates an edit instead of the code as it is. | -The documentation rules read the instruction files that coding agents load (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, and Claude, Cursor, Copilot, Windsurf, Cline, Kiro, Junie and Roo Code rules), even when hidden or gitignored. Each section is asked whether it only restates the stack, the manifest's commands, generic advice or a configured linter's rules, and whether text loaded in every session applies to only one directory. Project documentation in Markdown, MDX, reStructuredText or AsciiDoc of 300 or more lines is judged from its headings alone. Code finds staleness and duplication candidates: named paths or scripts that no longer exist, release tags, deleted files, and shared wording outside code examples. Jev then judges each candidate. Code comments and docstrings of application code are judged one at a time with the code they are about (the declaration they document, the lines below them or the line they end): whether they only repeat that code, hold sentences that add nothing, narrate an edit instead of the code as it is, or are code turned off. License headers, tool directives, type annotations, authorship tags and Sphinx version notes are left out; documentation that only repeats its declaration or says it at length (framework section banners included) is at most a `note`, as is every such comment in a project whose README says its code is written for learners; and the comments of one definition that span fewer than three lines in all are a `note`. Documentation findings are at most `consider`. The run also estimates the tokens each harness loads at session start; these estimates are evidence and never fail the gate. +The documentation rules read the instruction files that coding agents load (`AGENTS.md`, `CLAUDE.md`, `GEMINI.md`, and Claude, Cursor, Copilot, Windsurf, Cline, Kiro, Junie and Roo Code rules), even when hidden or gitignored. Each section is asked whether it only restates the stack, the manifest's commands, generic advice or a configured linter's rules, and whether text loaded in every session applies to only one directory. Project documentation in Markdown, MDX, reStructuredText or AsciiDoc of 300 or more lines is judged from its headings alone. Code finds staleness and duplication candidates: named paths or scripts that no longer exist, release tags, deleted files, and shared wording outside code examples. Jev then judges each candidate. Code comments and docstrings of application code are judged one at a time with the code they are about (the declaration they document, the lines below them or the line they end): whether they only repeat that code, hold sentences that add nothing, narrate an edit instead of the code as it is, or are code turned off. License headers, tool directives, type annotations, authorship tags and Sphinx version notes are left out; documentation that only repeats its declaration or says it at length (framework section banners included) is at most a `note`, as is every such comment in a project whose README says its code is written for learners; and the comments of one definition that span fewer than three lines in all are a `note`. Documentation findings are at most `consider`, and agent-context considers are the one documentation level that fails the check by default once these rules run: 22 of 24 were right on projects JevGate was never tuned on. The run also estimates the tokens each harness loads at session start; these estimates are evidence and never fail the gate. -`jevgate rules` prints every rule with its question and default; the [rules reference](reference/rules.md) lists them with what each one looks at. +`jevgate rules` prints every rule with its question, its default, and how often its reviews and considers were right on projects JevGate was never tuned on; by default only the levels right at least 80% of the time there fail the check. The [rules reference](reference/rules.md) lists them with what each one looks at. diff --git a/src/auth/file.rs b/src/auth/file.rs index b50b48b..51db188 100644 --- a/src/auth/file.rs +++ b/src/auth/file.rs @@ -1,11 +1,14 @@ -//! The fallback is confined to an owner-only directory and never writes a repository .env. -use super::secret::{MAX_KEY_BYTES, Secret}; +//! The fallback is confined to an owner-only directory and never writes a +//! repository .env. Beside it, a plain file names the saved key's provider. +use super::secret::MAX_STORED_BYTES; +use crate::provider::Provider; use anyhow::{Context, Result, ensure}; use std::{ fs, io::Read, path::{Path, PathBuf}, }; +use zeroize::Zeroizing; pub fn credential_path() -> Result { if let Some(path) = std::env::var_os("JEVGATE_CONFIG_DIR") { @@ -76,14 +79,15 @@ fn private_metadata(_path: &Path, _directory: bool) -> Result { ) } -pub fn load(path: &Path) -> Result> { +/// The saved credential's text, as `store` wrote it. +pub fn load(path: &Path) -> Result>> { if !path.try_exists()? && !path.is_symlink() { return Ok(None); } private_metadata(path.parent().context("Missing credential directory")?, true)?; let metadata = private_metadata(path, false)?; ensure!( - metadata.len() <= MAX_KEY_BYTES as u64, + metadata.len() <= MAX_STORED_BYTES as u64, "Saved credential file is too large" ); let mut options = fs::OpenOptions::new(); @@ -93,23 +97,26 @@ pub fn load(path: &Path) -> Result> { use std::os::unix::fs::OpenOptionsExt; options.custom_flags(libc::O_NOFOLLOW); } - let mut value = zeroize::Zeroizing::new(String::new()); + let mut value = Zeroizing::new(String::new()); options .open(path)? - .take((MAX_KEY_BYTES + 1) as u64) + .take((MAX_STORED_BYTES + 1) as u64) .read_to_string(&mut value) .map_err(|_| anyhow::anyhow!("Cannot read saved credential; run jevgate auth login"))?; ensure!( - value.len() <= MAX_KEY_BYTES, + value.len() <= MAX_STORED_BYTES, "Saved credential file is too large" ); - Secret::parse(std::mem::take(&mut *value)).map(Some) + Ok(Some(value)) } -pub fn save(path: &Path, secret: &Secret) -> Result<()> { +/// Write the saved credential's text, as `store` made it, to the owner-only +/// file at `path`, replacing it whole. It holds an API key, which is sent to +/// its provider as written, so it is kept as is: a hash could not be sent. +pub fn save(path: &Path, text: &str) -> Result<()> { #[cfg(not(unix))] { - let _ = (path, secret); + let _ = (path, text); anyhow::bail!( "Protected-file storage is unavailable; use the system credential store or TYPESAFE_API_KEY" ); @@ -117,13 +124,9 @@ pub fn save(path: &Path, secret: &Secret) -> Result<()> { #[cfg(unix)] { use std::io::Write; - use std::os::unix::fs::{DirBuilderExt, OpenOptionsExt}; + use std::os::unix::fs::OpenOptionsExt; let parent = path.parent().context("Missing credential directory")?; - fs::DirBuilder::new() - .recursive(true) - .mode(0o700) - .create(parent)?; - private_metadata(parent, true)?; + private_directory(parent)?; if path.exists() || path.is_symlink() { private_metadata(path, false)?; } @@ -137,7 +140,7 @@ pub fn save(path: &Path, secret: &Secret) -> Result<()> { .mode(0o600) .open(&temporary)?; let result = (|| -> Result<()> { - file.write_all(secret.expose().as_bytes())?; + file.write_all(text.as_bytes())?; file.sync_all()?; fs::rename(&temporary, path)?; Ok(()) @@ -158,3 +161,61 @@ pub fn remove(path: &Path) -> Result { fs::remove_file(path)?; Ok(true) } + +/// The directory of saved credentials, created owner-only. +fn private_directory(directory: &Path) -> Result<()> { + #[cfg(unix)] + { + use std::os::unix::fs::DirBuilderExt; + fs::DirBuilder::new() + .recursive(true) + .mode(0o700) + .create(directory)?; + private_metadata(directory, true)?; + } + #[cfg(not(unix))] + fs::create_dir_all(directory)?; + Ok(()) +} + +/// The file beside the saved credential that names its provider, so a check +/// can choose its model before it reads the credential itself. +fn provider_path(credential: &Path) -> PathBuf { + credential.with_file_name("provider") +} + +/// The most of that file read: a provider's name is a short word. +const MAX_PROVIDER_RECORD_BYTES: u64 = 64; + +/// The provider recorded beside the saved credential; none when no key was +/// saved with one, as before 0.26. +pub fn recorded_provider(credential: &Path) -> Option { + let path = provider_path(credential); + if path.is_symlink() { + return None; + } + let mut name = String::new(); + fs::File::open(path) + .ok()? + .take(MAX_PROVIDER_RECORD_BYTES) + .read_to_string(&mut name) + .ok()?; + Provider::named(name.trim()) +} + +pub fn record_provider(credential: &Path, provider: Provider) -> Result<()> { + let path = provider_path(credential); + private_directory(path.parent().context("Missing credential directory")?)?; + ensure!( + !path.is_symlink(), + "The saved key's provider record must be a regular file" + ); + fs::write(&path, provider.name()).context("Could not record the saved key's provider") +} + +pub fn forget_provider(credential: &Path) -> Result<()> { + match fs::remove_file(provider_path(credential)) { + Err(error) if error.kind() != std::io::ErrorKind::NotFound => Err(error.into()), + _ => Ok(()), + } +} diff --git a/src/auth/mod.rs b/src/auth/mod.rs index 5bcf34a..b6f4c37 100644 --- a/src/auth/mod.rs +++ b/src/auth/mod.rs @@ -4,33 +4,38 @@ mod file; not(any(target_os = "macos", target_os = "ios", target_os = "android")) ))] mod native_unix; -mod provider; mod secret; pub(crate) mod sources; mod store; +mod verify; -use anyhow::{Result, ensure}; +use crate::provider::{Endpoint, Provider}; +use anyhow::{Result, bail, ensure}; use clap::{Args, Subcommand}; -use provider::{TypeSafe, Verifier}; use secret::Secret; -use std::{io::IsTerminal, path::PathBuf}; +use std::{ + io::{BufRead, IsTerminal, Write}, + path::PathBuf, +}; use store::{Backend, NativeBackend, SavedCredentials, StorageMode}; +use verify::Verifier; #[derive(Subcommand)] pub enum AuthCommand { - /// Validate a TypeSafe API key and save it for every repository + /// Validate an API key from TypeSafe, OpenRouter or Vercel AI Gateway and save it for every repository /// - /// Prompts without echo, checks the key with TypeSafe (no source is sent), - /// and saves it in the OS credential store, or in an owner-only file where - /// no store is available. Create a key at - /// https://console.typesafe.ai/settings/keys. + /// Asks which kind of key it is, prompts without echo, checks the key + /// with its provider (no source is sent), and saves it with its provider + /// in the OS credential store, or in an owner-only file where no store is + /// available. Create a key at https://console.typesafe.ai/settings/keys, + /// https://openrouter.ai/settings/keys or in the Vercel dashboard. Login(LoginArgs), /// Show which credential a check would use; exit 0 when it works, 2 otherwise /// - /// Verifies the key with TypeSafe unless --offline. No source is sent and - /// the key is never printed. + /// Verifies the key with its provider unless --offline. No source is sent + /// and the key is never printed. Status(StatusArgs), - /// Remove saved credentials; TYPESAFE_API_KEY and repository .env files are left alone + /// Remove saved credentials; environment variables and repository .env files are left alone Logout, } @@ -39,6 +44,9 @@ pub struct LoginArgs { /// Read one key from stdin instead of prompting, for scripts #[arg(long)] with_key: bool, + /// The key's provider [default: asked on a terminal; typesafe with --with-key] + #[arg(long, value_enum)] + provider: Option, /// Where to save the key [default: JEVGATE_CREDENTIAL_STORE, else auto] /// /// `auto` uses the OS credential store and, on Unix, falls back to an @@ -50,13 +58,13 @@ pub struct LoginArgs { #[derive(Args)] pub struct StatusArgs { - /// Inspect this credential file instead of the repository .env; TYPESAFE_API_KEY still wins + /// Inspect this credential file instead of the repository .env; keys in the environment still win #[arg(long, value_name = "FILE")] env_file: Option, - /// Report the credential source without contacting TypeSafe + /// Report the credential source without contacting the provider #[arg(long)] offline: bool, - /// Print source, configured, connection_checked, authenticated and error as JSON (never the key) + /// Print source, provider, endpoint, configured, connection_checked, authenticated, error and unused as JSON (never the key) #[arg(long)] json: bool, } @@ -70,96 +78,196 @@ pub fn run(command: AuthCommand) -> Result { } fn login(args: LoginArgs) -> Result { - let key = if args.with_key { + let (provider, key) = if args.with_key { ensure!( !std::io::stdin().is_terminal(), "Pipe the key into jevgate auth login --with-key, or omit --with-key for hidden terminal entry" ); - secret::read_stdin(std::io::stdin().lock())? + let provider = args.provider.unwrap_or_default(); + (provider, secret::read_stdin(std::io::stdin().lock())?) } else { - ensure!( - std::io::stdin().is_terminal() && std::io::stderr().is_terminal(), - "Interactive login requires a terminal. For automation use jevgate auth login --with-key < key-file, or set TYPESAFE_API_KEY" - ); - note!("Create an API key at https://console.typesafe.ai/settings/keys"); - let value = rpassword::prompt_password("TypeSafe API key (hidden): ").map_err(|_| { - anyhow::anyhow!("Could not read hidden input; use --with-key to read from stdin") - })?; - Secret::parse(value)? + prompt(args.provider)? }; + let service = provider.service(); + sources::refuse_foreign(provider, &key, "jevgate auth login")?; + let endpoint = Endpoint::new(provider)?; let mode = args .storage .map(Ok) .unwrap_or_else(StorageMode::configured)?; let store = SavedCredentials::native(mode)?; - note!("Validating with TypeSafe; no source code is uploaded."); - let saved = validate_and_save(&TypeSafe, &store, &key)?; + note!( + "Validating with {}; no source code is uploaded.", + service.label + ); + let saved = validate_and_save(&endpoint, &store, provider, &key)?; if saved.fallback { note!("System credential store unavailable; using owner-only file storage (unencrypted)."); } - say!("API key verified and saved in {}.", saved.description); + say!( + "{} API key verified and saved in {}.", + service.label, + saved.description + ); say!("Ready: jevgate check . --dry-run"); report_override(); Ok(0) } +/// Ask on the terminal which kind of key it is, unless `--provider` said, +/// then read the key without echo. +fn prompt(chosen: Option) -> Result<(Provider, Secret)> { + ensure!( + std::io::stdin().is_terminal() && std::io::stderr().is_terminal(), + "Interactive login requires a terminal. For automation use jevgate auth login --with-key [--provider NAME] < key-file, or set TYPESAFE_API_KEY, OPENROUTER_API_KEY or AI_GATEWAY_API_KEY" + ); + let provider = match chosen { + Some(provider) => provider, + None => ask_provider(&mut std::io::stdin().lock(), &mut std::io::stderr())?, + }; + let service = provider.service(); + note!("Create an API key at {}", service.keys_page); + let value = rpassword::prompt_password(format!("{} API key (hidden): ", service.label)) + .map_err(|_| { + anyhow::anyhow!("Could not read hidden input; use --with-key to read from stdin") + })?; + Ok((provider, Secret::parse(value)?)) +} + +/// Answers `ask_provider` takes before it gives up. +const MAX_ANSWERS: usize = 3; + +/// Ask which kind of key it is until the answer names one: its number or its +/// name, or nothing for TypeSafe. +fn ask_provider(input: &mut impl BufRead, output: &mut impl Write) -> Result { + let menu: Vec = Provider::ALL + .iter() + .enumerate() + .map(|(i, provider)| format!("{} {}", i + 1, provider.service().label)) + .collect(); + for _ in 0..MAX_ANSWERS { + write!(output, "Key kind: {} [1]: ", menu.join(", "))?; + output.flush()?; + let mut answer = String::new(); + if input.read_line(&mut answer)? == 0 { + break; + } + if let Some(provider) = choice(answer.trim()) { + return Ok(provider); + } + writeln!( + output, + "Answer 1, 2 or 3, or typesafe, openrouter or vercel." + )?; + } + bail!("No key kind given; pass --provider typesafe, openrouter or vercel") +} + +/// The provider an answer names by number or name; empty means TypeSafe. +fn choice(answer: &str) -> Option { + if answer.is_empty() { + return Some(Provider::Typesafe); + } + let by_number = answer + .parse::() + .ok() + .and_then(|n| Provider::ALL.get(n.checked_sub(1)?).copied()); + by_number.or_else(|| Provider::named(&answer.to_ascii_lowercase())) +} + fn validate_and_save( verifier: &impl Verifier, store: &SavedCredentials, + provider: Provider, key: &Secret, ) -> Result { verifier.verify(key)?; - store.save(key) + store.save(provider, key) +} + +/// What `auth status` found: where the key is, its provider and endpoint, +/// the keys set besides it, and whether it works. +struct Status { + source: Option, + provider: Option, + endpoint: Option, + unused: Vec, + result: Result<()>, + checked: bool, } fn status(args: StatusArgs) -> Result { let path = environment_path(args.env_file.as_ref())?; - let credential = sources::resolve(&path, args.env_file.is_some()); - let (source, result) = match credential { - Ok(credential) => { - let result = if args.offline { + let explicit = args.env_file.is_some(); + let unused = sources::unused(&path, explicit).unwrap_or_default(); + let found = sources::resolve(&path, explicit) + .and_then(|credential| Ok((Endpoint::new(credential.provider)?, credential))); + let status = match found { + Ok((endpoint, credential)) => Status { + source: Some(credential.source), + provider: Some(credential.provider), + endpoint: Some(endpoint.describe()), + unused, + result: if args.offline { Ok(()) } else { - TypeSafe.verify(&credential.key) - }; - (Some(credential.source), result) - } - Err(error) => (None, Err(error)), + endpoint.verify(&credential.key) + }, + checked: !args.offline, + }, + Err(error) => Status { + source: None, + provider: None, + endpoint: None, + unused, + result: Err(error), + checked: false, + }, }; - let checked = !args.offline && source.is_some(); - let code = if result.is_ok() { 0 } else { 2 }; - print_status(&args, source, result, checked)?; + let code = if status.result.is_ok() { 0 } else { 2 }; + print_status(&args, status)?; Ok(code) } -fn print_status( - args: &StatusArgs, - source: Option, - result: Result<()>, - checked: bool, -) -> Result<()> { +fn print_status(args: &StatusArgs, status: Status) -> Result<()> { if args.json { say!( "{}", serde_json::to_string_pretty(&serde_json::json!({ - "source":source,"configured":source.is_some(),"connection_checked":checked, - "authenticated":provider::authenticated(&result, checked), - "error":result.err().map(|e|format!("{e:#}")), + "source": status.source, + "provider": status.provider.map(Provider::name), + "endpoint": status.endpoint, + "configured": status.source.is_some(), + "connection_checked": status.checked, + "authenticated": verify::authenticated(&status.result, status.checked), + "error": status.result.as_ref().err().map(|e| format!("{e:#}")), + "unused": status.unused, }))? ); return Ok(()); } - if let Some(source) = source { + if let (Some(source), Some(provider), Some(endpoint)) = + (&status.source, status.provider, &status.endpoint) + { + let service = provider.service(); say!("Credential source: {source}"); + say!( + "Provider: {} at {endpoint}; default model {}", + service.label, + service.default_model + ); + } + if !status.unused.is_empty() { + say!( + "Also set, not used: {} (the first key found is used)", + status.unused.join("; ") + ); } - match result { + match status.result { + Ok(()) if args.offline => say!("Connection: not checked (--offline)."), Ok(()) => say!( - "{}", - if args.offline { - "Connection: not checked (--offline)." - } else { - "Connection: authenticated with TypeSafe. No source code was uploaded." - } + "Connection: authenticated with {}. No source code was uploaded.", + status.provider.unwrap_or_default().service().label ), Err(error) => note!("Authentication: {error:#}"), } @@ -195,14 +303,7 @@ fn environment_path(selected: Option<&PathBuf>) -> Result { } fn report_override() { - let override_source = (|| -> Result> { - if sources::environment()?.is_some() { - return Ok(Some("TYPESAFE_API_KEY environment variable".into())); - } - let path = environment_path(None)?; - Ok(sources::key_from_file(&path)?.map(|_| format!("repository .env: {}", path.display()))) - })(); - match override_source { + match environment_path(None).and_then(|path| sources::override_source(&path)) { Ok(Some(source)) => say!( "Current override: {source}. Saved credentials are used when this override is absent." ), diff --git a/src/auth/native_unix.rs b/src/auth/native_unix.rs index beff136..f9f56bb 100644 --- a/src/auth/native_unix.rs +++ b/src/auth/native_unix.rs @@ -1,10 +1,10 @@ //! Secret Service reads and deletion never invoke its Unlock or Prompt methods. //! The higher-level keyring library can open a dialog even when just reading. -use super::secret::Secret; use anyhow::{Result, ensure}; use secret_service::{EncryptionType, blocking::SecretService}; use std::collections::HashMap; use zbus::blocking::{Connection, Proxy}; +use zeroize::Zeroizing; fn attributes() -> HashMap<&'static str, &'static str> { HashMap::from([("service", "jevgate"), ("username", "typesafe-api-key")]) @@ -19,15 +19,16 @@ fn unlocked_item<'a>( Ok(items.unlocked.pop()) } -pub fn get() -> Result> { - (|| -> Result> { +/// The saved credential's text, as `store` wrote it. +pub fn get() -> Result>> { + (|| -> Result>> { let service = SecretService::connect(EncryptionType::Dh)?; let Some(item) = unlocked_item(&service)? else { return Ok(None); }; - let bytes = zeroize::Zeroizing::new(item.get_secret()?); + let bytes = Zeroizing::new(item.get_secret()?); let value = std::str::from_utf8(&bytes)?; - Secret::parse(value.to_owned()).map(Some) + Ok(Some(Zeroizing::new(value.to_owned()))) })() .map_err(|_| anyhow::anyhow!( "Cannot read the system credential store; unlock it and retry, or use TYPESAFE_API_KEY or an --env-file. Run jevgate auth login to configure credentials" diff --git a/src/auth/provider.rs b/src/auth/provider.rs deleted file mode 100644 index 3997b54..0000000 --- a/src/auth/provider.rs +++ /dev/null @@ -1,74 +0,0 @@ -use super::secret::Secret; -use anyhow::{Result, bail, ensure}; -use serde_json::Value; -use std::time::Duration; - -pub trait Verifier { - fn verify(&self, key: &Secret) -> Result<()>; -} -pub struct TypeSafe; -#[derive(Debug)] -pub struct RejectedKey(u16); -impl std::fmt::Display for RejectedKey { - fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - write!( - formatter, - "TypeSafe rejected this API key (HTTP {}); create a key at https://console.typesafe.ai/settings/keys and run jevgate auth login", - self.0 - ) - } -} -impl std::error::Error for RejectedKey {} - -pub fn authenticated(result: &Result<()>, checked: bool) -> Option { - if !checked { - return None; - } - match result { - Ok(()) => Some(true), - Err(error) if error.is::() => Some(false), - Err(_) => None, - } -} -impl Verifier for TypeSafe { - fn verify(&self, key: &Secret) -> Result<()> { - let agent: ureq::Agent = ureq::Agent::config_builder() - .timeout_global(Some(Duration::from_secs(15))) - .max_redirects(0) - .build() - .into(); - let response = agent - .get("https://api.typesafe.ai/v1/models") - .header("Authorization", format!("Bearer {}", key.expose())) - .call(); - let mut response = match response { - Ok(response) => response, - Err(ureq::Error::StatusCode(code)) => return http_error(code), - Err(_) => bail!( - "Could not reach TypeSafe or the request timed out; check the connection and retry. The credential was not verified" - ), - }; - let body: Value = response.body_mut().with_config().limit(1_048_576).read_json() - .map_err(|_| anyhow::anyhow!("TypeSafe returned invalid or oversized model-list JSON; credential was not verified"))?; - validate_models(&body) - } -} - -pub fn http_error(code: u16) -> Result<()> { - match code { - 401 | 403 => Err(RejectedKey(code).into()), - _ => bail!( - "TypeSafe returned HTTP {code}; credential was not verified and the request was not retried" - ), - } -} - -pub fn validate_models(body: &Value) -> Result<()> { - ensure!( - body["models"].as_array().is_some_and(|models| models - .iter() - .all(|model| model["name"].as_str().is_some_and(|name| !name.is_empty()))), - "TypeSafe returned an unexpected model list; credential was not verified" - ); - Ok(()) -} diff --git a/src/auth/secret.rs b/src/auth/secret.rs index 3e6aa86..242d978 100644 --- a/src/auth/secret.rs +++ b/src/auth/secret.rs @@ -3,6 +3,8 @@ use std::io::Read; use zeroize::Zeroizing; pub const MAX_KEY_BYTES: usize = 4096; +/// A saved credential holds a key and, for a gateway, its provider's name. +pub const MAX_STORED_BYTES: usize = MAX_KEY_BYTES + 32; // Intentionally no Debug/Display/Serialize implementation. pub struct Secret(Zeroizing); @@ -14,7 +16,7 @@ impl Secret { !trimmed.is_empty() && trimmed.len() <= MAX_KEY_BYTES && trimmed.bytes().all(|b| b.is_ascii_graphic()), - "Invalid TYPESAFE_API_KEY: provide one nonempty key without spaces or embedded newlines" + "Invalid API key: provide one nonempty key without spaces or embedded newlines" ); Ok(Self(Zeroizing::new(trimmed.to_owned()))) } diff --git a/src/auth/sources.rs b/src/auth/sources.rs index d6d6847..e3795fe 100644 --- a/src/auth/sources.rs +++ b/src/auth/sources.rs @@ -1,76 +1,315 @@ +//! Where a check finds its key: TYPESAFE_API_KEY in the environment, then a +//! credential file, then the key saved by `jevgate auth login`, then a +//! gateway's variable in the environment. The first key found is used, and it +//! goes only to the provider that issued it. use super::{ secret::Secret, - store::{NativeBackend, SavedCredentials, StorageMode}, + store::{self, NativeBackend, SavedCredentials, SavedKey, StorageMode}, }; +use crate::provider::Provider; use anyhow::{Context, Result, bail, ensure}; -use std::{io::Read, path::Path}; +use std::{cell::OnceCell, io::Read, path::Path}; pub struct Credential { pub key: Secret, + pub provider: Provider, pub source: String, } -pub fn resolve(path: &Path, explicit: bool) -> Result { - let environment = environment()?; - resolve_with(environment, path, explicit, || { - SavedCredentials::::native(StorageMode::configured()?)?.get() - }) +/// The credential file a check reads: the one `--env-file` names, else the +/// repository's `.env`. +#[derive(Clone, Copy)] +pub struct CredentialFile<'a> { + pub path: &'a Path, + pub explicit: bool, } -pub fn environment() -> Result> { - match std::env::var("TYPESAFE_API_KEY") { - Ok(value) if value.trim().is_empty() => Ok(None), - Ok(value) => Ok(Some(value)), - Err(std::env::VarError::NotPresent) => Ok(None), - Err(_) => bail!("TYPESAFE_API_KEY must contain UTF-8 text"), +impl CredentialFile<'_> { + /// The providers whose keys this file may hold: all three in a file named + /// with `--env-file`; only TYPESAFE_API_KEY in the repository's `.env`, + /// where a gateway's key is usually the application's own and would bill + /// its account. + fn providers(self) -> &'static [Provider] { + if self.explicit { + &Provider::ALL + } else { + &Provider::ALL[..1] + } } -} -pub fn resolve_with( - environment: Option, - path: &Path, - explicit: bool, - saved: impl FnOnce() -> Result>, -) -> Result { - if let Some(value) = environment { - return Ok(Credential { - key: Secret::parse(value)?, - source: "TYPESAFE_API_KEY environment variable".into(), - }); + /// Its keys for those providers, in the order a check reads them. + fn keys(self) -> Result> { + match read_limited(self.path)? { + Some(text) => parse_keys(&text, self.providers()), + None => Ok(Vec::new()), + } } - if let Some(key) = key_from_file(path)? { - return Ok(Credential { - key, - source: format!( - "{}: {}", - if explicit { - "--env-file" - } else { - "repository .env" - }, - path.display() + + fn source(self, provider: Provider) -> String { + let kind = if self.explicit { + "--env-file" + } else { + "repository .env" + }; + match provider { + Provider::Typesafe => format!("{kind}: {}", self.path.display()), + gateway => format!( + "{kind}: {} ({})", + self.path.display(), + gateway.service().variable ), - }); + } + } + + /// For the repository's `.env`, the gateway keys it holds that a check + /// does not read, as a hint after "No API key configured". + fn unread_hint(self) -> String { + if self.explicit { + return String::new(); + } + let unread: Vec<&str> = read_limited(self.path) + .ok() + .flatten() + .and_then(|text| parse_keys(&text, &Provider::ALL[1..]).ok()) + .into_iter() + .flatten() + .map(|(provider, _)| provider.service().variable) + .collect(); + if unread.is_empty() { + return String::new(); + } + format!( + ". The repository .env's {} is read only with --env-file .env", + unread.join(" and ") + ) + } +} + +/// Where the key a check will use is, found without reading the saved key +/// where its provider is recorded. +pub enum Located { + Key(Credential), + /// The key saved by `jevgate auth login`, of the provider recorded beside + /// it; a key saved before 0.26 recorded none and was TypeSafe's. + Saved(Provider), +} + +impl Located { + pub fn provider(&self) -> Provider { + match self { + Self::Key(credential) => credential.provider, + Self::Saved(provider) => *provider, + } + } +} + +/// Keys set in the environment, in the order a check reads them; an empty +/// variable counts as unset. +pub fn environment() -> Result> { + let mut keys = Vec::new(); + for provider in Provider::ALL { + let name = provider.service().variable; + match std::env::var(name) { + Ok(value) if value.trim().is_empty() => {} + Ok(value) => keys.push((provider, value)), + Err(std::env::VarError::NotPresent) => {} + Err(_) => bail!("{name} must contain UTF-8 text"), + } + } + Ok(keys) +} + +fn environment_source(provider: Provider) -> String { + format!("{} environment variable", provider.service().variable) +} + +/// Whether a key in the environment is read before the credential file and +/// the saved key: only TypeSafe's. Other tools read OPENROUTER_API_KEY and +/// AI_GATEWAY_API_KEY too, so one exported for them is read last: it must not +/// move a check off the key it was given, to another account's credits, +/// another data processor and a model name no cached answer was asked with. +fn read_first((provider, _): &(Provider, String)) -> bool { + *provider == Provider::Typesafe +} + +/// The first key in a check's order: TYPESAFE_API_KEY in the environment, +/// the credential file's keys, the saved key, then the gateways' variables. +/// `recorded` is the provider recorded beside the saved key; `stored` reads +/// the provider of a key saved before 0.26, which recorded none, and is +/// called only when a gateway's variable would otherwise be used. +pub fn locate( + environment: Vec<(Provider, String)>, + file: CredentialFile<'_>, + recorded: Option, + stored: impl FnOnce() -> Option, +) -> Result { + let (first, last): (Vec<_>, Vec<_>) = environment.into_iter().partition(read_first); + if let Some(found) = first.into_iter().next() { + return environment_key(found); + } + if let Some((provider, key)) = file.keys()?.into_iter().next() { + return credential(provider, key, file.source(provider)).map(Located::Key); } ensure!( - !explicit, - "Selected --env-file is missing or has no TYPESAFE_API_KEY; correct the path or run jevgate auth login without --env-file" + !file.explicit, + "Selected --env-file is missing or has no TYPESAFE_API_KEY, OPENROUTER_API_KEY or AI_GATEWAY_API_KEY; correct the path or run jevgate auth login without --env-file" ); - match saved() - .context("Saved credential unavailable; run jevgate auth login or set TYPESAFE_API_KEY")? - { - Some((key, source)) => Ok(Credential { key, source }), - None => bail!( - "No API key configured. Run jevgate auth login, set TYPESAFE_API_KEY, or provide --env-file PATH" - ), + let Some(gateway) = last.into_iter().next() else { + return Ok(Located::Saved(recorded.unwrap_or_default())); + }; + match recorded.or_else(stored) { + Some(provider) => Ok(Located::Saved(provider)), + None => environment_key(gateway), + } +} + +fn environment_key((provider, value): (Provider, String)) -> Result { + let key = Secret::parse(value)?; + credential(provider, key, environment_source(provider)).map(Located::Key) +} + +/// A key found for `provider`, unless its prefix shows another provider +/// issued it: sent to the wrong host, it would fail there and hand that host +/// the key. +fn credential(provider: Provider, key: Secret, source: String) -> Result { + refuse_foreign(provider, &key, &source)?; + Ok(Credential { + key, + provider, + source, + }) +} + +/// Fail when `key` starts as another provider's keys do; `holder` names where it was found. +pub fn refuse_foreign(provider: Provider, key: &Secret, holder: &str) -> Result<()> { + if let Some(issuer) = Provider::issuer(key.expose()).filter(|issuer| *issuer != provider) { + let issued = issuer.service(); + bail!( + "{holder}: the key was issued by {} (it starts with {}), not by {}; set it as {}, or save it with jevgate auth login --provider {}", + issued.label, + issued.key_prefix.unwrap_or_default(), + provider.service().label, + issued.variable, + issuer.name() + ); + } + Ok(()) +} + +/// The key saved by `jevgate auth login`, from the store +/// `JEVGATE_CREDENTIAL_STORE` names. +fn read_saved() -> Result> { + SavedCredentials::::native(StorageMode::configured()?)?.get() +} + +/// The provider of a key saved before 0.26, which recorded none beside it, +/// read from the credential store; none when no key is saved or the store +/// cannot be read, as in CI, where the system store needs a terminal. +fn stored_provider() -> Option { + read_saved().ok().flatten().map(|saved| saved.provider) +} + +/// The provider of the key a check will use, since a check's default model +/// depends on it: found without reading the credential store unless a +/// gateway's variable is set and no provider is recorded. TypeSafe when no +/// key is found or its source cannot be read, which then fails where it is +/// used. +pub fn planned_provider(path: &Path, explicit: bool) -> Provider { + let file = CredentialFile { path, explicit }; + environment() + .and_then(|environment| { + locate( + environment, + file, + store::recorded_provider(), + stored_provider, + ) + }) + .map_or(Provider::Typesafe, |located| located.provider()) +} + +/// The key a check uses; the saved key is read only when nothing before it +/// holds one, and at most once. +pub fn resolve(path: &Path, explicit: bool) -> Result { + let file = CredentialFile { path, explicit }; + let read = OnceCell::new(); + let stored = || { + read.get_or_init(read_saved) + .as_ref() + .ok() + .and_then(Option::as_ref) + .map(|saved| saved.provider) + }; + match locate(environment()?, file, store::recorded_provider(), stored)? { + Located::Key(credential) => Ok(credential), + Located::Saved(provider) => saved(provider, file, || { + read.into_inner().unwrap_or_else(read_saved) + }), } } -pub fn key_from_file(path: &Path) -> Result> { - match read_limited(path)? { - Some(text) => parse_key(&text), - None => Ok(None), +/// The saved key, which must be of the provider `recorded` beside it (or +/// read from it): a check planned its model for that provider. +pub fn saved( + recorded: Provider, + file: CredentialFile<'_>, + read: impl FnOnce() -> Result>, +) -> Result { + let saved = read() + .context("Saved credential unavailable; run jevgate auth login or set TYPESAFE_API_KEY")? + .with_context(|| { + format!( + "No API key configured. Run jevgate auth login, set TYPESAFE_API_KEY (or OPENROUTER_API_KEY, AI_GATEWAY_API_KEY), or provide --env-file PATH{}", + file.unread_hint() + ) + })?; + ensure!( + saved.provider == recorded, + "The saved key is for {}, but the provider recorded beside it is {}; run jevgate auth login again", + saved.provider.service().label, + recorded.service().label + ); + // A key saved before 0.26 was saved as TypeSafe's, whatever it was. + credential(saved.provider, saved.key, saved.description) +} + +/// The keys set besides the one a check uses, in the order a check reads them. +pub fn unused(path: &Path, explicit: bool) -> Result> { + let file = CredentialFile { path, explicit }; + let (first, last): (Vec<_>, Vec<_>) = environment()?.into_iter().partition(read_first); + let source = |(provider, _): (Provider, String)| environment_source(provider); + let mut found: Vec = first.into_iter().map(source).collect(); + found.extend( + file.keys()? + .into_iter() + .map(|(provider, _)| file.source(provider)), + ); + let saved = + store::recorded_provider().or_else(|| (!last.is_empty()).then(stored_provider).flatten()); + if let Some(provider) = saved { + found.push(format!( + "the {} key saved by jevgate auth login", + provider.service().label + )); } + found.extend(last.into_iter().map(source)); + Ok(found.into_iter().skip(1).collect()) +} + +/// The key that a check uses instead of the saved one, when there is one: +/// TYPESAFE_API_KEY in the environment or the repository's `.env`. A +/// gateway's variable is read after the saved key, so it overrides nothing. +pub fn override_source(path: &Path) -> Result> { + let file = CredentialFile { + path, + explicit: false, + }; + // Located as if a key were saved, so only the keys read before it count. + let as_if_saved = Some(Provider::default()); + Ok(match locate(environment()?, file, as_if_saved, || None)? { + Located::Key(credential) => Some(credential.source), + Located::Saved(_) => None, + }) } /// The file's text, at most 64 KiB and zeroed on drop; `None` when it does not exist. @@ -97,23 +336,35 @@ fn read_limited(path: &Path) -> Result>> { Ok(Some(text)) } -/// The single `TYPESAFE_API_KEY=` definition, optionally exported or quoted. -fn parse_key(text: &str) -> Result> { - let mut key = None; +/// The keys a credential file defines for `providers`, in the order a check +/// reads them: each variable once, optionally exported or quoted. An empty +/// value counts as unset, as in the environment, so a template's +/// `OPENROUTER_API_KEY=` beside a real key does not fail the file. +fn parse_keys(text: &str, providers: &[Provider]) -> Result> { + let mut keys: Vec<(Provider, Secret)> = Vec::new(); for line in text.lines() { let line = line.trim().strip_prefix("export ").unwrap_or(line.trim()); let Some((name, value)) = line.split_once('=') else { continue; }; - if name.trim() != "TYPESAFE_API_KEY" { + let Some(provider) = providers + .iter() + .copied() + .find(|provider| provider.service().variable == name.trim()) + else { + continue; + }; + let value = value.trim().trim_matches(['\'', '"']); + if value.is_empty() { continue; } ensure!( - key.is_none(), - "Credential file contains duplicate TYPESAFE_API_KEY definitions" + !keys.iter().any(|(found, _)| *found == provider), + "Credential file contains duplicate {} definitions", + provider.service().variable ); - let value = value.trim().trim_matches(['\'', '"']); - key = Some(Secret::parse(value.to_owned())?); + keys.push((provider, Secret::parse(value.to_owned())?)); } - Ok(key) + keys.sort_by_key(|(provider, _)| Provider::ALL.iter().position(|p| p == provider)); + Ok(keys) } diff --git a/src/auth/store.rs b/src/auth/store.rs index c65fef5..b91d55e 100644 --- a/src/auth/store.rs +++ b/src/auth/store.rs @@ -1,7 +1,9 @@ use super::{file, secret::Secret}; +use crate::provider::Provider; use anyhow::{Context, Result, bail, ensure}; use clap::ValueEnum; use std::{io::IsTerminal, path::PathBuf}; +use zeroize::Zeroizing; #[derive(Clone, Copy, Debug, PartialEq, Eq, ValueEnum)] pub enum StorageMode { @@ -23,9 +25,10 @@ impl StorageMode { } } +/// Where the saved credential's text lives: the OS store, or a fake in tests. pub trait Backend { - fn get(&self) -> Result>; - fn set(&self, secret: &Secret) -> Result<()>; + fn get(&self) -> Result>>; + fn set(&self, text: &str) -> Result<()>; fn delete(&self) -> Result; } @@ -43,7 +46,7 @@ impl NativeBackend { } } impl Backend for NativeBackend { - fn get(&self) -> Result> { + fn get(&self) -> Result>> { #[cfg(all( unix, not(any(target_os = "macos", target_os = "ios", target_os = "android")) @@ -54,20 +57,19 @@ impl Backend for NativeBackend { not(any(target_os = "macos", target_os = "ios", target_os = "android")) )))] match self.entry()?.get_password() { - Ok(value) => Secret::parse(value).map(Some), + Ok(value) => Ok(Some(Zeroizing::new(value))), Err(keyring::Error::NoEntry) => Ok(None), Err(_) => bail!( "Cannot read the system credential store; unlock it or run jevgate auth login" ), } } - fn set(&self, secret: &Secret) -> Result<()> { + fn set(&self, text: &str) -> Result<()> { self.entry()? - .set_password(secret.expose()) + .set_password(text) .map_err(|_| anyhow::anyhow!("Cannot save to the system credential store"))?; ensure!( - self.get()? - .is_some_and(|saved| saved.expose() == secret.expose()), + self.get()?.is_some_and(|saved| *saved == text), "System credential store did not retain the credential" ); Ok(()) @@ -92,6 +94,35 @@ impl Backend for NativeBackend { } } +/// The text a saved key is stored as: the bare key for TypeSafe, as before +/// 0.26, else the provider's name, a space and the key. Versions before 0.26 +/// refuse a value with a space as an invalid key instead of sending a +/// gateway's key to TypeSafe. +fn stored_text(provider: Provider, key: &Secret) -> Zeroizing { + Zeroizing::new(match provider { + Provider::Typesafe => key.expose().to_owned(), + gateway => format!("{} {}", gateway.name(), key.expose()), + }) +} + +/// A saved key and its provider, from the text `stored_text` wrote. +pub(super) fn saved_key(text: &str) -> Result<(Provider, Secret)> { + let text = text.trim(); + let Some((name, key)) = text.split_once(' ') else { + return Ok((Provider::Typesafe, Secret::parse(text.to_owned())?)); + }; + let provider = Provider::named(name) + .context("The saved credential names an unknown provider; run jevgate auth login")?; + Ok((provider, Secret::parse(key.to_owned())?)) +} + +/// A key saved by `jevgate auth login`, and where it was found. +pub struct SavedKey { + pub provider: Provider, + pub key: Secret, + pub description: String, +} + pub struct SavedCredentials { pub backend: B, pub path: PathBuf, @@ -115,28 +146,53 @@ impl SavedCredentials { }) } } + +/// The provider recorded beside the saved key, read without the key itself; +/// none when no key was saved with one, as before 0.26. +pub fn recorded_provider() -> Option { + file::credential_path() + .ok() + .and_then(|path| file::recorded_provider(&path)) +} + impl SavedCredentials { - pub fn get(&self) -> Result> { + pub fn get(&self) -> Result> { // A fallback written during a keyring outage must not later expose an older keyring key. if self.mode != StorageMode::Keyring - && let Some(key) = file::load(&self.path)? + && let Some(text) = file::load(&self.path)? { - return Ok(Some(( + let (provider, key) = saved_key(&text)?; + let description = format!("protected file: {}", self.path.display()); + return Ok(Some(SavedKey { + provider, key, - format!("protected file: {}", self.path.display()), - ))); + description, + })); } if self.mode == StorageMode::File { return Ok(None); } - Ok(self - .backend - .get()? - .map(|key| (key, "system credential store".into()))) + let Some(text) = self.backend.get()? else { + return Ok(None); + }; + let (provider, key) = saved_key(&text)?; + Ok(Some(SavedKey { + provider, + key, + description: "system credential store".into(), + })) } - pub fn save(&self, secret: &Secret) -> Result { + + /// Save the key with its provider, then record the provider beside it. + pub fn save(&self, provider: Provider, secret: &Secret) -> Result { + let location = self.save_text(&stored_text(provider, secret))?; + file::record_provider(&self.path, provider)?; + Ok(location) + } + + fn save_text(&self, text: &str) -> Result { if self.mode != StorageMode::File { - match self.backend.set(secret) { + match self.backend.set(text) { Ok(()) => { file::remove(&self.path).context("Key saved in system credential store, but an older fallback file could not be removed")?; return Ok(SavedLocation { @@ -148,12 +204,13 @@ impl SavedCredentials { Err(_) => (), } } - file::save(&self.path, secret)?; + file::save(&self.path, text)?; Ok(SavedLocation { description: format!("protected file: {}", self.path.display()), fallback: self.mode == StorageMode::Auto, }) } + pub fn remove(&self) -> Result { let native = if self.mode == StorageMode::File { Ok(false) @@ -161,6 +218,7 @@ impl SavedCredentials { self.backend.delete() }; let removed_file = file::remove(&self.path)?; + file::forget_provider(&self.path)?; Ok(Removal { file: removed_file, keyring: native.as_ref().copied().unwrap_or(false), diff --git a/src/auth/tests.rs b/src/auth/tests.rs index db6ef5f..88ac75c 100644 --- a/src/auth/tests.rs +++ b/src/auth/tests.rs @@ -1,8 +1,15 @@ +//! Saving, finding and checking keys: the credential store and its file +//! fallback, the order a check reads keys in, each provider's key check, and +//! the login question. use super::*; +use crate::provider::{OPENROUTER, TYPESAFE, VERCEL}; +use sources::{CredentialFile, Located}; use std::{ cell::{Cell, RefCell}, path::Path, }; +use store::SavedKey; +use zeroize::Zeroizing; #[derive(Default)] struct FakeStore { @@ -11,14 +18,14 @@ struct FakeStore { writes: Cell, } impl Backend for FakeStore { - fn get(&self) -> Result> { + fn get(&self) -> Result>> { ensure!(!self.unavailable.get(), "unavailable"); - self.secret.borrow().clone().map(Secret::parse).transpose() + Ok(self.secret.borrow().clone().map(Zeroizing::new)) } - fn set(&self, key: &Secret) -> Result<()> { + fn set(&self, text: &str) -> Result<()> { ensure!(!self.unavailable.get(), "unavailable"); self.writes.set(self.writes.get() + 1); - *self.secret.borrow_mut() = Some(key.expose().into()); + *self.secret.borrow_mut() = Some(text.into()); Ok(()) } fn delete(&self) -> Result { @@ -41,20 +48,30 @@ fn store(root: &Path, mode: StorageMode) -> SavedCredentials { } } +fn key(value: &str) -> Secret { + Secret::parse(value.into()).unwrap() +} + +/// The saved key's value. +fn saved_value(saved: &SavedCredentials) -> String { + saved.get().unwrap().unwrap().key.expose().to_owned() +} + /// A store that already holds `old-key`, and the `new-key` meant to replace it. fn replacing_old_key(root: &Path) -> (SavedCredentials, Secret) { let saved = store(root, StorageMode::Auto); *saved.backend.secret.borrow_mut() = Some("old-key".into()); - (saved, Secret::parse("new-key".into()).unwrap()) + (saved, key("new-key")) } #[test] fn validation_failure_preserves_the_previous_credential() { let project = crate::tests::Project::new(); let (saved, key) = replacing_old_key(&project.0); - assert!(validate_and_save(&Verification(false), &saved, &key).is_err()); + let provider = Provider::Typesafe; + assert!(validate_and_save(&Verification(false), &saved, provider, &key).is_err()); assert_eq!(saved.backend.writes.get(), 0); - assert_eq!(saved.get().unwrap().unwrap().0.expose(), "old-key"); + assert_eq!(saved_value(&saved), "old-key"); assert!(!saved.path.exists()); } @@ -62,16 +79,56 @@ fn validation_failure_preserves_the_previous_credential() { fn successful_login_and_logout_use_the_system_store() { let project = crate::tests::Project::new(); let saved = store(&project.0, StorageMode::Auto); - let key = Secret::parse("new-key".into()).unwrap(); - let location = validate_and_save(&Verification(true), &saved, &key).unwrap(); + let provider = Provider::Typesafe; + let location = validate_and_save(&Verification(true), &saved, provider, &key("new-key")); + let location = location.unwrap(); assert_eq!(location.description, "system credential store"); assert!(!location.fallback && !saved.path.exists()); - assert_eq!(saved.get().unwrap().unwrap().0.expose(), "new-key"); + assert_eq!(saved_value(&saved), "new-key"); let removed = saved.remove().unwrap(); assert!(removed.keyring && !removed.keyring_error); assert!(saved.get().unwrap().is_none()); } +#[test] +fn a_gateway_key_is_saved_with_its_provider_where_older_versions_refuse_it() { + let project = crate::tests::Project::new(); + let saved = store(&project.0, StorageMode::Auto); + let gateway_key = key("sk-or-v1-private"); + saved.save(Provider::Openrouter, &gateway_key).unwrap(); + let text = saved.backend.secret.borrow().clone().unwrap(); + assert_eq!(text, "openrouter sk-or-v1-private"); + assert!( + Secret::parse(text).is_err(), + "0.25 parses the stored value as a bare key" + ); + let found = saved.get().unwrap().unwrap(); + assert_eq!( + (found.provider, found.key.expose()), + (Provider::Openrouter, "sk-or-v1-private") + ); + assert_eq!( + file::recorded_provider(&saved.path), + Some(Provider::Openrouter) + ); + saved + .save(Provider::Typesafe, &key("typesafe-key")) + .unwrap(); + assert_eq!( + saved.backend.secret.borrow().as_deref(), + Some("typesafe-key") + ); + assert_eq!( + file::recorded_provider(&saved.path), + Some(Provider::Typesafe) + ); + saved.remove().unwrap(); + assert_eq!(file::recorded_provider(&saved.path), None); + for invalid in ["gateway sk-or-v1-x", "openrouter two words"] { + assert!(store::saved_key(invalid).is_err(), "{invalid}"); + } +} + #[cfg(unix)] #[test] fn fallback_is_private_survives_store_recovery_and_can_migrate_back() { @@ -79,8 +136,8 @@ fn fallback_is_private_survives_store_recovery_and_can_migrate_back() { let project = crate::tests::Project::new(); let (saved, key) = replacing_old_key(&project.0); saved.backend.unavailable.set(true); - let location = validate_and_save(&Verification(true), &saved, &key).unwrap(); - assert!(location.fallback); + let location = validate_and_save(&Verification(true), &saved, Provider::Typesafe, &key); + assert!(location.unwrap().fallback); assert_eq!( std::fs::metadata(&saved.path).unwrap().permissions().mode() & 0o777, 0o600 @@ -94,10 +151,10 @@ fn fallback_is_private_survives_store_recovery_and_can_migrate_back() { 0o700 ); saved.backend.unavailable.set(false); - assert_eq!(saved.get().unwrap().unwrap().0.expose(), "new-key"); - saved.save(&key).unwrap(); + assert_eq!(saved_value(&saved), "new-key"); + saved.save(Provider::Typesafe, &key).unwrap(); assert!(!saved.path.exists()); - assert_eq!(saved.backend.get().unwrap().unwrap().expose(), "new-key"); + assert_eq!(*saved.backend.get().unwrap().unwrap(), "new-key"); } #[cfg(unix)] @@ -106,53 +163,233 @@ fn keyring_only_mode_never_falls_back_and_partial_logout_is_reported() { let project = crate::tests::Project::new(); let mut saved = store(&project.0, StorageMode::Keyring); saved.backend.unavailable.set(true); - let key = Secret::parse("test-key".into()).unwrap(); - assert!(saved.save(&key).is_err()); + let key = key("test-key"); + assert!(saved.save(Provider::Typesafe, &key).is_err()); assert!(!saved.path.exists()); saved.mode = StorageMode::File; - saved.save(&key).unwrap(); + saved.save(Provider::Typesafe, &key).unwrap(); saved.mode = StorageMode::Auto; let removed = saved.remove().unwrap(); assert!(removed.file && removed.keyring_error); assert!(!saved.path.exists()); } +/// What `jevgate auth login` saved, as a check finds it before reading the +/// key: the provider recorded beside it, and the provider a key saved before +/// 0.26, which recorded none, turns out to have when the store is read. +#[derive(Clone, Copy, Default)] +struct Saved { + recorded: Option, + stored: Option, +} + +/// The first key a check finds with `environment` set, `file` as its +/// credential file and `saved` saved; `read` notes whether the credential +/// store was read. +fn located_with( + environment: &[(Provider, &str)], + file: CredentialFile<'_>, + saved: Saved, + read: &Cell, +) -> Result { + let environment = environment + .iter() + .map(|(provider, value)| (*provider, value.to_string())) + .collect(); + sources::locate(environment, file, saved.recorded, || { + read.set(true); + saved.stored + }) +} + +/// The first key a check finds when no key is saved. +fn located(environment: &[(Provider, &str)], file: CredentialFile<'_>) -> Result { + located_with(environment, file, Saved::default(), &Cell::new(false)) +} + +/// The found key's provider, value and source. +fn found(located: Result) -> (Provider, String, String) { + match located.unwrap() { + Located::Key(credential) => ( + credential.provider, + credential.key.expose().to_owned(), + credential.source, + ), + Located::Saved(provider) => (provider, String::new(), "saved".into()), + } +} + #[test] -fn precedence_is_environment_then_selected_file_then_saved_credentials() { +fn a_gateways_variable_is_read_after_every_key_given_to_jevgate() { let project = crate::tests::Project::new(); - let file = project.0.join(".env"); - project.write(".env", "TYPESAFE_API_KEY=repo-key\n"); - let env = sources::resolve_with(Some("environment-key".into()), &file, true, || { - panic!("must not read saved credentials") - }) - .unwrap(); - assert_eq!(env.key.expose(), "environment-key"); - let repo = sources::resolve_with(None, &file, false, || { - panic!("must not read saved credentials") - }) - .unwrap(); - assert_eq!(repo.key.expose(), "repo-key"); + let path = project.0.join(".env"); + let repository = CredentialFile { + path: &path, + explicit: false, + }; + let selected = CredentialFile { + path: &path, + explicit: true, + }; + let gateways = [ + (Provider::Openrouter, "sk-or-environment"), + (Provider::Vercel, "vck_environment"), + ]; + project.write( + ".env", + "OPENROUTER_API_KEY=sk-or-file\nTYPESAFE_API_KEY=file-key\n", + ); + let everything = [(Provider::Typesafe, "environment-key"), gateways[0]]; + let (provider, value, source) = found(located(&everything, selected)); + assert_eq!( + (provider, value.as_str(), source.as_str()), + ( + Provider::Typesafe, + "environment-key", + "TYPESAFE_API_KEY environment variable" + ) + ); + for file in [repository, selected] { + let (provider, value, _) = found(located(&gateways, file)); + assert_eq!( + (provider, value.as_str()), + (Provider::Typesafe, "file-key"), + "the credential file's TypeSafe key comes before a gateway's variable" + ); + } + project.write( + ".env", + "TYPESAFE_API_KEY=\nOPENROUTER_API_KEY=sk-or-file\nAI_GATEWAY_API_KEY=''\nUNRELATED=keep-me\n", + ); + let (provider, value, source) = found(located(&gateways[1..], selected)); + assert_eq!( + (provider, value.as_str()), + (Provider::Openrouter, "sk-or-file") + ); + assert!(source.starts_with("--env-file:") && source.ends_with("(OPENROUTER_API_KEY)")); + let (provider, value, source) = found(located(&gateways, repository)); + assert_eq!( + (provider, value.as_str(), source.as_str()), + ( + Provider::Openrouter, + "sk-or-environment", + "OPENROUTER_API_KEY environment variable" + ), + "the repository .env is read only for TYPESAFE_API_KEY, and nothing is saved" + ); project.write(".env", "UNRELATED=keep-me\n"); - let saved = sources::resolve_with(None, &file, false, || { - Ok(Some(( - Secret::parse("saved-key".into())?, - "system credential store".into(), - ))) - }) - .unwrap(); - assert_eq!(saved.key.expose(), "saved-key"); assert!( - sources::resolve_with(None, &file, true, || panic!( - "explicit missing key must fail" - )) - .is_err() + located(&gateways, selected).is_err(), + "a selected file must hold a key" ); assert_eq!( - std::fs::read_to_string(file).unwrap(), + std::fs::read_to_string(&path).unwrap(), "UNRELATED=keep-me\n" ); } +#[test] +fn the_saved_key_comes_before_a_gateways_variable_and_the_store_is_read_only_to_decide_that() { + let project = crate::tests::Project::new(); + let path = project.0.join("absent.env"); + let file = CredentialFile { + path: &path, + explicit: false, + }; + let gateway = [(Provider::Openrouter, "sk-or-environment")]; + let recorded = Saved { + recorded: Some(Provider::Vercel), + stored: None, + }; + let before_0_26 = Saved { + recorded: None, + stored: Some(Provider::Typesafe), + }; + for (saved, reads_the_store) in [(recorded, false), (before_0_26, true)] { + let read = Cell::new(false); + let (provider, _, source) = found(located_with(&gateway, file, saved, &read)); + assert_eq!(source, "saved"); + assert_eq!(Some(provider), saved.recorded.or(saved.stored)); + assert_eq!(read.get(), reads_the_store); + } + let read = Cell::new(false); + let (provider, _, source) = found(located_with(&[], file, before_0_26, &read)); + assert_eq!( + (provider, source.as_str(), read.get()), + (Provider::Typesafe, "saved", false), + "without a gateway's variable, a key saved before 0.26 is TypeSafe's, as it was then" + ); +} + +#[test] +fn a_key_issued_by_another_provider_is_refused_without_being_shown() { + let project = crate::tests::Project::new(); + let path = project.0.join("absent.env"); + let file = CredentialFile { + path: &path, + explicit: false, + }; + for (provider, value, variable) in [ + (Provider::Typesafe, "sk-or-v1-private", "OPENROUTER_API_KEY"), + (Provider::Openrouter, "vck_private", "AI_GATEWAY_API_KEY"), + ] { + let error = located(&[(provider, value)], file) + .err() + .unwrap() + .to_string(); + assert!(error.contains(variable), "{error}"); + assert!(!error.contains("private"), "{error}"); + } + let (provider, _, _) = found(located(&[(Provider::Vercel, "vck_ok")], file)); + assert_eq!(provider, Provider::Vercel); +} + +#[test] +fn the_saved_key_must_be_of_the_provider_recorded_beside_it() { + let project = crate::tests::Project::new(); + project.write(".env", "AI_GATEWAY_API_KEY=vck_app\n"); + let path = project.0.join(".env"); + let file = CredentialFile { + path: &path, + explicit: false, + }; + let openrouter = || { + Ok(Some(SavedKey { + provider: Provider::Openrouter, + key: key("sk-or-saved"), + description: "system credential store".into(), + })) + }; + let credential = sources::saved(Provider::Openrouter, file, openrouter).unwrap(); + assert_eq!(credential.key.expose(), "sk-or-saved"); + let error = sources::saved(Provider::Typesafe, file, openrouter) + .err() + .unwrap() + .to_string(); + assert!(error.contains("run jevgate auth login again"), "{error}"); + let missing = sources::saved(Provider::Typesafe, file, || Ok(None)) + .err() + .unwrap() + .to_string(); + assert!(missing.starts_with("No API key configured"), "{missing}"); + assert!( + missing.contains("AI_GATEWAY_API_KEY is read only with --env-file .env"), + "{missing}" + ); + let saved_by_0_25 = || { + Ok(Some(SavedKey { + provider: Provider::Typesafe, + key: key("sk-or-v1-private"), + description: "system credential store".into(), + })) + }; + let refused = sources::saved(Provider::Typesafe, file, saved_by_0_25) + .err() + .unwrap() + .to_string(); + assert!(refused.contains("OPENROUTER_API_KEY") && !refused.contains("private")); +} + #[test] fn invalid_or_duplicate_keys_never_fall_through_or_appear_in_errors() { let project = crate::tests::Project::new(); @@ -160,23 +397,36 @@ fn invalid_or_duplicate_keys_never_fall_through_or_appear_in_errors() { ".env", "TYPESAFE_API_KEY=private-key\nTYPESAFE_API_KEY=another-private-key\n", ); - let error = sources::key_from_file(&project.0.join(".env")) - .err() - .unwrap() - .to_string(); + let path = project.0.join(".env"); + let file = CredentialFile { + path: &path, + explicit: false, + }; + let error = located(&[], file).err().unwrap().to_string(); assert!(!error.contains("private-key")); for value in ["", "private\nkey", "private key", "private\u{7f}key"] { assert!(Secret::parse(value.into()).is_err()); } - assert!( - sources::resolve_with( - Some("bad key".into()), - &project.0.join(".env"), - false, - || panic!("invalid override must not fall through") - ) - .is_err() + let error = located(&[(Provider::Typesafe, "bad key")], file) + .err() + .unwrap() + .to_string(); + assert!(!error.contains("bad key")); +} + +#[test] +fn credential_parser_does_not_execute_shell() { + let project = crate::tests::Project::new(); + project.write( + ".env", + "export TYPESAFE_API_KEY='literal$(do-not-execute)'\n", ); + let path = project.0.join(".env"); + let file = CredentialFile { + path: &path, + explicit: false, + }; + assert_eq!(found(located(&[], file)).1, "literal$(do-not-execute)"); } #[test] @@ -189,25 +439,47 @@ fn stdin_supports_one_key_with_a_trailing_newline_and_rejects_unbounded_input() assert!(secret::read_stdin(&vec![b'x'; secret::MAX_KEY_BYTES + 1][..]).is_err()); } +#[test] +fn login_asks_for_the_kind_of_key_by_number_or_name() { + let ask = |answers: &str| { + let mut output = Vec::new(); + let chosen = ask_provider(&mut answers.as_bytes(), &mut output); + (chosen.ok(), String::from_utf8(output).unwrap()) + }; + let (chosen, prompt) = ask("\n"); + assert_eq!(chosen, Some(Provider::Typesafe)); + assert_eq!( + prompt, + "Key kind: 1 TypeSafe, 2 OpenRouter, 3 Vercel AI Gateway [1]: " + ); + assert_eq!(ask("2\n").0, Some(Provider::Openrouter)); + assert_eq!(ask(" Vercel \n").0, Some(Provider::Vercel)); + let (chosen, prompt) = ask("4\nopenrouter\n"); + assert_eq!(chosen, Some(Provider::Openrouter)); + assert!(prompt.contains("Answer 1, 2 or 3")); + assert_eq!(ask("0\nx\ny\n").0, None, "three wrong answers"); + assert_eq!(ask("").0, None, "no answer"); +} + #[cfg(unix)] #[test] fn fallback_rejects_symlinks_hardlinks_and_broad_permissions() { use std::os::unix::fs::{PermissionsExt, symlink}; let project = crate::tests::Project::new(); let saved = store(&project.0, StorageMode::File); - let key = Secret::parse("test-key".into()).unwrap(); - saved.save(&key).unwrap(); + let key = key("test-key"); + saved.save(Provider::Typesafe, &key).unwrap(); let another = project.0.join("linked-secret"); std::fs::hard_link(&saved.path, &another).unwrap(); assert!(saved.get().is_err()); - assert!(saved.save(&key).is_err()); + assert!(saved.save(Provider::Typesafe, &key).is_err()); std::fs::remove_file(another).unwrap(); std::fs::set_permissions(&saved.path, std::fs::Permissions::from_mode(0o644)).unwrap(); assert!(saved.get().is_err()); std::fs::remove_file(&saved.path).unwrap(); project.write("external", "do-not-touch"); symlink(project.0.join("external"), &saved.path).unwrap(); - assert!(saved.save(&key).is_err()); + assert!(saved.save(Provider::Typesafe, &key).is_err()); assert!(saved.remove().is_err()); assert_eq!( std::fs::read_to_string(project.0.join("external")).unwrap(), @@ -217,39 +489,59 @@ fn fallback_rejects_symlinks_hardlinks_and_broad_permissions() { #[test] fn authentication_is_known_only_after_a_checked_connection() { - assert_eq!(provider::authenticated(&Ok(()), false), None); - assert_eq!(provider::authenticated(&Ok(()), true), Some(true)); + assert_eq!(verify::authenticated(&Ok(()), false), None); + assert_eq!(verify::authenticated(&Ok(()), true), Some(true)); assert_eq!( - provider::authenticated(&provider::http_error(403), true), + verify::authenticated(&verify::http_error(&TYPESAFE, 403), true), Some(false) ); assert_eq!( - provider::authenticated(&provider::http_error(429), true), + verify::authenticated(&verify::http_error(&TYPESAFE, 429), true), None ); } #[test] -fn model_listing_errors_never_echo_provider_text() { - assert!( - provider::validate_models(&serde_json::json!({"models":[{"name":"jev-latest"}]})).is_ok() - ); - let error = provider::validate_models(&serde_json::json!({"error":"secret-do-not-echo"})) +fn each_key_check_answer_is_validated_without_echoing_provider_text() { + use crate::provider::KeyAnswer; + for (service, valid) in [ + ( + &TYPESAFE, + serde_json::json!({"models": [{"name": "jev-latest"}]}), + ), + ( + &OPENROUTER, + serde_json::json!({"data": {"label": "k", "limit": null}}), + ), + ( + &VERCEL, + serde_json::json!({"balance": "95.50", "total_used": "4.50"}), + ), + ] { + let answer = service.key_check.answer; + assert!(verify::valid_answer(service, answer, &valid).is_ok()); + let error = verify::valid_answer( + service, + answer, + &serde_json::json!({"error": "secret-do-not-echo"}), + ) .unwrap_err() .to_string(); - assert!(!error.contains("secret-do-not-echo")); + assert!(!error.contains("secret-do-not-echo")); + assert!(error.starts_with(service.label)); + } + assert_eq!(OPENROUTER.key_check.answer, KeyAnswer::Key); } #[test] fn http_errors_explain_rejection_and_retry() { + let rejected = verify::http_error(&OPENROUTER, 401) + .unwrap_err() + .to_string(); + assert!(rejected.starts_with("OpenRouter rejected this API key")); + assert!(rejected.contains(OPENROUTER.keys_page)); assert!( - provider::http_error(403) - .unwrap_err() - .to_string() - .contains("rejected") - ); - assert!( - provider::http_error(429) + verify::http_error(&TYPESAFE, 429) .unwrap_err() .to_string() .contains("not retried") diff --git a/src/auth/verify.rs b/src/auth/verify.rs new file mode 100644 index 0000000..caaa303 --- /dev/null +++ b/src/auth/verify.rs @@ -0,0 +1,111 @@ +//! Checking a key with its provider: one free request that sends no source, +//! whose answer says whether the key is valid. +use super::secret::Secret; +use crate::provider::{Endpoint, KeyAnswer, Service}; +use anyhow::{Result, bail, ensure}; +use serde_json::Value; +use std::time::Duration; + +/// How long a key check may take. +const CHECK_TIMEOUT: Duration = Duration::from_secs(15); +/// The largest key-check answer read. +const MAX_ANSWER_BYTES: u64 = 1_048_576; + +pub trait Verifier { + fn verify(&self, key: &Secret) -> Result<()>; +} + +#[derive(Debug)] +pub struct RejectedKey { + status: u16, + service: &'static Service, +} +impl std::fmt::Display for RejectedKey { + fn fmt(&self, formatter: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { + write!( + formatter, + "{} rejected this API key (HTTP {}); create a key at {} and run jevgate auth login", + self.service.label, self.status, self.service.keys_page + ) + } +} +impl std::error::Error for RejectedKey {} + +pub fn authenticated(result: &Result<()>, checked: bool) -> Option { + if !checked { + return None; + } + match result { + Ok(()) => Some(true), + Err(error) if error.is::() => Some(false), + Err(_) => None, + } +} + +impl Verifier for Endpoint { + fn verify(&self, key: &Secret) -> Result<()> { + let label = self.service.label; + let (url, answer) = self.key_check(); + let agent: ureq::Agent = ureq::Agent::config_builder() + .timeout_global(Some(CHECK_TIMEOUT)) + .max_redirects(0) + .build() + .into(); + let response = agent + .get(&url) + .header("Authorization", format!("Bearer {}", key.expose())) + .call(); + let mut response = match response { + Ok(response) => response, + Err(ureq::Error::StatusCode(code)) => return http_error(self.service, code), + Err(_) => bail!( + "Could not reach {label} or the request timed out; check the connection and retry. The credential was not verified" + ), + }; + let body: Value = response + .body_mut() + .with_config() + .limit(MAX_ANSWER_BYTES) + .read_json() + .map_err(|_| { + anyhow::anyhow!( + "{label} returned invalid or oversized JSON; credential was not verified" + ) + })?; + valid_answer(self.service, answer, &body) + } +} + +pub fn http_error(service: &'static Service, code: u16) -> Result<()> { + match code { + 401 | 403 => Err(RejectedKey { + status: code, + service, + } + .into()), + _ => bail!( + "{} returned HTTP {code}; credential was not verified and the request was not retried", + service.label + ), + } +} + +/// Whether a key check's answer has the shape a valid key gets; its text is +/// never echoed. +pub fn valid_answer(service: &Service, answer: KeyAnswer, body: &Value) -> Result<()> { + let valid = match answer { + KeyAnswer::Models => body["models"].as_array().is_some_and(|models| { + models + .iter() + .all(|model| model["name"].as_str().is_some_and(|name| !name.is_empty())) + }), + KeyAnswer::Key => body["data"].is_object(), + KeyAnswer::Credits => body["balance"].is_string() || body["balance"].is_number(), + }; + ensure!( + valid, + "{} returned an unexpected answer; credential was not verified", + service.label + ); + Ok(()) +} diff --git a/src/baseline.rs b/src/baseline.rs index a53e2af..fd89618 100644 --- a/src/baseline.rs +++ b/src/baseline.rs @@ -2,7 +2,7 @@ //! why findings were accepted, and counting those reasons per rule. use crate::{ options::Disposition, - schema::{Report, Strength}, + schema::{Report, Scope, Strength}, }; use anyhow::{Context, Result, ensure}; use serde::{Deserialize, Serialize}; @@ -76,7 +76,9 @@ pub struct Written { /// Accept every finding of the last complete check. No source is read or sent. /// With `merge`, earlier entries stay for files the check did not cover, such /// as unchanged files of a `--base` run; entries for checked or deleted files -/// are replaced by what the check found. A finding accepted before keeps its +/// are replaced by what the check found. A check of changed lines judged only +/// what its change touched, so the entries of the files it checked stay too, +/// and only deleted files' entries go. A finding accepted before keeps its /// reason; the others get `reason`. pub fn write(root: &Path, merge: bool, reason: Option) -> Result { let report = crate::storage::read_latest(root) @@ -85,56 +87,18 @@ pub fn write(root: &Path, merge: bool, reason: Option) -> Result = report - .files - .iter() - .flat_map(|file| { - // A suppressed finding is accepted where its comment is; removing - // the comment brings it back. - file.findings - .iter() - .filter(|f| f.suppressed.is_none()) - .map(|f| Accepted { - fingerprint: f.fingerprint.clone(), - rule: f.rule.clone(), - path: file.path.clone(), - line: Some(f.line), - strength: Some(f.strength), - message: f.message.clone(), - reason, - }) - }) - .collect(); + let mut findings = to_accept(&report, reason); let accepted = findings.len(); - let mut kept = 0; let previous = read_baseline(root)?; if let Some(previous) = &previous { - let reasons: BTreeMap<&str, Disposition> = previous - .findings - .iter() - .filter_map(|f| Some((f.fingerprint.as_str(), f.reason?))) - .collect(); - for finding in &mut findings { - if let Some(earlier) = reasons.get(finding.fingerprint.as_str()) { - finding.reason = Some(*earlier); - } - } - } - if merge && let Some(previous) = previous { - let covered: BTreeSet<&Path> = report - .files - .iter() - .map(|f| f.path.as_path()) - .chain(report.deleted_files.iter().map(|p| p.as_path())) - .collect(); - let earlier: Vec = previous - .findings - .into_iter() - .filter(|f| !covered.contains(f.path.as_path())) - .collect(); - kept = earlier.len(); - findings.extend(earlier); + keep_reasons(&mut findings, previous); } + let earlier = match previous { + Some(previous) if merge => uncovered(&report, previous), + _ => Vec::new(), + }; + let kept = earlier.len(); + findings.extend(earlier); findings.sort_by(|a, b| (&a.path, &a.fingerprint).cmp(&(&b.path, &b.fingerprint))); findings.dedup_by(|a, b| a.fingerprint == b.fingerprint); let path = save_baseline( @@ -152,6 +116,61 @@ pub fn write(root: &Path, merge: bool, reason: Option) -> Result) -> Vec { + report + .files + .iter() + .flat_map(|file| { + file.findings + .iter() + .filter(|f| f.suppressed.is_none()) + .map(|f| Accepted { + fingerprint: f.fingerprint.clone(), + rule: f.rule.clone(), + path: file.path.clone(), + line: Some(f.line), + strength: Some(f.strength), + message: f.message.clone(), + reason, + }) + }) + .collect() +} + +/// Give each finding accepted before the reason it was accepted with. +fn keep_reasons(findings: &mut [Accepted], previous: &Baseline) { + let reasons: BTreeMap<&str, Disposition> = previous + .findings + .iter() + .filter_map(|f| Some((f.fingerprint.as_str(), f.reason?))) + .collect(); + for finding in findings { + if let Some(earlier) = reasons.get(finding.fingerprint.as_str()) { + finding.reason = Some(*earlier); + } + } +} + +/// The earlier entries a merge keeps: those of files the check did not +/// cover. A check of changed lines covers only the files it deleted. +fn uncovered(report: &Report, previous: Baseline) -> Vec { + let whole = report.scope == Scope::WholeFiles; + let covered: BTreeSet<&Path> = report + .files + .iter() + .filter(|_| whole) + .map(|f| f.path.as_path()) + .chain(report.deleted_files.iter().map(|p| p.as_path())) + .collect(); + previous + .findings + .into_iter() + .filter(|f| !covered.contains(f.path.as_path())) + .collect() +} + fn save_baseline(root: &Path, baseline: &Baseline) -> Result { let path = root.join(BASELINE_FILE); let mut bytes = serde_json::to_vec_pretty(baseline)?; diff --git a/src/catalog.rs b/src/catalog.rs index 3432a6c..59709d1 100644 --- a/src/catalog.rs +++ b/src/catalog.rs @@ -89,7 +89,10 @@ pub fn rules() -> Vec { Rule { id: "maintainability/hardcoded-values", group: "maintainability", - default_enabled: true, + // Opt-in: 6 of its 37 labeled reviews and considers were right on + // projects JevGate was never tuned on (16%), against 47 of 85 on + // the projects it was tuned on (55%). + default_enabled: false, key: HARDCODED_VALUES, version: rule_version(HARDCODED_VALUES), scope: "application functions and module constants that use literal values other than 0, 1, 2 or one-character strings", @@ -416,37 +419,69 @@ pub fn policy() -> BTreeMap { } /// One line per rule: ID, whether it runs by default, whether it needs -/// `--include-tests`, and its question; groups and selection follow. +/// `--include-tests`, the levels that fail the default gate, how often its +/// reviews and considers were right on unseen projects, and its question; +/// what the columns mean, groups and selection follow. pub fn table() -> String { let rules = rules(); let width = rules.iter().map(|r| r.id.len()).max().unwrap_or(0); - let mut lines = vec![format!("{:width$} DEFAULT QUESTION", "RULE")]; - for rule in &rules { - let default = match (rule.default_enabled, rule.requires_tests) { - (false, _) => "opt-in", - (true, true) => "tests", - (true, false) => "yes", - }; - lines.push(format!( - "{:width$} {default:7} {}", - rule.id, rule.inspection - )); - } + let mut lines = vec![format!( + "{:width$} DEFAULT BLOCKS REVIEWS RIGHT CONSIDERS RIGHT QUESTION", + "RULE" + )]; + lines.extend(rules.iter().map(|rule| table_row(rule, width))); lines.push(String::new()); + lines.push(format!( + "BLOCKS: the levels that fail the check by default, right at least {}% of the time over at least {} labeled findings on projects JevGate was never tuned on; an opt-in rule's levels fail it once the rule is selected. The rest are reported without failing it until they measure up. REVIEWS RIGHT and CONSIDERS RIGHT: the share of labeled findings right on those projects, a debatable one counting as not right; tests/laws is labeled only on Bend 2 projects, which these numbers leave out.", + crate::maturity::MIN_PERCENT_RIGHT, + crate::maturity::MIN_LABELS + )); lines.push(format!( "Groups: {}, {DEFAULT_GROUP} (every rule marked yes or tests), {ALL_GROUP}.", groups().join(", ") )); - lines.push("Select with --rule and --skip-rule, or [rules] in jevgate.toml; `tests` rules need --include-tests.".into()); + lines.push("Select with --rule and --skip-rule, or [rules] in jevgate.toml; `tests` rules need --include-tests. --fail-on and [rules] levels replace the default gate.".into()); lines.join("\n") } +fn table_row(rule: &Rule, width: usize) -> String { + use crate::{maturity, schema::Strength}; + let default = match (rule.default_enabled, rule.requires_tests) { + (false, _) => "opt-in", + (true, true) => "tests", + (true, false) => "yes", + }; + let mature: Vec = maturity::mature_levels(rule.key) + .iter() + .map(crate::output::label) + .collect(); + let blocks = if mature.is_empty() { + "-".to_string() + } else { + mature.join(", ") + }; + let right = |level| { + maturity::measure(rule.key, level) + .and_then(|m| m.unseen.summary()) + .unwrap_or_else(|| "-".into()) + }; + format!( + "{:width$} {default:7} {blocks:8} {:13} {:15} {}", + rule.id, + right(Strength::Review), + right(Strength::Consider), + rule.inspection + ) +} + pub fn describe() -> Value { Value::Array( rules() .into_iter() .map(|r| { + let maturity = crate::maturity::describe(r.key); let mut value = serde_json::to_value(r).unwrap(); + value["maturity"] = maturity; value["decision_policy"] = serde_json::json!(policy()); value }) diff --git a/src/changes.rs b/src/changes.rs index 0ddf4e8..6798cd8 100644 --- a/src/changes.rs +++ b/src/changes.rs @@ -1,5 +1,5 @@ //! Finding lineage between snapshots, by fingerprint: introduced, persistent or resolved. -use crate::schema::{Change, Report, Status}; +use crate::schema::{Change, Report, Scope, Status}; use std::{ collections::{BTreeMap, BTreeSet}, path::Path, @@ -75,15 +75,20 @@ fn lineage(previous: &Report, report: &Report) -> Vec { if current.contains(fingerprint) { continue; } - let (state, reason) = if comparable && judged.contains(path) { + let (state, reason) = if !comparable || !judged.contains(path) { ( - "resolved", - "The finding no longer triggers; correctness is not certified", + "non-comparable", + "The file was not judged in this snapshot, or rubric or model changed", ) - } else { + } else if report.scope == Scope::ChangedLines { ( "non-comparable", - "The file was not judged in this snapshot, or rubric or model changed", + "Only what the change touched was judged in this snapshot", + ) + } else { + ( + "resolved", + "The finding no longer triggers; correctness is not certified", ) }; changes.push(change(rule, path, fingerprint, generation, state, reason)); @@ -158,6 +163,7 @@ mod tests { rank: 1.0, baselined: false, suppressed: None, + gate: None, } } diff --git a/src/check.rs b/src/check.rs index 4a228fa..08960c9 100644 --- a/src/check.rs +++ b/src/check.rs @@ -30,7 +30,7 @@ fn validate(args: &CheckArgs) -> Result<()> { } /// The credential file: `--env-file` from the invocation directory, else the root `.env`. -fn credential_path(args: &CheckArgs, context: &ConfigContext) -> std::path::PathBuf { +pub(crate) fn credential_path(args: &CheckArgs, context: &ConfigContext) -> std::path::PathBuf { args.env_file .as_ref() .map(|p| context.input_path(p)) @@ -81,16 +81,18 @@ pub fn run(args: &CheckArgs, context: &ConfigContext) -> Result { return Ok(0); } let store = store.unwrap(); - let mut client = - transport::Client::new(&credential_path(args, context), args.env_file.is_some()); + let mut client = transport::Client::new( + &credential_path(args, context), + args.env_file.is_some(), + args.provider, + )?; let mut session = evaluate::Session { args, context, store: &store, evaluator: &mut client, requests: 0, - paid_input_tokens: 0, - paid_output_tokens: 0, + paid: Default::default(), budget: token_budget::TokenBudget::load(&context.root), observed: (0, 0), }; diff --git a/src/command.rs b/src/command.rs index 9f14700..57a5437 100644 --- a/src/command.rs +++ b/src/command.rs @@ -51,6 +51,10 @@ fn configured(command: JevCommand) -> Result { } JevCommand::Check(mut args) => { context.configure(&mut args)?; + args.provider = crate::auth::sources::planned_provider( + &crate::check::credential_path(&args, &context), + args.env_file.is_some(), + ); if let Some(base) = &args.base { args.base = Some(revision::resolve(&context.root, base)?); } diff --git a/src/config.rs b/src/config.rs index 2e64abc..33861a0 100644 --- a/src/config.rs +++ b/src/config.rs @@ -27,17 +27,17 @@ pub struct Config { pub rules: Rules, /// Ceiling on API attempts per invocation; flags can only lower it. Default: unlimited. pub max_requests: Option, - /// Ceiling on simultaneous requests (1-8). Default: 6. + /// Most simultaneous requests; flags can only lower it. JevGate sends at most 6 at once, so a higher value means 6. Default: 6 with a TypeSafe key, 3 with an OpenRouter or Vercel AI Gateway key. pub concurrency: Option, /// Files larger than this are reported as needs-context, never truncated. Default: 262144. pub max_file_bytes: Option, /// Ceiling on context bytes per request. Default: 32768. pub max_context_bytes: Option, - /// The level for rules without their own, like `--fail-on`. Default: ["review"]. + /// The level for rules without their own, like `--fail-on`. Default: ["mature"], which fails only on the levels of a rule measured right at least 80% of the time on projects JevGate was never tuned on; `jevgate rules` shows them. pub fail_on: Vec, - /// TypeSafe model; a pinned version keeps results repeatable. `--model` overrides it. + /// Model, as the key's provider names it; a pinned version keeps results repeatable. `--model` overrides it. Default: `jev-1.13.0` with a TypeSafe key, `typesafe/jev-1.13` with an OpenRouter key, `typesafe-ai/jev` with a Vercel AI Gateway key. pub model: Option, - /// Cache lifetime in seconds for the `jev-latest` and `jev-preview` aliases; pinned versions never expire. Default: 3600. + /// Cache lifetime in seconds for an alias, a model name without an x.y.z version such as `jev-latest`; pinned versions never expire. Default: 3600. pub cache_ttl_secs: Option, /// Judge tests, like `--include-tests`. Default: false. pub include_tests: bool, @@ -104,7 +104,7 @@ impl Level { .into_iter() .map(|name| { FailOn::parse(name).ok_or_else(|| { - anyhow!("Unknown level {name:?} for {target}; use review, consider, uncertain, report or off") + anyhow!("Unknown level {name:?} for {target}; use review, consider, mature, uncertain, report or off") }) }) .collect() @@ -132,7 +132,12 @@ impl ConfigContext { let config = if required || file.exists() { let text = std::fs::read_to_string(&file) .with_context(|| format!("Cannot read {}", file.display()))?; - toml::from_str(&text).with_context(|| format!("Invalid {}", file.display()))? + let config = + toml::from_str(&text).with_context(|| format!("Invalid {}", file.display()))?; + if let Some(notice) = written_before_mature(&text) { + note!("jevgate: {notice}"); + } + config } else { Config::default() }; @@ -202,7 +207,7 @@ impl ConfigContext { /// Each enabled rule's gate levels. The command line wins over the file; /// within each, a rule's own entry wins over its group's, then over the - /// levels for every rule, then `review`. + /// levels for every rule, then `mature`. fn configure_gate(&self, args: &mut CheckArgs) -> Result<()> { let cli = Levels::from_cli(&args.fail_on_specs)?; let file = self.file_levels()?; @@ -210,7 +215,7 @@ impl ConfigContext { .into_iter() .find(|levels| !levels.is_empty()) .cloned() - .unwrap_or_else(|| vec![FailOn::Review]); + .unwrap_or_else(|| vec![FailOn::Mature]); args.fail_on = fallback.clone(); args.rule_fail_on.clear(); for rule in catalog::rules() { @@ -280,13 +285,11 @@ impl ConfigContext { args.max_requests = Some(args.max_requests.map_or(n, |limit| limit.min(n))); } if let Some(n) = self.config.concurrency { - ensure!( - (1..=crate::options::MAX_CONCURRENCY).contains(&n), - "Concurrency must be between 1 and {}", - crate::options::MAX_CONCURRENCY - ); - args.concurrency = args.concurrency.min(n); + ensure!(n > 0, "Concurrency must be at least 1"); + let most = crate::options::MAX_CONCURRENCY; + args.concurrency = Some(args.concurrency.map_or(n.min(most), |flag| flag.min(n))); } + cap_concurrency(args); if let Some(n) = self.config.max_file_bytes { args.max_file_bytes = args.max_file_bytes.min(n); } @@ -301,6 +304,22 @@ impl ConfigContext { } } +/// Lower a `--concurrency` above [`MAX_CONCURRENCY`] to it, saying so on +/// stderr: 0.25 accepted it up to 8, and a script valid then keeps working. +/// A higher `concurrency` in jevgate.toml, which 0.25 accepted too, is +/// lowered without a word when the file is read. +/// +/// [`MAX_CONCURRENCY`]: crate::options::MAX_CONCURRENCY +fn cap_concurrency(args: &mut CheckArgs) { + let most = crate::options::MAX_CONCURRENCY; + if let Some(asked) = args.concurrency.filter(|n| *n > most) { + note!( + "jevgate: concurrency {asked} lowered to {most}, the most requests JevGate sends at once" + ); + args.concurrency = Some(most); + } +} + /// The levels one `[[scope]]` sets. `off` is not a gate level there: a rule /// is judged for every file or none, and `upload_deny` keeps files out. fn scope_levels(scope: &Scope) -> Result { @@ -382,6 +401,49 @@ fn most_specific<'a, T>(entries: &'a BTreeMap, rule: &catalog::Rule) .map(|(_, value)| value) } +/// The groups `jevgate init` gave a `review` level before 0.26: every group +/// whose rules all ran by default. +const INIT_REVIEW_GROUPS: [&str; 2] = ["maintainability", "tests"]; + +/// What to tell a user whose jevgate.toml still holds the `[rules]` lines +/// `jevgate init` wrote before 0.26, such as `maintainability = "review" # +/// file-organization, …`. They keep every review of those groups failing +/// the check and judge hardcoded values, which 0.26's measured default gate +/// and rules leave out, and most configurations were written that way. The +/// comment listing the group's rules tells them from a level set by hand, +/// so deleting it keeps the level without the notice. +fn written_before_mature(text: &str) -> Option { + let groups: Vec<&str> = INIT_REVIEW_GROUPS + .into_iter() + .filter(|group| { + text.lines().any(|line| { + line.trim() + .strip_prefix(group) + .is_some_and(|rest| rest.starts_with(" = \"review\" # ")) + }) + }) + .collect(); + let (them, their) = match groups.len() { + 0 => return None, + 1 => ("it", "its comment"), + _ => ("them", "their comments"), + }; + let lines: Vec = groups + .iter() + .map(|group| format!("`{group} = \"review\"`")) + .collect(); + let hardcoded = if groups.contains(&"maintainability") { + ", and hardcoded values is judged" + } else { + "" + }; + Some(format!( + "jevgate.toml keeps {} as `jevgate init` wrote {them} before 0.26: every {} review fails the check{hardcoded}. Delete {them} for the default rules and gate, which fails only on rule levels measured right at least 80% of the time; to keep {them}, delete {their} and this notice stops.", + lines.join(" and "), + groups.join(" and "), + )) +} + pub fn repository_root(invocation_dir: &Path) -> PathBuf { invocation_dir .ancestors() @@ -399,16 +461,21 @@ mod tests { use super::*; use crate::options::FailOnSpec; + /// The configuration `toml_text` holds, in the working directory. + fn context(toml_text: &str) -> Result { + Ok(ConfigContext { + invocation_dir: PathBuf::from("."), + root: PathBuf::from("."), + config: toml::from_str(toml_text)?, + }) + } + fn configured( toml_text: &str, rules: &[&str], specs: &[(Option<&str>, FailOn)], ) -> Result { - let context = ConfigContext { - invocation_dir: PathBuf::from("."), - root: PathBuf::from("."), - config: toml::from_str(toml_text)?, - }; + let context = context(toml_text)?; let mut args = crate::tests::args(); args.rules = rules.iter().map(|r| r.to_string()).collect(); args.fail_on.clear(); @@ -440,11 +507,42 @@ mod tests { assert!(error.to_string().contains("`jevgate rules`"), "{error}"); } + #[test] + fn the_levels_init_wrote_before_0_26_are_named_until_their_comments_go() { + // What `jevgate init` wrote from 0.3 to 0.25, less its comments. + let before = "[rules]\n\ + maintainability = \"review\" # file-organization, function-simplification, hardcoded-values, shared-logic\n\ + tests = \"review\" # test-value, test-redundancy\n\ + # security = \"consider\" # injection, sensitive-data (opt-in)\n"; + let notice = written_before_mature(before).unwrap(); + assert!( + notice.starts_with( + "jevgate.toml keeps `maintainability = \"review\"` and `tests = \"review\"` as `jevgate init` wrote them before 0.26: every maintainability and tests review fails the check, and hardcoded values is judged. Delete them" + ), + "{notice}" + ); + let tests_only = written_before_mature("tests = \"review\" # test-value\n").unwrap(); + assert!( + tests_only.contains("every tests review fails the check. Delete it"), + "{tests_only}" + ); + let project = crate::tests::Project::new(); + let (written, _) = crate::init::run(&project.0, false).unwrap(); + for kept in [ + "[rules]\nmaintainability = \"review\"\ntests = \"review\"\n".to_string(), + "[rules]\nmaintainability = \"consider\" # file-organization\n".to_string(), + std::fs::read_to_string(written).unwrap(), + ] { + assert_eq!(written_before_mature(&kept), None, "{kept}"); + } + } + #[test] fn default_group_runs_when_nothing_is_configured() { let args = configured("", &[], &[]).unwrap(); assert_eq!(args.rules, catalog::select(catalog::DEFAULT_GROUP).unwrap()); - assert_eq!(args.fail_on, [FailOn::Review]); + assert!(!args.rules.iter().any(|r| r == catalog::HARDCODED_VALUES)); + assert_eq!(args.fail_on, [FailOn::Mature]); assert!(args.rule_fail_on.is_empty()); } @@ -544,6 +642,49 @@ mod tests { } } + #[test] + fn mature_is_the_default_and_a_level_like_the_others() { + let args = configured("", &["default", "documentation"], &[]).unwrap(); + assert_eq!(args.levels(catalog::COMMENTS), [FailOn::Mature]); + let names = |pairs: &[(&str, &str)]| -> BTreeMap> { + pairs + .iter() + .map(|(rule, level)| (rule.to_string(), vec![level.to_string()])) + .collect() + }; + assert_eq!( + args.mature_level_names(), + names(&[ + ("maintainability/function-simplification", "review"), + ("documentation/agent-context", "consider") + ]) + ); + let file = r#" + fail_on = ["consider"] + [rules] + security = "mature" + [[scope]] + paths = ["scripts/**"] + fail_on = ["mature", "uncertain"] + "#; + let args = configured(file, &["default", "security"], &[]).unwrap(); + assert_eq!(args.levels(catalog::SHARED_LOGIC), [FailOn::Consider]); + assert_eq!(args.levels(catalog::INJECTION), [FailOn::Mature]); + assert_eq!( + args.levels_at(catalog::SHARED_LOGIC, Path::new("scripts/a.py")), + [FailOn::Mature, FailOn::Uncertain] + ); + assert_eq!( + args.mature_level_names(), + names(&[("maintainability/function-simplification", "review")]), + "mature in a scope; injection has no mature level" + ); + let flagged = configured(file, &["default"], &[(None, FailOn::Mature)]).unwrap(); + assert_eq!(flagged.levels(catalog::SHARED_LOGIC), [FailOn::Mature]); + let explicit = configured("fail_on = [\"review\"]", &[], &[]).unwrap(); + assert!(explicit.mature_level_names().is_empty()); + } + #[test] fn file_settings_apply_unless_a_flag_sets_them() { let file = "model = \"jev-latest\"\ncache_ttl_secs = 60\ninclude_tests = true\n"; @@ -552,15 +693,10 @@ mod tests { (args.model(), args.cache_ttl_secs(), args.include_tests), ("jev-latest", 60, true) ); - let context = ConfigContext { - invocation_dir: PathBuf::from("."), - root: PathBuf::from("."), - config: toml::from_str(file).unwrap(), - }; let mut args = crate::tests::args(); args.model = Some("jev-preview".into()); args.cache_ttl_secs = Some(5); - context.configure(&mut args).unwrap(); + context(file).unwrap().configure(&mut args).unwrap(); assert_eq!((args.model(), args.cache_ttl_secs()), ("jev-preview", 5)); let defaults = configured("", &[], &[]).unwrap(); assert_eq!(defaults.model(), crate::options::DEFAULT_MODEL); @@ -579,4 +715,37 @@ mod tests { assert!(configured("[rules]\nmaintainability = \"sometimes\"\n", &[], &[]).is_err()); assert!(configured("[rules]\nnothing = \"review\"\n", &[], &[]).is_err()); } + + #[test] + fn concurrency_follows_the_key_unless_set_and_is_at_most_six() { + use crate::provider::Provider; + // As a check runs: the configuration first, then the key's provider. + let concurrency = |file: &str, flag: Option, provider| -> Result { + let mut args = crate::tests::args(); + args.concurrency = flag; + context(file)?.configure(&mut args)?; + args.provider = provider; + Ok(args.concurrency()) + }; + for (provider, default) in [ + (Provider::Typesafe, 6), + (Provider::Openrouter, 3), + (Provider::Vercel, 3), + ] { + let set = |file, flag| concurrency(file, flag, provider).unwrap(); + assert_eq!(set("", None), default, "{provider:?}"); + assert_eq!(set("concurrency = 5", None), 5, "the file sets it"); + assert_eq!(set("", Some(4)), 4, "the flag sets it"); + assert_eq!(set("concurrency = 2", Some(5)), 2, "the file caps the flag"); + for valid_in_0_25 in ["concurrency = 7", "concurrency = 8"] { + assert_eq!(set(valid_in_0_25, None), 6, "{valid_in_0_25} means 6"); + } + assert_eq!( + set("concurrency = 8", Some(8)), + 6, + "--concurrency 8 is lowered" + ); + } + assert!(concurrency("concurrency = 0", None, Provider::Typesafe).is_err()); + } } diff --git a/src/config_schema.rs b/src/config_schema.rs index 59a504c..0d5ae85 100644 --- a/src/config_schema.rs +++ b/src/config_schema.rs @@ -7,7 +7,14 @@ use serde_json::{Value, json}; const ID: &str = "https://raw.githubusercontent.com/Tech-Byte-Frontier/jevgate/main/jevgate.schema.json"; /// Levels `fail_on` accepts; `[rules]` also accepts `off`. -const LEVELS: [&str; 5] = ["review", "consider", "uncertain", "report", "none"]; +const LEVELS: [&str; 6] = [ + "review", + "consider", + "mature", + "uncertain", + "report", + "none", +]; pub fn schema() -> Value { let mut schema = serde_json::to_value(schemars::schema_for!(Config)).expect("a schema is JSON"); diff --git a/src/docs/history.rs b/src/docs/history.rs index e45f759..eda4b72 100644 --- a/src/docs/history.rs +++ b/src/docs/history.rs @@ -1,6 +1,6 @@ //! What Git holds about a repository's paths: tracked files, release tags, //! and paths since deleted or renamed. Evidence for staleness candidates. -use crate::revision::git; +use crate::revision::{git, git_in}; use std::{ collections::{BTreeMap, BTreeSet}, path::{Path, PathBuf}, @@ -29,11 +29,8 @@ pub fn ignored(root: &Path, paths: &[PathBuf]) -> BTreeSet { if input.is_empty() { return BTreeSet::new(); } - let child = std::process::Command::new("git") - .arg("-C") - .arg(root) + let child = git_in(root) .args(["check-ignore", "--no-index", "--stdin", "-z"]) - .env("GIT_OPTIONAL_LOCKS", "0") .stdin(std::process::Stdio::piped()) .stdout(std::process::Stdio::piped()) .stderr(std::process::Stdio::null()) diff --git a/src/docs/references.rs b/src/docs/references.rs index 6e476ae..87f97eb 100644 --- a/src/docs/references.rs +++ b/src/docs/references.rs @@ -98,6 +98,16 @@ const BUILTIN: &[&str] = &[ "recursive", ]; +/// Whether `name`, as a document in `base` writes it, is one of `paths`: +/// read from the repository root or from the document's directory, as the +/// missing names are. +pub fn names_one_of(base: &Path, name: &str, paths: &BTreeSet) -> bool { + let trimmed = name.trim_end_matches('/'); + [normal(Path::new(trimmed)), normal(&base.join(trimmed))] + .iter() + .any(|path| paths.contains(path)) +} + /// Everything `text`, a section of `doc`, names that the repository lacks. pub fn missing( root: &Path, @@ -767,14 +777,7 @@ mod tests { fn ignored_paths_are_local_files() { let project = crate::tests::Project::new(); project.write(".gitignore", "out/\n*.local\nweb/dist/\n"); - let git = |args: &[&str]| { - std::process::Command::new("git") - .args(args) - .current_dir(&*project.0) - .output() - .unwrap() - }; - assert!(git(&["init", "-q"]).status.success()); + project.git(&["init", "-q"]); // `web/` is tracked, so its missing entries are candidates. let history = History { tracked: [PathBuf::from("web/index.html")].into(), diff --git a/src/evaluate.rs b/src/evaluate.rs index 994b4dd..00d3d62 100644 --- a/src/evaluate.rs +++ b/src/evaluate.rs @@ -17,8 +17,8 @@ pub struct Session<'a> { pub store: &'a Store, pub evaluator: &'a mut dyn Evaluator, pub requests: u32, - pub paid_input_tokens: u64, - pub paid_output_tokens: u64, + /// What this invocation's requests were billed. + pub paid: crate::requests::Usage, pub budget: TokenBudget, /// Uploaded bytes and billed input tokens of fresh requests, for calibration. pub observed: (u64, u64), @@ -85,6 +85,11 @@ fn empty_report(args: &CheckArgs, current: &SnapshotContext<'_>, files: Vec, files: Vec, files: Vec { fn progress(&self, report: &mut Report) -> Result<()> { report.api_requests = self.requests; - report.paid_input_tokens = self.paid_input_tokens; - report.paid_output_tokens = self.paid_output_tokens; + report.paid_input_tokens = self.paid.input_tokens; + report.paid_output_tokens = self.paid.output_tokens; + report.paid_models = self.paid.models.clone(); + report.unmetered_requests = self.paid.unmetered; + report.estimated_usd = self.paid.usd(); self.verify_current(report); report.update_status(); self.publish(report) diff --git a/src/gate.rs b/src/gate.rs index 860f24e..3e37c7b 100644 --- a/src/gate.rs +++ b/src/gate.rs @@ -1,8 +1,9 @@ //! The configurable quality gate: which results fail a check. Accepted -//! findings are in `baseline`. Classification never depends on this policy. +//! findings are in `baseline`, and which rules and levels fail by default in +//! `maturity`. Classification never depends on this policy. use crate::{ options::{CheckArgs, FailOn}, - schema::{Finding, Report, Status, Strength}, + schema::{Finding, Gating, Report, Status, Strength}, }; use anyhow::Result; use serde::{Deserialize, Serialize}; @@ -29,19 +30,25 @@ pub fn exit_code(report: &Report) -> u8 { } } +/// Record how the gate counts each finding, then decide it for a complete run. pub fn evaluate(report: &mut Report, args: &CheckArgs) { + for file in &mut report.files { + for finding in &mut file.findings { + finding.gate = gating(finding, &file.path, args); + } + } // Notes are optional improvements; no gate counts them. let findings = report .files .iter() - .flat_map(|f| f.findings.iter().map(|finding| (f.path.as_path(), finding))) - .filter(|(_, f)| f.strength != Strength::Note); - let baselined = findings.clone().filter(|(_, f)| f.baselined).count(); + .flat_map(|f| &f.findings) + .filter(|f| f.strength != Strength::Note); + let baselined = findings.clone().filter(|f| f.baselined).count(); let suppressed = findings .clone() - .filter(|(_, f)| !f.baselined && f.suppressed.is_some()) + .filter(|f| !f.baselined && f.suppressed.is_some()) .count(); - let new: Vec<_> = findings.filter(|(_, f)| !f.accepted()).collect(); + let new: Vec<_> = findings.filter(|f| !f.accepted()).collect(); let reasons = failures(report, &new, args); report.gate = report.complete.then_some(Gate { passed: reasons.is_empty(), @@ -52,28 +59,36 @@ pub fn evaluate(report: &mut Report, args: &CheckArgs) { }); } -/// Whether a finding in `path` fails the gate: new, not a note, and at its -/// rule's level for that path. Consider is the lower bar, so it also fails -/// on review findings. -pub fn fails(finding: &Finding, path: &Path, args: &CheckArgs) -> bool { +/// How the gate counts a finding in `path`, at its rule's levels for that +/// path: none for notes and accepted findings. Consider counts every finding +/// and review only reviews; `mature` counts the rule's mature levels, and a +/// finding it leaves out is still being measured. +fn gating(finding: &Finding, path: &Path, args: &CheckArgs) -> Option { + if finding.accepted() || finding.strength == Strength::Note { + return None; + } let levels = args.levels_at(&finding.rule, path); - !finding.accepted() - && finding.strength != Strength::Note - && (levels.contains(&FailOn::Consider) - || (finding.strength == Strength::Review && levels.contains(&FailOn::Review))) + let mature = levels.contains(&FailOn::Mature); + let counted = levels.contains(&FailOn::Consider) + || (finding.strength == Strength::Review && levels.contains(&FailOn::Review)) + || (mature && crate::maturity::mature(&finding.rule, finding.strength)); + Some(if counted { + Gating::Fails + } else if mature { + Gating::Measuring + } else { + Gating::Advisory + }) } /// Why the gate fails: new findings at their rule's level, or undecided /// results of a rule whose level includes `uncertain`. -fn failures(report: &Report, new: &[(&Path, &Finding)], args: &CheckArgs) -> Vec { +fn failures(report: &Report, new: &[&Finding], args: &CheckArgs) -> Vec { let mut reasons = Vec::new(); - let failing: Vec<_> = new - .iter() - .filter(|(path, f)| fails(f, path, args)) - .collect(); + let failing: Vec<_> = new.iter().filter(|f| f.fails_gate()).collect(); let review = failing .iter() - .filter(|(_, f)| f.strength == Strength::Review) + .filter(|f| f.strength == Strength::Review) .count(); let consider = failing.len() - review; if review > 0 { diff --git a/src/github.rs b/src/github.rs index f0fd098..77bb0b1 100644 --- a/src/github.rs +++ b/src/github.rs @@ -12,8 +12,9 @@ use std::{io::Write, path::Path}; const SUMMARY_ROWS: usize = 50; /// Annotations for run errors, failed files and every new finding that is not -/// a note (an error when it fails the gate, else a warning), then the agent -/// text. The summary goes to `$GITHUB_STEP_SUMMARY` when the runner sets it. +/// a note (an error when it fails the gate, else a warning, which says so when +/// its rule and level are still being measured), then the agent text. The +/// summary goes to `$GITHUB_STEP_SUMMARY` when the runner sets it. pub fn emit(out: &mut impl Write, report: &Report, args: &CheckArgs) -> Result<()> { for error in &report.errors { writeln!(out, "::error title=JevGate run incomplete::{}", data(error))?; @@ -27,23 +28,19 @@ pub fn emit(out: &mut impl Write, report: &Report, args: &CheckArgs) -> Result<( data(error) )?; } - let shown: Vec<(&Path, &Finding)> = output::ranked(report) + let shown: Vec<(&Path, &Finding)> = output::failing_first(report) .into_iter() .filter(|(_, f)| f.strength != Strength::Note && !f.accepted()) .collect(); for (path, finding) in &shown { - writeln!( - out, - "{}", - annotation(path, finding, crate::gate::fails(finding, path, args)) - )?; + writeln!(out, "{}", annotation(path, finding))?; } if let Some(file) = std::env::var_os("GITHUB_STEP_SUMMARY") { let written = std::fs::OpenOptions::new() .append(true) .create(true) .open(&file) - .and_then(|mut f| f.write_all(summary(report, &shown, args).as_bytes())); + .and_then(|mut f| f.write_all(summary(report, &shown).as_bytes())); if let Err(error) = written { note!("jevgate: cannot write the job summary: {error}"); } @@ -51,19 +48,27 @@ pub fn emit(out: &mut impl Write, report: &Report, args: &CheckArgs) -> Result<( output::agent(out, report, args.verbose, output::Style::PLAIN) } -fn annotation(path: &Path, finding: &Finding, fails: bool) -> String { +fn annotation(path: &Path, finding: &Finding) -> String { let end = finding .locations .iter() .find(|l| l.path == path && l.start_line == finding.line) .map_or(String::new(), |l| format!(",endLine={}", l.end_line)); + let mut message = format!("{}\n→ {}", finding.message, finding.action); + if let Some(note) = output::measuring_note(finding) { + message.push_str(&format!("\n{note}")); + } format!( "::{} file={},line={}{end},title={}::{}", - if fails { "error" } else { "warning" }, + if finding.fails_gate() { + "error" + } else { + "warning" + }, property(&path.to_string_lossy()), finding.line, property(&format!("JevGate {} [{}]", label(finding), finding.rule)), - data(&format!("{}\n→ {}", finding.message, finding.action)), + data(&message), ) } @@ -71,8 +76,9 @@ fn label(finding: &Finding) -> String { output::label(&finding.strength) } -/// The Markdown job summary: the headline, then a table of findings. -fn summary(report: &Report, shown: &[(&Path, &Finding)], args: &CheckArgs) -> String { +/// The Markdown job summary: the headline, then a table of findings, with +/// the ones that fail the gate in bold. +fn summary(report: &Report, shown: &[(&Path, &Finding)]) -> String { let mut text = format!("### {}\n\n", output::headline(report)); for error in &report.errors { text.push_str(&format!("- **Error:** {}\n", cell(error))); @@ -94,7 +100,7 @@ fn summary(report: &Report, shown: &[(&Path, &Finding)], args: &CheckArgs) -> St } text.push_str("| | Location | Rule | Finding |\n|---|---|---|---|\n"); for (path, finding) in shown.iter().take(SUMMARY_ROWS) { - let level = if crate::gate::fails(finding, path, args) { + let level = if finding.fails_gate() { format!("**{}**", label(finding)) } else { label(finding) @@ -114,6 +120,9 @@ fn summary(report: &Report, shown: &[(&Path, &Finding)], args: &CheckArgs) -> St shown.len() - SUMMARY_ROWS )); } + if let Some(line) = output::measuring(report) { + text.push_str(&format!("\n{line}\n")); + } text.push('\n'); text } @@ -138,25 +147,96 @@ fn cell(text: &str) -> String { #[cfg(test)] mod tests { use super::*; - use crate::tests::finding; + use crate::{schema::Gating, tests::finding}; + + fn counted(strength: Strength, gate: Gating) -> Finding { + Finding { + gate: Some(gate), + ..finding(strength) + } + } + + #[test] + fn a_review_that_does_not_fail_is_annotated_before_higher_ranked_considers() { + use crate::tests::finding_of; + let ranked = |rule: &str, strength: Strength, rank: f64| Finding { + rank, + ..finding_of(rule, strength) + }; + let considers = + (0..11).map(|_| ranked("maintainability/shared-logic", Strength::Consider, 0.9)); + let measured = ranked("maintainability/shared-logic", Strength::Review, 0.5); + let failing = ranked( + "maintainability/function-simplification", + Strength::Review, + 0.1, + ); + let report = crate::tests::gated( + considers.chain([measured, failing]).collect(), + &crate::tests::args(), + ); + let order: Vec<(Strength, bool)> = output::failing_first(&report) + .iter() + .map(|(_, f)| (f.strength, f.fails_gate())) + .collect(); + assert_eq!( + order[..3], + [ + (Strength::Review, true), + (Strength::Review, false), + (Strength::Consider, false) + ], + "the review still being measured is the first warning, not the twelfth" + ); + } #[test] fn annotations_escape_commands_and_mark_what_fails_the_gate() { - let line = annotation(Path::new("src/a,b.rs"), &finding(Strength::Review), true); + let review = counted(Strength::Review, Gating::Fails); + let line = annotation(Path::new("src/a,b.rs"), &review); assert_eq!( line, "::error file=src/a%2Cb.rs,line=12,endLine=20,title=JevGate review [maintainability/shared-logic]::Copies: 50%25 alike,%0Asee `b`%0A→ Share one | implementation" ); assert!(!line.contains('\n')); - let consider = annotation(Path::new("x.rs"), &finding(Strength::Consider), false); + let consider = annotation(Path::new("x.rs"), &finding(Strength::Consider)); assert!(consider.starts_with("::warning file=x.rs,line=12,title=")); } + #[test] + fn a_review_still_being_measured_is_a_warning_that_says_why() { + let review = counted(Strength::Review, Gating::Measuring); + let line = annotation(Path::new("x.rs"), &review); + assert!(line.starts_with("::warning file=x.rs,"), "{line}"); + assert!( + line.ends_with("%0ADoes not fail the gate: maintainability/shared-logic reviews are still being measured (54%25 of 85 right on projects JevGate was never tuned on)."), + "{line}" + ); + } + + #[test] + fn the_summary_says_why_reviews_still_being_measured_did_not_fail() { + let report = crate::tests::gated( + vec![crate::tests::finding_of( + "maintainability/shared-logic", + Strength::Review, + )], + &crate::tests::args(), + ); + let shown = output::failing_first(&report); + let text = summary(&report, &shown); + assert!(text.contains("| review | `src/lib.rs:12`"), "{text}"); + assert!( + text.contains("\n1 review did not fail the gate: by default only rules and levels"), + "{text}" + ); + } + #[test] fn summary_rows_stay_on_one_line_and_bold_gate_failures() { let args = crate::tests::args(); - let review = finding(Strength::Review); - let consider = finding(Strength::Consider); + let review = counted(Strength::Review, Gating::Fails); + let consider = counted(Strength::Consider, Gating::Measuring); let path = Path::new("src/a.rs"); let report = crate::evaluate::snapshot( &[], @@ -168,7 +248,7 @@ mod tests { requests: 0, }, ); - let text = summary(&report, &[(path, &review), (path, &consider)], &args); + let text = summary(&report, &[(path, &review), (path, &consider)]); let rows: Vec<&str> = text.lines().filter(|l| l.starts_with("| ")).collect(); assert_eq!(rows.len(), 3, "{text}"); assert!( diff --git a/src/gitlab.rs b/src/gitlab.rs index 560b6d4..44b2ae7 100644 --- a/src/gitlab.rs +++ b/src/gitlab.rs @@ -1,7 +1,6 @@ //! `--format gitlab`: the findings as a GitLab Code Quality report, which //! merge requests show as a widget and on the changed lines. use crate::{ - options::CheckArgs, output, schema::{Finding, Report, Strength}, }; @@ -10,29 +9,34 @@ use serde_json::{Value, json}; use std::{io::Write, path::Path}; /// The findings the GitHub annotations show: `major` when a finding fails the -/// gate, `minor` otherwise. -pub fn emit(out: &mut impl Write, report: &Report, args: &CheckArgs) -> Result<()> { +/// gate, `minor` otherwise, and the description says when its rule and level +/// are still being measured. +pub fn emit(out: &mut impl Write, report: &Report) -> Result<()> { let issues: Vec = output::ranked(report) .into_iter() .filter(|(_, f)| f.strength != Strength::Note && !f.accepted()) - .map(|(path, finding)| issue(path, finding, crate::gate::fails(finding, path, args))) + .map(|(path, finding)| issue(path, finding)) .collect(); serde_json::to_writer_pretty(&mut *out, &issues)?; writeln!(out)?; Ok(()) } -fn issue(path: &Path, finding: &Finding, fails: bool) -> Value { +fn issue(path: &Path, finding: &Finding) -> Value { let end = finding .locations .iter() .find(|l| l.path == path && l.start_line == finding.line) .map_or(finding.line, |l| l.end_line.max(finding.line)); + let mut description = format!("{} Next step: {}", finding.message, finding.action); + if let Some(note) = output::measuring_note(finding) { + description.push_str(&format!(" {note}")); + } json!({ - "description": format!("{} Next step: {}", finding.message, finding.action), + "description": description, "check_name": finding.rule, "fingerprint": fingerprint(path, finding), - "severity": if fails { "major" } else { "minor" }, + "severity": if finding.fails_gate() { "major" } else { "minor" }, "location": { "path": path.to_string_lossy(), "lines": {"begin": finding.line.max(1), "end": end.max(1)}, @@ -58,12 +62,16 @@ fn fingerprint(path: &Path, finding: &Finding) -> String { #[cfg(test)] mod tests { use super::*; - use crate::tests::finding; + use crate::{schema::Gating, tests::finding}; #[test] fn issues_carry_severity_location_and_a_fingerprint() { let path = Path::new("src/a,b.rs"); - let review = issue(path, &finding(Strength::Review), true); + let failing = Finding { + gate: Some(Gating::Fails), + ..finding(Strength::Review) + }; + let review = issue(path, &failing); assert_eq!(review["severity"], "major"); assert_eq!(review["check_name"], "maintainability/shared-logic"); assert_eq!(review["location"]["path"], "src/a,b.rs"); @@ -75,10 +83,22 @@ mod tests { .ends_with("Next step: Share one | implementation") ); assert_eq!(review["fingerprint"].as_str().unwrap().len(), 64); - let consider = issue(path, &finding(Strength::Consider), false); + let consider = issue(path, &finding(Strength::Consider)); assert_eq!(consider["severity"], "minor"); let mut named = finding(Strength::Consider); named.fingerprint = "abc".into(); - assert_eq!(issue(path, &named, false)["fingerprint"], "abc"); + assert_eq!(issue(path, &named)["fingerprint"], "abc"); + let measuring = Finding { + gate: Some(Gating::Measuring), + ..finding(Strength::Review) + }; + let measured = issue(path, &measuring); + assert_eq!(measured["severity"], "minor"); + assert!( + measured["description"] + .as_str() + .unwrap() + .ends_with("Next step: Share one | implementation Does not fail the gate: maintainability/shared-logic reviews are still being measured (54% of 85 right on projects JevGate was never tuned on).") + ); } } diff --git a/src/html_report.rs b/src/html_report.rs index fb5d4f9..e9aad58 100644 --- a/src/html_report.rs +++ b/src/html_report.rs @@ -13,12 +13,12 @@ fn dimension(rule: &str, d: &Dimension) -> Value { } fn batch_cost(report: &Report) -> Option { - crate::output::estimated_usd(report).map(|usd| { + report.estimated_usd.map(|usd| { json!({ "estimated_usd": usd, - "input_per_million": crate::output::INPUT_USD_PER_MILLION, + "input_per_million": crate::model::INPUT_USD_PER_MILLION, "output_per_million": 0.0, - "checked_at": crate::output::PRICE_CHECKED + "checked_at": crate::model::PRICE_CHECKED }) }) } @@ -43,7 +43,11 @@ pub fn render(report: &Report) -> Result { .collect(); let rules: Vec<_> = crate::catalog::rules() .into_iter() - .map(|r| json!({"key":r.key,"id":r.id,"description":r.inspection})) + .map(|r| { + json!({"key":r.key,"id":r.id,"description":r.inspection, + "maturity":crate::maturity::describe(r.key), + "unmeasured":crate::maturity::unmeasured(r.key)}) + }) .collect(); // A run that leaves out default rules lists fewer files; say so. let partial = crate::catalog::rules() @@ -53,8 +57,11 @@ pub fn render(report: &Report) -> Result { "status":report.status,"complete":report.complete,"settled":report.settled, "refresh":report.watcher_pid.is_some(),"model":report.requested_model, "requests":report.api_requests,"tokens":report.paid_input_tokens, - "cost":batch_cost(report),"gate":report.gate,"fail_on":report.fail_on, - "errors":report.errors,"deleted":report.deleted_files,"files":files,"rules":rules, + "cost":batch_cost(report),"unmetered":report.unmetered_requests, + "gate":report.gate,"fail_on":report.fail_on, + "fail_on_mature":report.fail_on_mature, + "errors":report.errors,"deleted":report.deleted_files,"base":report.base_revision, + "scope":report.scope,"files":files,"rules":rules, "selected":report.rules,"partial":partial}); // Even a filename or analyzer message may contain . Never let data // terminate the JSON element, and insert all displayed strings with textContent. @@ -148,15 +155,24 @@ mod tests { .unwrap() .is_empty() ); - assert_eq!(decoded["fail_on"], serde_json::json!(["review"])); + assert_eq!(decoded["fail_on"], serde_json::json!(["mature"])); + assert_eq!( + decoded["fail_on_mature"], + serde_json::json!({"maintainability/function-simplification": ["review"], "documentation/agent-context": ["consider"]}) + ); + assert_eq!( + decoded["rules"][1]["maturity"]["review"]["unseen"], + serde_json::json!({"right": 20, "labeled": 23}) + ); + assert_eq!(decoded["scope"], "whole-files"); + assert_eq!(decoded["unmetered"], 0); assert!(!html.contains("id=\"root\"")); - report.paid_input_tokens = 1_000_000; - report.paid_output_tokens = 12_345; - let cost = batch_cost(&report).unwrap(); - assert!((cost["estimated_usd"].as_f64().unwrap() - 0.042).abs() < 1e-12); - report.paid_input_tokens = 0; assert_eq!(batch_cost(&report).unwrap()["estimated_usd"], 0.0); - report.requested_model = "unknown-model".into(); + report.estimated_usd = Some(0.042); + let cost = batch_cost(&report).unwrap(); + assert_eq!(cost["estimated_usd"], 0.042); + assert_eq!(cost["checked_at"], crate::model::PRICE_CHECKED); + report.estimated_usd = None; assert!(batch_cost(&report).is_none()); assert!(!html.contains("def value()")); } diff --git a/src/init.rs b/src/init.rs index 35e5c7b..569a998 100644 --- a/src/init.rs +++ b/src/init.rs @@ -90,23 +90,7 @@ fn render(allow: &[String]) -> String { } else { format!("upload_allow = {}\n", list(allow)) }; - let mut rules = String::new(); - for group in catalog::groups() { - let members: Vec<_> = catalog::rules() - .into_iter() - .filter(|r| r.group == group) - .collect(); - let names: Vec<&str> = members - .iter() - .map(|r| r.id.trim_start_matches(&format!("{group}/")[..])) - .collect(); - let comment = format!("# {}", names.join(", ")); - if members.iter().all(|r| r.default_enabled) { - rules.push_str(&format!("{group} = \"review\" {comment}\n")); - } else { - rules.push_str(&format!("# {group} = \"consider\" {comment} (opt-in)\n")); - } - } + let rules: String = catalog::groups().into_iter().map(group_example).collect(); format!( r#"#:schema https://raw.githubusercontent.com/Tech-Byte-Frontier/jevgate/v{version}/jevgate.schema.json # JevGate configuration, written by `jevgate init`. Unknown keys are errors. @@ -122,17 +106,23 @@ upload_deny = ["**/.env*", "**/*.pem", "**/*.key"] # File organization judges test files either way. # include_tests = true -# The model, pinned so results stay repeatable; --model overrides it. +# The model, pinned so results stay repeatable; --model overrides it. The +# default follows the key: this one for TypeSafe, typesafe/jev-1.13 for +# OpenRouter, typesafe-ai/jev for Vercel AI Gateway. # model = "{model}" # Budgets for one invocation; flags can only lower them. # max_requests = 200 # concurrency = 4 -# Each group or rule ID set to a level is judged and fails the check at that -# level: "review", "consider" (also fails on review), "uncertain", "report" -# (judge, never fail) or "off". A rule's own entry wins over its group's. -# Test rules also need include_tests or --include-tests. +# Unset, the default rules run and only rule levels measured right at least +# 80% of the time on projects JevGate was never tuned on fail the check +# ("mature"; `jevgate rules` shows them); other findings are reported without +# failing it. A group or rule ID set to a level is judged, every rule of a +# group included, and fails the check at exactly that level: "review", +# "consider" (also fails on review), "mature", "uncertain", "report" (judge, +# never fail) or "off". A rule's own entry wins over its group's. Test rules +# also need include_tests or --include-tests. [rules] {rules} # Levels for the files some paths match, such as report-only tooling. The last @@ -146,6 +136,33 @@ upload_deny = ["**/.env*", "**/*.pem", "**/*.key"] ) } +/// A commented `[rules]` line for one group: a level to set, and its rules, +/// with the ones that do not run by default marked opt-in. +fn group_example(group: &str) -> String { + let members: Vec<_> = catalog::rules() + .into_iter() + .filter(|r| r.group == group) + .collect(); + let opt_in = members.iter().all(|r| !r.default_enabled); + let names: Vec = members + .iter() + .map(|r| { + let name = r.id.trim_start_matches(&format!("{group}/")[..]); + if r.default_enabled || opt_in { + name.to_string() + } else { + format!("{name} (opt-in)") + } + }) + .collect(); + let (level, suffix) = if opt_in { + ("consider", " (opt-in)") + } else { + ("review", "") + }; + format!("# {group} = \"{level}\" # {}{suffix}\n", names.join(", ")) +} + #[cfg(test)] mod tests { use super::*; @@ -181,12 +198,26 @@ mod tests { ".cursor/rules/**" ] ); - let config: Config = toml::from_str(&std::fs::read_to_string(&path).unwrap()).unwrap(); + let text = std::fs::read_to_string(&path).unwrap(); + let config: Config = toml::from_str(&text).unwrap(); assert_eq!(config.upload_allow, allow); - let Rules::Levels(levels) = config.rules else { - panic!("rules is a table of levels"); + let levels = |config: Config| match config.rules { + Rules::Levels(levels) => levels, + Rules::List(_) => panic!("rules is a table of levels"), }; - assert!(levels.contains_key("maintainability") && levels.contains_key("tests")); + assert!(levels(config).is_empty(), "the default rules and gate"); + assert!(text.contains("hardcoded-values (opt-in)"), "{text}"); + let uncommented: Vec<&str> = text + .lines() + .map(|line| { + let example = catalog::groups() + .into_iter() + .any(|g| line.starts_with(&format!("# {g} = "))); + if example { &line[2..] } else { line } + }) + .collect(); + let config: Config = toml::from_str(&uncommented.join("\n")).unwrap(); + assert_eq!(levels(config).len(), catalog::groups().len()); assert!(run(&dir, false).is_err(), "an existing file is kept"); assert!(run(&dir, true).is_ok()); } diff --git a/src/inventory/documents.rs b/src/inventory/documents.rs index fc557b6..41328e1 100644 --- a/src/inventory/documents.rs +++ b/src/inventory/documents.rs @@ -2,15 +2,22 @@ //! not they are hidden or ignored. use super::*; +/// Which documents a check reads: those `selected`, and with `broken` the +/// others in its scope that name one of the paths a change removed. +pub(super) type Selection<'a> = ( + &'a dyn Fn(&Path) -> bool, + Option<(&'a dyn Fn(&Path) -> bool, &'a Arc>)>, +); + /// Agent instruction files and project docs the selected rules judge. pub(super) fn add_documents( args: &CheckArgs, context: &ConfigContext, boundary: &Boundary, - selected: &dyn Fn(&Path) -> bool, + (selected, broken): Selection<'_>, inputs: &mut Vec, ) -> Result<()> { - let repository = std::sync::Arc::new(crate::docs::scan(&context.root)?); + let repository = Arc::new(crate::docs::scan(&context.root)?); let instructions = repository .readers .keys() @@ -22,20 +29,58 @@ pub(super) fn add_documents( .filter(|_| args.enabled(crate::catalog::LARGE_DOCS) || cross_document(args)) .map(|p| (p, DOCS)); for (relative, role) in instructions.chain(docs) { - let wanted = selected(relative) - && (role == DOCS || repository.judged(relative)) + let judged = (role == DOCS || repository.judged(relative)) && !inputs.iter().any(|i| i.result.path == *relative); - if wanted { - let mut input = bounded(relative, role, args, context, boundary)?; - if input.result.status != Status::Skipped { - input.repository = Some(repository.clone()); + if !judged { + continue; + } + let change = match broken { + _ if selected(relative) => None, + Some((in_scope, removed)) if in_scope(relative) => { + Some(crate::revision::FileChange::unchanged(removed.clone())) + } + _ => continue, + }; + let mut input = bounded(relative, role, args, context, boundary)?; + if let Some(change) = change { + let text = input.source.as_deref().unwrap_or(""); + if !names_removed(relative, text, &change.removed) { + continue; } - inputs.push(input); + input.changed = Some(change); } + if input.result.status != Status::Skipped { + input.repository = Some(repository.clone()); + } + inputs.push(input); } Ok(()) } +/// Whether `text`, the document at `doc`, may name one of `removed`: the +/// path from the repository root, or from a directory holding the document, +/// as a relative link climbing out of it ends. Only the top of each removed +/// tree is looked for, since a name inside it holds the tree's own path: a +/// deleted directory of vendored files is one search. The staleness rule +/// then decides which sections name a removed path. Both paths are written +/// with `/`, as Git and `discovery::relative` write them. +fn names_removed(doc: &Path, text: &str, removed: &BTreeSet) -> bool { + let folders: Vec<&Path> = doc + .ancestors() + .skip(1) + .filter(|dir| !dir.as_os_str().is_empty()) + .collect(); + removed + .iter() + .filter(|path| path.parent().is_none_or(|dir| !removed.contains(dir))) + .any(|path| { + let below = folders.iter().filter_map(|dir| path.strip_prefix(dir).ok()); + std::iter::once(path.as_path()) + .chain(below) + .any(|name| text.contains(name.to_string_lossy().as_ref())) + }) +} + /// A documentation file, read whole; hidden and ignored paths are allowed. pub(super) fn load_document( relative: &std::path::Path, @@ -60,6 +105,7 @@ pub(super) fn load_document( package: None, settings_selected_by: Vec::new(), templates: Vec::new(), + changed: None, } } // dvja's docs hold a Markdown file with NUL bytes, which made the run incomplete. @@ -71,3 +117,24 @@ pub(super) fn load_document( fn cross_document(args: &CheckArgs) -> bool { args.enabled(crate::catalog::DOC_STALENESS) || args.enabled(crate::catalog::DOC_DUPLICATION) } + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn a_document_may_name_a_removed_path_from_the_root_or_a_folder_holding_it() { + let removed: BTreeSet = [ + "docs/api/old.md", + "vendor/lib", + "vendor/lib/a.js", + "vendor/lib/b.js", + ] + .map(PathBuf::from) + .into(); + let names = |text: &str| names_removed(Path::new("docs/guide/intro.md"), text, &removed); + assert!(names("See [the API](../api/old.md).")); + assert!(names("Built from `vendor/lib/a.js`.")); + assert!(!names("See old.md and lib.")); + } +} diff --git a/src/inventory/mod.rs b/src/inventory/mod.rs index 56a6664..f19a01d 100644 --- a/src/inventory/mod.rs +++ b/src/inventory/mod.rs @@ -6,14 +6,15 @@ use super::{ options::CheckArgs, schema::{FileResult, Status, hash}, }; -use crate::{boundary::Boundary, config::ConfigContext, discovery}; +use crate::{boundary::Boundary, config::ConfigContext, discovery, revision::Changes}; use anyhow::{Context, Result, ensure}; use django::{select_settings, unescaped_templates}; use documents::{add_documents, load_document}; use spacetimedb::{keep_module_packages, spacetimedb_package}; use std::{ - collections::BTreeMap, + collections::{BTreeMap, BTreeSet}, path::{Path, PathBuf}, + sync::Arc, }; #[derive(Clone)] @@ -35,6 +36,10 @@ pub struct Input { /// For a Python file, with injection judged: the Django templates it /// names that write values without escaping them. pub templates: Vec, + /// With `--base` judging what the change touched, what it did to this + /// file; none when the file is judged whole: without a base, with + /// `--whole-files`, or for a file the change added. + pub changed: Option, } /// A SpacetimeDB module's package: the directory of the `package.json` that @@ -79,11 +84,12 @@ pub(crate) fn walker(root: &std::path::Path) -> ignore::Walk { } pub fn collect(args: &CheckArgs, context: &ConfigContext, scope: &[PathBuf]) -> Result> { - let changes = args - .base + let changes = load_changes(args, context)?; + // The paths a change removed, when only what it touched is judged. + let removed = changes .as_ref() - .map(|b| crate::revision::Changes::load(&context.root, b)) - .transpose()?; + .filter(|_| args.changed_lines()) + .map(|c| Arc::new(c.removed(&context.root))); let extra = super::context::collect(args, context)?; let boundary = Boundary::new(&context.config)?; let in_scope = |relative: &Path| { @@ -113,7 +119,12 @@ pub fn collect(args: &CheckArgs, context: &ConfigContext, scope: &[PathBuf]) -> unescaped_templates(context, &boundary, &mut inputs); } if args.documentation() { - add_documents(args, context, &boundary, &selected, &mut inputs)?; + // A document the change left alone is judged for the paths it removed. + let broken = removed + .as_ref() + .filter(|r| !r.is_empty() && args.enabled(crate::catalog::DOC_STALENESS)) + .map(|r| (&in_scope as &dyn Fn(&Path) -> bool, r)); + add_documents(args, context, &boundary, (&selected, broken), &mut inputs)?; } // SQL is also a source extension, so the code rules' walk may have listed // the file already; the configuration rule's role replaces that entry. @@ -124,9 +135,41 @@ pub fn collect(args: &CheckArgs, context: &ConfigContext, scope: &[PathBuf]) -> None => inputs.push(input), } } + if let (Some(changes), Some(removed)) = (changes, &removed) { + mark_changes(&mut inputs, changes, &context.root, removed)?; + } Ok(inputs) } +/// With `--base`, what changed since the fork point with it. +fn load_changes(args: &CheckArgs, context: &ConfigContext) -> Result> { + args.base + .as_ref() + .map(|base| Changes::load(&context.root, base)) + .transpose() +} + +/// Record on each input what the change did to it, so only the units it +/// touched are judged. The lines are read only for the files read to be +/// judged, once they are known. A file the change added stays whole, and a +/// document it left alone keeps the mark it was selected with. +fn mark_changes( + inputs: &mut [Input], + changes: Changes, + root: &Path, + removed: &Arc>, +) -> Result<()> { + let judged = inputs + .iter() + .filter(|i| i.changed.is_none() && i.source.is_some()) + .map(|i| i.result.path.as_path()); + let changes = changes.with_lines(root, judged)?; + for input in inputs.iter_mut().filter(|i| i.changed.is_none()) { + input.changed = changes.file(root, &input.result.path, removed); + } + Ok(()) +} + /// Application source and tests in scope, with their roles. fn source_paths( args: &CheckArgs, @@ -312,6 +355,7 @@ fn source_input( package: crate::packages::package(&context.root, relative), settings_selected_by: Vec::new(), templates: Vec::new(), + changed: None, } } @@ -409,6 +453,7 @@ fn bare_input(result: FileResult) -> Input { package: None, settings_selected_by: Vec::new(), templates: Vec::new(), + changed: None, } } @@ -497,6 +542,7 @@ pub fn fingerprint(inputs: &[Input]) -> String { &i.result.context_complete, &i.result.context_limitations, &i.result.error, + i.changed.as_ref().map(|c| &c.lines), ) }) .collect(); diff --git a/src/main.rs b/src/main.rs index 4a71eb5..0e28b7b 100644 --- a/src/main.rs +++ b/src/main.rs @@ -43,14 +43,18 @@ mod inventory; mod line_ranges; mod locations; mod manual; +mod maturity; mod mcp; +mod model; mod options; mod output; mod packages; mod policy; +mod provider; mod provider_error; mod requests; mod response; +mod response_headers; mod revision; mod sarif; mod schema; @@ -76,8 +80,10 @@ use options::JevCommand; /// has a location, a probability and a next step, and undecided answers are /// reported as uncertain instead of hidden. /// -/// Rule groups: maintainability (on by default), tests (with -/// --include-tests), and the opt-in security and documentation groups. +/// Rule groups: maintainability (on by default, except hardcoded values), +/// tests (with --include-tests), and the opt-in security and documentation +/// groups. By default only rules and levels measured right at least 80% of +/// the time on projects JevGate was never tuned on fail the check. #[derive(Parser)] #[command(version, after_long_help = options::OVERVIEW)] pub struct Cli { diff --git a/src/maturity.rs b/src/maturity.rs new file mode 100644 index 0000000..cc8ead5 --- /dev/null +++ b/src/maturity.rs @@ -0,0 +1,277 @@ +//! Which rules and levels fail the gate by default. A rule and level is +//! mature when its findings were right at least 80% of the time on projects +//! JevGate was never tuned on, over at least 20 findings labeled by hand from +//! the code. The default gate level, `mature`, fails only on those; every +//! other finding is reported without failing the check, until its rule and +//! level measure up. The concern probability could not decide this: on those +//! projects, reviews were right 55%, 46%, 56% and 61% of the time with a +//! probability below 0.90, below 0.95, below 0.98 and above. +use crate::{ + catalog, + schema::Strength::{self, Consider, Review}, +}; +use serde_json::{Value, json}; + +/// Labeled findings a rule and level needs on unseen projects to be mature. +pub const MIN_LABELS: u32 = 20; +/// The share of those findings, in percent, that must be right. +pub const MIN_PERCENT_RIGHT: u32 = 80; + +/// Findings of one rule and level labeled by hand: how many were right, of +/// how many labeled. A debatable label counts as not right. +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub struct Labels { + pub right: u32, + pub labeled: u32, +} + +impl Labels { + /// The share right in whole percent, half rounded up as the HTML report + /// and the site round it; none without labels. + pub fn percent(self) -> Option { + (self.labeled > 0).then(|| (200 * self.right + self.labeled) / (2 * self.labeled)) + } + + /// "87% of 23"; none without labels. + pub fn summary(self) -> Option { + self.percent().map(|p| format!("{p}% of {}", self.labeled)) + } + + fn describe(self) -> Value { + json!({"right": self.right, "labeled": self.labeled}) + } +} + +/// One rule and level's labels on the projects never used for tuning and on +/// the ones the rules were tuned on. +pub struct Measure { + /// The rule's catalog key. + pub rule: &'static str, + pub level: Strength, + pub unseen: Labels, + pub tuned: Labels, +} + +impl Measure { + /// Right at least [`MIN_PERCENT_RIGHT`] of the time over at least + /// [`MIN_LABELS`] labels on unseen projects. + pub fn mature(&self) -> bool { + let Labels { right, labeled } = self.unseen; + labeled >= MIN_LABELS && 100 * right >= MIN_PERCENT_RIGHT * labeled + } +} + +const fn row(rule: &'static str, level: Strength, unseen: [u32; 2], tuned: [u32; 2]) -> Measure { + Measure { + rule, + level, + unseen: Labels { + right: unseen[0], + labeled: unseen[1], + }, + tuned: Labels { + right: tuned[0], + labeled: tuned[1], + }, + } +} + +/// Measured on 2026-09-28 from the corpus's hand labels +/// (`evaluation/labels`, joined by fingerprint) and JevGate 0.25.0's findings, +/// replayed from the answer cache over the 94 labeled projects outside Bend 2; +/// the few files whose 0.25.0 requests the cache lacked keep their 0.24.1 +/// findings. Unseen: [right, labeled] on the 11 held-out and 14 fresh projects +/// never used for tuning (22 of them have findings); tuned: on the other 72. +/// `docs/research/2026-09-28/scripts/maturity.py` in the maintainer's clone +/// prints these rows. Two are mature: function-simplification reviews (20 of +/// 23) and agent-context considers (22 of 24). The gap between the columns is +/// why only unseen projects count: shared-logic reviews were right 75% of the +/// time on tuned projects and 54% on unseen ones. +const TABLE: [Measure; 26] = [ + row(catalog::FILE_ORGANIZATION, Review, [2, 5], [15, 30]), + row(catalog::FILE_ORGANIZATION, Consider, [17, 29], [19, 43]), + row(catalog::FUNCTION_SIMPLIFICATION, Review, [20, 23], [57, 69]), + row( + catalog::FUNCTION_SIMPLIFICATION, + Consider, + [85, 126], + [147, 197], + ), + row(catalog::SHARED_LOGIC, Review, [46, 85], [181, 240]), + row(catalog::SHARED_LOGIC, Consider, [85, 158], [146, 244]), + row(catalog::HARDCODED_VALUES, Review, [1, 8], [15, 28]), + row(catalog::HARDCODED_VALUES, Consider, [5, 29], [32, 57]), + row(catalog::INJECTION, Review, [3, 4], [81, 96]), + row(catalog::INJECTION, Consider, [5, 13], [27, 47]), + row(catalog::SENSITIVE_DATA, Review, [10, 24], [41, 64]), + row(catalog::SENSITIVE_DATA, Consider, [0, 5], [12, 15]), + row(catalog::UNSAFE_SETTINGS, Review, [2, 4], [53, 72]), + row(catalog::UNSAFE_SETTINGS, Consider, [0, 4], [16, 21]), + row(catalog::ACCESS_CONTROL, Review, [0, 0], [2, 5]), + row(catalog::ACCESS_CONTROL, Consider, [0, 0], [5, 14]), + row(catalog::WORKFLOWS, Review, [1, 1], [0, 1]), + row(catalog::TEST_VALUE, Review, [3, 5], [5, 11]), + row(catalog::TEST_VALUE, Consider, [1, 2], [20, 27]), + row(catalog::TEST_REDUNDANCY, Review, [1, 1], [4, 4]), + row(catalog::TEST_REDUNDANCY, Consider, [27, 45], [27, 34]), + row(catalog::AGENT_CONTEXT, Consider, [22, 24], [64, 68]), + row(catalog::LARGE_DOCS, Consider, [1, 1], [1, 5]), + row(catalog::DOC_STALENESS, Consider, [2, 2], [14, 15]), + row(catalog::DOC_DUPLICATION, Consider, [3, 20], [5, 22]), + row(catalog::COMMENTS, Consider, [39, 72], [97, 150]), +]; + +/// The labels of a rule, by ID, name or key, at one level. +pub fn measure(rule: &str, level: Strength) -> Option<&'static Measure> { + let key = catalog::find(rule)?.key; + TABLE.iter().find(|m| m.rule == key && m.level == level) +} + +/// Why `rule` has no share of right findings on unseen projects at a level: +/// the laws rule judges Bend 2 code, whose labeled projects this table +/// leaves out; any other has no labeled finding there yet. +pub fn unmeasured(rule: &str) -> &'static str { + if catalog::find(rule).is_some_and(|found| found.key == catalog::LAWS) { + "labeled only on Bend 2 projects, which the maturity table leaves out" + } else { + "none labeled yet on projects JevGate was never tuned on" + } +} + +/// How often findings of `rule` at `level` were right on unseen projects, +/// "54% of 85 right on projects JevGate was never tuned on", or why that +/// is unknown. +pub fn unseen_share(rule: &str, level: Strength) -> String { + measure(rule, level) + .and_then(|m| m.unseen.summary()) + .map_or_else( + || unmeasured(rule).into(), + |share| format!("{share} right on projects JevGate was never tuned on"), + ) +} + +/// Whether the default gate fails on findings of `rule` at `level`. +pub fn mature(rule: &str, level: Strength) -> bool { + measure(rule, level).is_some_and(Measure::mature) +} + +/// A rule's mature levels, review first. +pub fn mature_levels(rule: &str) -> Vec { + [Review, Consider] + .into_iter() + .filter(|level| mature(rule, *level)) + .collect() +} + +/// A rule's measured levels for `jevgate rules --format json`: labels on +/// unseen and tuned projects, and whether each level is mature. +pub fn describe(rule: &str) -> Value { + let levels = [Review, Consider].into_iter().filter_map(|level| { + measure(rule, level).map(|m| { + let value = json!({ + "unseen": m.unseen.describe(), + "tuned": m.tuned.describe(), + "mature": m.mature(), + }); + (crate::output::label(&level), value) + }) + }); + Value::Object(levels.collect()) +} + +#[cfg(test)] +mod tests { + use super::*; + + fn unseen(right: u32, labeled: u32) -> Measure { + row(catalog::SHARED_LOGIC, Review, [right, labeled], [0, 0]) + } + + #[test] + fn a_level_is_mature_at_eighty_percent_over_twenty_labels() { + assert!(unseen(16, 20).mature(), "exactly 80% of 20"); + assert!(!unseen(15, 20).mature()); + assert!(!unseen(19, 19).mature(), "too few labels"); + assert!(unseen(20, 23).mature()); + assert!(!unseen(0, 0).mature()); + } + + #[test] + fn the_table_names_catalog_rules_once_per_level() { + let keys = catalog::keys(); + for (i, m) in TABLE.iter().enumerate() { + assert!(keys.contains(&m.rule), "{}", m.rule); + assert!(m.level != Strength::Note, "{}", m.rule); + for labels in [m.unseen, m.tuned] { + assert!(labels.right <= labels.labeled, "{}", m.rule); + } + assert!( + !TABLE[..i] + .iter() + .any(|other| other.rule == m.rule && other.level == m.level), + "{} twice", + m.rule + ); + } + } + + #[test] + fn function_simplification_reviews_and_agent_context_considers_are_mature() { + let mature: Vec<(&str, Strength)> = TABLE + .iter() + .filter(|m| m.mature()) + .map(|m| (m.rule, m.level)) + .collect(); + assert_eq!( + mature, + [ + (catalog::FUNCTION_SIMPLIFICATION, Review), + (catalog::AGENT_CONTEXT, Consider) + ] + ); + assert!(self::mature( + "maintainability/function-simplification", + Review + )); + assert!(!self::mature("function-simplification", Consider)); + assert!(!self::mature(catalog::LAWS, Review), "never measured"); + assert_eq!(mature_levels(catalog::AGENT_CONTEXT), [Consider]); + assert!(mature_levels("nothing").is_empty()); + } + + #[test] + fn measured_levels_describe_their_labels() { + let value = describe(catalog::FUNCTION_SIMPLIFICATION); + assert_eq!( + value["review"], + json!({"unseen": {"right": 20, "labeled": 23}, "tuned": {"right": 57, "labeled": 69}, "mature": true}) + ); + assert_eq!(value["consider"]["mature"], false); + assert_eq!(describe(catalog::LAWS), json!({})); + let unseen = |rule| measure(rule, Review).unwrap().unseen.summary(); + assert_eq!(unseen(catalog::SHARED_LOGIC).as_deref(), Some("54% of 85")); + assert_eq!( + unseen(catalog::HARDCODED_VALUES).as_deref(), + Some("13% of 8"), + "12.5% rounds up" + ); + assert_eq!(unseen(catalog::ACCESS_CONTROL), None); + } + + #[test] + fn a_level_without_a_share_says_why() { + assert_eq!( + unseen_share(catalog::SHARED_LOGIC, Review), + "54% of 85 right on projects JevGate was never tuned on" + ); + assert_eq!( + unseen_share(catalog::ACCESS_CONTROL, Review), + "none labeled yet on projects JevGate was never tuned on" + ); + assert_eq!( + unseen_share("tests/laws", Review), + "labeled only on Bend 2 projects, which the maturity table leaves out", + "law findings were labeled, on the Bend 2 projects kept apart" + ); + } +} diff --git a/src/mcp.rs b/src/mcp.rs index b2b9814..80ae2ca 100644 --- a/src/mcp.rs +++ b/src/mcp.rs @@ -17,8 +17,10 @@ const VERSIONS: [&str; 4] = ["2025-06-18", "2025-11-25", "2025-03-26", "2024-11- const MAX_FINDINGS: usize = 50; const INSTRUCTIONS: &str = "JevGate reviews code by asking TypeSafe Jev small questions about functions, files, tests and docs. \ -Call jevgate_check with `base` (such as origin/main) to review what changed; it uses the repository's jevgate.toml and TYPESAFE_API_KEY, and paid requests only for code the answer cache lacks. \ -Fix each `review` finding; for a `consider`, fix it or explain why the code should stay. Exit code 2 means the run could not finish: report it, never treat it as a pass. \ +Call jevgate_check with `base` (such as origin/main) to review what changed; it uses the repository's jevgate.toml and the API key `jevgate auth status` shows, and paid requests only for code the answer cache lacks. \ +Fix each finding marked to fail the gate (`gate: fails` in jevgate_findings): they decide the exit code, and by default only rules and levels measured right at least 80% of the time on projects JevGate was never tuned on fail it. \ +Weigh the other `review` and `consider` findings: fix one when it is right, or say why the code should stay. \ +Exit code 2 means the run could not finish: report it, never treat it as a pass. \ jevgate_findings reads the last report without running anything."; pub fn run() -> Result<()> { @@ -111,7 +113,8 @@ impl Server { Ok(format!("{text}\n\n({meaning})")) } - /// Findings of the last report, ranked, optionally for one path prefix. + /// Findings of the last report, those that fail the gate first, then + /// reviews before considers, each by rank, optionally for one path prefix. fn findings(&self, arguments: &Value) -> Result { let report = crate::storage::read_latest(&self.root) .context("No report yet; call jevgate_check first")?; @@ -139,6 +142,7 @@ fn check_arguments(arguments: &Value) -> Result> { args.push(format!("--rule={rule}")); } for (flag, name) in [ + ("--whole-files", "whole_files"), ("--include-tests", "include_tests"), ("--dry-run", "dry_run"), ("--verbose", "verbose"), @@ -170,8 +174,11 @@ fn strings(value: &Value, name: &str) -> Result> { } } +/// The findings of `report`, those that fail the gate first and then +/// reviews, so the cap never leaves one out for a finding that only warns or +/// for a consider. fn findings(report: &Report, prefix: Option<&str>, include_notes: bool) -> Value { - let all: Vec = output::ranked(report) + let all: Vec = output::failing_first(report) .into_iter() .filter(|(path, _)| prefix.is_none_or(|p| path.starts_with(p))) .filter(|(_, f)| include_notes || f.strength != crate::schema::Strength::Note) @@ -186,6 +193,7 @@ fn findings(report: &Report, prefix: Option<&str>, include_notes: bool) -> Value "probability": f.concern_probability, "baselined": f.baselined, "suppressed": f.suppressed, + "gate": f.gate, }) }) .collect(); @@ -218,11 +226,12 @@ fn tools() -> Value { { "name": "jevgate_check", "title": "Review code with JevGate", - "description": "Run `jevgate check` in the repository and return its ranked findings, each with a location, probability and next step. Uses jevgate.toml and TYPESAFE_API_KEY; unchanged code is answered from the cache for free, and dry_run costs nothing. Can take minutes on a large change.", + "description": "Run `jevgate check` in the repository and return its ranked findings, each with a location, probability and next step. Uses jevgate.toml and the API key `jevgate auth status` shows; unchanged code is answered from the cache for free, and dry_run costs nothing. Can take minutes on a large change.", "inputSchema": { "type": "object", "properties": { - "base": {"type": "string", "description": "Review only files changed since this Git revision, such as origin/main"}, + "base": {"type": "string", "description": "Review only what changed since this Git revision, such as origin/main: the functions, tests and comments on changed lines, and copies where either copy changed"}, + "whole_files": {"type": "boolean", "description": "With base, judge each changed file whole instead of only what the change touched"}, "paths": {"type": "array", "items": {"type": "string"}, "description": "Files or directories to review instead of the discovered source"}, "rules": {"type": "array", "items": {"type": "string"}, "description": "Rule IDs, names, keys or groups, such as security or file-organization; replaces the configured selection"}, "include_tests": {"type": "boolean", "description": "Also judge tests"}, @@ -319,6 +328,7 @@ mod tests { "base": "--config=/etc/passwd", "rules": ["security"], "paths": ["--refresh", "src"], + "whole_files": true, "dry_run": true, })) .unwrap(); @@ -330,6 +340,7 @@ mod tests { "--color=never", "--base=--config=/etc/passwd", "--rule=security", + "--whole-files", "--dry-run", "--", "--refresh", @@ -339,6 +350,35 @@ mod tests { assert!(check_arguments(&json!({"paths": "src"})).is_err()); } + #[test] + fn findings_say_how_the_gate_counted_them() { + use crate::{schema::Strength, tests::finding_of}; + let lower = crate::schema::Finding { + rank: 0.5, + ..finding_of("maintainability/function-simplification", Strength::Review) + }; + let report = crate::tests::gated( + vec![ + finding_of("maintainability/shared-logic", Strength::Review), + lower, + ], + &crate::tests::args(), + ); + let value = findings(&report, None, false); + let gates: Vec<&Value> = value["findings"] + .as_array() + .unwrap() + .iter() + .map(|f| &f["gate"]) + .collect(); + assert_eq!( + gates, + [&json!("fails"), &json!("measuring")], + "a failure first, whatever its rank" + ); + assert_eq!(value["gate"]["passed"], false); + } + #[test] fn a_failed_tool_call_is_a_tool_error_not_a_protocol_error() { let reply = server() diff --git a/src/model.rs b/src/model.rs new file mode 100644 index 0000000..6c4346c --- /dev/null +++ b/src/model.rs @@ -0,0 +1,153 @@ +//! What a model name says: a pinned version, whose answers never change, or an +//! alias that can move to a new version; and what its input costs. + +/// The longest model name accepted from a provider. +const MAX_NAME_BYTES: usize = 128; + +/// Jev 1.13's published price in dollars per million input tokens; output +/// tokens are free. TypeSafe's models page (https://docs.typesafe.ai/models), +/// checked on `PRICE_CHECKED`; OpenRouter's and Vercel AI Gateway's listings +/// give the same price. +pub const INPUT_USD_PER_MILLION: f64 = 0.042; +pub const PRICE_CHECKED: &str = "2026-09-28"; + +/// Model lines with a published price. A name is priced when its base name is +/// the line, one of its versions or a dated snapshot of it: `jev-1.13`, +/// `jev-1.13.0`, `typesafe/jev-1.13`, and `typesafe/jev-1.13-20260917`, the +/// endpoint OpenRouter lists for `typesafe/jev-1.13` at the same price +/// (checked on `PRICE_CHECKED`). An alias that names no version, such as +/// `typesafe-ai/jev`, is not: its price follows whatever it points to. +const PRICED_LINES: [&str; 1] = ["jev-1.13"]; + +/// Digits in a snapshot's date, as in `-20260917`. +const SNAPSHOT_DATE_DIGITS: usize = 8; + +/// A name a provider may return: letters, digits and `-_.`, with `/` between a +/// gateway's namespace and the model (`typesafe/jev-1.13`) and `~` for +/// OpenRouter's moving aliases (`~typesafe/jev-latest`). +pub fn valid_name(name: &str) -> bool { + !name.is_empty() + && name.len() <= MAX_NAME_BYTES + && name + .bytes() + .all(|c| c.is_ascii_alphanumeric() || b"-_.~/".contains(&c)) +} + +/// The name without a gateway's namespace: `jev-1.13` of `typesafe/jev-1.13`. +pub fn base_name(name: &str) -> &str { + name.rsplit_once('/').map_or(name, |(_, base)| base) +} + +/// A pinned version: the base name ends in an `x.y.z` version, as in +/// `jev-1.13.0` or `typesafe/jev-1.13.0`, and no `~` marks it as moving. Every +/// other name is an alias that can move to a new version: `jev`, `jev-1.13` +/// and `jev-latest` in TypeSafe's docs, and the gateways' `typesafe-ai/jev` and +/// `~typesafe/jev-latest`. +pub fn pinned(name: &str) -> bool { + let version = base_name(name).rsplit('-').next().unwrap_or_default(); + !name.contains('~') && dotted_numbers(version, 3) +} + +/// Estimated dollars for `tokens` input tokens answered by `name`; none when +/// its price is unknown. No tokens cost nothing, whatever the model. +pub fn usd(name: &str, tokens: u64) -> Option { + let base = base_name(name); + let priced = PRICED_LINES.iter().any(|line| { + base.strip_prefix(line).is_some_and(|rest| { + rest.is_empty() + || rest + .strip_prefix('.') + .is_some_and(|patch| dotted_numbers(patch, 1)) + || rest.strip_prefix('-').is_some_and(|date| { + date.len() == SNAPSHOT_DATE_DIGITS && date.bytes().all(|c| c.is_ascii_digit()) + }) + }) + }); + (tokens == 0 || priced).then(|| tokens as f64 * INPUT_USD_PER_MILLION / 1_000_000.0) +} + +/// Whether `text` is `count` runs of digits joined by dots: `1.13.0` for three. +fn dotted_numbers(text: &str, count: usize) -> bool { + let parts: Vec<&str> = text.split('.').collect(); + parts.len() == count + && parts + .iter() + .all(|part| !part.is_empty() && part.bytes().all(|c| c.is_ascii_digit())) +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn only_names_ending_in_an_x_y_z_version_are_pinned() { + for name in ["jev-1.13.0", "typesafe/jev-1.13.0", "typesafe-ai/jev-2.0.1"] { + assert!(pinned(name), "{name}"); + } + for name in [ + "jev", + "jev-1.13", + "jev-latest", + "jev-preview", + "typesafe/jev-1.13", + "typesafe-ai/jev", + "~typesafe/jev-latest", + "~typesafe/jev-1.13.0", + "jev-1.13.0-rc1", + "jev-1..0", + "other-version", + ] { + assert!(!pinned(name), "{name}"); + } + } + + #[test] + fn gateway_names_are_valid_and_their_base_is_the_model() { + for name in [ + "jev-1.13.0", + "typesafe/jev-1.13", + "~typesafe/jev-latest", + "typesafe-ai/jev", + ] { + assert!(valid_name(name), "{name}"); + } + for name in [ + "", + "jev 1", + "jev\n", + "jev@1", + &"j".repeat(MAX_NAME_BYTES + 1), + ] { + assert!(!valid_name(name), "{name:?}"); + } + assert_eq!(base_name("typesafe/jev-1.13"), "jev-1.13"); + assert_eq!(base_name("jev-1.13.0"), "jev-1.13.0"); + } + + #[test] + fn the_jev_1_13_line_is_priced_under_any_namespace_and_aliases_are_not() { + for name in [ + "jev-1.13.0", + "jev-1.13", + "typesafe/jev-1.13", + "typesafe-ai/jev-1.13.2", + // OpenRouter's endpoint for typesafe/jev-1.13, at $0.000000042 a token. + "typesafe/jev-1.13-20260917", + ] { + let usd = usd(name, 1_000_000).unwrap_or_else(|| panic!("{name}")); + assert!((usd - INPUT_USD_PER_MILLION).abs() < 1e-12, "{name}"); + } + for name in [ + "jev-latest", + "typesafe-ai/jev", + "~typesafe/jev-latest", + "jev-1.130", + "jev-1.13.x", + "jev-1.13-2026091", + "jev-1.13-rc1", + ] { + assert_eq!(usd(name, 1_000), None, "{name}"); + } + assert_eq!(usd("typesafe-ai/jev", 0), Some(0.0)); + } +} diff --git a/src/options/commands.rs b/src/options/commands.rs index 586fd7e..3cbf059 100644 --- a/src/options/commands.rs +++ b/src/options/commands.rs @@ -4,12 +4,15 @@ use clap::{Subcommand, ValueEnum}; #[derive(Subcommand)] pub enum JevCommand { - /// Save, inspect or remove your TypeSafe API credential + /// Save, inspect or remove your API key: TypeSafe, OpenRouter or Vercel AI Gateway /// - /// A check finds its key in this order: the TYPESAFE_API_KEY environment - /// variable, then the file named by `check --env-file` (by default the - /// repository's `.env`), then the key saved by `jevgate auth login`. In CI, set TYPESAFE_API_KEY - /// from a secret; nothing needs to be saved. + /// A check uses the first key it finds: TYPESAFE_API_KEY in the + /// environment; then the file named by `check --env-file` (TYPESAFE_API_KEY, + /// OPENROUTER_API_KEY or AI_GATEWAY_API_KEY), else TYPESAFE_API_KEY in the + /// repository's `.env`; then the key saved by `jevgate auth login`; then + /// OPENROUTER_API_KEY or AI_GATEWAY_API_KEY in the environment, which + /// other tools read too. The key goes only to its own provider. In CI, set + /// one variable from a secret; nothing needs to be saved. #[command(after_long_help = AUTH_EXAMPLES)] Auth { #[command(subcommand)] @@ -55,7 +58,9 @@ pub enum JevCommand { /// Without it, the file is replaced, so after a `--base` or path-limited /// check the findings accepted for every other file are dropped. With it, /// entries for files the check covered, or that were deleted, are replaced - /// by what the check found, and the rest are kept. + /// by what the check found, and the rest are kept. A `--base` check that + /// judged only what its change touched covers only the deleted files: + /// the other entries of the files it checked stay. #[arg(long)] merge: bool, /// Record this reason on findings accepted now without one @@ -66,14 +71,20 @@ pub enum JevCommand { #[command(subcommand)] action: Option, }, - /// List every rule with its group, default and the question it asks + /// List every rule with its default, the levels that fail the check by default, and its question /// /// A rule is named by its ID (`maintainability/shared-logic`), its key /// (`shared_logic`) or its group (`maintainability`, `tests`, `security`, /// `documentation`, plus `default` and `all`) anywhere a rule is accepted: /// `--rule`, `--skip-rule`, `--fail-on TARGET=LEVEL` and `[rules]`. + /// + /// Each rule shows how often its reviews and considers were right on + /// projects JevGate was never tuned on, from findings labeled by hand. The + /// levels right at least 80% of the time over at least 20 labels are + /// mature: by default only they fail the check (`--fail-on mature`), and + /// the other findings are reported without failing it. Rules { - /// `table` for people; `json` adds scope, evidence unit, version and decision policy + /// `table` for people; `json` adds scope, evidence unit, version, labels per level and decision policy #[arg(long, value_enum, default_value_t = RulesFormat::Table)] format: RulesFormat, }, @@ -177,7 +188,7 @@ Examples: pub const OVERVIEW: &str = "\ Workflow: jevgate init Write jevgate.toml: upload scope, rules and gate - jevgate auth login Save an API key (or set TYPESAFE_API_KEY) + jevgate auth login Save an API key: TypeSafe, OpenRouter or Vercel AI Gateway jevgate check --dry-run --show-requests Print every request body; no key, no network jevgate check Review and apply the gate jevgate baseline Accept current findings; later checks fail only on new ones @@ -185,7 +196,7 @@ Workflow: jevgate baseline mark wrong PATH[:LINE] Record why a finding was accepted; `baseline stats` counts them For agents and CI: - jevgate check --base origin/main Only files changed since a revision + jevgate check --base origin/main Only what changed since a revision jevgate check --base origin/main --format json The full report, raw probabilities included jevgate check --base origin/main --format github Annotations and a job summary on GitHub jevgate rules --format json Every rule and the question it asks @@ -205,7 +216,11 @@ Files (at the repository root): .jevgate/report.html HTML dashboard, with --report Environment: - TYPESAFE_API_KEY API key; takes precedence over every saved credential + TYPESAFE_API_KEY A TypeSafe key; wins over --env-file, .env and the saved key + OPENROUTER_API_KEY An OpenRouter key, used when no key comes from those + AI_GATEWAY_API_KEY A Vercel AI Gateway key, used when neither comes first + JEVGATE_BASE_URL Send requests to this API root instead (https, or http to localhost), + for a self-hosted proxy; never read from jevgate.toml or .env JEVGATE_CREDENTIAL_STORE Where `auth login` saves: auto, keyring or file JEVGATE_CONFIG_DIR Absolute directory for file-stored credentials CI When set, --report writes the dashboard without opening a browser @@ -218,26 +233,35 @@ const CHECK_EXAMPLES: &str = "\ Examples: jevgate check Discovered application source, default rules jevgate check src/billing --verbose One directory, with notes and per-file detail - jevgate check --base origin/main --format json Changed files only, machine-readable + jevgate check --base origin/main --format json Only what changed, machine-readable + jevgate check --base origin/main --whole-files Every unit of each changed file jevgate check --rule default --rule security Add the opt-in security group jevgate check --rule documentation Agent instruction files, project docs and code comments jevgate check --rule comments Only code comments: repeated code, filler, narrated edits jevgate check --include-tests Also judge test value and redundancy jevgate check --fail-on none Advisory: never exits 1; exits 2 when incomplete + jevgate check --fail-on review Fail on every review, not only on mature rules jevgate check --fail-on review --fail-on security=consider jevgate check --dry-run --show-requests Exactly what would be uploaded, offline - jevgate check --cache-only Replay cached answers; never contact TypeSafe + jevgate check --cache-only Replay cached answers; never contact the provider Reading the JSON report (--format json or .jevgate/latest.json): complete false when any selected file was not judged; the exit code is then 2 + scope whole-files, or changed-lines when --base judged what changed gate passed, reasons, new_findings, baselined_findings + fail_on the gate levels; fail_on_mature says what `mature` stands for files[].status clear, note, consider, review, uncertain, needs-context, not-applicable, skipped or error files[].findings rule, strength, line, message, action, locations, - concern_probability, fingerprint, baselined + concern_probability, fingerprint, baselined, and gate: + fails, measuring (its rule and level are still being + measured) or advisory (below the level in force) files[].dimensions per rule: status, unit counts and the units left undecided files[].judgments every raw answer, first pass and follow-ups - api_requests, paid_input_tokens, paid_output_tokens this run's usage"; + api_requests, paid_input_tokens, paid_output_tokens this run's usage + paid_models input tokens by the model that answered them + estimated_usd this run's cost; null when a response reported no usage + (unmetered_requests) or the model has no known price"; const COMPLETIONS_EXAMPLES: &str = "\ Examples: @@ -260,10 +284,11 @@ Examples: const AUTH_EXAMPLES: &str = "\ Examples: - jevgate auth login Hidden prompt; saved in the OS credential store - jevgate auth login --with-key < key.txt Read the key from stdin + jevgate auth login Asks the kind of key, then a hidden prompt + jevgate auth login --with-key < key.txt Read a TypeSafe key from stdin + jevgate auth login --with-key --provider openrouter < key.txt jevgate auth status Show which key a check would use and verify it - jevgate auth status --offline --json Same, without contacting TypeSafe + jevgate auth status --offline --json Same, without contacting the provider jevgate auth logout"; #[derive(Clone, Copy, Debug, ValueEnum, PartialEq, Eq)] diff --git a/src/options/mod.rs b/src/options/mod.rs index 447ba96..8a55603 100644 --- a/src/options/mod.rs +++ b/src/options/mod.rs @@ -39,6 +39,9 @@ pub enum FailOn { Review, /// New review or consider findings Consider, + /// New findings of the rule's mature levels, measured right at least 80% + /// of the time on projects JevGate was never tuned on (the default) + Mature, /// Files whose answers stayed undecided or that need context Uncertain, /// Nothing; findings are advisory and only an incomplete run exits 2 @@ -50,6 +53,7 @@ impl FailOn { match self { Self::Review => "review", Self::Consider => "consider", + Self::Mature => "mature", Self::Uncertain => "uncertain", Self::None => "none", } @@ -78,7 +82,7 @@ fn fail_on_spec(value: &str) -> Result { None => (None, value.trim()), }; let level = FailOn::parse(level).ok_or_else(|| { - format!("Unknown level {level:?}; use review, consider, uncertain or none") + format!("Unknown level {level:?}; use review, consider, mature, uncertain or none") })?; Ok(FailOnSpec { target, level }) } @@ -89,9 +93,11 @@ const OUTPUT: &str = "Output"; const BUDGETS: &str = "Model, budgets and cache"; const WATCH: &str = "Watch"; -/// The model used when neither `--model` nor `model` in jevgate.toml names one. +/// The model asked with a TypeSafe key when neither `--model` nor `model` in +/// jevgate.toml names one. pub const DEFAULT_MODEL: &str = "jev-1.13.0"; -/// Cache lifetime for the `jev-latest` and `jev-preview` aliases, in seconds. +/// Cache lifetime for an alias (a model name without an `x.y.z` version, such +/// as `jev-latest`), in seconds. pub const DEFAULT_CACHE_TTL_SECS: u64 = 3600; #[derive(Args, Debug)] @@ -106,14 +112,22 @@ pub struct CheckArgs { /// and vendored files are classified and skipped with a reason. /// `upload_allow`/`upload_deny` in jevgate.toml still bound what is sent. pub paths: Vec, - /// Review only files changed against this Git revision (commit, branch or tag) + /// Review only what changed against this Git revision (commit, branch or tag) /// - /// Includes committed, staged, unstaged and untracked changes. Deleted - /// files are listed in the report. The revision must exist locally: in CI, + /// Compares with the fork point, as a pull request diff does, and includes + /// committed, staged, unstaged and untracked changes. Only what the change + /// touches is asked about and reported: functions, tests, comments and + /// values on changed lines, copies where either copy changed, a file's + /// outline when the change adds members to it, and documents naming a + /// path it deleted or renamed. A new file is judged whole. Deleted files + /// are listed in the report. The revision must exist locally: in CI, /// check out with full history (for example `fetch-depth: 0`). When no /// supported file changed, the run is complete and exits 0. #[arg(long, value_name = "REVISION", help_heading = SCOPE)] pub base: Option, + /// With --base, judge each changed file whole, not only what the change touches + #[arg(long, requires = "base", help_heading = SCOPE)] + pub whole_files: bool, /// Also judge tests: test value, redundancy, and shared logic among tests /// /// Without it, test files are judged only for file organization. Also set @@ -147,14 +161,17 @@ pub struct CheckArgs { /// Deselect a rule ID, name, key or group (repeatable); applied after --rule and jevgate.toml #[arg(long = "skip-rule", value_name = "RULE", help_heading = RULES)] pub skip_rules: Vec, - /// What fails the gate: LEVEL for every rule, or TARGET=LEVEL (repeatable) [default: review] + /// What fails the gate: LEVEL for every rule, or TARGET=LEVEL (repeatable) [default: mature] /// - /// LEVEL is review, consider (also fails on review), uncertain, or none - /// (advisory; `report` is accepted as a synonym). TARGET is a rule ID, key - /// or group, for example `security=consider`; the most specific target - /// wins. Flags replace `fail_on` and `[rules]` levels from jevgate.toml - /// for the rules they address. Notes and baselined findings never fail the - /// gate. An incomplete run exits 2 regardless of the gate. + /// LEVEL is review, consider (also fails on review), mature, uncertain, or + /// none (advisory; `report` is accepted as a synonym). `mature` fails only + /// on the levels of a rule measured right at least 80% of the time on + /// projects JevGate was never tuned on (`jevgate rules` shows them); other + /// findings are reported without failing. TARGET is a rule ID, key or + /// group, for example `security=consider`; the most specific target wins. + /// Flags replace `fail_on` and `[rules]` levels from jevgate.toml for the + /// rules they address. Notes and baselined findings never fail the gate. + /// An incomplete run exits 2 regardless of the gate. #[arg(long = "fail-on", value_name = "[TARGET=]LEVEL", value_parser = fail_on_spec, help_heading = RULES)] pub fail_on_specs: Vec, /// The resolved levels for rules without their own: from --fail-on, else configuration. @@ -171,6 +188,10 @@ pub struct CheckArgs { /// who reads an error-detail finding's responses. #[arg(skip)] pub project: Option, + /// The provider of the key the check will use, found before planning so + /// that its default model is the one asked; never from jevgate.toml. + #[arg(skip)] + pub provider: crate::provider::Provider, /// Output format [default: agent; jsonl with --watch; json with --show-requests] #[arg(long, value_enum, help_heading = OUTPUT)] pub format: Option, @@ -198,10 +219,12 @@ pub struct CheckArgs { /// Follow-up requests depend on answers and are not known in advance. #[arg(long, requires = "dry_run", help_heading = OUTPUT)] pub show_requests: bool, - /// TypeSafe model; pin a version for repeatable results [default: jev-1.13.0] + /// Model, as the key's provider names it; pin a version for repeatable results [default: the key's provider's model] /// - /// Also set by `model` in jevgate.toml. Answers are cached per model, so - /// changing it re-asks every unit. + /// The default follows the key: jev-1.13.0 with a TypeSafe key, + /// typesafe/jev-1.13 with an OpenRouter key, typesafe-ai/jev with a Vercel + /// AI Gateway key. Also set by `model` in jevgate.toml. Answers are cached + /// per model, so changing it re-asks every unit. #[arg(long, help_heading = BUDGETS)] pub model: Option, /// Stop after this many API attempts in this invocation, watch updates included @@ -211,30 +234,37 @@ pub struct CheckArgs { /// ceiling this flag can only lower. #[arg(long, value_name = "N", value_parser = clap::value_parser!(u32).range(1..=1000000), help_heading = BUDGETS)] pub max_requests: Option, - /// Maximum simultaneous TypeSafe requests (1-8) - #[arg(long, value_name = "N", default_value_t = 6, value_parser = clap::value_parser!(u32).range(1..=MAX_CONCURRENCY as i64), help_heading = BUDGETS)] - pub concurrency: u32, + /// Maximum simultaneous requests, at most 6; a higher value is lowered to 6 [default: 6, or 3 with a gateway's key] + /// + /// The default follows the key: 6 with a TypeSafe key, 3 with an + /// OpenRouter or Vercel AI Gateway key. Also set by `concurrency` in + /// jevgate.toml, which this flag can only lower. + #[arg(long, value_name = "N", value_parser = concurrency, help_heading = BUDGETS)] + pub concurrency: Option, /// Per-file read limit; a larger file is reported as needs-context, never truncated #[arg(long, value_name = "BYTES", default_value_t = DEFAULT_MAX_FILE_BYTES, value_parser = clap::value_parser!(u64).range(1..=1048576), help_heading = BUDGETS)] pub max_file_bytes: u64, /// Total bytes of --context files per request; context is never truncated #[arg(long, value_name = "BYTES", default_value_t = 32768, value_parser = clap::value_parser!(u64).range(1..=1048576), help_heading = BUDGETS)] pub max_context_bytes: u64, - /// Cache lifetime for the jev-latest and jev-preview aliases [default: 3600] + /// Cache lifetime for an alias, a model name without an x.y.z version such as jev-latest [default: 3600] /// - /// Answers from a pinned model version never expire. Also set by - /// `cache_ttl_secs` in jevgate.toml. + /// Answers from a pinned model version, such as jev-1.13.0, never + /// expire. Also set by `cache_ttl_secs` in jevgate.toml. #[arg(long, value_name = "SECONDS", help_heading = BUDGETS)] pub cache_ttl_secs: Option, /// Ignore cached answers for this invocation and ask again #[arg(long, help_heading = BUDGETS)] pub refresh: bool, - /// Use cached answers only and never contact TypeSafe; unanswered units leave the run incomplete + /// Use cached answers only and never contact the provider; unanswered units leave the run incomplete #[arg(long, conflicts_with = "refresh", help_heading = BUDGETS)] pub cache_only: bool, - /// Credential file holding TYPESAFE_API_KEY [default: /.env] + /// Credential file holding TYPESAFE_API_KEY, OPENROUTER_API_KEY or AI_GATEWAY_API_KEY [default: /.env] /// - /// The TYPESAFE_API_KEY environment variable takes precedence. + /// TYPESAFE_API_KEY in the environment takes precedence; the file comes + /// before the saved key and a gateway's variable in the environment. The + /// repository's .env is read only for TYPESAFE_API_KEY: a gateway's key + /// there is usually the application's own. #[arg(long, value_name = "FILE", help_heading = BUDGETS)] pub env_file: Option, /// Keep running and re-check the selected files after each save @@ -264,8 +294,13 @@ pub struct PathLevels { pub rules: BTreeMap>, } -/// Upper bound on simultaneous requests; rate-limit retries share one cooldown. -pub const MAX_CONCURRENCY: u32 = 8; +/// Upper bound on simultaneous requests, and the default with a TypeSafe +/// key: six workers made 18 to 20 requests a second on the corpus's largest +/// runs (0.3 s a request), just under TypeSafe's limit of 1,200 a minute; +/// eight would make about 27. Rate-limit retries share one cooldown. A +/// higher `--concurrency` is lowered to it with a notice, since 0.25 +/// accepted up to 8; a higher `concurrency` in jevgate.toml means it. +pub const MAX_CONCURRENCY: u32 = 6; /// Default read limit per file. Units are sent separately, so this bounds /// local reading rather than one request. Configuration and @@ -276,6 +311,17 @@ fn names(levels: &[FailOn]) -> Vec { levels.iter().map(|f| f.name().to_string()).collect() } +/// `--concurrency`: at least 1. A higher value than [`MAX_CONCURRENCY`] is +/// accepted here and lowered to it with a notice when the check starts. +fn concurrency(value: &str) -> Result { + match value.parse::() { + Ok(n) if n > 0 => Ok(n), + _ => Err(format!( + "Use a whole number from 1 to {MAX_CONCURRENCY}; a higher one is lowered to {MAX_CONCURRENCY}" + )), + } +} + fn source_extension(value: &str) -> Result { if value.is_empty() || !value.bytes().all(|c| c.is_ascii_alphanumeric()) { return Err("Use an extension without a dot, for example: --source-extension zig".into()); @@ -304,6 +350,11 @@ impl CheckArgs { !crate::catalog::DOCUMENTATION.contains(&key) && key != crate::catalog::WORKFLOWS } + /// Whether a `--base` check judges only what its change touches. + pub fn changed_lines(&self) -> bool { + self.base.is_some() && !self.whole_files + } + /// Whether any documentation rule is selected, so instruction files are found. pub fn documentation(&self) -> bool { crate::catalog::DOCUMENTATION @@ -359,15 +410,49 @@ impl CheckArgs { .collect() } - /// The model to ask: `--model`, else configuration, else [`DEFAULT_MODEL`]. + /// What `mature` stands for, for the report: the mature levels of each + /// selected rule that has some and whose levels include `mature` outside + /// scopes or in one, by rule ID. + pub fn mature_level_names(&self) -> BTreeMap> { + let uses_mature = |key: &str| { + self.levels(key).contains(&FailOn::Mature) + || self.path_fail_on.iter().any(|scope| { + scope + .rules + .get(key) + .is_some_and(|l| l.contains(&FailOn::Mature)) + }) + }; + self.rules + .iter() + .filter(|key| uses_mature(key)) + .filter_map(|key| { + let levels = crate::maturity::mature_levels(key); + let names = levels.iter().map(crate::output::label).collect::>(); + (!names.is_empty()).then(|| (crate::catalog::id(key).to_string(), names)) + }) + .collect() + } + + /// The model to ask: `--model`, else configuration, else the default of + /// the key's provider ([`DEFAULT_MODEL`] for TypeSafe). pub fn model(&self) -> &str { - self.model.as_deref().unwrap_or(DEFAULT_MODEL) + self.model + .as_deref() + .unwrap_or(self.provider.service().default_model) } pub fn cache_ttl_secs(&self) -> u64 { self.cache_ttl_secs.unwrap_or(DEFAULT_CACHE_TTL_SECS) } + /// The most requests sent at once: `--concurrency` or `concurrency` in + /// jevgate.toml, else the default of the key's provider. + pub fn concurrency(&self) -> u32 { + self.concurrency + .unwrap_or(self.provider.service().default_concurrency) + } + pub fn output_format(&self) -> Format { self.format.unwrap_or(if self.show_requests { Format::Json diff --git a/src/output.rs b/src/output.rs index 18055c9..4a7fc2a 100644 --- a/src/output.rs +++ b/src/output.rs @@ -1,6 +1,6 @@ use crate::{ options::{CheckArgs, ColorChoice, Format}, - schema::{FileResult, Finding, Report, Status, Strength}, + schema::{FileResult, Finding, Gating, Report, Scope, Status, Strength}, }; use anyhow::Result; use std::{ @@ -12,24 +12,14 @@ use std::{ /// Consider findings shown by default; `--verbose` shows all. const TOP_CONSIDER: usize = 10; -// Published Jev rate, checked 2026-09-18: -// https://typesafe.ai/blog/introducing-system-one-models-and-jev -pub const INPUT_USD_PER_MILLION: f64 = 0.042; -pub const PRICE_CHECKED: &str = "2026-09-18"; - /// `n` and a noun, plural unless `n` is one: "1 finding", "2 findings". pub fn count(n: usize, noun: &str) -> String { format!("{n} {noun}{}", if n == 1 { "" } else { "s" }) } -/// Estimated dollars for this invocation's paid input tokens, for a priced model. -pub fn estimated_usd(report: &Report) -> Option { - usd(&report.requested_model, report.paid_input_tokens) -} - -/// Estimated dollars for `tokens` input tokens of `model`, when it is priced. -fn usd(model: &str, tokens: u64) -> Option { - (model == "jev-1.13.0").then(|| tokens as f64 * INPUT_USD_PER_MILLION / 1_000_000.0) +/// The headline's cost: estimated dollars, or unknown, never a guessed $0. +fn cost(usd: Option) -> String { + usd.map_or(" · cost unknown".into(), |usd| format!(" · ~${usd:.4}")) } // ANSI select-graphic-rendition codes. @@ -101,8 +91,8 @@ pub fn emit(report: &Report, args: &CheckArgs) -> Result<()> { Style::for_stdout(args.color), ), Format::Github => crate::github::emit(&mut out, report, args), - Format::Sarif => crate::sarif::emit(&mut out, report, args), - Format::Gitlab => crate::gitlab::emit(&mut out, report, args), + Format::Sarif => crate::sarif::emit(&mut out, report), + Format::Gitlab => crate::gitlab::emit(&mut out, report), }; match written { Err(error) if broken_pipe(&error) => Ok(()), @@ -137,6 +127,9 @@ pub(super) fn agent( ) -> Result<()> { emit_header(out, report, style)?; emit_findings(out, report, verbose, style)?; + if let Some(line) = measuring(report) { + writeln!(out, "\n{line}")?; + } emit_summary(out, report)?; if let Some(load) = &report.context_load { emit_context_load(out, load)?; @@ -158,11 +151,11 @@ pub(crate) fn headline(report: &Report) -> String { let planned: u64 = stages.clone().map(|s| s.planned_requests).sum(); let cached: u64 = stages.clone().map(|s| s.planned_cached).sum(); let tokens: u64 = stages.map(|s| s.planned_tokens).sum(); - let cost = usd(&report.requested_model, tokens) - .map_or(String::new(), |usd| format!(" · ~${usd:.4}")); + let cost = cost(crate::model::usd(&report.requested_model, tokens)); return format!( - "JevGate: dry run · {} files · {planned} first-pass requests, {cached} answered by the cache · ~{tokens} new input tokens{cost}; follow-ups depend on the answers", - report.files.len() + "JevGate: dry run · {} files{} · {planned} first-pass requests, {cached} answered by the cache · ~{tokens} new input tokens{cost}; follow-ups depend on the answers", + report.files.len(), + since(report) ); } let gate = match &report.gate { @@ -170,16 +163,47 @@ pub(crate) fn headline(report: &Report) -> String { Some(gate) => format!("gate failed: {}", gate.reasons.join("; ")), None => "gate not evaluated".to_string(), }; - let cost = estimated_usd(report).map_or(String::new(), |usd| format!(" · ~${usd:.4}")); + let cost = cost(report.estimated_usd); format!( - "JevGate: {} · {gate} · {} files · {} API requests · {} input tokens{cost}", + "JevGate: {} · {gate} · {} files{} · {} API requests{} · {} input tokens{cost}", report.status, report.files.len(), + since(report), report.api_requests, + via(report), report.paid_input_tokens ) } +/// ` via OpenRouter` when a gateway answered, so a key found in the +/// environment never bills another account unseen; nothing for TypeSafe. +fn via(report: &Report) -> String { + crate::provider::Provider::named(&report.provider) + .filter(|provider| *provider != crate::provider::Provider::Typesafe) + .map_or(String::new(), |provider| { + format!(" via {}", provider.service().label) + }) +} + +/// Characters of a commit id shown, as Git abbreviates it. +const SHORT_COMMIT: usize = 7; + +/// With a base revision, what the check judged since it: ` · changed lines +/// since 1a2b3c4` or ` · whole files changed since 1a2b3c4`. +fn since(report: &Report) -> String { + let Some(base) = &report.base_revision else { + return String::new(); + }; + let judged = match report.scope { + Scope::ChangedLines => "changed lines", + Scope::WholeFiles => "whole files changed", + }; + format!( + " · {judged} since {}", + base.get(..SHORT_COMMIT).unwrap_or(base) + ) +} + /// The headline, green when the gate passed and red when it failed, then run errors. fn emit_header(out: &mut impl Write, report: &Report, style: Style) -> Result<()> { let code = match &report.gate { @@ -209,10 +233,26 @@ pub(crate) fn ranked(report: &Report) -> Vec<(&Path, &Finding)> { findings } -/// Every review, then the top-ranked considers (all with `verbose`). Notes -/// are listed only with `verbose`; otherwise just counted. +/// Every finding with its file's path: those that fail the gate first, then +/// the rest by level, reviews first, each highest rank first. A capped list +/// never leaves out a failure for a finding that only warns, nor a review +/// still being measured for a higher-ranked consider. GitHub shows 10 +/// warning annotations a step: in whole-repository runs of 94 corpus +/// projects with the default rules, 267 of 424 such reviews fell past the +/// tenth when ranked with considers, and 109 with reviews first, all in the +/// 10 projects holding more than ten of them. +pub(crate) fn failing_first(report: &Report) -> Vec<(&Path, &Finding)> { + let mut findings = ranked(report); + // A stable sort keeps the rank order within each part. + findings.sort_by_key(|(_, f)| (!f.fails_gate(), std::cmp::Reverse(f.strength))); + findings +} + +/// Every review, then the top considers (all with `verbose`), those that +/// fail the gate first. Notes are listed only with `verbose`; otherwise just +/// counted. fn emit_findings(out: &mut impl Write, report: &Report, verbose: bool, style: Style) -> Result<()> { - let findings = ranked(report); + let findings = failing_first(report); let of = |strength: Strength| -> Vec<(&Path, &Finding)> { findings .iter() @@ -230,24 +270,7 @@ fn emit_findings(out: &mut impl Write, report: &Report, verbose: bool, style: St emit_section(out, &heading, BOLD_RED, &review, style)?; } if !consider.is_empty() { - let shown = if verbose { - consider.len() - } else { - TOP_CONSIDER - }; - let more = if consider.len() > shown { - format!(", top {shown}; --verbose shows all") - } else { - String::new() - }; - let heading = format!("Consider ({}{more}):", consider.len()); - emit_section( - out, - &heading, - BOLD_YELLOW, - &consider[..shown.min(consider.len())], - style, - )?; + emit_considers(out, &consider, verbose, style)?; } if notes.is_empty() { return Ok(()); @@ -264,6 +287,33 @@ fn emit_findings(out: &mut impl Write, report: &Report, verbose: bool, style: St emit_section(out, &heading, BOLD, ¬es, style) } +/// The top considers (all with `verbose`), under a heading that says how +/// many there are and which are shown. +fn emit_considers( + out: &mut impl Write, + consider: &[(&Path, &Finding)], + verbose: bool, + style: Style, +) -> Result<()> { + let shown = if verbose { + consider.len() + } else { + TOP_CONSIDER.min(consider.len()) + }; + let more = if consider.len() > shown { + let order = if consider.iter().any(|(_, f)| f.fails_gate()) { + ", those that fail the gate first" + } else { + "" + }; + format!(", top {shown}{order}; --verbose shows all") + } else { + String::new() + }; + let heading = format!("Consider ({}{more}):", consider.len()); + emit_section(out, &heading, BOLD_YELLOW, &consider[..shown], style) +} + /// A blank line, a heading, then its findings. fn emit_section( out: &mut impl Write, @@ -279,35 +329,93 @@ fn emit_section( Ok(()) } -/// Counts of undecided, unsent and failed files, and skip reasons. +/// Why reviews did not fail the gate when their rules and levels are still +/// being measured, with the considers beside them and how to make every +/// review fail it; none when no review is left out that way. +pub(crate) fn measuring(report: &Report) -> Option { + let left_out = |strength: Strength| { + report + .files + .iter() + .flat_map(|f| &f.findings) + .filter(|f| f.strength == strength && f.gate == Some(Gating::Measuring)) + .count() + }; + let (reviews, considers) = (left_out(Strength::Review), left_out(Strength::Consider)); + if reviews == 0 { + return None; + } + let considers = if considers > 0 { + format!(" and {}", count(considers, "consider")) + } else { + String::new() + }; + Some(format!( + "{}{considers} did not fail the gate: by default only rules and levels right at least {}% of the time on projects JevGate was never tuned on fail it, and theirs are still being measured. `jevgate rules` shows each one's precision; `--fail-on review` makes every review fail the gate.", + count(reviews, "review"), + crate::maturity::MIN_PERCENT_RIGHT + )) +} + +/// Why a finding still being measured does not fail the gate, with its rule +/// and level's precision on unseen projects; none for any other finding. +pub(crate) fn measuring_note(finding: &Finding) -> Option { + if finding.gate != Some(Gating::Measuring) { + return None; + } + Some(format!( + "Does not fail the gate: {} {}s are still being measured ({}).", + finding.rule, + label(&finding.strength), + crate::maturity::unseen_share(&finding.rule, finding.strength) + )) +} + +/// Counts of undecided, unsent and failed files, then why files failed and +/// why they were skipped. fn emit_summary(out: &mut impl Write, report: &Report) -> Result<()> { - let count = |status: Status| report.files.iter().filter(|f| f.status == status).count(); + let files = |status: Status| report.files.iter().filter(|f| f.status == status).count(); let undecided = report .files .iter() .filter(|f| f.dimensions.values().any(|d| d.status == Status::Uncertain)) .count(); let summary = [ - (undecided, "with uncertain units"), - (count(Status::NeedsContext), "need context"), - (count(Status::Error), "failed"), + (undecided, "with uncertain units", "with uncertain units"), + (files(Status::NeedsContext), "needs context", "need context"), + (files(Status::Error), "failed", "failed"), ]; let lines: Vec<_> = summary .iter() - .filter(|(n, _)| *n > 0) - .map(|(n, text)| format!("{n} files {text}")) + .filter(|(n, ..)| *n > 0) + .map(|&(n, one, many)| format!("{} {}", count(n, "file"), if n == 1 { one } else { many })) .collect(); if !lines.is_empty() { writeln!(out, "\n{}.", lines.join(" · "))?; } - let mut skipped = BTreeMap::::new(); - for file in report.files.iter().filter(|f| f.status == Status::Skipped) { - *skipped - .entry(file.error.clone().unwrap_or_else(|| "Skipped".into())) + emit_reasons(out, report, Status::Error, "Failed", "Not judged")?; + emit_reasons(out, report, Status::Skipped, "Skipped", "Skipped") +} + +/// Each reason the files of `status` give, with how many give it, as +/// `Failed 1: TypeSafe HTTP 402 (credits exhausted; …)`: a run that could +/// not finish says why without --verbose, and the MCP tool, which returns +/// this text, can tell exhausted credits from a missing key. +fn emit_reasons( + out: &mut impl Write, + report: &Report, + status: Status, + verb: &str, + unknown: &str, +) -> Result<()> { + let mut reasons = BTreeMap::<&str, usize>::new(); + for file in report.files.iter().filter(|f| f.status == status) { + *reasons + .entry(file.error.as_deref().unwrap_or(unknown)) .or_default() += 1; } - for (reason, n) in skipped { - writeln!(out, "Skipped {n}: {reason}")?; + for (reason, n) in reasons { + writeln!(out, "{verb} {n}: {reason}")?; } Ok(()) } @@ -353,8 +461,8 @@ fn emit_context_load(out: &mut impl Write, load: &crate::docs::load::ContextLoad Ok(()) } -/// `path:line [rule] message`, then the next step; the location is bold and -/// the rule dim. +/// `path:line [rule] message`, then the next step; the location is bold, +/// the rule dim, and a finding that fails the gate says so in red. fn emit_finding(out: &mut impl Write, path: &Path, finding: &Finding, style: Style) -> Result<()> { let location = format!("{}:{}", path.display(), finding.line); let accepted = match (&finding.suppressed, finding.baselined) { @@ -363,9 +471,14 @@ fn emit_finding(out: &mut impl Write, path: &Path, finding: &Finding, style: Sty (None, false) => String::new(), }; let rule = format!("[{}]{accepted}", finding.rule); + let fails = if finding.fails_gate() { + format!("{} ", style.paint(RED, "(fails the gate)")) + } else { + String::new() + }; writeln!( out, - " {} {} {}", + " {} {} {fails}{}", style.paint(BOLD, &location), style.paint(DIM, &rule), finding.message diff --git a/src/provider.rs b/src/provider.rs new file mode 100644 index 0000000..1c84e3f --- /dev/null +++ b/src/provider.rs @@ -0,0 +1,315 @@ +//! The services that answer Jev's questions: TypeSafe itself, and the +//! gateways that serve its API under their own keys, OpenRouter and Vercel AI +//! Gateway, at the same price. A key goes only to the host of the provider +//! that issued it, or to `JEVGATE_BASE_URL`, which only the environment sets: +//! nothing in a repository, which the change under review can edit, chooses +//! where a key is sent. +use anyhow::{Result, bail, ensure}; +use clap::ValueEnum; + +/// Whose key a check uses. +#[derive(Clone, Copy, Debug, Default, PartialEq, Eq, ValueEnum)] +pub enum Provider { + /// A TypeSafe key + #[default] + Typesafe, + /// An OpenRouter key + Openrouter, + /// A Vercel AI Gateway key + Vercel, +} + +impl Provider { + /// Every provider, in the order a check reads their keys. + pub const ALL: [Self; 3] = [Self::Typesafe, Self::Openrouter, Self::Vercel]; + + pub fn service(self) -> &'static Service { + match self { + Self::Typesafe => &TYPESAFE, + Self::Openrouter => &OPENROUTER, + Self::Vercel => &VERCEL, + } + } + + /// The name `--provider`, the saved credential and the report use. + pub fn name(self) -> &'static str { + self.service().name + } + + pub fn named(name: &str) -> Option { + Self::ALL + .into_iter() + .find(|provider| provider.name() == name) + } + + /// The provider whose keys start the way `key` does, when that is known. + pub fn issuer(key: &str) -> Option { + Self::ALL.into_iter().find(|provider| { + provider + .service() + .key_prefix + .is_some_and(|prefix| key.starts_with(prefix)) + }) + } +} + +/// One provider of TypeSafe's API. +#[derive(Debug)] +pub struct Service { + pub name: &'static str, + /// The name people read in messages. + pub label: &'static str, + /// The environment variable that holds its key. + pub variable: &'static str, + /// The API root; requests go to `/v1/systemone`. + pub api_root: &'static str, + /// The model asked when neither `--model` nor `model` names one. + pub default_model: &'static str, + /// The most requests sent at once when neither `--concurrency` nor + /// `concurrency` sets it. + pub default_concurrency: u32, + /// A free request that succeeds only with a valid key and sends no source. + pub key_check: KeyCheck, + /// Where to create a key. + pub keys_page: &'static str, + /// What to do when a request is refused for want of credits (HTTP 402). + pub credits: &'static str, + /// How its keys start, when that is known: a key with another + /// provider's prefix is refused rather than sent to the wrong host. + pub key_prefix: Option<&'static str>, +} + +/// Where a key is checked, and what a valid answer holds. +#[derive(Debug)] +pub struct KeyCheck { + pub url: &'static str, + pub answer: KeyAnswer, +} + +#[derive(Clone, Copy, Debug, PartialEq, Eq)] +pub enum KeyAnswer { + /// TypeSafe's model list: `{"models": [{"name": …}]}`. + Models, + /// OpenRouter's key: `{"data": {…}}`. + Key, + /// Vercel's credit balance: `{"balance": "95.50", …}`. + Credits, +} + +/// TypeSafe itself. Its API documents no 402; its terms bill prepaid credits +/// that can refill automatically (MCA §8.2), managed in the console. +pub const TYPESAFE: Service = Service { + name: "typesafe", + label: "TypeSafe", + variable: "TYPESAFE_API_KEY", + api_root: "https://api.typesafe.ai", + default_model: crate::options::DEFAULT_MODEL, + default_concurrency: crate::options::MAX_CONCURRENCY, + key_check: KeyCheck { + url: "https://api.typesafe.ai/v1/models", + answer: KeyAnswer::Models, + }, + keys_page: "https://console.typesafe.ai/settings/keys", + credits: "add credits or turn on auto-refill at https://console.typesafe.ai", + key_prefix: None, +}; + +/// Requests sent at once by default with a gateway's key, half of TypeSafe's +/// 6. Six workers at JevGate's pacing make up to 1,200 requests a minute, +/// TypeSafe's limit for an account; through a gateway the account is the +/// gateway's, shared with its other customers. A precaution rather than a +/// measured fix: on 2026-09-28 OpenRouter answered 503 to about as many +/// attempts (63%) as TypeSafe's own endpoint did at the time (65%), and its +/// rounds of a few requests at once fared only a little better (50%, within +/// noise). +pub const GATEWAY_CONCURRENCY: u32 = 3; + +/// OpenRouter serves TypeSafe's API at `/api/v1/systemone`. `typesafe/jev-1.13` +/// is its name for the Jev 1.13 line, the nearest to the `jev-1.13.0` whose +/// answers JevGate's thresholds were tuned on; `~typesafe/jev-latest` would +/// move to a new major version. It lists no pinned version. +pub const OPENROUTER: Service = Service { + name: "openrouter", + label: "OpenRouter", + variable: "OPENROUTER_API_KEY", + api_root: "https://openrouter.ai/api", + default_model: "typesafe/jev-1.13", + default_concurrency: GATEWAY_CONCURRENCY, + key_check: KeyCheck { + url: "https://openrouter.ai/api/v1/key", + answer: KeyAnswer::Key, + }, + keys_page: "https://openrouter.ai/settings/keys", + credits: "add credits at https://openrouter.ai/settings/credits, or raise the key's limit", + key_prefix: Some("sk-or-"), +}; + +/// Vercel AI Gateway serves TypeSafe's API under `/typesafe`, for Jev only +/// under the name `typesafe-ai/jev`. +pub const VERCEL: Service = Service { + name: "vercel", + label: "Vercel AI Gateway", + variable: "AI_GATEWAY_API_KEY", + api_root: "https://ai-gateway.vercel.sh/typesafe", + default_model: "typesafe-ai/jev", + default_concurrency: GATEWAY_CONCURRENCY, + key_check: KeyCheck { + url: "https://ai-gateway.vercel.sh/v1/credits", + answer: KeyAnswer::Credits, + }, + keys_page: "https://vercel.com/docs/ai-gateway/authentication-and-byok/api-keys", + credits: "add AI Gateway credits in the Vercel dashboard, or raise its budget", + key_prefix: Some("vck_"), +}; + +/// The environment variable that sends requests to another API root, for a +/// self-hosted proxy or a test server. +pub const BASE_URL: &str = "JEVGATE_BASE_URL"; + +/// Where a check sends its requests, and the provider that answers there. +#[derive(Debug)] +pub struct Endpoint { + pub service: &'static Service, + root: String, + /// The root came from `JEVGATE_BASE_URL`. + custom: bool, +} + +impl Endpoint { + /// The provider's API root, or `JEVGATE_BASE_URL` when the environment sets it. + pub fn new(provider: Provider) -> Result { + let service = provider.service(); + match std::env::var(BASE_URL) { + Ok(value) if !value.trim().is_empty() => Self::custom(service, &value), + Ok(_) | Err(std::env::VarError::NotPresent) => Ok(Self { + service, + root: service.api_root.into(), + custom: false, + }), + Err(_) => bail!("{BASE_URL} must contain UTF-8 text"), + } + } + + /// Another API root: `https://` to any host, or `http://` to this machine + /// only, so a key never crosses a network in clear text; no user, query + /// or fragment. + pub fn custom(service: &'static Service, value: &str) -> Result { + let root = value.trim().trim_end_matches('/'); + let rest = root + .strip_prefix("https://") + .or_else(|| root.strip_prefix("http://").filter(|rest| loopback(rest))); + ensure!( + rest.is_some_and(|rest| !host(rest).is_empty() + && rest + .bytes() + .all(|c| c.is_ascii_graphic() && !b"@?#\\".contains(&c))), + "{BASE_URL} must be an https:// URL, or http:// to localhost, 127.0.0.1 or [::1], without a user, query or fragment" + ); + Ok(Self { + service, + root: root.into(), + custom: true, + }) + } + + /// The URL questions are posted to. + pub fn systemone(&self) -> String { + format!("{}/v1/systemone", self.root) + } + + /// Where the key is checked: the provider's own check, or at another + /// root the TypeSafe model list that a proxy of TypeSafe's API serves. + pub fn key_check(&self) -> (String, KeyAnswer) { + if self.custom { + (format!("{}/v1/models", self.root), KeyAnswer::Models) + } else { + let check = &self.service.key_check; + (check.url.into(), check.answer) + } + } + + /// The API root, and whether `JEVGATE_BASE_URL` set it. + pub fn describe(&self) -> String { + if self.custom { + format!("{} ({BASE_URL})", self.root) + } else { + self.root.clone() + } + } + + pub fn is_custom(&self) -> bool { + self.custom + } +} + +/// The host of a URL after its scheme: `127.0.0.1` of `127.0.0.1:4010/api`. +fn host(rest: &str) -> &str { + let authority = rest.split('/').next().unwrap_or_default(); + if authority.starts_with('[') { + authority.split_inclusive(']').next().unwrap_or_default() + } else { + authority.split(':').next().unwrap_or_default() + } +} + +/// Whether a URL's host is this machine. +fn loopback(rest: &str) -> bool { + matches!(host(rest), "localhost" | "127.0.0.1" | "[::1]") +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn a_custom_root_is_https_or_this_machine_and_carries_no_credentials() { + for root in [ + "https://proxy.example.com/typesafe/", + "http://127.0.0.1:4010/api", + "http://localhost:8080", + "http://[::1]:9000", + ] { + let endpoint = Endpoint::custom(&TYPESAFE, root).unwrap(); + assert!(!endpoint.systemone().contains("//v1"), "{root}"); + assert!(endpoint.describe().ends_with("(JEVGATE_BASE_URL)")); + } + for root in [ + "http://proxy.example.com", + "http://127.0.0.1.example.com", + "http://localhost@example.com", + "https://user:pass@example.com", + "https://example.com/?key=1", + "https://example.com/#x", + "https://exa mple.com", + "ftp://127.0.0.1", + "https://", + "127.0.0.1:4010", + ] { + assert!(Endpoint::custom(&TYPESAFE, root).is_err(), "{root}"); + } + let custom = Endpoint::custom(&OPENROUTER, "http://127.0.0.1:1/api").unwrap(); + assert_eq!(custom.systemone(), "http://127.0.0.1:1/api/v1/systemone"); + assert_eq!(custom.key_check().0, "http://127.0.0.1:1/api/v1/models"); + } + + #[test] + fn each_provider_has_its_own_names_and_keys() { + for provider in Provider::ALL { + let value = provider.to_possible_value().unwrap(); + assert_eq!(value.get_name(), provider.name()); + assert_eq!(Provider::named(provider.name()), Some(provider)); + if let Some(prefix) = provider.service().key_prefix { + assert_eq!(Provider::issuer(&format!("{prefix}abc")), Some(provider)); + } + } + assert_eq!(Provider::issuer("tsk-abc"), None); + assert_eq!( + (OPENROUTER.api_root, OPENROUTER.default_model), + ("https://openrouter.ai/api", "typesafe/jev-1.13") + ); + assert_eq!( + (VERCEL.api_root, VERCEL.default_model), + ("https://ai-gateway.vercel.sh/typesafe", "typesafe-ai/jev") + ); + } +} diff --git a/src/provider_error.rs b/src/provider_error.rs index ebb7d17..8331204 100644 --- a/src/provider_error.rs +++ b/src/provider_error.rs @@ -1,28 +1,39 @@ //! What a failed provider response means: an unsent request, a context -//! limit, an edge-firewall block, or another HTTP status. Only verified -//! machine codes are recognized; provider text is never echoed. +//! limit, an edge-firewall block, exhausted credits, an invalid request, or +//! another HTTP status. Only verified machine codes are recognized; provider +//! text is never echoed. +use crate::provider::Service; use serde_json::Value; +use std::time::Duration; /// The connection failed before any request bytes were sent. #[derive(Debug)] -pub(crate) struct Unsent; +pub(crate) struct Unsent(pub &'static Service); impl std::error::Error for Unsent {} impl std::fmt::Display for Unsent { fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - write!(f, "Cannot connect to TypeSafe; request was not sent") + write!( + f, + "Cannot connect to {}; request was not sent", + self.0.label + ) } } /// The request was sent but no answer arrived: it timed out or the connection /// dropped. The provider may have run it, so it is retried only once. #[derive(Debug)] -pub(crate) struct Interrupted; +pub(crate) struct Interrupted(pub &'static Service); impl std::error::Error for Interrupted {} impl std::fmt::Display for Interrupted { fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - write!(f, "TypeSafe request timed out or its connection dropped") + write!( + f, + "{} request timed out or its connection dropped", + self.0.label + ) } } @@ -33,50 +44,231 @@ pub(crate) fn retryable(status: u16) -> bool { matches!(status, 408 | 429 | 500 | 502 | 503 | 504 | 520..=524 | 529) } +/// A failed response as it arrived: its status, body and headers. +#[derive(Default)] +pub(crate) struct Failure<'a> { + pub status: u16, + pub body: Option<&'a str>, + pub retry_after: Option, + pub request_id: Option, +} + #[derive(Debug)] pub(crate) struct ProviderError { + pub service: &'static Service, pub status: u16, pub context_limit: bool, /// A Cloudflare `error code: 10xx` page: the edge refused the client. pub edge_block: bool, - pub retry_after: Option, + pub retry_after: Option, + pub request_id: Option, + /// Where a 422 found the request invalid: each field's path and error type. + pub invalid: Vec, + /// The provider does not know the model asked for. + pub unknown_model: bool, +} + +impl ProviderError { + /// What the status means, when JevGate knows: the bracketed part of the message. + fn meaning(&self) -> Option { + if self.context_limit { + return Some("model context limit exceeded".into()); + } + if self.edge_block { + return Some("blocked by the provider's edge protection".into()); + } + if self.unknown_model { + return Some("unknown model; check the model name".into()); + } + match self.status { + 402 => Some(format!("credits exhausted; {}", self.service.credits)), + 404 => Some("not found; check the model name".into()), + 422 if !self.invalid.is_empty() => { + Some(format!("invalid request: {}", self.invalid.join(", "))) + } + _ => None, + } + } } impl std::error::Error for ProviderError {} impl std::fmt::Display for ProviderError { fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - let detail = if self.context_limit { - " (model context limit exceeded)" - } else if self.edge_block { - " (blocked by the provider's edge protection)" - } else { - "" - }; - let retried = if retryable(self.status) { - "" - } else { - "; request was not retried" - }; - write!(f, "TypeSafe HTTP {}{detail}{retried}", self.status) + write!(f, "{} HTTP {}", self.service.label, self.status)?; + if let Some(meaning) = self.meaning() { + write!(f, " ({meaning})")?; + } + if !retryable(self.status) { + write!(f, "; request was not retried")?; + } + if let Some(id) = &self.request_id { + write!(f, "; request id {id}")?; + } + Ok(()) } } -pub(crate) fn provider_error( - status: u16, - body: Option<&str>, - retry_after: Option, -) -> ProviderError { +pub(crate) fn provider_error(service: &'static Service, failure: Failure<'_>) -> ProviderError { // Recognize only verified machine codes; do not echo arbitrary provider text. - let json = body.and_then(|text| serde_json::from_str::(text).ok()); + let status = failure.status; + let json = failure + .body + .and_then(|text| serde_json::from_str::(text).ok()); ProviderError { + service, status, - context_limit: status == 400 - && json.is_some_and(|body| body["detail"]["error_type"] == "max_tokens_exceeded"), + // A gateway refuses an oversized request with 413 before the model sees it. + context_limit: status == 413 + || (status == 400 + && json + .as_ref() + .is_some_and(|body| body["detail"]["error_type"] == "max_tokens_exceeded")), edge_block: status == 403 - && body.is_some_and(|text| { + && failure.body.is_some_and(|text| { text.trim_start().starts_with("error code: 10") || text.contains("Attention Required! | Cloudflare") }), - retry_after, + retry_after: failure.retry_after, + request_id: failure.request_id, + invalid: if status == 422 { + invalid_fields(json.as_ref()) + } else { + Vec::new() + }, + // TypeSafe answered `jev-1.13`, a name its docs use, with this on + // 2026-09-28; only the message's fixed opening is read. + unknown_model: status == 400 + && json.as_ref().is_some_and(|body| { + body["detail"]["error_type"] == "api_usage_error" + && body["detail"]["message"] + .as_str() + .is_some_and(|message| message.starts_with("Unknown model:")) + }), + } +} + +/// Validation failures a 422 names, at most this many. +const MAX_INVALID: usize = 3; +/// A path segment or type longer than this is not a field name. +const MAX_FIELD_BYTES: usize = 64; + +/// Each `detail[]` entry of a 422 as its `loc` joined by dots and its `type`: +/// `body.questions.q1.criteria missing`. `msg` and `input` are never read, +/// since they can echo the request's source. +fn invalid_fields(body: Option<&Value>) -> Vec { + let Some(details) = body.and_then(|body| body["detail"].as_array()) else { + return Vec::new(); + }; + details + .iter() + .take(MAX_INVALID) + .map(|detail| { + let location: Vec = detail["loc"] + .as_array() + .into_iter() + .flatten() + .map(|part| match part { + Value::Number(index) => index.to_string(), + other => field_name(other.as_str()), + }) + .collect(); + format!( + "{} {}", + location.join("."), + field_name(detail["type"].as_str()) + ) + }) + .collect() +} + +/// A field name or error type fit to print, else `?`. +fn field_name(text: Option<&str>) -> String { + text.filter(|text| { + !text.is_empty() + && text.len() <= MAX_FIELD_BYTES + && text + .bytes() + .all(|c| c.is_ascii_alphanumeric() || b"_-".contains(&c)) + }) + .unwrap_or("?") + .to_owned() +} + +#[cfg(test)] +mod tests { + use super::*; + use crate::provider::TYPESAFE; + use serde_json::json; + + fn message(status: u16, body: Option<&str>, request_id: Option<&str>) -> String { + let failure = Failure { + status, + body, + request_id: request_id.map(Into::into), + ..Default::default() + }; + provider_error(&TYPESAFE, failure).to_string() + } + + #[test] + fn a_422_names_the_invalid_fields_but_never_their_message_or_input() { + let body = json!({"detail": [ + {"loc": ["body", "questions", "simplify_0", "criteria"], "msg": "private source", + "type": "missing", "input": {"source": "private source"}}, + {"loc": ["body", "state", 3], "msg": "private", "type": "string_too_long"}, + {"loc": ["body", "private source text"], "msg": "private", "type": "x"}, + {"loc": ["body", "model"], "msg": "private", "type": "fourth"}, + ]}) + .to_string(); + let text = message(422, Some(&body), Some("req_7")); + assert_eq!( + text, + "TypeSafe HTTP 422 (invalid request: body.questions.simplify_0.criteria missing, body.state.3 string_too_long, body.? x); request was not retried; request id req_7" + ); + assert!(!text.contains("private")); + assert_eq!( + message(422, Some("{\"detail\":\"private\"}"), None), + "TypeSafe HTTP 422; request was not retried" + ); + // TypeSafe's answer to a request without `state`, as received on + // 2026-09-28: its `input` echoes the whole request. + let real = r#"{"detail":[{"type":"missing","loc":["body","state"],"msg":"Field required","input":{"questions":{"q":{"type":"noul","instructions":"private question"}},"model":"jev-1.13.0"}}]}"#; + assert_eq!( + message(422, Some(real), None), + "TypeSafe HTTP 422 (invalid request: body.state missing); request was not retried" + ); + } + + #[test] + fn credits_not_found_and_oversized_requests_say_what_to_do() { + assert_eq!( + message(402, Some("{\"error\":\"private\"}"), None), + "TypeSafe HTTP 402 (credits exhausted; add credits or turn on auto-refill at https://console.typesafe.ai); request was not retried" + ); + assert!(message(404, None, None).contains("(not found; check the model name)")); + // TypeSafe's answer to `"model": "jev-1.13"`, as received on 2026-09-28. + let unknown = + r#"{"detail":{"error_type":"api_usage_error","message":"Unknown model: jev-1.13"}}"#; + assert_eq!( + message(400, Some(unknown), Some("req_01")), + "TypeSafe HTTP 400 (unknown model; check the model name); request was not retried; request id req_01" + ); + let other = r#"{"detail":{"error_type":"api_usage_error","message":"private"}}"#; + assert_eq!( + message(400, Some(other), None), + "TypeSafe HTTP 400; request was not retried" + ); + let oversized = provider_error( + &TYPESAFE, + Failure { + status: 413, + ..Default::default() + }, + ); + assert!(oversized.context_limit); + assert_eq!( + message(503, None, Some("abc")), + "TypeSafe HTTP 503; request id abc" + ); } } diff --git a/src/report.html b/src/report.html index 6754228..fa0642e 100644 --- a/src/report.html +++ b/src/report.html @@ -17,6 +17,7 @@

+

@@ -49,6 +50,7 @@ p($('meta'),new Date(data.generated_at*1000).toLocaleString()); p($('meta'),'Generation '+data.generation+' / '+data.model); p($('meta'),data.refresh?'Auto-refresh every 4 seconds':'Saved snapshot'); +if(data.base)p($('meta'),(data.scope==='changed-lines'?'Changed lines':'Whole files changed')+' since '+data.base.slice(0,7)); // Alert only on failure: a settled run that is incomplete, or a failed gate. // Run errors are listed separately. const failures=[];if(data.settled&&!data.complete)failures.push('Review execution incomplete.');if(data.gate&&!data.gate.passed)failures.push('Gate failed: '+data.gate.reasons.join('; ')+' (fail on '+data.fail_on.join(', ')+').'); @@ -57,8 +59,10 @@ const statItems=[['files',data.files.length],['review findings',strength('review')],['consider findings',strength('consider')],['notes',strength('note')],['unresolved',data.files.filter(f=>['uncertain','needs-context'].includes(f.status)).length],['API requests',data.requests],['input tokens',data.tokens]]; for(const [label,value] of statItems){const item=el('div');item.append(el('strong',value.toLocaleString()),el('span',label));$('stats').append(item)} const cost=el('div');cost.append(el('strong',data.cost?.estimated_usd!==undefined?'$'+data.cost.estimated_usd.toFixed(4):'Unavailable'),el('span','estimated batch cost'));$('stats').append(cost); -$('cost-basis').textContent=data.cost?'USD estimate for this invocation: $'+data.cost.input_per_million+' per million input tokens; output tokens free. Cached results incur no new inference cost. Pricing checked '+data.cost.checked_at+'.':'No published rate configured for this model.'; +$('cost-basis').textContent=data.cost?'USD estimate for this invocation: $'+data.cost.input_per_million+' per million input tokens; output tokens free. Cached results incur no new inference cost. Pricing checked '+data.cost.checked_at+'.':data.unmetered?'Cost unknown: '+data.unmetered+(data.unmetered===1?' response':' responses')+' reported no token usage.':'Cost unknown: no published rate for the model that answered.'; if(data.selected.length)$('selection').textContent='Rules checked: '+data.selected.map(name).join(', ')+'.'+(data.partial?' Only these rules were selected (--rule, --skip-rule or [rules]), so files that only other rules judge are not listed.':''); +const mature=Object.entries(data.fail_on_mature||{}),measuring=findings.filter(f=>f.gate==='measuring').length; +if(data.fail_on.includes('mature')||mature.length)$('gate-policy').textContent='Gate: by default only rules and levels right at least 80% of the time on projects JevGate was never tuned on fail it'+(mature.length?' ('+mature.map(([rule,levels])=>name(rule)+' '+levels.map(l=>l+'s').join(' and ')).join(', ')+')':'')+'.'+(measuring?' '+measuring+(measuring===1?' finding':' findings')+' of rules still being measured '+(measuring===1?'is':'are')+' reported without failing it; --fail-on review makes every review fail it.':''); for(const error of data.errors)p($('errors'),error); if(data.deleted.length)p($('errors'),'Deleted files (no current code reviewed): '+data.deleted.join(', '),'small'); let saved={};try{saved=JSON.parse(sessionStorage.getItem('jevgate:'+data.root)||'{}')}catch{} @@ -70,12 +74,16 @@ function save(){try{sessionStorage.setItem('jevgate:'+data.root,JSON.stringify({query:$('search').value,filter:$('filter').value,open:[...openPaths],page,pageSize:Number($('page-size').value),scroll:window.scrollY}))}catch{}} const order=['review','consider','note','error','needs-context','uncertain','pending','clear','not-applicable','skipped']; const files=[...data.files].sort((a,b)=>order.indexOf(a.status)-order.indexOf(b.status)||a.path.localeCompare(b.path)); +function stillMeasured(finding){const rule=rules.get(finding.rule);const unseen=rule?.maturity?.[finding.strength]?.unseen;const share=unseen&&unseen.labeled?Math.round(100*unseen.right/unseen.labeled)+'% of '+unseen.labeled+' right on projects JevGate was never tuned on':rule?.unmeasured;return 'Does not fail the gate: '+name(finding.rule)+' '+finding.strength+'s are still being measured ('+share+').'} function findingRow(finding){ const row=el('details',null,'row finding-row'),summary=el('summary'),main=el('span',null,'main'); main.append(el('span','Line '+finding.line+(finding.baselined?' · baselined':finding.suppressed?' · allowed: '+finding.suppressed:''),'path'),el('span',finding.message,'lead')); - summary.append(main,el('span',name(finding.rule),'count'),badge(finding.strength));row.append(summary); + summary.append(main,el('span',name(finding.rule),'count')); + if(finding.gate==='fails')summary.append(el('span','fails the gate','badge review')); + summary.append(badge(finding.strength));row.append(summary); const body=el('div',null,'detail');row.append(body); p(body,finding.action); + if(finding.gate==='measuring')p(body,stillMeasured(finding),'small'); p(body,finding.category,'small'); if(finding.locations&&finding.locations.length>1)p(body,'Locations: '+finding.locations.map(l=>l.path+':'+l.start_line+'–'+l.end_line).join(', '),'small'); p(body,'Concern '+pct(finding.concern_probability),'small'); diff --git a/src/requests.rs b/src/requests.rs index 8c80916..59bd9f2 100644 --- a/src/requests.rs +++ b/src/requests.rs @@ -74,7 +74,7 @@ pub(super) fn judgment_key(request: &Value) -> String { /// Aliases move to new model versions, so their answers expire. A pinned /// version answers the same request the same way; its entries never expire. fn cache_ttl(model: &str, ttl: u64) -> Option { - matches!(model, "jev-latest" | "jev-preview").then_some(ttl) + (!crate::model::pinned(model)).then_some(ttl) } /// A valid, unexpired cached answer to `request`, read through `load`; none @@ -171,20 +171,24 @@ impl Session<'_> { let batch: Vec<&Value> = pending.iter().map(|(_, r)| *r).collect(); let store = self.store; let requests_count = &mut self.requests; - let paid = (&mut self.paid_input_tokens, &mut self.paid_output_tokens); + let paid = &mut self.paid; let observed = &mut self.observed; self.evaluator.evaluate_queue( &batch, - self.args.concurrency as usize, + self.args.concurrency() as usize, &before, &mut |index, outcome| { let (i, request) = pending[index]; *requests_count += u32::from(outcome.attempted); let receipt = &mut receipts[i]; - record(store, request, outcome, receipt); - *paid.0 += receipt.metrics.input_tokens; - *paid.1 += receipt.metrics.output_tokens; - if receipt.metrics.evaluated_judgments > 0 { + let billed = record(store, request, outcome, receipt); + // An answer without usage says nothing of its tokens, so it + // stays out of the bytes-per-token calibration. + let metered = billed.as_ref().is_some_and(|b| b.input_tokens.is_some()); + if let Some(billed) = billed { + paid.bill(billed); + } + if metered && receipt.metrics.evaluated_judgments > 0 { observed.0 += serde_json::to_vec(&provider_request(request)) .map_or(0, |v| v.len() as u64); observed.1 += receipt.metrics.input_tokens; @@ -194,13 +198,60 @@ impl Session<'_> { } } -/// Record one outcome: timing, token usage, and a validated answer saved to the cache. +/// What one answered request was billed: the model that answered, and the +/// tokens its response reported. +pub(super) struct Billed { + model: String, + /// None when the response reported no usage. + input_tokens: Option, + output_tokens: u64, +} + +/// What this invocation's requests were billed, and by which models. +#[derive(Default)] +pub struct Usage { + pub input_tokens: u64, + pub output_tokens: u64, + /// Input tokens by the model that answered them. + pub models: BTreeMap, + /// Answers whose response reported no usage. + pub unmetered: u32, +} + +impl Usage { + fn bill(&mut self, billed: Billed) { + self.output_tokens += billed.output_tokens; + match billed.input_tokens { + Some(tokens) => { + self.input_tokens += tokens; + *self.models.entry(billed.model).or_default() += tokens; + } + None => self.unmetered += 1, + } + } + + /// Dollars, priced by the model that answered each request; unknown when + /// an answer reported no usage or a model has no published price. The + /// fold starts at 0.0: a float `sum` of nothing is -0.0, shown as "$-0.0000". + pub fn usd(&self) -> Option { + if self.unmetered > 0 { + return None; + } + self.models.iter().try_fold(0.0, |total, (model, tokens)| { + Some(total + crate::model::usd(model, *tokens)?) + }) + } +} + +/// Record one outcome: timing, token usage, and a validated answer saved to +/// the cache. Returns what an answered request was billed, even when its +/// answer failed validation. fn record( store: &crate::storage::Store, request: &Value, outcome: crate::transport::Outcome, receipt: &mut Receipt, -) { +) -> Option { receipt.metrics.service_ms = outcome.elapsed_ms; receipt.metrics.queue_wait_ms = outcome.started_ms; receipt.metrics.evidence_bytes = if outcome.attempted { @@ -208,9 +259,22 @@ fn record( } else { 0 }; + let mut billed = None; receipt.result = outcome.result.and_then(|body| { - receipt.metrics.input_tokens += usage(&body, "input_tokens"); - receipt.metrics.output_tokens += usage(&body, "output_tokens"); + let input_tokens = response::input_tokens(&body); + let output_tokens = response::output_tokens(&body); + receipt.metrics.input_tokens += input_tokens.unwrap_or(0); + receipt.metrics.output_tokens += output_tokens; + // A name that fails validation is billed to "unknown", which has no price. + let model = body["model"] + .as_str() + .filter(|name| crate::model::valid_name(name)) + .unwrap_or("unknown"); + billed = Some(Billed { + model: model.to_owned(), + input_tokens, + output_tokens, + }); response::validate(&body, request)?; let timestamp = schema::now(); store.save( @@ -229,6 +293,7 @@ fn record( receipt.metrics.failed_attempts = 1; } } + billed } pub(super) fn require_current( @@ -275,16 +340,6 @@ fn require_paths( Ok(()) } -/// A usage count above this is corrupt, not a real count, and is ignored. -const MAX_REPORTED_TOKENS: u64 = 1_000_000_000; - -fn usage(body: &Value, field: &str) -> u64 { - body["usage"][field] - .as_u64() - .filter(|n| *n <= MAX_REPORTED_TOKENS) - .unwrap_or(0) -} - #[cfg(test)] mod tests { use super::*; @@ -297,8 +352,12 @@ mod tests { store.save("old", &body, schema::now() - 7200).unwrap(); for (model, kept) in [ ("jev-1.13.0", true), + ("typesafe/jev-1.13.0", true), ("jev-latest", false), ("jev-preview", false), + ("jev-1.13", false), + ("typesafe/jev-1.13", false), + ("typesafe-ai/jev", false), ] { let ttl = cache_ttl(model, 3600); assert_eq!(store.load("old", ttl).is_some(), kept, "{model}"); diff --git a/src/response.rs b/src/response.rs index f80a9b6..0796115 100644 --- a/src/response.rs +++ b/src/response.rs @@ -7,7 +7,12 @@ use serde_json::{Map, Value}; const ROUNDING_PER_VALUE: f64 = 0.005; const MAX_MASS_ERROR: f64 = 0.05; const FLOAT_NOISE: f64 = 1e-9; +/// A usage count above this is corrupt, not a real count, and is ignored. +const MAX_REPORTED_TOKENS: u64 = 1_000_000_000; +/// A response whose answers match the request's questions. Its `usage` is not +/// required: a gateway need not pass TypeSafe's through, and such an answer +/// counts as unmetered rather than as free. pub fn validate(response: &Value, request: &Value) -> Result<()> { validate_model(response, request)?; let answers = response["answers"].as_object().context("Missing answers")?; @@ -27,33 +32,40 @@ pub fn validate(response: &Value, request: &Value) -> Result<()> { validate_distribution(answer, question)?; } } - for field in ["input_tokens", "output_tokens"] { - ensure!( - response["usage"][field] - .as_u64() - .is_some_and(|n| n <= 1_000_000_000), - "Missing token usage" - ); - } Ok(()) } -/// A well-formed model name, equal to the pinned model when one was requested. +/// The billed input tokens a response reports; none when it reports no usage. +pub fn input_tokens(response: &Value) -> Option { + token_count(response, "input_tokens") +} + +/// The output tokens a response reports, zero when it reports none: they are free. +pub fn output_tokens(response: &Value) -> u64 { + token_count(response, "output_tokens").unwrap_or(0) +} + +fn token_count(response: &Value, field: &str) -> Option { + response["usage"][field] + .as_u64() + .filter(|n| *n <= MAX_REPORTED_TOKENS) +} + +/// A well-formed model name; when a pinned version was requested, that +/// version, with or without a gateway's namespace (`typesafe-ai/jev-1.13.0` +/// asked, `jev-1.13.0` answered). An alias answers with whatever version it +/// points to. fn validate_model(response: &Value, request: &Value) -> Result<()> { let model = response["model"] .as_str() .context("Missing model identity")?; - ensure!( - !model.is_empty() - && model.len() <= 128 - && model - .bytes() - .all(|c| c.is_ascii_alphanumeric() || b"-_.".contains(&c)), - "Invalid model identity" - ); - if let Some(requested) = request["model"].as_str() { + ensure!(crate::model::valid_name(model), "Invalid model identity"); + if let Some(requested) = request["model"] + .as_str() + .filter(|name| crate::model::pinned(name)) + { ensure!( - matches!(requested, "jev-latest" | "jev-preview") || model == requested, + crate::model::base_name(model) == crate::model::base_name(requested), "Provider returned a different pinned model" ); } @@ -151,14 +163,23 @@ fn validate_choice(answer: &Value, probabilities: &Map) -> Result Ok(()) } +/// The cached form of a validated response: its model, the typed answers and, +/// when it reported them, its usage and the provider's request id. pub fn cache_value(response: &Value, request: &Value) -> Value { let mut answers = serde_json::Map::new(); for (key, question) in request["questions"].as_object().unwrap() { let kind = question["type"].as_str().unwrap(); answers.insert(key.clone(), typed_fields(&response["answers"][key], kind)); } - serde_json::json!({"model":response["model"], "answers":answers, - "usage":{"input_tokens":response["usage"]["input_tokens"], "output_tokens":response["usage"]["output_tokens"]}}) + let mut value = serde_json::json!({"model": response["model"], "answers": answers}); + if let Some(input) = input_tokens(response) { + value["usage"] = + serde_json::json!({"input_tokens": input, "output_tokens": output_tokens(response)}); + } + if let Some(id) = crate::response_headers::request_id(response["request_id"].as_str()) { + value["request_id"] = Value::String(id); + } + value } /// Only the fields a typed answer defines; anything else the provider sent is dropped. @@ -175,3 +196,71 @@ fn typed_fields(answer: &Value, kind: &str) -> Value { .collect(), ) } + +#[cfg(test)] +mod tests { + use super::*; + use serde_json::json; + + /// A one-question request for `requested`, and a valid answer from `answered`. + fn exchange(requested: &str, answered: &str) -> (Value, Value) { + let request = json!({"model": requested, "state": "x", + "questions": {"q": {"type": "noul", "instructions": "?"}}}); + let response = json!({"model": answered, "answers": {"q": {"type": "noul", "noul": 0.1}}, + "usage": {"input_tokens": 10, "output_tokens": 1}}); + (request, response) + } + + #[test] + fn an_alias_accepts_any_version_and_a_pinned_name_only_its_own() { + for (requested, answered) in [ + ("jev-1.13.0", "jev-1.13.0"), + ("jev-latest", "jev-1.13.0"), + ("jev-1.13", "jev-1.13.0"), + ("~typesafe/jev-latest", "typesafe/jev-1.13"), + ("typesafe/jev-1.13", "typesafe/jev-1.13"), + ("typesafe-ai/jev", "typesafe-ai/jev"), + ("typesafe-ai/jev-1.13.0", "jev-1.13.0"), + ] { + let (request, response) = exchange(requested, answered); + assert!( + validate(&response, &request).is_ok(), + "{requested} {answered}" + ); + } + for (requested, answered) in [ + ("jev-1.13.0", "jev-1.14.0"), + ("jev-1.13.0", "typesafe/jev-1.13"), + ("jev-latest", "jev 1.13.0"), + ("jev-latest", ""), + ] { + let (request, response) = exchange(requested, answered); + assert!( + validate(&response, &request).is_err(), + "{requested} {answered}" + ); + } + } + + #[test] + fn an_answer_without_usage_is_accepted_as_unmetered_and_cached_without_it() { + let (request, mut response) = exchange("typesafe-ai/jev", "typesafe-ai/jev"); + assert_eq!(input_tokens(&response), Some(10)); + let cached = cache_value(&response, &request); + assert_eq!( + cached["usage"], + json!({"input_tokens": 10, "output_tokens": 1}) + ); + for usage in [ + Value::Null, + json!({"inputTokens": 10}), + json!({"input_tokens": -1}), + ] { + response["usage"] = usage; + assert!(validate(&response, &request).is_ok(), "{response}"); + assert_eq!(input_tokens(&response), None); + assert!(cache_value(&response, &request).get("usage").is_none()); + assert!(validate(&cache_value(&response, &request), &request).is_ok()); + } + } +} diff --git a/src/response_headers.rs b/src/response_headers.rs new file mode 100644 index 0000000..a1e2f58 --- /dev/null +++ b/src/response_headers.rs @@ -0,0 +1,165 @@ +//! What a provider's response headers say: the id support needs to find a +//! request, and how long to wait before sending it again. +use std::time::{Duration, SystemTime, UNIX_EPOCH}; + +/// The header TypeSafe names each request by. +pub const REQUEST_ID: &str = "x-typesafe-request-id"; +/// The longest request id kept; a longer or stranger value is not an id. +const MAX_REQUEST_ID_BYTES: usize = 128; + +/// A request id fit to print: letters, digits and `._:-`. Anything else is +/// dropped, so a header cannot carry text into messages or logs. +pub fn request_id(value: Option<&str>) -> Option { + let value = value?.trim(); + (!value.is_empty() + && value.len() <= MAX_REQUEST_ID_BYTES + && value + .bytes() + .all(|c| c.is_ascii_alphanumeric() || b"._:-".contains(&c))) + .then(|| value.to_owned()) +} + +/// How long the provider asks to wait before a retry: `retry-after-ms`, which +/// TypeSafe's SDKs honor first, else `Retry-After` in seconds or as an HTTP +/// date. A date in the past asks for no wait. +pub fn retry_after(ms: Option<&str>, value: Option<&str>, now: SystemTime) -> Option { + if let Some(wait) = ms + .and_then(|ms| ms.trim().parse::().ok()) + .and_then(|ms| Duration::try_from_secs_f64(ms / 1000.0).ok()) + { + return Some(wait); + } + let value = value?.trim(); + if let Ok(seconds) = value.parse::() { + return Some(Duration::from_secs(seconds)); + } + http_date(value).map(|date| date.duration_since(now).unwrap_or_default()) +} + +const MONTHS: [&str; 12] = [ + "Jan", "Feb", "Mar", "Apr", "May", "Jun", "Jul", "Aug", "Sep", "Oct", "Nov", "Dec", +]; +const SECONDS_PER_DAY: u64 = 86_400; +/// Days from 0000-03-01 to 1970-01-01 in the proleptic Gregorian calendar. +const EPOCH_DAYS: u64 = 719_468; +/// Days in 400 Gregorian years, after which the calendar repeats. +const DAYS_PER_ERA: u64 = 146_097; + +/// An HTTP date in the form HTTP requires senders to use, as in +/// `Sun, 06 Nov 1994 08:49:37 GMT`. +fn http_date(text: &str) -> Option { + let parts: Vec<&str> = text.split_ascii_whitespace().collect(); + if parts.len() != 6 || parts[5] != "GMT" { + return None; + } + let day: u64 = parts[1].parse().ok()?; + let month = MONTHS.iter().position(|m| *m == parts[2])? as u64 + 1; + // The form's year is four digits: a longer one would overflow the time. + let year: u64 = Some(parts[3]) + .filter(|year| year.len() == 4 && year.bytes().all(|b| b.is_ascii_digit()))? + .parse() + .ok() + .filter(|year| *year >= 1970)?; + let clock: Vec = parts[4] + .split(':') + .map(|part| part.parse().ok()) + .collect::>()?; + let [hour, minute, second] = clock[..] else { + return None; + }; + if !(1..=31).contains(&day) || hour > 23 || minute > 59 || second > 60 { + return None; + } + let days = days_since_epoch(year, month, day); + UNIX_EPOCH.checked_add(Duration::from_secs( + days * SECONDS_PER_DAY + hour * 3600 + minute * 60 + second, + )) +} + +/// Days from 1970-01-01 to a date from 1970 on: Howard Hinnant's +/// `days_from_civil`, counting years from March so leap days fall last. +fn days_since_epoch(year: u64, month: u64, day: u64) -> u64 { + let year = if month <= 2 { year - 1 } else { year }; + let (era, year_of_era) = (year / 400, year % 400); + let day_of_year = (153 * ((month + 9) % 12) + 2) / 5 + day - 1; + let day_of_era = year_of_era * 365 + year_of_era / 4 - year_of_era / 100 + day_of_year; + era * DAYS_PER_ERA + day_of_era - EPOCH_DAYS +} + +#[cfg(test)] +mod tests { + use super::*; + + #[test] + fn a_request_id_is_kept_only_when_it_is_safe_to_print() { + assert_eq!( + request_id(Some(" req_01J9-abc.1:2 ")).as_deref(), + Some("req_01J9-abc.1:2") + ); + for value in [ + "", + "id with space", + "id\r\nX-Injected: 1", + "\u{1b}[31m", + &"a".repeat(129), + ] { + assert_eq!(request_id(Some(value)), None, "{value:?}"); + } + assert_eq!(request_id(None), None); + } + + #[test] + fn retry_after_ms_wins_then_seconds_then_an_http_date() { + let now = UNIX_EPOCH + Duration::from_secs(784_111_777); + let wait = |ms, value| retry_after(ms, value, now); + assert_eq!( + wait(Some("1500"), Some("9")), + Some(Duration::from_millis(1500)) + ); + assert_eq!( + wait(Some("12.5"), None), + Some(Duration::from_micros(12_500)) + ); + assert_eq!(wait(Some("-1"), Some("9")), Some(Duration::from_secs(9))); + assert_eq!( + wait(Some("soon"), Some(" 2 ")), + Some(Duration::from_secs(2)) + ); + // 784111777 is Sun, 06 Nov 1994 08:49:37 GMT. + assert_eq!( + wait(None, Some("Sun, 06 Nov 1994 08:50:07 GMT")), + Some(Duration::from_secs(30)) + ); + assert_eq!( + wait(None, Some("Sun, 06 Nov 1994 08:49:00 GMT")), + Some(Duration::ZERO) + ); + for value in [ + "Sunday, 06-Nov-94 08:49:37 GMT", + "Sun, 06 Nov 1994 08:49:37 UTC", + "Sun, 32 Nov 1994 08:49:37 GMT", + "Sun, 06 Nov 500000000000 08:49:37 GMT", + "Sun, 06 Nov 18446744073709551615 08:49:37 GMT", + "Sun, 06 Nov +994 08:49:37 GMT", + "tomorrow", + ] { + assert_eq!(wait(None, Some(value)), None, "{value}"); + } + assert_eq!(wait(None, None), None); + } + + #[test] + fn http_dates_count_leap_days() { + let seconds = |text| { + http_date(text) + .unwrap() + .duration_since(UNIX_EPOCH) + .unwrap() + .as_secs() + }; + assert_eq!(seconds("Thu, 01 Jan 1970 00:00:00 GMT"), 0); + assert_eq!(seconds("Tue, 29 Feb 2000 00:00:00 GMT"), 951_782_400); + assert_eq!(seconds("Mon, 28 Sep 2026 12:00:00 GMT"), 1_790_596_800); + assert_eq!(seconds("Fri, 31 Dec 9999 23:59:59 GMT"), 253_402_300_799); + } +} diff --git a/src/revision.rs b/src/revision.rs index 00d8273..6357f65 100644 --- a/src/revision.rs +++ b/src/revision.rs @@ -1,9 +1,14 @@ -//! Changed-file selection. Git never executes external diff helpers. -use anyhow::{Context, Result, ensure}; +//! What a change against a Git revision holds: the changed and deleted +//! files, and the lines of each file the change touched. Git never executes +//! external diff helpers. +use anyhow::{Context, Result, bail, ensure}; +use serde::Serialize; use std::{ - collections::BTreeMap, + collections::{BTreeMap, BTreeSet}, + io::{BufRead, BufReader, Read}, path::{Path, PathBuf}, - process::Command, + process::{Command, Stdio}, + sync::{Arc, OnceLock}, }; pub struct Changes { @@ -11,6 +16,121 @@ pub struct Changes { /// Current path -> previous path; None denotes a new or untracked file. pub paths: BTreeMap>, pub deleted: Vec, + /// The lines changed in each judged file that existed at the revision, + /// once read by [`Changes::with_lines`]. + pub lines: BTreeMap, +} + +/// The lines of one file a change touched: those it added or modified, and +/// the places it removed lines from without adding any. +#[derive(Clone, Debug, Default, PartialEq, Eq, Serialize)] +pub struct Lines { + /// Added or modified lines of the current file, as inclusive ranges in order. + pub changed: Vec<(usize, usize)>, + /// Lines removed with nothing in their place sit after each of these + /// current lines (0: before the first line). + pub removed_after: Vec, +} + +impl Lines { + /// Whether the change touched lines `start..=end`: it changed one of + /// them, or removed lines between two of them. A removal just before + /// the first line or after the last is not counted: without the old + /// file it belongs as much to the code beside it, and counting it would + /// judge the neighbours of every deleted function. + pub fn touch(&self, start: usize, end: usize) -> bool { + self.changed.iter().any(|&(a, b)| a <= end && start <= b) + || self + .removed_after + .iter() + .any(|&after| start <= after && after < end) + } + + fn add_hunk(&mut self, start: usize, count: usize) { + if count == 0 { + self.removed_after.push(start); + } else { + self.changed.push((start, start + count - 1)); + } + } +} + +/// What a change did to one file, when only what it touched is judged. +#[derive(Clone, Debug)] +pub struct FileChange { + pub lines: Lines, + /// The file as the base revision holds it; none for a document the + /// change left alone, judged only for the paths it removed. + before: Option, + /// Paths the change deleted or renamed away. + pub removed: Arc>, +} + +/// A file at the base revision, read from Git the first time a file-level +/// rule asks for it: a watch poll collects files every 250 ms, and only an +/// evaluation needs the old text. +#[derive(Clone, Debug)] +struct Before { + root: PathBuf, + revision: String, + path: PathBuf, + text: OnceLock>, +} + +impl FileChange { + /// A document the change left alone, judged for the paths it removed. + pub fn unchanged(removed: Arc>) -> Self { + Self { + lines: Lines::default(), + before: None, + removed, + } + } + + /// Whether the change left this document alone: it was selected only + /// for naming a path the change removed. + pub fn left_alone(&self) -> bool { + self.before.is_none() + } + + /// Whether the change adds one of `members`, each a name and the line it + /// starts on, that the file at the base revision lacks; `known` gives that + /// version's names from its path and text. A member the change added + /// starts on a line it added, so the base version is read and parsed only + /// when one does: most changes edit bodies. A base version Git cannot + /// give as text counts as lacking every name, so the rule is asked. + pub fn adds<'a>( + &self, + members: impl Iterator, + known: impl FnOnce(&Path, &str) -> BTreeSet, + ) -> bool { + let starts_changed = |line: usize| { + self.lines + .changed + .iter() + .any(|&(a, b)| a <= line && line <= b) + }; + let fresh: Vec<&str> = members + .filter(|&(_, line)| starts_changed(line)) + .map(|(name, _)| name) + .collect(); + if fresh.is_empty() { + return false; + } + let Some((path, text)) = self.before() else { + return true; + }; + let known = known(path, text); + fresh.iter().any(|name| !known.contains(*name)) + } + + fn before(&self) -> Option<(&Path, &str)> { + let before = self.before.as_ref()?; + let text = before + .text + .get_or_init(|| show(&before.root, &before.revision, &before.path).ok()); + Some((before.path.as_path(), text.as_deref()?)) + } } fn git_path<'a>(fields: &mut impl Iterator, missing: &str) -> Result { @@ -19,15 +139,42 @@ fn git_path<'a>(fields: &mut impl Iterator, missing: &str) -> R )?)) } -pub(crate) fn git(root: &Path, args: &[&str]) -> Result> { - let output = Command::new("git") - .arg("--literal-pathspecs") +/// Git in `root`, without taking optional locks. GIT_DIFF_OPTS is dropped: +/// Git lets it outrank `-U0`, and its context lines would count as changed. +/// +/// The variables that point Git at a repository, such as `GIT_DIR` and +/// `GIT_INDEX_FILE`, are honored, as by any Git command. Git sets them for +/// the hooks, `rebase --exec` and `bisect run` it starts, naming the +/// repository `root` is in: a pre-commit hook's `GIT_INDEX_FILE` is the index +/// being committed, a temporary one for `commit -a` or a partial commit, and +/// a repository kept apart from its work tree is found only through +/// `GIT_DIR`. Every call only reads, except `git diff` refreshing the index's +/// cached file stats, which Git 2.51 does despite GIT_OPTIONAL_LOCKS; what is +/// staged stays as it was. Unit tests drop them here, as the tests' own Git +/// does (`tests/support/git.rs`): inherited from a hook, they would point a +/// test's check at the repository running the tests instead of the one the +/// test built. +pub(crate) fn git_in(root: &Path) -> Command { + let mut command = Command::new("git"); + command .arg("-C") .arg(root) - .args(args) .env("GIT_OPTIONAL_LOCKS", "0") - .output() - .context("Cannot run Git")?; + .env_remove("GIT_DIFF_OPTS"); + #[cfg(test)] + crate::tests::git::isolate(&mut command); + command +} + +/// Git in `root` with `args`, as [`git_in`], taking pathspecs literally. +fn git_command(root: &Path, args: &[&str]) -> Command { + let mut command = git_in(root); + command.arg("--literal-pathspecs").args(args); + command +} + +pub(crate) fn git(root: &Path, args: &[&str]) -> Result> { + let output = git_command(root, args).output().context("Cannot run Git")?; ensure!( output.status.success(), "Git {} failed: {}", @@ -44,7 +191,9 @@ pub(crate) fn git(root: &Path, args: &[&str]) -> Result> { type ChangedPaths = (BTreeMap>, Vec); /// Changed paths since `revision`, each with its previous path (none when -/// added), and deleted paths. Renames keep their source; conflicts stop the review. +/// added), and deleted paths, relative to the root: a `jevgate.toml` in a +/// subdirectory of a Git repository sees the paths below it, as +/// `ls-files` does. Renames keep their source; conflicts stop the review. fn tracked_changes(root: &Path, revision: &str) -> Result { let bytes = git( root, @@ -53,6 +202,7 @@ fn tracked_changes(root: &Path, revision: &str) -> Result { "--no-ext-diff", "--no-textconv", "--find-renames", + "--relative", "--name-status", "-z", revision, @@ -84,6 +234,243 @@ fn tracked_changes(root: &Path, revision: &str) -> Result { Ok((paths, deleted)) } +/// Bytes of paths given to one `git diff`, well under the 32,767 +/// characters a Windows command line holds. +const PATHSPEC_BYTES: usize = 16 * 1024; + +/// `groups` of paths joined in order into batches of about `limit` bytes, +/// a group never split: a renamed file's two names must be diffed together. +fn batches<'a>(groups: &[Vec<&'a str>], limit: usize) -> Vec> { + let mut batches: Vec> = Vec::new(); + let mut bytes = 0; + for group in groups { + let size: usize = group.iter().map(|path| path.len() + 1).sum(); + if batches.is_empty() || bytes + size > limit { + batches.push(Vec::new()); + bytes = 0; + } + bytes += size; + batches.last_mut().unwrap().extend(group); + } + batches +} + +/// The lines each of `paths` changed since `revision`, when it was modified +/// or renamed, from one `git diff -U0` read as it streams: only hunk +/// headers and paths are kept, since a patch can run to many megabytes. +/// Added and deleted files are left out; an added one changed throughout. +/// `--text` gives the lines of a file `.gitattributes` marks `-diff` or +/// `binary`, which Git would otherwise report only as changed. +fn changed_lines(root: &Path, revision: &str, paths: &[&str]) -> Result> { + let mut args = vec![ + "diff", + "-U0", + "--text", + "--inter-hunk-context=0", + "--no-color", + "--no-ext-diff", + "--no-textconv", + "--find-renames", + "--relative", + "--src-prefix=a/", + "--dst-prefix=b/", + "--diff-filter=MRT", + revision, + "--", + ]; + args.extend_from_slice(paths); + let mut child = git_command(root, &args) + .stdin(Stdio::null()) + .stdout(Stdio::piped()) + .stderr(Stdio::piped()) + .spawn() + .context("Cannot run Git")?; + // Read apart from the patch, so a long warning cannot block Git. + let errors = child.stderr.take().map(|mut stderr| { + std::thread::spawn(move || { + let mut text = Vec::new(); + let _ = stderr.read_to_end(&mut text); + text + }) + }); + let stdout = child.stdout.take().context("Git gave no output")?; + let parsed = parse_diff(BufReader::new(stdout)); + if parsed.is_err() { + // Git may still be writing; it must not wait on a full pipe. + let _ = child.kill(); + } + let status = child.wait().context("Cannot run Git")?; + let errors = errors + .and_then(|thread| thread.join().ok()) + .unwrap_or_default(); + ensure!( + status.success() || parsed.is_err(), + "Git diff failed: {}", + String::from_utf8_lossy(&errors).trim() + ); + parsed +} + +/// Bytes of one patch line kept: enough for a header naming any path. +const KEPT_LINE_BYTES: usize = 64 * 1024; + +/// The changed lines per new path of a `-U0` patch. Content lines are +/// counted off each hunk's header, so a line starting `+++ ` inside a hunk +/// is never read as a header. +fn parse_diff(mut reader: impl BufRead) -> Result> { + let mut files = BTreeMap::::new(); + let mut current: Option = None; + // Old and new lines still to come in the current hunk. + let (mut old, mut new) = (0usize, 0usize); + let mut line = Vec::new(); + while read_line(&mut reader, &mut line, KEPT_LINE_BYTES)? { + if old + new > 0 { + match line.first() { + Some(b'-') => old = old.saturating_sub(1), + Some(b'+') => new = new.saturating_sub(1), + Some(b' ') => { + old = old.saturating_sub(1); + new = new.saturating_sub(1); + } + Some(b'\\') => {} + _ => bail!("Git diff ended a hunk early"), + } + continue; + } + if line.starts_with(b"diff ") { + current = None; + } else if let Some(name) = line.strip_prefix(b"+++ ") { + current = new_path(name)?; + if let Some(path) = ¤t { + files.entry(path.clone()).or_default(); + } + } else if line.starts_with(b"@@ ") { + let hunk = hunk_header(&line).context("Git diff has an unreadable hunk header")?; + (old, new) = (hunk.old_count, hunk.count); + if let Some(path) = ¤t { + files + .entry(path.clone()) + .or_default() + .add_hunk(hunk.start, hunk.count); + } + } + } + Ok(files) +} + +/// Read one line into `line`, keeping at most `keep` bytes of it and +/// skipping the rest; false at the end of the input. +fn read_line(reader: &mut impl BufRead, line: &mut Vec, keep: usize) -> Result { + line.clear(); + let mut read = false; + loop { + let buffer = reader.fill_buf()?; + if buffer.is_empty() { + return Ok(read); + } + read = true; + let end = buffer.iter().position(|&b| b == b'\n'); + let chunk = &buffer[..end.unwrap_or(buffer.len())]; + let room = keep.saturating_sub(line.len()); + line.extend_from_slice(&chunk[..chunk.len().min(room)]); + let used = end.map_or(buffer.len(), |end| end + 1); + reader.consume(used); + if end.is_some() { + return Ok(true); + } + } +} + +/// One hunk's counts: old lines replaced, and where its new lines start and +/// how many there are. +struct Hunk { + old_count: usize, + start: usize, + count: usize, +} + +/// `@@ -a[,b] +c[,d] @@ …`, a missing count meaning one line. +fn hunk_header(line: &[u8]) -> Option { + let text = std::str::from_utf8(line.get(3..)?).ok()?; + let mut ranges = text.split(' '); + let range = |text: &str| -> Option<(usize, usize)> { + let (start, count) = text.split_once(',').unwrap_or((text, "1")); + Some((start.parse().ok()?, count.parse().ok()?)) + }; + let (_, old_count) = range(ranges.next()?.strip_prefix('-')?)?; + let (start, count) = range(ranges.next()?.strip_prefix('+')?)?; + Some(Hunk { + old_count, + start, + count, + }) +} + +/// The path of a `+++` line, none for `/dev/null`. Git quotes a name with +/// unusual bytes C-style and puts a tab after a name holding a space. +fn new_path(name: &[u8]) -> Result> { + let name = name.strip_suffix(b"\t").unwrap_or(name); + if name == b"/dev/null" { + return Ok(None); + } + let name = match name.strip_prefix(b"\"") { + Some(quoted) => unquote(quoted.strip_suffix(b"\"").unwrap_or(quoted)), + None => name.to_vec(), + }; + let name = name + .strip_prefix(b"b/") + .context("Git diff path lacks its prefix")?; + Ok(Some(PathBuf::from( + String::from_utf8(name.to_vec()).context("Git path is not UTF-8")?, + ))) +} + +/// A C-quoted name without its quotes: `\t`, `\n`, `\"`, `\\` and octal +/// bytes such as `\303\251`. +fn unquote(quoted: &[u8]) -> Vec { + let mut out = Vec::with_capacity(quoted.len()); + let mut bytes = quoted.iter().copied().peekable(); + while let Some(byte) = bytes.next() { + if byte != b'\\' { + out.push(byte); + continue; + } + let Some(escaped) = bytes.next() else { break }; + out.push(match escaped { + b'a' => b'\x07', + b'b' => b'\x08', + b't' => b'\t', + b'n' => b'\n', + b'v' => b'\x0b', + b'f' => b'\x0c', + b'r' => b'\r', + b'0'..=b'7' => { + let mut value = u32::from(escaped - b'0'); + for _ in 0..2 { + if let Some(digit) = bytes.next_if(|b| (b'0'..=b'7').contains(b)) { + value = value * 8 + u32::from(digit - b'0'); + } + } + value as u8 + } + other => other, + }); + } + out +} + +/// Bytes of a file's base version read to compare its members; a larger one +/// is unknown, like other files too large to parse locally. +const BEFORE_BYTES: usize = 1_048_576; + +/// `path`, as Git writes it, at `revision`, as UTF-8 text. +fn show(root: &Path, revision: &str, path: &Path) -> Result { + let path = path.to_str().context("Path is not UTF-8")?; + let bytes = git(root, &["cat-file", "blob", &format!("{revision}:./{path}")])?; + ensure!(bytes.len() <= BEFORE_BYTES, "Base version is too large"); + String::from_utf8(bytes).context("Base version is not UTF-8") +} + /// The commit to compare with: where `revision` and HEAD diverged, as a pull /// request diff does, so changes made only on the base branch are not reviewed. pub fn resolve(root: &Path, revision: &str) -> Result { @@ -131,6 +518,85 @@ impl Changes { revision, paths, deleted, + lines: BTreeMap::new(), + }) + } + + /// The same changes with the lines each of the `judged` files changed, + /// when it existed at the revision. Only those files are diffed, each + /// with its name before a rename: a changed binary or a regenerated + /// lockfile the check never reads costs nothing, which matters to + /// `--watch`, whose every poll collects the files anew. A 150 MB binary + /// diffed as text took 0.59 s and 460 MB of memory in Git. + pub fn with_lines<'a>( + mut self, + root: &Path, + judged: impl IntoIterator, + ) -> Result { + let mut groups: Vec> = Vec::new(); + for path in judged { + let Some(Some(previous)) = self.paths.get(path) else { + continue; + }; + let mut names: Vec<&str> = [path, previous.as_path()] + .iter() + .filter_map(|name| name.to_str()) + .collect(); + names.dedup(); + groups.push(names); + } + for batch in batches(&groups, PATHSPEC_BYTES) { + self.lines + .extend(changed_lines(root, &self.revision, &batch)?); + } + Ok(self) + } + + /// Paths the change deleted or renamed away, with the directories it + /// emptied, which no longer exist under `root`: a document names + /// `src/legacy/` as readily as a file in it. + pub fn removed(&self, root: &Path) -> BTreeSet { + let renamed = self + .paths + .iter() + .filter_map(|(path, previous)| previous.as_ref().filter(|p| *p != path)); + let mut removed: BTreeSet = self.deleted.iter().chain(renamed).cloned().collect(); + let folders: BTreeSet<&Path> = removed + .iter() + .flat_map(|path| path.ancestors().skip(1)) + .filter(|dir| !dir.as_os_str().is_empty()) + .collect(); + let emptied: Vec = folders + .into_iter() + .filter(|dir| !root.join(dir).exists()) + .map(Path::to_path_buf) + .collect(); + removed.extend(emptied); + removed + } + + /// What the change did to `path`, judged by the lines it touched; none + /// for a file it added, judged whole, or one it left alone. `root` is + /// where Git reads the file's base version. + pub fn file( + &self, + root: &Path, + path: &Path, + removed: &Arc>, + ) -> Option { + let previous = self.paths.get(path)?.as_ref()?; + Some(FileChange { + lines: self.lines.get(path).cloned().unwrap_or_default(), + before: Some(Before { + root: root.to_path_buf(), + revision: self.revision.clone(), + path: previous.clone(), + text: OnceLock::new(), + }), + removed: removed.clone(), }) } } + +#[cfg(test)] +mod tests; diff --git a/src/revision/tests.rs b/src/revision/tests.rs new file mode 100644 index 0000000..12dd06b --- /dev/null +++ b/src/revision/tests.rs @@ -0,0 +1,218 @@ +use super::*; +use crate::tests::Project; + +fn lines(changed: &[(usize, usize)], removed_after: &[usize]) -> Lines { + Lines { + changed: changed.to_vec(), + removed_after: removed_after.to_vec(), + } +} + +#[test] +fn a_patch_gives_each_files_changed_lines_and_removals() { + let patch = concat!( + "diff --git a/src/lib.rs b/src/lib.rs\n", + "index 1111111..2222222 100644\n", + "--- a/src/lib.rs\n", + "+++ b/src/lib.rs\n", + "@@ -3 +3 @@ fn a() {\n", + "- old();\n", + "+ new();\n", + "@@ -10,2 +9,0 @@ fn b() {\n", + "- gone();\n", + "- gone_too();\n", + "@@ -20,0 +19,3 @@ fn c() {\n", + "++++ b/not/a/header.rs\n", + "+@@ -1 +1 @@\n", + "+diff --git a/x b/x\n", + "diff --git a/my file.rs b/my file.rs\n", + "--- a/my file.rs\t\n", + "+++ b/my file.rs\t\n", + "@@ -1 +1,2 @@\n", + "-x\n", + "\\ No newline at end of file\n", + "+x\n", + "+y\n", + "\\ No newline at end of file\n", + "diff --git \"a/caf\\303\\251 \\\"q\\\".rs\" \"b/caf\\303\\251 \\\"q\\\".rs\"\n", + "--- \"a/caf\\303\\251 \\\"q\\\".rs\"\t\n", + "+++ \"b/caf\\303\\251 \\\"q\\\".rs\"\t\n", + "@@ -2 +2 @@\n", + "-a\n", + "+b\n", + "diff --git a/old.rs b/new.rs\n", + "similarity index 90%\n", + "rename from old.rs\n", + "rename to new.rs\n", + "--- a/old.rs\n", + "+++ b/new.rs\n", + "@@ -5 +5 @@\n", + "-a\n", + "+b\n", + "diff --git a/pure.rs b/moved.rs\n", + "similarity index 100%\n", + "rename from pure.rs\n", + "rename to moved.rs\n", + "diff --git a/logo.png b/logo.png\n", + "Binary files a/logo.png and b/logo.png differ\n", + ); + let files = parse_diff(patch.as_bytes()).unwrap(); + let expected: BTreeMap = [ + ("src/lib.rs", lines(&[(3, 3), (19, 21)], &[9])), + ("my file.rs", lines(&[(1, 2)], &[])), + ("café \"q\".rs", lines(&[(2, 2)], &[])), + ("new.rs", lines(&[(5, 5)], &[])), + ] + .into_iter() + .map(|(path, lines)| (PathBuf::from(path), lines)) + .collect(); + assert_eq!(files, expected); + assert!(parse_diff(&b"+++ b/a.rs\n@@ -1,2 +1,2 @@\n-a\n@@ -9 +9 @@\n"[..]).is_err()); +} + +#[test] +fn a_change_touches_the_spans_holding_its_lines_or_its_removals() { + let change = lines(&[(10, 12)], &[20]); + assert!(change.touch(12, 15) && change.touch(1, 10) && change.touch(11, 11)); + assert!(!change.touch(13, 19) && !change.touch(1, 9)); + // Lines removed after line 20 sit inside 18..=25, and at the edge of the + // spans that end at 20 or start at 21. + assert!(change.touch(18, 25)); + assert!(!change.touch(15, 20) && !change.touch(21, 30)); +} + +#[test] +fn long_lines_are_cut_without_losing_the_next() { + let text = format!("{}\nnext\n", "x".repeat(100)); + let mut reader = text.as_bytes(); + let mut line = Vec::new(); + assert!(read_line(&mut reader, &mut line, 8).unwrap()); + assert_eq!(line, b"xxxxxxxx"); + assert!(read_line(&mut reader, &mut line, 8).unwrap()); + assert_eq!(line, b"next"); + assert!(!read_line(&mut reader, &mut line, 8).unwrap()); +} + +const LIB: &str = "fn a() {\n one();\n}\n\nfn b() {\n two();\n three();\n four();\n}\n"; + +#[test] +fn changes_from_git_hold_each_files_lines_and_its_base_text() { + let project = Project::new(); + project.write("src/lib.rs", LIB); + project.write("old.rs", "fn kept() {\n stays();\n}\n"); + project.write("notes.md", "# Notes\n"); + project.write("legacy/one.rs", "fn one() {}\n"); + project.write(".gitattributes", "marked.rs -diff\n"); + project.write("marked.rs", "fn marked() {\n one();\n}\n"); + project.write("Cargo.lock", "version = 3\n"); + project.commit_all(); + std::fs::remove_dir_all(project.0.join("legacy")).unwrap(); + project.write("marked.rs", "fn marked() {\n two();\n}\n"); + project.write("Cargo.lock", "version = 4\n"); + let edited = LIB + .replace(" one();", " uno();") + .replace(" three();\n", ""); + project.write("src/lib.rs", &format!("{edited}fn c() {{ added(); }}\n")); + project.git(&["mv", "old.rs", "renamed.rs"]); + project.write("renamed.rs", "fn kept() {\n stays();\n more();\n}\n"); + std::fs::remove_file(project.0.join("notes.md")).unwrap(); + project.write("new.rs", "fn fresh() {}\n"); + let judged = ["src/lib.rs", "renamed.rs", "marked.rs", "new.rs"].map(Path::new); + let changes = Changes::load(&project.0, "HEAD") + .unwrap() + .with_lines(&project.0, judged) + .unwrap(); + assert!(changes.paths.contains_key(Path::new("Cargo.lock"))); + assert_eq!( + changes.lines.keys().collect::>(), + ["marked.rs", "renamed.rs", "src/lib.rs"].map(Path::new), + "only the judged files are diffed, and a renamed one with its old name" + ); + assert_eq!(changes.paths[Path::new("new.rs")], None); + assert_eq!( + changes.paths[Path::new("renamed.rs")].as_deref(), + Some(Path::new("old.rs")) + ); + let removed = changes.removed(&project.0); + let expected = ["legacy", "legacy/one.rs", "notes.md", "old.rs"]; + assert_eq!(removed, expected.into_iter().map(PathBuf::from).collect()); + let removed = Arc::new(removed); + let lib = changes + .file(&project.0, Path::new("src/lib.rs"), &removed) + .unwrap(); + // `uno` on line 2, `three` removed after line 6 and `c` on line 9. + assert_eq!(lib.lines, lines(&[(2, 2), (9, 9)], &[6])); + let names = |_: &Path, text: &str| -> BTreeSet { + ["a", "b", "c"] + .into_iter() + .filter(|name| text.contains(&format!("fn {name}("))) + .map(String::from) + .collect() + }; + // `c` starts on an added line and is new; `a` and `b` start on lines + // the change left, and `a` pretending to start on line 2 is known. + assert!(lib.adds([("a", 1), ("b", 5), ("c", 9)].into_iter(), names)); + assert!(!lib.adds([("a", 1), ("b", 5)].into_iter(), names)); + assert!(!lib.adds([("a", 2)].into_iter(), names)); + let renamed = changes + .file(&project.0, Path::new("renamed.rs"), &removed) + .unwrap(); + assert_eq!(renamed.lines, lines(&[(3, 3)], &[])); + assert!( + changes + .file(&project.0, Path::new("new.rs"), &removed) + .is_none() + ); + let marked = changes + .file(&project.0, Path::new("marked.rs"), &removed) + .unwrap(); + assert_eq!( + marked.lines, + lines(&[(2, 2)], &[]), + "`-diff` keeps the lines" + ); + assert!(!FileChange::unchanged(removed).adds([("a", 1)].into_iter(), names)); +} + +#[test] +fn a_root_below_the_git_top_level_sees_its_own_paths() { + let project = Project::new(); + project.write("pkg/lib.rs", LIB); + project.write("other/lib.rs", LIB); + project.commit_all(); + project.write("pkg/lib.rs", &LIB.replace("one", "uno")); + project.write("other/lib.rs", &LIB.replace("one", "uno")); + let root = project.0.join("pkg"); + let changes = Changes::load(&root, "HEAD") + .unwrap() + .with_lines(&root, [Path::new("lib.rs")]) + .unwrap(); + assert_eq!( + changes.paths.keys().collect::>(), + [Path::new("lib.rs")] + ); + let file = changes + .file(&root, Path::new("lib.rs"), &Arc::default()) + .unwrap(); + assert_eq!(file.lines, lines(&[(2, 2)], &[])); + assert_eq!(file.before().unwrap().1, LIB); +} + +#[test] +fn paths_are_diffed_in_batches_that_keep_a_renamed_files_names_together() { + let groups = [vec!["a.rs"], vec!["b.rs", "old/b.rs"], vec!["c.rs"]]; + // Each name counts with the space that follows it: 5, 14 and 5 bytes. + assert_eq!( + batches(&groups, 19), + [vec!["a.rs", "b.rs", "old/b.rs"], vec!["c.rs"]] + ); + assert_eq!( + batches(&groups, 10), + [vec!["a.rs"], vec!["b.rs", "old/b.rs"], vec!["c.rs"]], + "a group longer than the limit is a batch of its own, never split" + ); + assert!( + batches(&[], 10).is_empty(), + "no path, no diff of everything" + ); +} diff --git a/src/sarif.rs b/src/sarif.rs index f9c1649..5b9f9f1 100644 --- a/src/sarif.rs +++ b/src/sarif.rs @@ -1,9 +1,7 @@ //! `--format sarif`: the findings as a SARIF 2.1.0 log, for GitHub code //! scanning, GitLab and editors that read static analysis results. use crate::{ - catalog, - options::CheckArgs, - output, + catalog, output, schema::{Finding, Report, Status, Strength}, }; use anyhow::Result; @@ -14,30 +12,26 @@ const SCHEMA: &str = "https://json.schemastore.org/sarif-2.1.0.json"; const HOME: &str = "https://github.com/Tech-Byte-Frontier/jevgate"; /// The same findings the GitHub annotations show: every new finding that is -/// not a note, an `error` when it fails the gate and a `warning` otherwise. -/// Run errors and files that could not be judged are tool notifications. -pub fn emit(out: &mut impl Write, report: &Report, args: &CheckArgs) -> Result<()> { +/// not a note, an `error` when it fails the gate and a `warning` otherwise, +/// with how the gate counted it as the `gate` property. Run errors and files +/// that could not be judged are tool notifications. +pub fn emit(out: &mut impl Write, report: &Report) -> Result<()> { let shown: Vec<(&Path, &Finding)> = output::ranked(report) .into_iter() .filter(|(_, f)| f.strength != Strength::Note && !f.accepted()) .collect(); - serde_json::to_writer_pretty(&mut *out, &document(report, &shown, args))?; + serde_json::to_writer_pretty(&mut *out, &document(report, &shown))?; writeln!(out)?; Ok(()) } -fn document(report: &Report, shown: &[(&Path, &Finding)], args: &CheckArgs) -> Value { +fn document(report: &Report, shown: &[(&Path, &Finding)]) -> Value { let rules = catalog::rules(); let results: Vec = shown .iter() .map(|(path, finding)| { let index = rules.iter().position(|r| r.id == finding.rule); - result( - path, - finding, - index, - crate::gate::fails(finding, path, args), - ) + result(path, finding, index) }) .collect(); let mut notifications: Vec = report @@ -101,7 +95,7 @@ fn title(id: &str) -> String { }) } -fn result(path: &Path, finding: &Finding, rule_index: Option, fails: bool) -> Value { +fn result(path: &Path, finding: &Finding, rule_index: Option) -> Value { let end = finding .locations .iter() @@ -122,10 +116,14 @@ fn result(path: &Path, finding: &Finding, rule_index: Option, fails: bool }) }) .collect(); + let mut text = format!("{}\n\nNext step: {}", finding.message, finding.action); + if let Some(note) = output::measuring_note(finding) { + text.push_str(&format!("\n\n{note}")); + } let mut value = json!({ "ruleId": finding.rule, - "level": if fails { "error" } else { "warning" }, - "message": {"text": format!("{}\n\nNext step: {}", finding.message, finding.action)}, + "level": if finding.fails_gate() { "error" } else { "warning" }, + "message": {"text": text}, "locations": [{"physicalLocation": { "artifactLocation": artifact(path), "region": {"startLine": finding.line.max(1), "endLine": end.max(1)}, @@ -145,6 +143,9 @@ fn result(path: &Path, finding: &Finding, rule_index: Option, fails: bool if let Some(category) = &finding.category { value["properties"]["category"] = json!(category); } + if let Some(gate) = finding.gate { + value["properties"]["gate"] = json!(gate); + } value } @@ -156,7 +157,7 @@ fn artifact(path: &Path) -> Value { #[cfg(test)] mod tests { use super::*; - use crate::tests::finding; + use crate::{options::CheckArgs, schema::Gating, tests::finding}; fn report(args: &CheckArgs) -> Report { crate::evaluate::snapshot( @@ -174,20 +175,31 @@ mod tests { #[test] fn results_name_their_rule_level_location_and_fingerprint() { let args = crate::tests::args(); - let review = finding(Strength::Review); - let consider = finding(Strength::Consider); + let review = Finding { + gate: Some(Gating::Fails), + ..finding(Strength::Review) + }; + let measuring = Finding { + gate: Some(Gating::Measuring), + ..finding(Strength::Review) + }; let path = Path::new("src/a,b.rs"); - let log = document(&report(&args), &[(path, &review), (path, &consider)], &args); + let log = document(&report(&args), &[(path, &review), (path, &measuring)]); assert_eq!(log["version"], "2.1.0"); let run = &log["runs"][0]; let rules = run["tool"]["driver"]["rules"].as_array().unwrap(); assert_eq!(rules.len(), catalog::rules().len()); let results = run["results"].as_array().unwrap(); - assert_eq!( - results[0]["level"], "error", - "a review fails the default gate" - ); + assert_eq!(results[0]["level"], "error", "a review that fails the gate"); + assert_eq!(results[0]["properties"]["gate"], "fails"); assert_eq!(results[1]["level"], "warning"); + assert_eq!(results[1]["properties"]["gate"], "measuring"); + assert!( + results[1]["message"]["text"] + .as_str() + .unwrap() + .ends_with("reviews are still being measured (54% of 85 right on projects JevGate was never tuned on).") + ); let first = &results[0]; let index = first["ruleIndex"].as_u64().unwrap() as usize; assert_eq!(rules[index]["id"], "maintainability/shared-logic"); @@ -209,7 +221,7 @@ mod tests { #[test] fn security_rules_are_tagged_for_code_scanning() { let args = crate::tests::args(); - let log = document(&report(&args), &[], &args); + let log = document(&report(&args), &[]); let rules = log["runs"][0]["tool"]["driver"]["rules"] .as_array() .unwrap(); diff --git a/src/schema/mod.rs b/src/schema/mod.rs index a2e76cd..e618622 100644 --- a/src/schema/mod.rs +++ b/src/schema/mod.rs @@ -70,6 +70,9 @@ pub struct Judgment { pub version: String, pub pass: Pass, pub answer: Answer, + /// The provider's id for the request that answered, to quote to its support. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub request_id: Option, } pub fn now() -> u64 { diff --git a/src/schema/report.rs b/src/schema/report.rs index d742dde..2ff34ac 100644 --- a/src/schema/report.rs +++ b/src/schema/report.rs @@ -122,6 +122,10 @@ pub struct Finding { /// finding is accepted as a baselined one is. #[serde(default, skip_serializing_if = "Option::is_none")] pub suppressed: Option, + /// How the gate counted it: none for notes and accepted findings, and + /// before the gate is applied. + #[serde(default, skip_serializing_if = "Option::is_none")] + pub gate: Option, } impl Finding { @@ -129,6 +133,26 @@ impl Finding { pub fn accepted(&self) -> bool { self.baselined || self.suppressed.is_some() } + + /// Whether the gate counted it as a failure. + pub fn fails_gate(&self) -> bool { + self.gate == Some(Gating::Fails) + } +} + +/// How the gate counted a new finding, at the level in force for its rule +/// and path. Set even when the run is incomplete and the gate not evaluated. +#[derive(Debug, Clone, Copy, Serialize, Deserialize, PartialEq, Eq)] +#[serde(rename_all = "kebab-case")] +pub enum Gating { + /// It fails the gate. + Fails, + /// Reported without failing: the level is `mature`, and this rule and + /// level is still being measured. + Measuring, + /// Reported without failing: the level does not count it, as `review` + /// does not count a consider and `none` counts nothing. + Advisory, } #[derive(Debug, Clone, Serialize, Deserialize)] @@ -193,6 +217,19 @@ pub struct StageMetrics { pub evidence_bytes: u64, } +/// What a check judged of its files. +#[derive(Debug, Clone, Copy, Default, Serialize, Deserialize, PartialEq, Eq)] +#[serde(rename_all = "kebab-case")] +pub enum Scope { + /// Every unit of each selected file. + #[default] + WholeFiles, + /// With a base revision, only what the change touched: units on changed + /// lines, copies where either copy changed, outlines the change added + /// members to, and documents naming a path it removed. + ChangedLines, +} + #[derive(Debug, Clone, Serialize, Deserialize)] pub struct Report { #[serde(default)] @@ -201,6 +238,8 @@ pub struct Report { pub base_revision: Option, #[serde(default)] pub deleted_files: Vec, + #[serde(default)] + pub scope: Scope, pub schema_version: u32, pub command: String, pub rubric_version: String, @@ -218,6 +257,9 @@ pub struct Report { #[serde(default, skip_serializing_if = "Vec::is_empty")] pub initial_requests: Vec, pub requested_model: String, + /// Whose key the requests use: `typesafe`, `openrouter` or `vercel`. + #[serde(default)] + pub provider: String, pub decision_policy: BTreeMap, /// The configured gate policy; classification does not depend on it. #[serde(default)] @@ -228,6 +270,11 @@ pub struct Report { /// Gate levels for the files `[[scope]]` paths match, in configuration order. #[serde(default, skip_serializing_if = "Vec::is_empty")] pub fail_on_paths: Vec, + /// What the `mature` level stands for: the levels of each selected rule + /// measured right at least 80% of the time on projects JevGate was never + /// tuned on, by rule ID, for the rules whose levels include `mature`. + #[serde(default, skip_serializing_if = "BTreeMap::is_empty")] + pub fail_on_mature: BTreeMap>, #[serde(default, skip_serializing_if = "Option::is_none")] pub gate: Option, pub api_requests: u32, @@ -235,6 +282,18 @@ pub struct Report { pub concurrency: u32, pub paid_input_tokens: u64, pub paid_output_tokens: u64, + /// This invocation's paid input tokens by the model that answered them, + /// which for an alias is the version it pointed to. + #[serde(default, skip_serializing_if = "BTreeMap::is_empty")] + pub paid_models: BTreeMap, + /// Paid requests whose response reported no token usage; their tokens are + /// not in `paid_input_tokens`, and the cost is unknown. + #[serde(default)] + pub unmetered_requests: u32, + /// Estimated dollars for this invocation's paid requests, priced by the + /// model that answered each; null when unknown. + #[serde(default)] + pub estimated_usd: Option, #[serde(default)] pub stages: BTreeMap, pub settled: bool, diff --git a/src/tests/gating.rs b/src/tests/gating.rs index af1ff39..dbeff74 100644 --- a/src/tests/gating.rs +++ b/src/tests/gating.rs @@ -15,7 +15,7 @@ fn gate_fails_only_on_the_configured_results() { ]; for (level, status, default, stricter, stricter_code) in cases { options.refresh = true; - options.fail_on = vec![options::FailOn::Review]; + options.fail_on = vec![options::FailOn::Mature]; let mut mock = Mock { level, ..Default::default() @@ -60,6 +60,18 @@ fn a_rule_level_fails_the_gate_only_for_that_rule() { } } +/// A `[[scope]]` that sets `levels` for every rule in the files `glob` matches. +fn every_rule_in(glob: &str, levels: Vec) -> options::PathLevels { + options::PathLevels { + paths: vec![glob.into()], + matcher: crate::boundary::globs(&[glob.into()]).unwrap(), + rules: crate::catalog::keys() + .into_iter() + .map(|key| (key.to_string(), levels.clone())) + .collect(), + } +} + #[test] fn a_scope_makes_its_paths_report_only_while_other_files_gate() { let project = Project::new(); @@ -71,14 +83,7 @@ fn a_scope_makes_its_paths_report_only_while_other_files_gate() { level: 4, ..Default::default() }; - let scripts = |levels: Vec| options::PathLevels { - paths: vec!["scripts/**".into()], - matcher: crate::boundary::globs(&["scripts/**".into()]).unwrap(), - rules: crate::catalog::keys() - .into_iter() - .map(|key| (key.to_string(), levels.clone())) - .collect(), - }; + let scripts = |levels| every_rule_in("scripts/**", levels); options.path_fail_on = vec![scripts(vec![options::FailOn::None])]; let report = run(&project, &options, &mut mock); let gate = report.gate.as_ref().unwrap(); @@ -163,6 +168,36 @@ fn a_merged_baseline_keeps_accepted_findings_for_files_the_check_did_not_cover() assert_eq!(gate::exit_code(&run(&project, &options, &mut review)), 1); } +#[test] +fn a_merged_baseline_after_a_check_of_changed_lines_keeps_the_rest_of_its_files() { + let project = Project::new(); + let source = format!("{}{}", function("a"), function("b")); + project.write("lib.rs", &source); + let mut options = args(); + let mut review = Mock { + level: 2, + ..Default::default() + }; + publish(&project, &run(&project, &options, &mut review)); + assert_eq!( + baseline::write(&project.0, false, None).unwrap().accepted, + 2 + ); + project.commit_all(); + // The change touches `a` alone: `b`'s accepted finding was not judged. + project.write("lib.rs", &source.replacen("doubled + 1", "doubled + 2", 1)); + options.base = Some("HEAD".into()); + let report = run(&project, &options, &mut review); + assert_eq!(report.scope, schema::Scope::ChangedLines); + publish(&project, &report); + let merged = baseline::write(&project.0, true, None).unwrap(); + assert_eq!((merged.accepted, merged.kept), (1, 2)); + options.base = None; + let report = run(&project, &options, &mut review); + assert!(report.files[0].findings.iter().all(|f| f.baselined)); + assert_eq!(gate::exit_code(&report), 0); +} + #[test] fn baseline_reasons_are_marked_counted_and_kept_across_rewrites() { use options::Disposition::{Later, Wrong}; @@ -241,3 +276,129 @@ fn an_allow_comment_accepts_a_finding_only_with_a_reason() { ); } } + +const SIMPLIFICATION: &str = "maintainability/function-simplification"; +const SHARED_LOGIC: &str = "maintainability/shared-logic"; + +/// How the gate counted each finding of a report, in order. +fn gates(report: &schema::Report) -> Vec> { + report.files[0].findings.iter().map(|f| f.gate).collect() +} + +#[test] +fn the_default_gate_fails_only_on_mature_rule_levels() { + use schema::{Gating::*, Strength::*}; + let accepted = schema::Finding { + baselined: true, + ..finding_of(SIMPLIFICATION, Review) + }; + let findings = vec![ + finding_of(SIMPLIFICATION, Review), + finding_of(SIMPLIFICATION, Consider), + finding_of(SHARED_LOGIC, Review), + finding_of(SIMPLIFICATION, Note), + accepted, + ]; + let report = gated(findings.clone(), &args()); + assert_eq!( + gates(&report), + [Some(Fails), Some(Measuring), Some(Measuring), None, None] + ); + assert_eq!( + report.gate.as_ref().unwrap().reasons, + ["1 new review finding"] + ); + assert_eq!(gate::exit_code(&report), 1); + let measured = gated(findings[1..].to_vec(), &args()); + let gate = measured.gate.as_ref().unwrap(); + assert!(gate.passed, "reviews still being measured are reported"); + assert_eq!((gate.new_findings, gate.baselined_findings), (2, 1)); + assert_eq!(gate::exit_code(&measured), 0); +} + +#[test] +fn an_explicit_level_replaces_the_default_exactly_as_it_says() { + use schema::{Gating::*, Strength::*}; + let findings = vec![ + finding_of(SIMPLIFICATION, Review), + finding_of(SHARED_LOGIC, Review), + finding_of(SHARED_LOGIC, Consider), + ]; + let mut options = args(); + options.fail_on = vec![options::FailOn::Review]; + let report = gated(findings.clone(), &options); + assert_eq!(gates(&report), [Some(Fails), Some(Fails), Some(Advisory)]); + assert_eq!(report.gate.unwrap().reasons, ["2 new review findings"]); + // A level for one rule leaves the others at the default. + let mut options = args(); + options.rule_fail_on = std::collections::BTreeMap::from([( + crate::catalog::SHARED_LOGIC.to_string(), + vec![options::FailOn::None], + )]); + let report = gated(findings.clone(), &options); + assert_eq!( + gates(&report), + [Some(Fails), Some(Advisory), Some(Advisory)] + ); + // A scope's report level covers the mature rule too. + let mut options = args(); + options.path_fail_on = vec![every_rule_in("src/**", vec![options::FailOn::None])]; + let report = gated(findings, &options); + assert_eq!(gates(&report), [Some(Advisory); 3]); + assert_eq!(gate::exit_code(&report), 0); +} + +#[test] +fn agent_text_marks_what_fails_and_says_why_the_rest_did_not() { + use schema::Strength::*; + let report = gated( + vec![ + finding_of(SIMPLIFICATION, Review), + finding_of(SHARED_LOGIC, Review), + finding_of(SHARED_LOGIC, Consider), + ], + &args(), + ); + let mut out = Vec::new(); + output::agent(&mut out, &report, false, output::Style::PLAIN).unwrap(); + let text = String::from_utf8(out).unwrap(); + assert!( + text.contains("[maintainability/function-simplification] (fails the gate) Copies"), + "{text}" + ); + assert!( + text.contains("[maintainability/shared-logic] Copies"), + "{text}" + ); + assert!( + text.contains("\n1 review and 1 consider did not fail the gate: by default only rules and levels right at least 80% of the time on projects JevGate was never tuned on fail it, and theirs are still being measured."), + "{text}" + ); + let considers_only = gated(vec![finding_of(SHARED_LOGIC, Consider)], &args()); + assert!( + output::measuring(&considers_only).is_none(), + "considers never failed by default" + ); +} + +#[test] +fn a_capped_list_shows_the_findings_that_fail_the_gate_first() { + use schema::Strength::*; + let failing = schema::Finding { + rank: 0.1, + ..finding_of("documentation/agent-context", Consider) + }; + let mut findings = vec![finding_of(SHARED_LOGIC, Consider); 11]; + findings.push(failing); + let report = gated(findings, &args()); + let mut out = Vec::new(); + output::agent(&mut out, &report, false, output::Style::PLAIN).unwrap(); + let text = String::from_utf8(out).unwrap(); + let section = text.split("Consider (").nth(1).unwrap(); + assert!( + section.starts_with( + "12, top 10, those that fail the gate first; --verbose shows all):\n src/lib.rs:12 [documentation/agent-context] (fails the gate)" + ), + "{text}" + ); +} diff --git a/src/tests/mod.rs b/src/tests/mod.rs index 278bb74..0d0d9a9 100644 --- a/src/tests/mod.rs +++ b/src/tests/mod.rs @@ -7,6 +7,11 @@ use clap::Parser; use serde_json::{Value, json}; use std::path::PathBuf; mod gating; +#[path = "../../tests/support/git.rs"] +pub(super) mod git; +#[path = "../../tests/support/mock_provider.rs"] +pub(super) mod mock_provider; +pub(super) use mock_provider::answer; mod scope; #[path = "../../tests/support/temp_dir.rs"] mod temp_dir; @@ -27,6 +32,17 @@ impl Project { config: Default::default(), } } + /// Run Git in the project, with a fixed identity and no signing, apart + /// from the repository running the tests. + pub(super) fn git(&self, args: &[&str]) { + git::run(&self.0, args); + } + /// A Git repository holding the project's files as its first commit. + pub(super) fn commit_all(&self) { + self.git(&["init", "-q"]); + self.git(&["add", "-A"]); + self.git(&["commit", "-qm", "base"]); + } } #[derive(Parser)] @@ -37,11 +53,10 @@ struct TestCli { pub(super) fn args() -> CheckArgs { let mut a = TestCli::parse_from(["test"]).args; a.rules = crate::catalog::keys().into_iter().map(Into::into).collect(); - a.fail_on = vec![options::FailOn::Review]; + a.fail_on = vec![options::FailOn::Mature]; a } -/// A function large enough to judge (five body lines). /// A shared-logic finding at `src/a,b.rs:12`, with text the output formats escape. pub(super) fn finding(strength: crate::schema::Strength) -> crate::schema::Finding { crate::schema::Finding { @@ -66,9 +81,31 @@ pub(super) fn finding(strength: crate::schema::Strength) -> crate::schema::Findi rank: 1.0, baselined: false, suppressed: None, + gate: None, + } +} + +/// A finding of `rule` (an ID) at `strength`, otherwise as [`finding`]. +pub(super) fn finding_of(rule: &str, strength: crate::schema::Strength) -> crate::schema::Finding { + crate::schema::Finding { + rule: rule.into(), + ..finding(strength) } } +/// A complete report of one judged file, `src/lib.rs`, holding `findings`, +/// with the gate applied as `options` set it. +pub(super) fn gated(findings: Vec, options: &CheckArgs) -> schema::Report { + let project = Project::new(); + project.write("src/lib.rs", &function("f")); + let (_, mut report) = snapshot(&project, options); + report.files[0].findings = findings; + report.complete = true; + gate::evaluate(&mut report, options); + report +} + +/// A function large enough to judge (five body lines). pub(super) fn function(name: &str) -> String { format!( "fn {name}(values: &[i32]) -> i32 {{\n let mut total = 0;\n for value in values {{\n total += value;\n }}\n let doubled = total * 2;\n doubled + 1\n}}\n" @@ -82,55 +119,6 @@ pub(super) fn long_function(name: &str) -> String { ) } -/// Levels: 0 answers the bottom of every scale (clear), 1 the middle (consider, -/// or a note where the middle says the code is fine), 2 the top (review), -/// 3 spreads probability (uncertain), 4 leans to the top without reaching review. -pub(super) fn answer(request: &Value, level: usize) -> Value { - let answers = request["questions"] - .as_object() - .unwrap() - .iter() - .map(|(name, q)| (name.clone(), typed_answer(q, level))) - .collect::>(); - json!({"model":request["model"],"answers":answers,"usage":{"input_tokens":10,"output_tokens":0}}) -} - -/// A valid answer of the question's type at `level` (see [`answer`]). -fn typed_answer(question: &Value, level: usize) -> Value { - match question["type"].as_str().unwrap() { - "noul" => { - let noul = [0.05, 0.5, 0.95, 0.5, 0.5][level]; - json!({"type":"noul","noul":noul}) - } - "score" => { - let p = [ - [1.0, 0.0, 0.0], - [0.0, 1.0, 0.0], - [0.0, 0.0, 1.0], - [0.4, 0.2, 0.4], - [0.1, 0.3, 0.6], - ][level]; - json!({"type":"score","score":p[1] + 2.0 * p[2],"confidence":1.0, - "probabilities":{"0":p[0],"1":p[1],"2":p[2]}}) - } - _ => choice_answer(question["criteria"].as_object().unwrap()), - } -} - -/// A certain Choice of `none` when offered, else the first option. -fn choice_answer(options: &serde_json::Map) -> Value { - let chosen = if options.contains_key("none") { - "none" - } else { - options.keys().next().unwrap() - }; - let probabilities: serde_json::Map<_, _> = options - .keys() - .map(|k| (k.clone(), json!(if k == chosen { 1.0 } else { 0.0 }))) - .collect(); - json!({"type":"choice","choice":chosen,"confidence":1.0,"probabilities":probabilities}) -} - #[derive(Default)] pub(super) struct Mock { pub(super) calls: usize, @@ -166,8 +154,7 @@ pub(super) fn session<'a>( store, evaluator, requests: 0, - paid_input_tokens: 0, - paid_output_tokens: 0, + paid: Default::default(), budget: token_budget::TokenBudget::default(), observed: (0, 0), } @@ -278,6 +265,57 @@ fn model_and_refresh_invalidate_cache() { assert_eq!(mock.calls, 3); } +/// Answers as a provider that names `model` and reports usage only when `metered`. +struct Answering { + model: &'static str, + metered: bool, +} +impl transport::Evaluator for Answering { + fn evaluate(&mut self, request: &Value) -> anyhow::Result { + let mut body = answer(request, 0); + body["model"] = json!(self.model); + if !self.metered { + body.as_object_mut().unwrap().remove("usage"); + } + Ok(body) + } +} + +#[test] +fn cost_is_priced_by_the_answering_model_and_unknown_without_usage() { + let project = Project::new(); + project.write("lib.rs", &function("f")); + let mut options = args(); + options.model = Some("jev-latest".into()); + let metered = Answering { + model: "jev-1.13.0", + metered: true, + }; + let report = run(&project, &options, &mut { metered }); + assert_eq!(report.paid_models, [("jev-1.13.0".to_string(), 10)].into()); + assert!((report.estimated_usd.unwrap() - 10.0 * 0.042 / 1e6).abs() < 1e-15); + assert!(output::headline(&report).ends_with("· 10 input tokens · ~$0.0000")); + options.model = Some("typesafe-ai/jev".into()); + let unmetered = Answering { + model: "typesafe-ai/jev", + metered: false, + }; + let report = run(&project, &options, &mut { unmetered }); + assert!(report.complete); + assert_eq!((report.api_requests, report.unmetered_requests), (1, 1)); + assert_eq!((report.paid_input_tokens, report.estimated_usd), (0, None)); + assert!(output::headline(&report).ends_with("· 0 input tokens · cost unknown")); + let replay = run(&project, &options, &mut Mock::default()); + assert_eq!(replay.api_requests, 0); + assert!( + replay + .estimated_usd + .is_some_and(|usd| usd.is_sign_positive() && usd == 0.0) + ); + assert!(output::headline(&replay).ends_with("· 0 input tokens · ~$0.0000")); + assert!(replay.files.iter().all(|file| file.cached)); +} + #[test] fn malformed_response_and_exhausted_budget_never_pass() { let project = two_files(); diff --git a/src/transport.rs b/src/transport.rs deleted file mode 100644 index 63d5a67..0000000 --- a/src/transport.rs +++ /dev/null @@ -1,915 +0,0 @@ -use crate::provider_error::{Interrupted, ProviderError, Unsent, provider_error, retryable}; -use anyhow::{Result, bail}; -use serde_json::Value; -use std::{ - path::Path, - sync::{ - Mutex, - atomic::{AtomicU16, Ordering}, - }, - time::{Duration, Instant}, -}; - -/// Attempts per request, including the first send. -const ATTEMPTS: u32 = 4; -/// Attempts after a timeout or dropped connection: the provider may have run -/// (and billed) the first send, so it is repeated only once. -const INTERRUPTED_ATTEMPTS: u32 = 2; -/// Longest provider-requested pause that is honored before a retry. -const RETRY_AFTER_CAP: Duration = Duration::from_secs(30); - -pub trait Evaluator { - /// A new snapshot may retry after account access has been restored. - fn begin_review(&mut self) {} - - fn evaluate(&mut self, request: &Value) -> Result; - - fn evaluate_batch(&mut self, requests: &[&Value]) -> Vec> { - requests - .iter() - .map(|request| self.evaluate(request)) - .collect() - } - - fn evaluate_queue( - &mut self, - requests: &[&Value], - concurrency: usize, - before: &(dyn Fn(&Value) -> Result<()> + Sync), - completed: &mut dyn FnMut(usize, Outcome), - ) { - let queue_start = std::time::Instant::now(); - let mut index = 0; - for chunk in requests.chunks(concurrency) { - let checks: Vec<_> = chunk.iter().map(|r| before(r)).collect(); - let valid: Vec<_> = chunk - .iter() - .zip(&checks) - .filter_map(|(r, c)| c.is_ok().then_some(*r)) - .collect(); - let started_ms = queue_start.elapsed().as_millis() as u64; - let start = std::time::Instant::now(); - let mut bodies = self.evaluate_batch(&valid).into_iter(); - for check in checks { - completed( - index, - match check { - Ok(()) => Outcome::attempted( - bodies.next().unwrap_or_else(|| { - Err(anyhow::anyhow!("Missing provider receipt")) - }), - start, - started_ms, - ), - Err(e) => Outcome::skipped(e), - }, - ); - index += 1; - } - } - } -} - -pub struct Outcome { - pub result: Result, - pub elapsed_ms: u64, - pub started_ms: u64, - pub attempted: bool, - /// Sends after the first one; zero when the request was not retried. - pub retries: u32, -} - -/// Workers claim the next item immediately after finishing, independent of the -/// slowest sibling. Completions carry their input index and are persisted immediately; -/// cancellation/freshness run at send. -pub(crate) fn work_queue( - items: &[T], - concurrency: usize, - work: impl Fn(&T) -> R + Sync, - mut completed: impl FnMut(usize, R), -) { - use std::sync::{ - atomic::{AtomicUsize, Ordering}, - mpsc, - }; - let next = AtomicUsize::new(0); - let (sender, receiver) = mpsc::channel(); - std::thread::scope(|scope| { - for _ in 0..concurrency.min(items.len()) { - let (next, work, sender) = (&next, &work, sender.clone()); - scope.spawn(move || { - loop { - let i = next.fetch_add(1, Ordering::Relaxed); - let Some(item) = items.get(i) else { - break; - }; - if sender.send((i, work(item))).is_err() { - break; - } - } - }); - } - drop(sender); - for (index, result) in receiver { - completed(index, result); - } - }); -} - -pub struct Client { - agent: ureq::Agent, - key_file: std::path::PathBuf, - key: Option, - explicit_file: bool, - access: ProviderAccess, -} - -impl Client { - pub fn new(key_file: &Path, explicit_file: bool) -> Self { - Self { - agent: ureq::Agent::config_builder() - .timeout_global(Some(Duration::from_secs(60))) - .max_redirects(0) - .http_status_as_error(false) - .build() - .into(), - key_file: key_file.into(), - key: None, - explicit_file, - access: ProviderAccess::default(), - } - } - - fn credential(&mut self) -> Result<&str> { - if self.key.is_none() { - self.key = Some(crate::auth::sources::resolve( - &self.key_file, - self.explicit_file, - )?); - } - Ok(self.key.as_ref().unwrap().key.expose()) - } -} - -impl Evaluator for Client { - fn begin_review(&mut self) { - if self.access.reset() { - // A rejected credential may have been replaced between snapshots. - self.key = None; - } - } - - fn evaluate(&mut self, request: &Value) -> Result { - self.access.check()?; - let agent = self.agent.clone(); - let result = send(&agent, self.credential()?, request); - self.access.observe(&result); - result - } - - fn evaluate_queue( - &mut self, - requests: &[&Value], - concurrency: usize, - before: &(dyn Fn(&Value) -> Result<()> + Sync), - completed: &mut dyn FnMut(usize, Outcome), - ) { - let agent = self.agent.clone(); - match self.credential() { - Ok(_) => {} - Err(error) => { - for i in 0..requests.len() { - completed( - i, - Outcome { - result: Err(anyhow::anyhow!(error.to_string())), - elapsed_ms: 0, - started_ms: 0, - attempted: false, - retries: 0, - }, - ); - } - return; - } - } - let key = self.key.as_ref().unwrap().key.expose(); - self.access.evaluate_queue( - requests, - concurrency, - before, - |request| send(&agent, key, request), - completed, - ); - } -} - -/// Consecutive edge blocks that stop further uploads. -const EDGE_BLOCKS: u16 = 3; - -/// Backoff jitter in thousandths of the delay: coprime steps spread requests -/// and retries over up to a quarter of it. -const JITTER_INDEX_STEP: u64 = 37; -const JITTER_RETRY_STEP: u64 = 101; -const JITTER_RANGE: u64 = 250; -const PER_MILLE: u32 = 1000; - -/// The first pause after a rate-limit or overload response; later ones grow from it. -const FIRST_BACKOFF: Duration = Duration::from_millis(500); - -/// Reject further uploads in this review only after a typed account/access failure. -/// Completed and in-flight requests keep their individual results; caches bypass this gate. -/// Rate-limit and overload responses pause every worker through one shared cooldown. -struct ProviderAccess { - rejected: AtomicU16, - /// The rejection came from the provider's edge protection, not the account. - edge: std::sync::atomic::AtomicBool, - /// Consecutive edge blocks: a request whose content trips a firewall rule - /// fails alone; blocks with no success between them stop further uploads. - edge_blocks: AtomicU16, - cooldown: Mutex>, - backoff: Duration, -} - -impl Default for ProviderAccess { - fn default() -> Self { - Self { - rejected: AtomicU16::new(0), - edge: std::sync::atomic::AtomicBool::new(false), - edge_blocks: AtomicU16::new(0), - cooldown: Mutex::new(None), - backoff: FIRST_BACKOFF, - } - } -} - -impl ProviderAccess { - fn reset(&mut self) -> bool { - *self.cooldown.lock().unwrap() = None; - self.edge_blocks.store(0, Ordering::Release); - self.edge.store(false, Ordering::Release); - self.rejected.swap(0, Ordering::AcqRel) != 0 - } - - fn check(&self) -> Result<()> { - let status = self.rejected.load(Ordering::Acquire); - if status != 0 && self.edge.load(Ordering::Acquire) { - bail!( - "TypeSafe request not sent after HTTP {status} from the provider's edge protection; wait before rerunning, and contact TypeSafe if it persists" - ); - } - if status != 0 { - bail!( - "TypeSafe request not sent after HTTP {status}; restore account access and rerun the review" - ); - } - Ok(()) - } - - fn observe(&self, result: &Result) { - if result.is_ok() { - self.edge_blocks.store(0, Ordering::Release); - } - if let Err(error) = result - && let Some(error) = error.downcast_ref::() - && matches!(error.status, 401..=403) - { - if error.edge_block { - if self.edge_blocks.fetch_add(1, Ordering::AcqRel) + 1 < EDGE_BLOCKS { - return; - } - self.edge.store(true, Ordering::Release); - } - let _ = self.rejected.compare_exchange( - 0, - error.status, - Ordering::AcqRel, - Ordering::Acquire, - ); - } - } - - /// Extend the shared pause; a shorter request never shortens another worker's wait. - fn pause(&self, delay: Duration) { - let until = Instant::now() + delay; - let mut cooldown = self.cooldown.lock().unwrap(); - if cooldown.is_none_or(|current| current < until) { - *cooldown = Some(until); - } - } - - fn wait(&self) { - let until = *self.cooldown.lock().unwrap(); - if let Some(remaining) = - until.and_then(|until| until.checked_duration_since(Instant::now())) - { - std::thread::sleep(remaining); - } - } - - /// Exponential backoff with deterministic jitter, so reruns are reproducible - /// while concurrent requests still spread out. - fn backoff(&self, index: usize, retry: u32) -> Duration { - let jitter = (index as u64 * JITTER_INDEX_STEP + u64::from(retry) * JITTER_RETRY_STEP) - % JITTER_RANGE; - self.backoff * 2u32.pow(retry) * (PER_MILLE + jitter as u32) / PER_MILLE - } - - fn send_with_retries( - &self, - index: usize, - request: &Value, - before: &(dyn Fn(&Value) -> Result<()> + Sync), - send: &(impl Fn(&Value) -> Result + Sync), - ) -> (Result, u32) { - let mut retry = 0; - loop { - self.wait(); - let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| send(request))) - .unwrap_or_else(|_| Err(anyhow::anyhow!("TypeSafe request worker failed"))); - self.observe(&result); - let Some((delay, attempts)) = result.as_ref().err().and_then(retry_delay) else { - return (result, retry); - }; - if retry + 1 >= attempts { - return ( - result.map_err(|error| { - anyhow::anyhow!("{error}; gave up after {attempts} attempts") - }), - retry, - ); - } - retry += 1; - self.pause(delay.unwrap_or_default().max(self.backoff(index, retry))); - // A rejected sibling or an edited source stops the retry like a first send. - if let Err(error) = self - .check() - .and_then(|()| before(request)) - .and_then(|()| self.check()) - { - return (Err(error), retry); - } - } - } - - fn evaluate_queue( - &self, - requests: &[&Value], - concurrency: usize, - before: &(dyn Fn(&Value) -> Result<()> + Sync), - send: impl Fn(&Value) -> Result + Sync, - completed: &mut dyn FnMut(usize, Outcome), - ) { - let queue_start = std::time::Instant::now(); - let indexed: Vec<_> = requests.iter().enumerate().collect(); - work_queue( - &indexed, - concurrency.clamp(1, crate::options::MAX_CONCURRENCY as usize), - |(index, request)| { - // Recheck after freshness work in case a sibling has since been rejected. - if let Err(error) = self - .check() - .and_then(|()| before(request)) - .and_then(|()| self.check()) - { - return Outcome::skipped(error); - } - let started_ms = queue_start.elapsed().as_millis() as u64; - let start = std::time::Instant::now(); - let (result, retries) = self.send_with_retries(*index, request, before, &send); - let mut outcome = Outcome::attempted(result, start, started_ms); - outcome.retries = retries; - outcome - }, - completed, - ); - } -} - -/// The pause and the attempt limit for a failure worth retrying: rate limits, -/// overload, server and gateway errors, and connections that failed before -/// the request was sent. A timeout or dropped connection is retried once, -/// since the request may have run. Validation and account errors are never -/// retried. -fn retry_delay(error: &anyhow::Error) -> Option<(Option, u32)> { - if let Some(error) = error.downcast_ref::() { - return retryable(error.status).then(|| { - let pause = error - .retry_after - .map(|s| Duration::from_secs(s).min(RETRY_AFTER_CAP)); - (pause, ATTEMPTS) - }); - } - if error.downcast_ref::().is_some() { - return Some((None, INTERRUPTED_ATTEMPTS)); - } - error.downcast_ref::().map(|_| (None, ATTEMPTS)) -} - -impl Outcome { - fn attempted(result: Result, start: std::time::Instant, started_ms: u64) -> Self { - Self { - result, - elapsed_ms: start.elapsed().as_millis() as u64, - started_ms, - attempted: true, - retries: 0, - } - } - - fn skipped(error: anyhow::Error) -> Self { - Self { - result: Err(error), - elapsed_ms: 0, - started_ms: 0, - attempted: false, - retries: 0, - } - } -} - -/// Compact JSON: ureq's `send_json` pretty-prints, and the provider's edge -/// blocks indented bodies carrying JSX that it accepts when compact. -fn request_body(request: &Value) -> Result> { - Ok(serde_json::to_vec( - crate::requests::provider_request(request).as_ref(), - )?) -} - -fn send(agent: &ureq::Agent, key: &str, request: &Value) -> Result { - let body = request_body(request)?; - let response = agent - .post("https://api.typesafe.ai/v1/systemone") - .header("Authorization", format!("Bearer {key}")) - .header( - "User-Agent", - concat!( - "jevgate/", - env!("CARGO_PKG_VERSION"), - " (+https://github.com/Tech-Byte-Frontier/jevgate)" - ), - ) - .content_type("application/json") - .send(&body[..]); - let mut response = match response { - Ok(response) => response, - Err(ureq::Error::StatusCode(status)) => { - return Err(provider_error(status, None, None).into()); - } - Err(ureq::Error::HostNotFound | ureq::Error::ConnectionFailed) => { - return Err(Unsent.into()); - } - Err(ureq::Error::Timeout(_) | ureq::Error::Io(_)) => return Err(Interrupted.into()), - Err(_) => bail!("TypeSafe transport failure; request was not retried"), - }; - if !response.status().is_success() { - let status = response.status().as_u16(); - let retry_after = response - .headers() - .get("retry-after") - .and_then(|value| value.to_str().ok()) - .and_then(|value| value.trim().parse::().ok()); - let body = response - .body_mut() - .with_config() - .limit(65_536) - .read_to_string() - .ok(); - return Err(provider_error(status, body.as_deref(), retry_after).into()); - } - // Error bodies and headers may echo credentials or source; never render them. - response - .body_mut() - .with_config() - .limit(1_048_576) - .read_json() - .map_err(|error| match error { - ureq::Error::Timeout(_) | ureq::Error::Io(_) => Interrupted.into(), - _ => anyhow::anyhow!("TypeSafe returned invalid or oversized JSON"), - }) -} - -#[cfg(test)] -pub(super) fn key_from_file(path: &Path) -> Result { - crate::auth::sources::key_from_file(path)? - .map(|key| key.expose().to_owned()) - .ok_or_else(|| anyhow::anyhow!("Credential file has no TYPESAFE_API_KEY")) -} - -#[cfg(test)] -mod tests { - use super::*; - use serde_json::json; - - fn fast() -> ProviderAccess { - ProviderAccess { - backoff: Duration::from_millis(1), - ..Default::default() - } - } - - fn sends(access: &ProviderAccess, results: Vec>) -> (Outcome, usize) { - let request = json!({"index":0}); - let calls = std::sync::atomic::AtomicUsize::new(0); - let results = Mutex::new(results.into_iter()); - let mut last = None; - access.evaluate_queue( - &[&request], - 1, - &|_| Ok(()), - |_| { - calls.fetch_add(1, Ordering::Relaxed); - results.lock().unwrap().next().unwrap() - }, - &mut |_, outcome| last = Some(outcome), - ); - (last.unwrap(), calls.load(Ordering::Relaxed)) - } - - #[test] - fn rate_limits_retry_after_the_requested_pause_and_count_retries() { - let access = fast(); - let start = Instant::now(); - let (outcome, calls) = sends( - &access, - vec![ - Err(provider_error(429, None, Some(1)).into()), - Ok(json!({"answers":{}})), - ], - ); - assert!(outcome.result.is_ok()); - assert_eq!((calls, outcome.retries), (2, 1)); - assert!( - start.elapsed() >= Duration::from_secs(1), - "retry-after is honored" - ); - for (status, error) in [ - (529, provider_error(529, None, None)), - (502, provider_error(502, None, None)), - ] { - let (outcome, calls) = sends(&access, vec![Err(error.into()), Ok(json!({}))]); - assert_eq!((calls, outcome.retries), (2, 1), "{status}"); - } - let (outcome, calls) = sends(&access, vec![Err(Unsent.into()), Ok(json!({}))]); - assert_eq!((calls, outcome.retries), (2, 1), "connection never opened"); - assert_eq!( - retry_delay(&provider_error(429, None, Some(3600)).into()), - Some((Some(RETRY_AFTER_CAP), ATTEMPTS)) - ); - } - - #[test] - fn server_errors_retry_and_an_interrupted_request_is_sent_twice_at_most() { - for status in [500, 520, 522, 524] { - let (outcome, calls) = sends( - &fast(), - vec![ - Err(provider_error(status, None, None).into()), - Ok(json!({})), - ], - ); - assert!(outcome.result.is_ok(), "{status}"); - assert_eq!((calls, outcome.retries), (2, 1), "{status}"); - } - let (outcome, calls) = sends(&fast(), vec![Err(Interrupted.into()), Ok(json!({}))]); - assert!(outcome.result.is_ok()); - assert_eq!(calls, 2, "a timeout passes on its second send"); - let failures = (0..4).map(|_| Err(Interrupted.into())).collect(); - let (outcome, calls) = sends(&fast(), failures); - assert_eq!(calls, INTERRUPTED_ATTEMPTS as usize); - let message = outcome.result.unwrap_err().to_string(); - assert!(message.contains("timed out") && message.contains("gave up after 2 attempts")); - } - - #[test] - fn validation_transport_and_account_errors_are_sent_once() { - for error in [ - anyhow::Error::from(provider_error(422, None, None)), - provider_error(400, None, Some(1)).into(), - provider_error(401, None, None).into(), - anyhow::anyhow!("TypeSafe transport failure; request was not retried"), - ] { - let text = error.to_string(); - let (outcome, calls) = sends(&fast(), vec![Err(error), Ok(json!({}))]); - assert_eq!((calls, outcome.retries), (1, 0), "{text}"); - assert!(outcome.result.is_err()); - } - } - - #[test] - fn persistent_overload_stops_after_the_attempt_limit() { - let failures = (0..ATTEMPTS + 2) - .map(|_| Err(provider_error(503, None, None).into())) - .collect(); - let (outcome, calls) = sends(&fast(), failures); - assert_eq!(calls, ATTEMPTS as usize); - assert_eq!(outcome.retries, ATTEMPTS - 1); - let message = outcome.result.unwrap_err().to_string(); - assert!(message.contains("HTTP 503") && message.contains("gave up after 4 attempts")); - } - - #[test] - fn account_rejections_stop_pending_uploads_but_keep_in_flight_successes() { - use std::sync::{Barrier, Condvar, Mutex, atomic::AtomicUsize}; - let access = fast(); - let requests: Vec<_> = (0..24).map(|i| json!({"index":i})).collect(); - let batch: Vec<_> = requests.iter().collect(); - let first_four = Barrier::new(4); - let released = (Mutex::new(false), Condvar::new()); - let calls = AtomicUsize::new(0); - let mut outcomes = Vec::new(); - access.evaluate_queue( - &batch, - 4, - &|_| Ok(()), - |request| { - calls.fetch_add(1, Ordering::Relaxed); - let index = request["index"].as_u64().unwrap(); - if index < 4 { - first_four.wait(); - if index == 0 { - return Err(provider_error(402, None, None).into()); - } - let (released, timeout) = released - .1 - .wait_timeout_while( - released.0.lock().unwrap(), - Duration::from_secs(5), - |done| !*done, - ) - .unwrap(); - assert!(*released && !timeout.timed_out()); - } - Ok(request.clone()) - }, - &mut |index, outcome| { - if index == 0 { - *released.0.lock().unwrap() = true; - released.1.notify_all(); - } - outcomes.push((index, outcome)); - }, - ); - outcomes.sort_by_key(|(index, _)| *index); - assert_eq!(calls.load(Ordering::Relaxed), 4); - assert_eq!(outcomes.len(), 24); - assert!(outcomes[0].1.attempted && outcomes[0].1.result.is_err()); - for (_, outcome) in &outcomes[1..4] { - assert!(outcome.attempted && outcome.result.is_ok()); - } - for (_, outcome) in &outcomes[4..] { - assert!(!outcome.attempted); - assert_eq!(outcome.elapsed_ms, 0); - assert!( - outcome - .result - .as_ref() - .unwrap_err() - .to_string() - .contains("not sent after HTTP 402") - ); - } - access.evaluate_queue( - &batch, - 4, - &|_| panic!("stopped before freshness work"), - |_| panic!("stopped across later stages"), - &mut |_, outcome| assert!(!outcome.attempted), - ); - } - - #[test] - fn only_typed_account_errors_stop_siblings_and_a_new_review_can_retry() { - let requests = [json!({"index":0}), json!({"index":1})]; - let batch: Vec<_> = requests.iter().collect(); - for status in [400, 401, 402, 403, 422, 429, 503, 529] { - let mut access = fast(); - let mut attempts = 0; - access.evaluate_queue( - &batch, - 1, - &|_| Ok(()), - |request| { - if request["index"] == 0 { - Err(anyhow::Error::new(provider_error(status, None, None)) - .context("provider response")) - } else { - Ok(request.clone()) - } - }, - &mut |_, outcome| attempts += usize::from(outcome.attempted), - ); - let rejected = matches!(status, 401..=403); - assert_eq!(attempts, if rejected { 1 } else { 2 }, "{status}"); - assert_eq!(access.reset(), rejected); - access.evaluate_queue( - &batch, - 1, - &|_| Ok(()), - |request| Ok(request.clone()), - &mut |_, outcome| assert!(outcome.attempted && outcome.result.is_ok()), - ); - } - let access = fast(); - let mut attempts = 0; - access.evaluate_queue( - &batch, - 1, - &|_| Ok(()), - |_| Err(anyhow::anyhow!("source text says TypeSafe HTTP 402")), - &mut |_, outcome| attempts += usize::from(outcome.attempted), - ); - assert_eq!(attempts, 2, "arbitrary error text cannot close the queue"); - access.evaluate_queue( - &batch, - 1, - &|request| { - if request["index"] == 0 { - bail!("stale source") - } - Ok(()) - }, - |request| Ok(request.clone()), - &mut |index, outcome| { - assert_eq!(outcome.attempted, index == 1); - }, - ); - } - - #[test] - fn rejected_review_keeps_cached_judgments_and_recovers_only_unfinished_work() { - use crate::tests::{Project, answer, args, run}; - struct Provider { - access: ProviderAccess, - reject: bool, - } - impl Evaluator for Provider { - fn begin_review(&mut self) { - self.access.reset(); - } - fn evaluate(&mut self, _: &Value) -> Result { - unreachable!("queue path") - } - fn evaluate_queue( - &mut self, - requests: &[&Value], - concurrency: usize, - before: &(dyn Fn(&Value) -> Result<()> + Sync), - completed: &mut dyn FnMut(usize, Outcome), - ) { - self.access.evaluate_queue( - requests, - concurrency, - before, - |request| { - if self.reject { - Err(provider_error(402, None, None).into()) - } else { - Ok(answer(request, 0)) - } - }, - completed, - ); - } - } - let project = Project::new(); - for name in ["a", "b", "c", "d"] { - project.write(&format!("{name}.py"), &format!("def {name}(fn):\n try:\n return fn()\n except OSError:\n log(fn)\n raise\n")); - } - project.write("b.py", "def b(fn):\n try:\n return fn()\n except OSError:\n log(fn)\n raise\n\ndef second(fn):\n try:\n return fn()\n except ValueError:\n log(fn)\n raise\n"); - let mut options = args(); - options.quick = true; - options.rules = vec!["function_simplification".into()]; - options.concurrency = 1; - options.paths = vec!["a.py".into()]; - let mut provider = Provider { - access: fast(), - reject: false, - }; - let warm = run(&project, &options, &mut provider); - assert!(warm.complete); - assert_eq!(warm.api_requests, 1); - options.paths.clear(); - provider.reject = true; - let rejected = run(&project, &options, &mut provider); - assert!(!rejected.complete); - assert_eq!(rejected.api_requests, 1); - assert_eq!(rejected.files[0].status, crate::schema::Status::Clear); - assert!(rejected.files[0].cached); - assert!( - rejected.files[1..] - .iter() - .all(|f| f.status == crate::schema::Status::Error) - ); - assert_eq!(rejected.stages["functions"].failed_attempts, 1); - assert_eq!(rejected.stages["functions"].cache_hits, 1); - assert_eq!( - rejected.files[1].error.as_deref(), - Some("TypeSafe HTTP 402; request was not retried"), - "later unsent work in the same file must not hide the original provider failure" - ); - assert!( - rejected.files[2..] - .iter() - .all(|f| f.error.as_ref().unwrap().contains("not sent")) - ); - let saved = crate::storage::read_latest(&project.0).unwrap(); - assert!(!saved.complete); - assert_eq!(saved.api_requests, 1); - provider.reject = false; - let recovered = run(&project, &options, &mut provider); - assert!(recovered.complete); - assert_eq!( - recovered.api_requests, 3, - "failed and unsent requests were not cached" - ); - options.cache_only = true; - let replay = run(&project, &options, &mut provider); - assert!(replay.complete); - assert_eq!(replay.api_requests, 0); - for (expected, actual) in recovered.files.iter().zip(replay.files) { - assert_eq!(json!(expected.dimensions), json!(actual.dimensions)); - } - } - - #[test] - fn a_context_limit_error_is_named_without_echoing_private_text() { - let body = json!({"detail":{"error_type":"max_tokens_exceeded","message":"private source and credentials"}}); - assert_eq!( - provider_error(400, Some(&body.to_string()), None).to_string(), - "TypeSafe HTTP 400 (model context limit exceeded); request was not retried" - ); - } - - #[test] - fn unknown_error_details_are_not_echoed() { - for body in [ - json!({"detail":"private source"}), - json!({"detail":{"error_type":"private credentials"}}), - ] { - assert_eq!( - provider_error(400, Some(&body.to_string()), None).to_string(), - "TypeSafe HTTP 400; request was not retried" - ); - } - assert_eq!( - provider_error(503, None, None).to_string(), - "TypeSafe HTTP 503" - ); - } - - const EDGE_PAGE: &str = - "Attention Required! | Cloudflare"; - - #[test] - fn request_bodies_are_compact_without_local_metadata() { - let request = json!({"model": "m", "state": {"source": ""}, "jevgate": {}}); - let body = String::from_utf8(request_body(&request).unwrap()).unwrap(); - assert_eq!(body, r#"{"model":"m","state":{"source":""}}"#); - } - - #[test] - fn edge_firewall_blocks_are_told_apart_from_account_rejections() { - let edge = provider_error(403, Some("error code: 1010\n"), None); - assert!(edge.edge_block); - assert_eq!( - edge.to_string(), - "TypeSafe HTTP 403 (blocked by the provider's edge protection); request was not retried" - ); - assert!(provider_error(403, Some(EDGE_PAGE), None).edge_block); - assert!(!provider_error(403, Some("{\"detail\":\"forbidden\"}"), None).edge_block); - } - - #[test] - fn isolated_edge_blocks_fail_alone_and_consecutive_blocks_stop_uploads() { - let access = ProviderAccess::default(); - let blocked: Result = Err(provider_error(403, Some(EDGE_PAGE), None).into()); - access.observe(&blocked); - access.observe(&blocked); - access.observe(&Ok(json!({}))); - access.observe(&blocked); - assert!(access.check().is_ok(), "a success resets the count"); - access.observe(&blocked); - access.observe(&blocked); - assert!( - access - .check() - .unwrap_err() - .to_string() - .contains("edge protection") - ); - } - - #[test] - fn credential_parser_does_not_execute_shell() { - let project = crate::tests::Project::new(); - project.write( - ".env", - "export TYPESAFE_API_KEY='literal$(do-not-execute)'\n", - ); - assert_eq!( - key_from_file(&project.0.join(".env")).unwrap(), - "literal$(do-not-execute)" - ); - } -} diff --git a/src/transport/mod.rs b/src/transport/mod.rs new file mode 100644 index 0000000..b179e3f --- /dev/null +++ b/src/transport/mod.rs @@ -0,0 +1,630 @@ +use crate::provider::{Endpoint, Provider, Service, TYPESAFE}; +use crate::provider_error::{ + Failure, Interrupted, ProviderError, Unsent, provider_error, retryable, +}; +use crate::response_headers::{REQUEST_ID, request_id, retry_after}; +use anyhow::{Result, bail, ensure}; +use serde_json::Value; +use std::{ + path::Path, + sync::{ + Mutex, + atomic::{AtomicU16, Ordering}, + }, + time::{Duration, Instant}, +}; + +/// Sends of a request the provider answered with a status worth retrying +/// (a rate limit, overload, or a server or gateway error), the first +/// included. On 2026-09-28 TypeSafe answered 503 to about two attempts in +/// three for at least ten minutes, directly (22 of 34) and through OpenRouter +/// (141 of 224), each failed attempt taking about 10 s: with 4 sends, 1 of +/// 13 TypeSafe requests and 16 of 99 OpenRouter ones gave up, and every run +/// ended incomplete. Simulating this queue, 6 sends leave 6% of such +/// requests unanswered instead of 16%; when 1 attempt in 5 fails, a +/// 1,000-request run completes 95% of the time instead of 21%; and a +/// provider failing every attempt is outlasted for 26 s instead of 8. It +/// costs time only while attempts fail: a hard outage takes up to three +/// times as long to end a run incomplete, 8 minutes instead of 3 for 100 +/// requests. +const ANSWERED_ATTEMPTS: u32 = 6; +/// Sends of a request whose connection failed before anything was sent, as +/// on a machine without a network: 4 as before, so such a run ends no later. +const UNSENT_ATTEMPTS: u32 = 4; +/// Attempts after a timeout or dropped connection: the provider may have run +/// (and billed) the first send, so it is repeated only once. +const INTERRUPTED_ATTEMPTS: u32 = 2; +/// Longest provider-requested pause that is honored before a retry. +const RETRY_AFTER_CAP: Duration = Duration::from_secs(30); +/// How long one attempt may take, from connecting to reading the answer. A +/// request takes about 0.3 s and TypeSafe's SDKs wait 10 s per attempt; one +/// that has not answered in 20 s is sent again once, rather than holding a +/// worker for a minute. +const ATTEMPT_TIMEOUT: Duration = Duration::from_secs(20); +/// Requests start at least this far apart: 1,200 a minute, TypeSafe's limit. +const REQUEST_INTERVAL: Duration = Duration::from_millis(50); + +pub trait Evaluator { + /// A new snapshot may retry after account access has been restored. + fn begin_review(&mut self) {} + + fn evaluate(&mut self, request: &Value) -> Result; + + fn evaluate_batch(&mut self, requests: &[&Value]) -> Vec> { + requests + .iter() + .map(|request| self.evaluate(request)) + .collect() + } + + fn evaluate_queue( + &mut self, + requests: &[&Value], + concurrency: usize, + before: &(dyn Fn(&Value) -> Result<()> + Sync), + completed: &mut dyn FnMut(usize, Outcome), + ) { + let queue_start = std::time::Instant::now(); + let mut index = 0; + for chunk in requests.chunks(concurrency) { + let checks: Vec<_> = chunk.iter().map(|r| before(r)).collect(); + let valid: Vec<_> = chunk + .iter() + .zip(&checks) + .filter_map(|(r, c)| c.is_ok().then_some(*r)) + .collect(); + let started_ms = queue_start.elapsed().as_millis() as u64; + let start = std::time::Instant::now(); + let mut bodies = self.evaluate_batch(&valid).into_iter(); + for check in checks { + completed( + index, + match check { + Ok(()) => Outcome::attempted( + bodies.next().unwrap_or_else(|| { + Err(anyhow::anyhow!("Missing provider receipt")) + }), + start, + started_ms, + ), + Err(e) => Outcome::skipped(e), + }, + ); + index += 1; + } + } + } +} + +pub struct Outcome { + pub result: Result, + pub elapsed_ms: u64, + pub started_ms: u64, + pub attempted: bool, + /// Sends after the first one; zero when the request was not retried. + pub retries: u32, +} + +/// Workers claim the next item immediately after finishing, independent of the +/// slowest sibling. Completions carry their input index and are persisted immediately; +/// cancellation/freshness run at send. +pub(crate) fn work_queue( + items: &[T], + concurrency: usize, + work: impl Fn(&T) -> R + Sync, + mut completed: impl FnMut(usize, R), +) { + use std::sync::{ + atomic::{AtomicUsize, Ordering}, + mpsc, + }; + let next = AtomicUsize::new(0); + let (sender, receiver) = mpsc::channel(); + std::thread::scope(|scope| { + for _ in 0..concurrency.min(items.len()) { + let (next, work, sender) = (&next, &work, sender.clone()); + scope.spawn(move || { + loop { + let i = next.fetch_add(1, Ordering::Relaxed); + let Some(item) = items.get(i) else { + break; + }; + if sender.send((i, work(item))).is_err() { + break; + } + } + }); + } + drop(sender); + for (index, result) in receiver { + completed(index, result); + } + }); +} + +pub struct Client { + agent: ureq::Agent, + /// The provider the check planned its model for; the key must be its. + provider: Provider, + endpoint: Endpoint, + key_file: std::path::PathBuf, + key: Option, + explicit_file: bool, + access: ProviderAccess, +} + +/// An HTTP client that gives up on an attempt after `timeout`, follows no +/// redirect (the key must not travel elsewhere) and returns error statuses as +/// responses, so their headers and body can be read. +fn agent(timeout: Duration) -> ureq::Agent { + ureq::Agent::config_builder() + .timeout_global(Some(timeout)) + .max_redirects(0) + .http_status_as_error(false) + .build() + .into() +} + +impl Client { + /// A client for `provider`'s key, sending to its endpoint; says on + /// stderr when `JEVGATE_BASE_URL` sends requests elsewhere. + pub fn new(key_file: &Path, explicit_file: bool, provider: Provider) -> Result { + let endpoint = Endpoint::new(provider)?; + if endpoint.is_custom() { + note!("Sending requests to {}", endpoint.describe()); + } + Ok(Self { + agent: agent(ATTEMPT_TIMEOUT), + provider, + access: ProviderAccess { + service: endpoint.service, + ..Default::default() + }, + endpoint, + key_file: key_file.into(), + key: None, + explicit_file, + }) + } + + fn credential(&mut self) -> Result<&str> { + if self.key.is_none() { + let credential = crate::auth::sources::resolve(&self.key_file, self.explicit_file)?; + ensure!( + credential.provider == self.provider, + "The key in {} is for {}, but this check planned its requests for {}; rerun the check", + credential.source, + credential.provider.service().label, + self.provider.service().label + ); + self.key = Some(credential); + } + Ok(self.key.as_ref().unwrap().key.expose()) + } +} + +impl Evaluator for Client { + fn begin_review(&mut self) { + if self.access.reset() { + // A rejected credential may have been replaced between snapshots. + self.key = None; + } + } + + fn evaluate(&mut self, request: &Value) -> Result { + self.access.check()?; + let agent = self.agent.clone(); + self.credential()?; + let key = self.key.as_ref().unwrap().key.expose(); + let result = send(&agent, &self.endpoint, key, request); + self.access.observe(&result); + result + } + + fn evaluate_queue( + &mut self, + requests: &[&Value], + concurrency: usize, + before: &(dyn Fn(&Value) -> Result<()> + Sync), + completed: &mut dyn FnMut(usize, Outcome), + ) { + let agent = self.agent.clone(); + match self.credential() { + Ok(_) => {} + Err(error) => { + for i in 0..requests.len() { + completed( + i, + Outcome { + result: Err(anyhow::anyhow!(error.to_string())), + elapsed_ms: 0, + started_ms: 0, + attempted: false, + retries: 0, + }, + ); + } + return; + } + } + let key = self.key.as_ref().unwrap().key.expose(); + let endpoint = &self.endpoint; + self.access.evaluate_queue( + requests, + concurrency, + before, + |request| send(&agent, endpoint, key, request), + completed, + ); + } +} + +/// Consecutive edge blocks that stop further uploads. +const EDGE_BLOCKS: u16 = 3; + +/// Backoff jitter in thousandths of the delay: coprime steps spread requests +/// and retries over up to a quarter of it. +const JITTER_INDEX_STEP: u64 = 37; +const JITTER_RETRY_STEP: u64 = 101; +const JITTER_RANGE: u64 = 250; +const PER_MILLE: u32 = 1000; + +/// The first pause after an answer worth retrying; each later one doubles +/// the one before, up to [`LONGEST_BACKOFF`]: 1, 2, 4, 8 and 8 s. +const FIRST_BACKOFF: Duration = Duration::from_secs(1); +/// The longest pause before its jitter. A pause holds every worker, so a +/// longer one would stall the whole run for one request's last send. +const LONGEST_BACKOFF: Duration = Duration::from_secs(8); + +/// Reject further uploads in this review only after a typed account/access failure. +/// Completed and in-flight requests keep their individual results; caches bypass this gate. +/// Rate-limit and overload responses pause every worker through one shared cooldown. +struct ProviderAccess { + rejected: AtomicU16, + /// The rejection came from the provider's edge protection, not the account. + edge: std::sync::atomic::AtomicBool, + /// Consecutive edge blocks: a request whose content trips a firewall rule + /// fails alone; blocks with no success between them stop further uploads. + edge_blocks: AtomicU16, + cooldown: Mutex>, + backoff: Duration, + /// The earliest time the next request may start, so that every worker's + /// sends together stay `interval` apart. + next_start: Mutex, + interval: Duration, + /// The provider the requests go to, named in the messages. + service: &'static Service, +} + +impl Default for ProviderAccess { + fn default() -> Self { + Self { + rejected: AtomicU16::new(0), + edge: std::sync::atomic::AtomicBool::new(false), + edge_blocks: AtomicU16::new(0), + cooldown: Mutex::new(None), + backoff: FIRST_BACKOFF, + next_start: Mutex::new(Instant::now()), + interval: REQUEST_INTERVAL, + service: &TYPESAFE, + } + } +} + +impl ProviderAccess { + fn reset(&mut self) -> bool { + *self.cooldown.lock().unwrap() = None; + self.edge_blocks.store(0, Ordering::Release); + self.edge.store(false, Ordering::Release); + self.rejected.swap(0, Ordering::AcqRel) != 0 + } + + fn check(&self) -> Result<()> { + let status = self.rejected.load(Ordering::Acquire); + let provider = self.service.label; + if status != 0 && self.edge.load(Ordering::Acquire) { + bail!( + "{provider} request not sent after HTTP {status} from the provider's edge protection; wait before rerunning, and contact {provider} if it persists" + ); + } + if status == 402 { + bail!( + "{provider} request not sent after HTTP 402 (credits exhausted); {}, then rerun the review", + self.service.credits + ); + } + if status != 0 { + bail!( + "{provider} request not sent after HTTP {status}; restore account access and rerun the review" + ); + } + Ok(()) + } + + fn observe(&self, result: &Result) { + if result.is_ok() { + self.edge_blocks.store(0, Ordering::Release); + } + if let Err(error) = result + && let Some(error) = error.downcast_ref::() + && matches!(error.status, 401..=403) + { + if error.edge_block { + if self.edge_blocks.fetch_add(1, Ordering::AcqRel) + 1 < EDGE_BLOCKS { + return; + } + self.edge.store(true, Ordering::Release); + } + let _ = self.rejected.compare_exchange( + 0, + error.status, + Ordering::AcqRel, + Ordering::Acquire, + ); + } + } + + /// Extend the shared pause; a shorter request never shortens another worker's wait. + fn pause(&self, delay: Duration) { + let until = Instant::now() + delay; + let mut cooldown = self.cooldown.lock().unwrap(); + if cooldown.is_none_or(|current| current < until) { + *cooldown = Some(until); + } + } + + fn wait(&self) { + let until = *self.cooldown.lock().unwrap(); + if let Some(remaining) = + until.and_then(|until| until.checked_duration_since(Instant::now())) + { + std::thread::sleep(remaining); + } + } + + /// Take the next start time, `interval` after the one before, and sleep until it. + fn pace(&self) { + let start = { + let mut next = self.next_start.lock().unwrap(); + let start = (*next).max(Instant::now()); + *next = start + self.interval; + start + }; + if let Some(remaining) = start.checked_duration_since(Instant::now()) { + std::thread::sleep(remaining); + } + } + + /// Exponential backoff with deterministic jitter, so reruns are reproducible + /// while concurrent requests still spread out; `retry` counts from 1. + fn backoff(&self, index: usize, retry: u32) -> Duration { + let jitter = (index as u64 * JITTER_INDEX_STEP + u64::from(retry) * JITTER_RETRY_STEP) + % JITTER_RANGE; + let pause = self + .backoff + .saturating_mul(2u32.saturating_pow(retry.saturating_sub(1))) + .min(LONGEST_BACKOFF); + pause * (PER_MILLE + jitter as u32) / PER_MILLE + } + + fn send_with_retries( + &self, + index: usize, + request: &Value, + before: &(dyn Fn(&Value) -> Result<()> + Sync), + send: &(impl Fn(&Value) -> Result + Sync), + ) -> (Result, u32) { + let mut retry = 0; + loop { + self.wait(); + self.pace(); + let result = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| send(request))) + .unwrap_or_else(|_| { + Err(anyhow::anyhow!( + "{} request worker failed", + self.service.label + )) + }); + self.observe(&result); + let Some((delay, attempts)) = result.as_ref().err().and_then(retry_delay) else { + return (result, retry); + }; + if retry + 1 >= attempts { + return ( + result.map_err(|error| { + anyhow::anyhow!("{error}; gave up after {attempts} attempts") + }), + retry, + ); + } + retry += 1; + self.pause(delay.unwrap_or_default().max(self.backoff(index, retry))); + // A rejected sibling or an edited source stops the retry like a first send. + if let Err(error) = self + .check() + .and_then(|()| before(request)) + .and_then(|()| self.check()) + { + return (Err(error), retry); + } + } + } + + fn evaluate_queue( + &self, + requests: &[&Value], + concurrency: usize, + before: &(dyn Fn(&Value) -> Result<()> + Sync), + send: impl Fn(&Value) -> Result + Sync, + completed: &mut dyn FnMut(usize, Outcome), + ) { + let queue_start = std::time::Instant::now(); + let indexed: Vec<_> = requests.iter().enumerate().collect(); + work_queue( + &indexed, + concurrency.clamp(1, crate::options::MAX_CONCURRENCY as usize), + |(index, request)| { + // Recheck after freshness work in case a sibling has since been rejected. + if let Err(error) = self + .check() + .and_then(|()| before(request)) + .and_then(|()| self.check()) + { + return Outcome::skipped(error); + } + let started_ms = queue_start.elapsed().as_millis() as u64; + let start = std::time::Instant::now(); + let (result, retries) = self.send_with_retries(*index, request, before, &send); + let mut outcome = Outcome::attempted(result, start, started_ms); + outcome.retries = retries; + outcome + }, + completed, + ); + } +} + +/// The pause and the attempt limit for a failure worth retrying: rate limits, +/// overload, server and gateway errors, and connections that failed before +/// the request was sent. A timeout or dropped connection is retried once, +/// since the request may have run. Validation and account errors are never +/// retried. +fn retry_delay(error: &anyhow::Error) -> Option<(Option, u32)> { + if let Some(error) = error.downcast_ref::() { + return retryable(error.status).then(|| { + ( + error.retry_after.map(|p| p.min(RETRY_AFTER_CAP)), + ANSWERED_ATTEMPTS, + ) + }); + } + if error.downcast_ref::().is_some() { + return Some((None, INTERRUPTED_ATTEMPTS)); + } + error + .downcast_ref::() + .map(|_| (None, UNSENT_ATTEMPTS)) +} + +impl Outcome { + fn attempted(result: Result, start: std::time::Instant, started_ms: u64) -> Self { + Self { + result, + elapsed_ms: start.elapsed().as_millis() as u64, + started_ms, + attempted: true, + retries: 0, + } + } + + fn skipped(error: anyhow::Error) -> Self { + Self { + result: Err(error), + elapsed_ms: 0, + started_ms: 0, + attempted: false, + retries: 0, + } + } +} + +/// Compact JSON: ureq's `send_json` pretty-prints, and the provider's edge +/// blocks indented bodies carrying JSX that it accepts when compact. +fn request_body(request: &Value) -> Result> { + Ok(serde_json::to_vec( + crate::requests::provider_request(request).as_ref(), + )?) +} + +/// Send one request and return the provider's answer, with the provider's +/// request id under `request_id`: TypeSafe's `x-typesafe-request-id` header, +/// else the response's own `id`, which OpenRouter sends. +fn send(agent: &ureq::Agent, endpoint: &Endpoint, key: &str, request: &Value) -> Result { + let service = endpoint.service; + let body = request_body(request)?; + let response = agent + .post(endpoint.systemone()) + .header("Authorization", format!("Bearer {key}")) + .header( + "User-Agent", + concat!( + "jevgate/", + env!("CARGO_PKG_VERSION"), + " (+https://github.com/Tech-Byte-Frontier/jevgate)" + ), + ) + .content_type("application/json") + .send(&body[..]); + let mut response = match response { + Ok(response) => response, + Err(ureq::Error::StatusCode(status)) => { + let failure = Failure { + status, + ..Default::default() + }; + return Err(provider_error(service, failure).into()); + } + Err(ureq::Error::HostNotFound | ureq::Error::ConnectionFailed) => { + return Err(Unsent(service).into()); + } + Err(ureq::Error::Timeout(_) | ureq::Error::Io(_)) => { + return Err(Interrupted(service).into()); + } + Err(_) => bail!( + "{} transport failure; request was not retried", + service.label + ), + }; + let header = |name: &str| { + response + .headers() + .get(name) + .and_then(|value| value.to_str().ok()) + .map(str::to_owned) + }; + let id = request_id(header(REQUEST_ID).as_deref()); + if !response.status().is_success() { + let wait = retry_after( + header("retry-after-ms").as_deref(), + header("retry-after").as_deref(), + std::time::SystemTime::now(), + ); + let status = response.status().as_u16(); + let body = response + .body_mut() + .with_config() + .limit(65_536) + .read_to_string() + .ok(); + let failure = Failure { + status, + body: body.as_deref(), + retry_after: wait, + request_id: id, + }; + return Err(provider_error(service, failure).into()); + } + // Error bodies and headers may echo credentials or source; never render them. + let mut answer: Value = response + .body_mut() + .with_config() + .limit(1_048_576) + .read_json() + .map_err(|error| match error { + ureq::Error::Timeout(_) | ureq::Error::Io(_) => Interrupted(service).into(), + _ => anyhow::anyhow!("{} returned invalid or oversized JSON", service.label), + })?; + // A `request_id` the body carries itself is replaced by the checked id, + // or dropped without one, as a cached answer drops it. + let id = id.or_else(|| request_id(answer["id"].as_str())); + if let Some(fields) = answer.as_object_mut() { + match id { + Some(id) => fields.insert("request_id".into(), Value::String(id)), + None => fields.remove("request_id"), + }; + } + Ok(answer) +} + +#[cfg(test)] +mod tests; diff --git a/src/transport/tests.rs b/src/transport/tests.rs new file mode 100644 index 0000000..bf1f0cf --- /dev/null +++ b/src/transport/tests.rs @@ -0,0 +1,666 @@ +//! The request queue, retries and access failures, and whole exchanges with a +//! mock provider over HTTP. +use super::*; +use crate::tests::mock_provider::{MockProvider, Received, Reply}; +use serde_json::json; + +/// A TypeSafe failure with `status`, `body` and a pause in seconds. +fn failed(status: u16, body: Option<&str>, pause: Option) -> ProviderError { + let failure = Failure { + status, + body, + retry_after: pause.map(Duration::from_secs), + request_id: None, + }; + provider_error(&TYPESAFE, failure) +} + +/// Retries and starts without the production pauses. +fn fast() -> ProviderAccess { + ProviderAccess { + backoff: Duration::from_millis(1), + interval: Duration::ZERO, + ..Default::default() + } +} + +fn sends(access: &ProviderAccess, results: Vec>) -> (Outcome, usize) { + let request = json!({"index":0}); + let calls = std::sync::atomic::AtomicUsize::new(0); + let results = Mutex::new(results.into_iter()); + let mut last = None; + access.evaluate_queue( + &[&request], + 1, + &|_| Ok(()), + |_| { + calls.fetch_add(1, Ordering::Relaxed); + results.lock().unwrap().next().unwrap() + }, + &mut |_, outcome| last = Some(outcome), + ); + (last.unwrap(), calls.load(Ordering::Relaxed)) +} + +#[test] +fn rate_limits_retry_after_the_requested_pause_and_count_retries() { + let access = fast(); + let start = Instant::now(); + let (outcome, calls) = sends( + &access, + vec![ + Err(failed(429, None, Some(1)).into()), + Ok(json!({"answers":{}})), + ], + ); + assert!(outcome.result.is_ok()); + assert_eq!((calls, outcome.retries), (2, 1)); + assert!( + start.elapsed() >= Duration::from_secs(1), + "retry-after is honored" + ); + for (status, error) in [ + (529, failed(529, None, None)), + (502, failed(502, None, None)), + ] { + let (outcome, calls) = sends(&access, vec![Err(error.into()), Ok(json!({}))]); + assert_eq!((calls, outcome.retries), (2, 1), "{status}"); + } + let (outcome, calls) = sends(&access, vec![Err(Unsent(&TYPESAFE).into()), Ok(json!({}))]); + assert_eq!((calls, outcome.retries), (2, 1), "connection never opened"); + assert_eq!( + retry_delay(&failed(429, None, Some(3600)).into()), + Some((Some(RETRY_AFTER_CAP), ANSWERED_ATTEMPTS)) + ); + assert_eq!( + retry_delay(&failed(503, None, Some(20)).into()), + Some((Some(Duration::from_secs(20)), ANSWERED_ATTEMPTS)), + "a pause longer than the longest backoff is still the provider's" + ); +} + +#[test] +fn server_errors_retry_and_an_interrupted_request_is_sent_twice_at_most() { + for status in [500, 520, 522, 524] { + let (outcome, calls) = sends( + &fast(), + vec![Err(failed(status, None, None).into()), Ok(json!({}))], + ); + assert!(outcome.result.is_ok(), "{status}"); + assert_eq!((calls, outcome.retries), (2, 1), "{status}"); + } + let (outcome, calls) = sends( + &fast(), + vec![Err(Interrupted(&TYPESAFE).into()), Ok(json!({}))], + ); + assert!(outcome.result.is_ok()); + assert_eq!(calls, 2, "a timeout passes on its second send"); + let failures = (0..4).map(|_| Err(Interrupted(&TYPESAFE).into())).collect(); + let (outcome, calls) = sends(&fast(), failures); + assert_eq!(calls, INTERRUPTED_ATTEMPTS as usize); + let message = outcome.result.unwrap_err().to_string(); + assert!(message.contains("timed out") && message.contains("gave up after 2 attempts")); +} + +#[test] +fn validation_transport_and_account_errors_are_sent_once() { + for error in [ + anyhow::Error::from(failed(422, None, None)), + failed(400, None, Some(1)).into(), + failed(401, None, None).into(), + anyhow::anyhow!("TypeSafe transport failure; request was not retried"), + ] { + let text = error.to_string(); + let (outcome, calls) = sends(&fast(), vec![Err(error), Ok(json!({}))]); + assert_eq!((calls, outcome.retries), (1, 0), "{text}"); + assert!(outcome.result.is_err()); + } +} + +#[test] +fn persistent_overload_stops_after_the_attempt_limit() { + let failures = (0..ANSWERED_ATTEMPTS + 2) + .map(|_| Err(failed(503, None, None).into())) + .collect(); + let (outcome, calls) = sends(&fast(), failures); + assert_eq!(calls, ANSWERED_ATTEMPTS as usize); + assert_eq!(outcome.retries, ANSWERED_ATTEMPTS - 1); + let message = outcome.result.unwrap_err().to_string(); + assert!(message.contains("HTTP 503") && message.contains("gave up after 6 attempts")); + let unsent = (0..ANSWERED_ATTEMPTS) + .map(|_| Err(Unsent(&TYPESAFE).into())) + .collect(); + let (outcome, calls) = sends(&fast(), unsent); + assert_eq!( + calls, UNSENT_ATTEMPTS as usize, + "an offline run ends as soon as before" + ); + let message = outcome.result.unwrap_err().to_string(); + assert!(message.contains("was not sent; gave up after 4 attempts")); +} + +#[test] +fn pauses_double_from_one_second_to_at_most_eight() { + let access = ProviderAccess::default(); + let pauses: Vec<_> = (1..ANSWERED_ATTEMPTS) + .map(|retry| access.backoff(0, retry).as_millis()) + .collect(); + // 1, 2, 4, 8 and 8 s, each lengthened by its jitter; the first three are + // the pauses 0.25 made. + assert_eq!(pauses, [1101, 2404, 4212, 9232, 8040]); + for (retry, seconds) in (1..ANSWERED_ATTEMPTS).zip([1, 2, 4, 8, 8]) { + let base = Duration::from_secs(seconds); + for index in 0..MAX_WORKERS * 4 { + let pause = access.backoff(index, retry); + assert!(pause >= base && pause < base * 5 / 4, "{retry} {index}"); + } + } +} + +#[test] +fn account_rejections_stop_pending_uploads_but_keep_in_flight_successes() { + use std::sync::{Barrier, Condvar, Mutex, atomic::AtomicUsize}; + let access = fast(); + let requests: Vec<_> = (0..24).map(|i| json!({"index":i})).collect(); + let batch: Vec<_> = requests.iter().collect(); + let first_four = Barrier::new(4); + let released = (Mutex::new(false), Condvar::new()); + let calls = AtomicUsize::new(0); + let mut outcomes = Vec::new(); + access.evaluate_queue( + &batch, + 4, + &|_| Ok(()), + |request| { + calls.fetch_add(1, Ordering::Relaxed); + let index = request["index"].as_u64().unwrap(); + if index < 4 { + first_four.wait(); + if index == 0 { + return Err(failed(402, None, None).into()); + } + let (released, timeout) = released + .1 + .wait_timeout_while( + released.0.lock().unwrap(), + Duration::from_secs(5), + |done| !*done, + ) + .unwrap(); + assert!(*released && !timeout.timed_out()); + } + Ok(request.clone()) + }, + &mut |index, outcome| { + if index == 0 { + *released.0.lock().unwrap() = true; + released.1.notify_all(); + } + outcomes.push((index, outcome)); + }, + ); + outcomes.sort_by_key(|(index, _)| *index); + assert_eq!(calls.load(Ordering::Relaxed), 4); + assert_eq!(outcomes.len(), 24); + assert!(outcomes[0].1.attempted && outcomes[0].1.result.is_err()); + for (_, outcome) in &outcomes[1..4] { + assert!(outcome.attempted && outcome.result.is_ok()); + } + for (_, outcome) in &outcomes[4..] { + assert!(!outcome.attempted); + assert_eq!(outcome.elapsed_ms, 0); + assert!( + outcome + .result + .as_ref() + .unwrap_err() + .to_string() + .contains("not sent after HTTP 402") + ); + } + access.evaluate_queue( + &batch, + 4, + &|_| panic!("stopped before freshness work"), + |_| panic!("stopped across later stages"), + &mut |_, outcome| assert!(!outcome.attempted), + ); +} + +#[test] +fn only_typed_account_errors_stop_siblings_and_a_new_review_can_retry() { + let requests = [json!({"index":0}), json!({"index":1})]; + let batch: Vec<_> = requests.iter().collect(); + for status in [400, 401, 402, 403, 422, 429, 503, 529] { + let mut access = fast(); + let mut attempts = 0; + access.evaluate_queue( + &batch, + 1, + &|_| Ok(()), + |request| { + if request["index"] == 0 { + Err(anyhow::Error::new(failed(status, None, None)).context("provider response")) + } else { + Ok(request.clone()) + } + }, + &mut |_, outcome| attempts += usize::from(outcome.attempted), + ); + let rejected = matches!(status, 401..=403); + assert_eq!(attempts, if rejected { 1 } else { 2 }, "{status}"); + assert_eq!(access.reset(), rejected); + access.evaluate_queue( + &batch, + 1, + &|_| Ok(()), + |request| Ok(request.clone()), + &mut |_, outcome| assert!(outcome.attempted && outcome.result.is_ok()), + ); + } + let access = fast(); + let mut attempts = 0; + access.evaluate_queue( + &batch, + 1, + &|_| Ok(()), + |_| Err(anyhow::anyhow!("source text says TypeSafe HTTP 402")), + &mut |_, outcome| attempts += usize::from(outcome.attempted), + ); + assert_eq!(attempts, 2, "arbitrary error text cannot close the queue"); + access.evaluate_queue( + &batch, + 1, + &|request| { + if request["index"] == 0 { + bail!("stale source") + } + Ok(()) + }, + |request| Ok(request.clone()), + &mut |index, outcome| { + assert_eq!(outcome.attempted, index == 1); + }, + ); +} + +#[test] +fn rejected_review_keeps_cached_judgments_and_recovers_only_unfinished_work() { + use crate::tests::{Project, answer, args, run}; + struct Provider { + access: ProviderAccess, + reject: bool, + } + impl Evaluator for Provider { + fn begin_review(&mut self) { + self.access.reset(); + } + fn evaluate(&mut self, _: &Value) -> Result { + unreachable!("queue path") + } + fn evaluate_queue( + &mut self, + requests: &[&Value], + concurrency: usize, + before: &(dyn Fn(&Value) -> Result<()> + Sync), + completed: &mut dyn FnMut(usize, Outcome), + ) { + self.access.evaluate_queue( + requests, + concurrency, + before, + |request| { + if self.reject { + Err(failed(402, None, None).into()) + } else { + Ok(answer(request, 0)) + } + }, + completed, + ); + } + } + let project = Project::new(); + for name in ["a", "b", "c", "d"] { + project.write(&format!("{name}.py"), &format!("def {name}(fn):\n try:\n return fn()\n except OSError:\n log(fn)\n raise\n")); + } + project.write("b.py", "def b(fn):\n try:\n return fn()\n except OSError:\n log(fn)\n raise\n\ndef second(fn):\n try:\n return fn()\n except ValueError:\n log(fn)\n raise\n"); + let mut options = args(); + options.quick = true; + options.rules = vec!["function_simplification".into()]; + options.concurrency = Some(1); + options.paths = vec!["a.py".into()]; + let mut provider = Provider { + access: fast(), + reject: false, + }; + let warm = run(&project, &options, &mut provider); + assert!(warm.complete); + assert_eq!(warm.api_requests, 1); + options.paths.clear(); + provider.reject = true; + let rejected = run(&project, &options, &mut provider); + assert!(!rejected.complete); + assert_eq!(rejected.api_requests, 1); + assert_eq!(rejected.files[0].status, crate::schema::Status::Clear); + assert!(rejected.files[0].cached); + assert!( + rejected.files[1..] + .iter() + .all(|f| f.status == crate::schema::Status::Error) + ); + assert_eq!(rejected.stages["functions"].failed_attempts, 1); + assert_eq!(rejected.stages["functions"].cache_hits, 1); + assert_eq!( + rejected.files[1].error.as_deref(), + Some( + "TypeSafe HTTP 402 (credits exhausted; add credits or turn on auto-refill at https://console.typesafe.ai); request was not retried" + ), + "later unsent work in the same file must not hide the original provider failure" + ); + assert!(rejected.files[2..].iter().all(|f| { + f.error + .as_ref() + .unwrap() + .contains("not sent after HTTP 402 (credits exhausted); add credits") + })); + let saved = crate::storage::read_latest(&project.0).unwrap(); + assert!(!saved.complete); + assert_eq!(saved.api_requests, 1); + provider.reject = false; + let recovered = run(&project, &options, &mut provider); + assert!(recovered.complete); + assert_eq!( + recovered.api_requests, 3, + "failed and unsent requests were not cached" + ); + options.cache_only = true; + let replay = run(&project, &options, &mut provider); + assert!(replay.complete); + assert_eq!(replay.api_requests, 0); + for (expected, actual) in recovered.files.iter().zip(replay.files) { + assert_eq!(json!(expected.dimensions), json!(actual.dimensions)); + } +} + +#[test] +fn a_context_limit_error_is_named_without_echoing_private_text() { + let body = json!({"detail":{"error_type":"max_tokens_exceeded","message":"private source and credentials"}}); + assert_eq!( + failed(400, Some(&body.to_string()), None).to_string(), + "TypeSafe HTTP 400 (model context limit exceeded); request was not retried" + ); +} + +#[test] +fn unknown_error_details_are_not_echoed() { + for body in [ + json!({"detail":"private source"}), + json!({"detail":{"error_type":"private credentials"}}), + ] { + assert_eq!( + failed(400, Some(&body.to_string()), None).to_string(), + "TypeSafe HTTP 400; request was not retried" + ); + } + assert_eq!(failed(503, None, None).to_string(), "TypeSafe HTTP 503"); +} + +const EDGE_PAGE: &str = + "Attention Required! | Cloudflare"; + +#[test] +fn request_bodies_are_compact_without_local_metadata() { + let request = json!({"model": "m", "state": {"source": ""}, "jevgate": {}}); + let body = String::from_utf8(request_body(&request).unwrap()).unwrap(); + assert_eq!(body, r#"{"model":"m","state":{"source":""}}"#); +} + +#[test] +fn edge_firewall_blocks_are_told_apart_from_account_rejections() { + let edge = failed(403, Some("error code: 1010\n"), None); + assert!(edge.edge_block); + assert_eq!( + edge.to_string(), + "TypeSafe HTTP 403 (blocked by the provider's edge protection); request was not retried" + ); + assert!(failed(403, Some(EDGE_PAGE), None).edge_block); + assert!(!failed(403, Some("{\"detail\":\"forbidden\"}"), None).edge_block); +} + +#[test] +fn isolated_edge_blocks_fail_alone_and_consecutive_blocks_stop_uploads() { + let access = ProviderAccess::default(); + let blocked: Result = Err(failed(403, Some(EDGE_PAGE), None).into()); + access.observe(&blocked); + access.observe(&blocked); + access.observe(&Ok(json!({}))); + access.observe(&blocked); + assert!(access.check().is_ok(), "a success resets the count"); + access.observe(&blocked); + access.observe(&blocked); + assert!( + access + .check() + .unwrap_err() + .to_string() + .contains("edge protection") + ); +} + +/// A mock provider answering with `respond`, and a TypeSafe endpoint at it. +fn mock(respond: impl Fn(&Received) -> Reply + Send + Sync + 'static) -> (MockProvider, Endpoint) { + let provider = MockProvider::start(respond); + let endpoint = Endpoint::custom(&TYPESAFE, &provider.url).unwrap(); + (provider, endpoint) +} + +/// A one-question request, with local metadata that must not be uploaded. +fn question() -> Value { + json!({"model": "jev-1.13.0", "state": "x", "jevgate": {"stage": "functions"}, + "questions": {"q": {"type": "noul", "instructions": "?"}}}) +} + +/// A valid answer to the request the provider received. +fn answered(received: &Received) -> Reply { + Reply::json(200, &crate::tests::answer(&received.json(), 0)) +} + +#[test] +fn an_exchange_sends_the_bearer_key_and_keeps_only_a_checked_request_id() { + let (provider, endpoint) = + mock(|received| answered(received).header(REQUEST_ID, "req_01J9-abc")); + let answer = send(&agent(ATTEMPT_TIMEOUT), &endpoint, "test-key", &question()).unwrap(); + assert_eq!(answer["request_id"], "req_01J9-abc"); + assert!(crate::response::validate(&answer, &question()).is_ok()); + let received = &provider.received()[0]; + assert_eq!( + (received.method.as_str(), received.path.as_str()), + ("POST", "/v1/systemone") + ); + assert_eq!(received.header("Authorization"), Some("Bearer test-key")); + assert!(received.json().get("jevgate").is_none()); + let (_, openrouter) = mock(|received| { + let mut body = crate::tests::answer(&received.json(), 0); + body["id"] = json!("gen-dec-1789738314-X5e5"); + Reply::json(200, &body) + }); + let answer = send(&agent(ATTEMPT_TIMEOUT), &openrouter, "k", &question()).unwrap(); + assert_eq!( + answer["request_id"], "gen-dec-1789738314-X5e5", + "a response's own id stands in" + ); + let (_, forged) = mock(|received| { + let mut body = crate::tests::answer(&received.json(), 0); + body["request_id"] = json!("not an id \u{1b}[31m