JevGate 0.26: gateway keys, judge the change, block only mature rules - #42
Closed
tauanbinato wants to merge 44 commits into
Closed
tauanbinato wants to merge 44 commits into
tauanbinato wants to merge 44 commits into
Conversation
A rule and level is mature when its findings were right at least 80% of the time on the 25 projects JevGate was never tuned on, over at least 20 hand labels. The default gate level is now `mature`, which fails only on those; every other finding is reported without failing the check. Any level set in jevgate.toml or with --fail-on replaces it as it says, and `mature` can be set by name. Measured from the corpus labels joined to 0.25.0's findings (replayed from the answer cache): function-simplification reviews (20 of 23) and agent-context considers (22 of 24) are mature. On the unseen projects' full runs the default gate failed on 122 findings, 57% of the labeled ones right, and on 17 of 22 projects; it now fails on 23, 87% right, and on 9. The concern probability does not separate right from wrong reviews there (55%, 46%, 56% and 61% across four bands), so the table in src/maturity.rs decides. Each finding records how the gate counted it (`gate`: fails, measuring or advisory), so the JSON report, annotations, SARIF, GitLab, the HTML report and the MCP tool agree, and the report names what `mature` stands for in `fail_on_mature`. The agent text marks findings that fail the gate and lists them first wherever a list is capped, and says when reviews did not fail it because their rules are still being measured. `jevgate rules` shows the levels that fail by default and each rule's precision on unseen projects. `jevgate init` writes the groups as commented examples, since `maintainability = "review"` would opt every maintainability review back in.
On the 25 projects JevGate was never tuned on, 6 of its 37 labeled reviews and considers were right (16%), against 47 of 85 on the projects it was tuned on (55%). It stays in the maintainability group and in `all`; `--rule default --rule hardcoded-values` adds it back. Without it, a default run plans 44% fewer first-pass requests on the 94 labeled corpus projects (22,370 to 12,623) and uploads 39% fewer bytes; with `--rule all` every request body is unchanged.
…ateway names Gateways name Jev `typesafe/jev-1.13`, `~typesafe/jev-latest` and `typesafe-ai/jev`. JevGate took any name but `jev-latest` and `jev-preview` for a pinned version, whose cached answers never expire and whose answer must name exactly that model, and its model check refused the `/` and `~` in an answering model's name: every gateway answer would have failed. A name is pinned only when its last segment ends in an x.y.z version and no `~` marks it as moving; a pinned request accepts its own version with or without a gateway's namespace. The default `jev-1.13.0` is unchanged, so nothing is asked again.
… usage
The cost was priced only when the requested name was exactly
`jev-1.13.0`, so a `jev-latest` run showed none although TypeSafe answers
it with `jev-1.13.0`, and a response without `usage` failed validation.
Gateways need not pass TypeSafe's `usage` through.
Each paid answer is now billed to the model it names; an answer without
usage is accepted, cached without it and counted as unmetered, which makes
the cost unknown ("cost unknown" in the headline, `estimated_usd: null`)
instead of $0, and keeps it out of the bytes-per-token calibration. The
price table moves to `model.rs`, by model line (`jev-1.13` and its
versions under any gateway namespace), checked against TypeSafe's models
page on 2026-09-28.
A pure move: the queue, retry and error tests leave transport.rs (915 lines, half of them tests) for transport/tests.rs, before the provider exchange tests join them.
…402 and 422 TypeSafe names each request in `x-typesafe-request-id` and its SDKs honor `retry-after-ms` before `Retry-After`; JevGate kept neither and read `Retry-After` only as whole seconds. A 402 said "request was not retried" and nothing about credits; a 422 said nothing about what was invalid. An error now ends with the request id when the provider sent one (else the response's own `id`, which OpenRouter sends), and each judgment and cached answer keeps the id of the request behind it. A retry waits for `retry-after-ms`, else `Retry-After` in seconds or as an HTTP date, still capped at 30 s. A 402 says the credits are exhausted and where to add them, also in the message of every request it stopped; a 422 names each `detail[].loc` and `type`, never `msg` or `input`, which can echo source; a 404 suggests the model name; a 413 counts as beyond the model's context. Header values are kept only when they are safe to print. Messages name the provider through a `Service`, so the gateways can share them; a mock provider on 127.0.0.1 (tests/support/mock_provider.rs) tests whole exchanges: the bearer key, the uploaded body without local metadata, the request id, a 429 with `retry-after-ms`, and a 422.
…pts after 20 s TypeSafe documents 1,200 requests a minute. On the corpus's largest runs six workers made 18 to 20 requests a second (0.3 s a request), just under it, and `--concurrency 8` would make about 27. Requests now start at least 50 ms apart across every worker and retry, and concurrency accepts 1 to 6. An attempt gave up only after 60 s; TypeSafe's SDKs wait 10 s. It now gives up after 20 s, and a timed-out request is still sent once more. A mock provider that answers late shows the retry.
People who already pay for OpenRouter or Vercel AI Gateway can run JevGate without a TypeSafe account, as they can Abide and Qlty: both gateways serve TypeSafe's API and the same model at the same price. JevGate read only TYPESAFE_API_KEY and sent every request to api.typesafe.ai. The provider follows the key. `jevgate auth login` asks which kind of key it is (`--provider` for scripts) and saves it with its provider; a check reads TYPESAFE_API_KEY, OPENROUTER_API_KEY or AI_GATEWAY_API_KEY from the environment, then `--env-file`, then the saved key. The repository's .env is read only for TYPESAFE_API_KEY: a gateway's key there is usually the application's own. `auth status` shows the provider, the endpoint and the keys set but not used; the headline says `via OpenRouter`. A key goes only to its own provider. jevgate.toml and .env cannot name a host; a key with another provider's prefix (`sk-or-`, `vck_`) is refused; JEVGATE_BASE_URL, from the environment only, takes https or loopback http for a self-hosted proxy and is announced on stderr. A gateway's saved key is stored as `<provider> <key>`, which 0.25 refuses as invalid rather than send to TypeSafe; a plain `provider` file beside it lets a check choose its default model (`typesafe/jev-1.13`, `typesafe-ai/jev`) before reading the key, and a mismatch stops the run instead of sending anything. Tested end to end against a local server in each gateway's shape (OpenRouter's `id`, `provider` and `usage.cost`; an answer without usage); no gateway has been called.
A float `sum` over no items is -0.0, so a fully cached run printed "~$-0.0000" and wrote `"estimated_usd": -0.0`; the canaries replayed from their caches showed it. The total now folds from 0.0.
TypeSafe answers `"model": "jev-1.13"`, a name its docs use, with HTTP 400
and `{"detail": {"error_type": "api_usage_error", "message": "Unknown
model: jev-1.13"}}` (asked on 2026-09-28; the request id header came with
it). JevGate said only "TypeSafe HTTP 400". It now says "unknown model;
check the model name", reading the message's fixed opening and never
echoing it.
The Models changelog entry said TypeSafe took `jev-1.13`; it does not, so
the entry now says what the alias rule changes: gateways' names and names
other than jev-latest and jev-preview.
…tate Asked on 2026-09-28: its detail[].input echoes the whole request, which is why the message reads only loc and type.
…ame two numbers A key saved by 0.25 was saved as TypeSafe's whatever it was; its prefix is now checked like the environment's and the files'. A paid answer whose model name fails validation is billed to "unknown" (no price) rather than to its raw text. JevGate's own check of this branch noted two numbers a reader must guess: the provider record's size limit and the days in 400 years.
An `--env-file` is now read for all three providers' variables, and a template's empty `OPENROUTER_API_KEY=` beside a real TypeSafe key failed the whole file as an invalid key. An empty value counts as unset, as an empty environment variable already does.
`--base` compared the files `git diff --name-status` printed from the Git top level with paths relative to JevGate's root, the nearest directory holding `jevgate.toml`. In a repository whose `jevgate.toml` is in a subdirectory, such as one package of a monorepo, no tracked change matched: the check selected only untracked files, which `ls-files` already prints from the root, and a pull request of edited files passed as "no-changed-source". `--relative` makes the diff print paths from the root and leave out changes outside it. The unit tests' Git helper moves to `Project::git`, shared by the documentation tests and the new one.
`--base` selected whole changed files, so a pull request check asked about, paid for and reported every unit of a touched file: on each corpus project's last commit, 54% of the review and consider findings in the changed files sat on lines the commit did not touch. A check with `--base` now asks about and reports only what the change touches: - functions, tests, comments, values and security units whose lines the change added, modified or removed lines inside (`git diff -U0`, streamed so a large patch never sits in memory, `--text` so a file marked `-diff` keeps its lines); a new or untracked file is judged whole; - copy pairs where either copy changed; - a file's outline, and a large document's, only when the change adds a member or heading its base version lacks (read from Git only then); - a document the change left alone when one of its sections names a path the change deleted or renamed, and only that section. Functions, values, comments, security units, instruction sections and laws are still packed within runs whose ends are computed over the whole file, so the changed functions of one run share a pack and the packs of other runs stay as they were. `--whole-files` keeps the whole-file check and asks exactly what 0.25.0 asked. The report says which scope a check used (`scope`: `changed-lines` or `whole-files`), as do the headline and the HTML report. `baseline --merge` after a check of changed lines keeps the other entries of the files it checked, and a finding such a check did not judge is `non-comparable` rather than `resolved`. The MCP tool takes `whole_files`. On the last commits of 118 corpus projects (all rules, tests included), 0.25.0 reported 311 review and consider findings, 167 of them off the changed lines; this reports 149, 9 off them (5 outlines the change added members to, 3 copy pairs whose other copy changed, 1 section with lines removed inside). Of the 144 findings on changed lines, 137 stay: 3 outlines of files the commit added no member to are not asked, a group of six comments of which the commit touched one is a note, and 3 function-simplification considers asked beside fewer functions are notes. With nothing cached, the first pass asks 2,259 requests and 4.55M input tokens instead of 4,828 and 11.05M. The corpus runs cost about $0.12.
0.26 caps concurrency at 6, and `concurrency = 7` or `8` in jevgate.toml, or `--concurrency 8`, became an error. 0.25 accepted up to 8, so a CI configuration valid yesterday would stop every run with exit 2 after an upgrade, which the stability page's rule for deprecated settings (keep working, with a warning on stderr) is meant to prevent. Nothing needs more than 6: requests start at least 50 ms apart whatever the concurrency, and the transport already clamped the workers to 6. A value above 6, from the flag or the file, is now lowered to 6 with one line on stderr; 0 is still an error. The schema keeps 6 as its maximum, so editors still point at the value to change.
Since 0.26 a run's cost is unknown when a response reports no token usage, which a gateway need not send, as well as when the model that answered has no published price. The headline says "cost unknown" for both, but the HTML report still explained every unknown cost as "No published rate configured for this model", wrong for the first case, which is the one a Vercel AI Gateway run is likely to hit. The page now gets the report's unmetered request count and says "Cost unknown: 3 responses reported no token usage", or that the model that answered has no published rate. Rendered in jsdom with both cases and a known cost.
The three 0.26 changes each updated the pages they touched; read together, a few places still missed the others: - ci.md: how the action takes a gateway key (`api-key-kind`, jevgate-action 1.2), and what the two gate changes add up to for a pull request check: replayed on the last commits of 118 corpus projects, the defaults fail 9 of them, all on function-simplification reviews, where 0.25.0 failed 19. - output.md: the headline now carries the scope a `--base` check judged, the gateway that answered and an unknown cost; say so in one place. - `jevgate rules`, the rules reference, what-it-finds and coding-agents: say plainly that an opt-in rule's mature levels fail the check once it is selected, as agent-context considers do with the documentation rules. - privacy-and-cost: Vercel AI Gateway's public model list marks Jev `zdr: none` (checked today), so it cannot apply zero data retention to these requests, as the page said it could. - `check --help`: `--cache-only` never contacts the provider, whichever it is. - "a key that starts as another provider's do" was missing a word.
…ck fails on The three changes were merged as eleven bullets in the order they landed. They are now one section: an opening paragraph with what they do together, then the default gate, `--base`, and gateway keys, each with its details under it. The opening gives the number the two gate changes decide together, which neither change measured: a pull request check with the defaults (`jevgate check --base`, as the action runs it), replayed from the answer cache on the last commit of the 118 corpus projects the scope change used. 0.25.0 fails 19 of them on 61 reviews, 31 of the 45 labeled right (9 wrong); 0.26 fails 9 on 11 function-simplification reviews, 8 of the 10 labeled right (none wrong, 2 debatable). Also: the `--base` table gives the findings lost on changed lines next to the ones off them, agent-context considers failing the check once the documentation rules run is said plainly, and the action's `api-key-kind` is named where gateway keys are.
JevGate's own check of the version (`check --base main --rule all --include-tests`) left two considers, both in the gateway change's tests: - `serve` in the mock provider read the request line, the headers and the body, logged the request and wrote the reply in one function (0.92). Reading the request, its headers and writing the reply are now their own functions. - The rate-limit and timeout tests repeated the same steps: a counter that treats the first request differently, then one request through a one-worker queue (0.86). `mock_first` and `queued` hold them. No behavior changes; the same tests pass.
The heading came right after the paragraph above it when the gate and key sections were merged.
The default gate and the changed-lines scope were each tested alone. A check with `--base` now runs both: of six functions a scripted provider calls review-worthy, only the one the change edited is judged, and its function-simplification review, a mature level, fails the gate on its own; with `--whole-files` all six do.
OPENROUTER_API_KEY and AI_GATEWAY_API_KEY in the environment came before --env-file, the repository's .env and the key saved by `jevgate auth login`. Other tools read both variables, so a user who exported one for them and upgraded had every check moved to the gateway: billed to that account, sent through a data processor they never chose for JevGate, and asked again from scratch, since the model name changed. The critique reproduced it with `--env-file ts.env` and an exported OpenRouter key: every request went out with the OpenRouter key. A check now reads TYPESAFE_API_KEY from the environment, then the credential file, then the saved key, and only then the gateways' variables. A key saved before 0.26 has no provider recorded beside it, so the credential store is read to find it, and only when a gateway's variable would otherwise be used; `resolve` reads it once. In CI the system store needs a terminal, so there the gateway's variable is used as before. `jevgate auth logout` no longer names a gateway's variable as an override of the saved key, since it is read after it.
A run stopped by exhausted credits printed its headline, "1 files failed." and exit 2: the reason (`TypeSafe HTTP 402 (credits exhausted; …)`) was only in the JSON report, with --verbose, or in the GitHub annotations. The MCP server's `jevgate_check` returns this text, so an agent could not tell a 402 from a missing key, and 0.26's new messages for 402, 422, an unknown model or a key issued by another provider went unseen by default. The summary now lists each distinct failure reason with how many files gave it, before the skip reasons it already listed, and counts one file in the singular. Tested end to end against a mock provider answering 402, in the agent text and in the MCP tool's result.
`Retry-After: Sun, 06 Nov 500000000000 08:49:37 GMT` panicked in the
request worker ("overflow when adding duration to instant"): the date
parser took any year from 1970 up, and the seconds since 1970 no longer
fit the system time. The request then failed as a worker failure without
its retry, and the run ended incomplete. For larger years the product
wrapped silently in release builds.
The form HTTP requires senders to use has a four-digit year, so any other
is not a date and the usual backoff applies; the sum is also added with
checked_add. 31 Dec 9999 still parses.
`send` put the checked request id (the header, else the body's `id`) under `request_id`, but when there was none it left a `request_id` the provider's body carried itself, and the report copied it unchecked: a mock answering `"request_id": "not an id \u001b[31m<script>…"` put the escape and newline into files[].judgments[].request_id on a live run, while the cache dropped it, so a cached replay's report disagreed. A body's own `request_id` is now replaced by the checked id, or dropped without one, as the cached answer drops it.
A changed-lines check ran `git diff -U0 --text` over the whole change. A branch that changes a 150 MB binary made Git write a 301 MB patch: a dry run took 0.87 s and 460 MB where 0.25.0 took 0.14 s, and `--base --watch` repeated that diff on every 250 ms poll, keeping Git busy with nothing edited. A regenerated 10 MB lockfile cost 0.14 s a poll. The lines are now read after the inputs are known, only for the files read to be judged, each with its name before a rename so Git still pairs the two. The paths go to Git in batches of about 16 KiB, well under a Windows command line, and no path means no diff. On the binary branch a dry run is back to 0.10 s, as with 0.25.0; on 16 corpus projects the changed-lines request bodies are identical (439 requests). Git also lets GIT_DIFF_OPTS outrank -U0: with `--unified=6` set, context lines counted as changed, and a function next to the edited one was judged and could fail the gate. The variable is now removed from Git's environment.
…in --concurrency 0 The CHANGELOG, the configuration page and the schema said a `concurrency` of 7 or 8 in jevgate.toml is lowered to 6 with a notice. It never is: the file's value is a ceiling on the flag, whose default is 6, so it stays 6 without a word, as in 0.25.0. Only `--concurrency` above 6 reaches the cap and its notice. The docs and the schema now say so, and the unit test lowers `--concurrency 8` through `configure`, so it fails without the cap; it passed without it before. `--concurrency 0` answered "0 is not in 1..=4294967295", contradicting "at most 6"; it now says "Use a whole number from 1 to 6; a higher one is lowered to 6".
Capped lists put the findings that fail the gate first, then the rest by rank alone, so a review still being measured competed with higher-ranked considers: GitHub shows 10 warning annotations a step, and in 0.25 reviews were errors with a cap of their own. Replaying the order over the whole-repository runs of 94 corpus projects with the default rules (results/mat-gate), 267 of 424 such reviews fell past the tenth warning, in 35 projects. With reviews first, 109 do, all in the 10 projects that hold more than ten of them. The same order serves the agent text, the job summary's 50 rows and the MCP server's 50 findings. A stable sort keeps the rank order within each part.
OpenRouter's public listing gives `typesafe/jev-1.13`, JevGate's default model with an OpenRouter key, the canonical slug `typesafe/jev-1.13-20260917` at $0.000000042 per prompt token (checked on 2026-09-28, no key needed). The price table knew only `jev-1.13` and its `.N` versions, so if OpenRouter's answers name the endpoint, every OpenRouter run would show "cost unknown" and a null `estimated_usd`. A dated snapshot of a priced line (eight digits after a hyphen) is now priced like the line. Vercel AI Gateway's `typesafe-ai/jev` stays unpriced: it names no version, so its price follows whatever it points to; the privacy page and the CHANGELOG now say so. The maintainer's probe script shows which name each gateway answers with.
…led yet" A law finding still being measured said its rule was "still being measured (none labeled yet on projects JevGate was never tuned on)". The maturity table leaves Bend 2's labels out, and `tests/laws` judges only Bend 2 code, so it has no row; but its findings were labeled, on Bend 2 projects (0.22.0 counted 15 right of 23 on projects then unseen), so the text was wrong for it. The maturity module now words why a level has no share: law findings were labeled only on Bend 2 projects, which the table leaves out; any other level has none labeled yet. GitHub, SARIF and GitLab messages and the HTML report use that text, `jevgate rules` and the rules reference say it in their legend, and the CHANGELOG says Bend 2's labels were kept apart.
`jevgate check -h`, the man page and the completions showed `--model … [default: jev-1.13.0]` whatever the key, while the default follows the key's provider (typesafe/jev-1.13 with an OpenRouter key, typesafe-ai/jev with a Vercel AI Gateway key), as the long help says. A gateway user reading -h could pass jev-1.13.0, which OpenRouter refuses. It now reads `[default: the key's provider's model]`, as --rule and --env-file describe theirs.
…rrect them ci.md called the 118 corpus projects of the last-commit replays open-source; 28 of them are the maintainer's private repositories, and 8 of the 9 projects the 0.26 default check fails are among them. ci.md and the CHANGELOG now say so, and the CHANGELOG adds that 10 of the 11 failing findings are on those repositories and the other, ky's `Ky` constructor, was labeled right (labeled from the code for this: the constructor carries `eslint-disable complexity` and folds URL resolution and a 60-line search-params merge into setup). With that label the 0.26 replay is 9 of 11 right, none wrong, 2 debatable, and 0.25.0's 32 of 46. 0.25.0 failed 18 of those projects, not 19: the 19th's replay lacked one cached answer and ended incomplete. The losses on changed lines now say the comments consider was labeled right, as the scope note has it. The CHANGELOG said twice that nothing is asked again on upgrade; that holds for the model. A --base check asks its touched units again once, in changed-lines packs: with a cache holding whole-file answers of the same commits, the corpus's first pass needed 716 new requests and 1.34M input tokens against 765 and 1.06M for 0.25.0. jevgate-action 1.1, today's @v1, has no `api-key-kind` input and passes `api-key` as TYPESAFE_API_KEY, which 0.26 refuses for an OpenRouter key. ci.md and the CHANGELOG now give the form that works with both: leave `api-key` out and set the gateway's variable in the step's `env`.
JevGate's own check (`--rule all --include-tests`) found two considers in code this version changed: - `baseline::write` mixed separate jobs (0.94): collecting the findings to accept, carrying earlier reasons, and keeping the entries a merge keeps, which 0.26 extended for changed-lines checks. Each is now a named function; behavior is unchanged. - The new gateway test's `save_as_0_25` repeated `save_credential` of the auth tests (0.99). Both now use one `Project::save_credential(key)`. `auth::file::save` gets the doc comment its sibling `load` has, saying the file holds an API key kept as written, since the provider needs it as is; the check read it as a password kept in plain text (0.83).
…xture The pull request's own review (JevGate 0.25.0, `--rule all --include-tests`, where every review fails the gate) failed on `src/units/tests/mod.rs:312`: "One of this file's constants fixes a value that differs between deployments (0.98). The constant is `HARDCODED`." `HARDCODED` is the source text the hardcoded-values tests check, `const REGION: &str = "eu-west-1"` included, so the finding is wrong (test data). The code is untouched by this version; its request changed because the file gained a line above it. Accepted with `jevgate baseline --merge`, keeping only this entry, and `jevgate baseline mark wrong src/units/tests/mod.rs:312`, as AGENTS.md asks for a mistaken finding.
A check with an OpenRouter or Vercel AI Gateway key now sends at most 3 requests at once unless --concurrency or concurrency in jevgate.toml says otherwise; a TypeSafe key keeps 6. Six workers at JevGate's pacing make up to TypeSafe's limit of 1,200 requests a minute for one account, and through a gateway that account is the gateway's, shared with its other customers. It is a precaution, not a measured fix. On 2026-09-28 the canaries through OpenRouter got HTTP 503 on 141 of 224 attempts (63%), but a self-check with a TypeSafe key at the same time got 22 of 34 (65%), and a 1,083-request TypeSafe run half an hour earlier, 6 at once, had no retry: the 503s were TypeSafe's. OpenRouter's rounds of a few requests at once fared only a little better (18 of 36 attempts, against 123 of 188 in first-pass rounds of up to six; z about 1.7). concurrency in jevgate.toml now applies without the flag and caps it, as max_requests does, so a repository can raise a gateway's default up to 6; above 6 it means 6, silently, as before. The report's concurrency is the value used.
An answer worth retrying (HTTP 408, 429, 500, 502-504, 520-524, 529) is now sent up to 6 times instead of 4, with every provider, pausing 1, 2, 4, 8 and 8 s, each up to a quarter longer; the first three pauses are 0.25's, and a longer pause the provider asks for is still honored up to 30 s. A timeout is still sent twice at most, and a connection that fails before sending 4 times, so an offline run ends no later. On 2026-09-28 TypeSafe answered 503 to about two attempts in three for at least ten minutes, directly (a self-check at 10:30: 22 of 34 attempts, 1 of 13 requests gave up) and through OpenRouter (the four canaries, 10:34 to 10:39: 141 of 224 attempts, 16 of 99 requests gave up), each failed attempt taking about 10 s. Every one of those runs ended incomplete. An independent failure of 63% per attempt predicts 15.6 of 99 requests exhausting 4 attempts; 16 did. A simulation of this queue (shared pause, 50 ms pacing, jitter) puts numbers on the change. In that brownout, 6 attempts leave 6.3% of requests unanswered instead of 16%, and a 13-request check completes 42% of the time instead of 9%. When 1 attempt in 5 fails, a 1,000-request run completes 95% of the time instead of 21%. A provider failing every attempt is outlasted for 26 s instead of 8.7. The cost comes only while attempts fail: a hard outage answered at once ends a 100-request run incomplete after 8.3 minutes instead of 2.7 at 6 at once (16 instead of 5.3 at a gateway's 3). The 0.27 hook is bounded by its own deadline. TypeSafe's SDKs retry twice; a gate's run that ends incomplete costs a rerun of the whole job, where one more pause of 8 s costs little. The pause stops doubling at 8 s because it holds every worker.
Git exports GIT_DIR to a hook, `git rebase --exec` or `git bisect run` in a linked worktree, and GIT_INDEX_FILE to a pre-commit hook (absolute for `commit -a` and a partial commit; probed on Git 2.51.0). The tests run `git init`, `add` and `commit` in temporary projects, and with those variables inherited they acted on the outer repository instead. `git init` with GIT_DIR set to a worktree's administrative directory guesses a bare repository and writes `core.bare = true` into the configuration the worktree shares with its main clone; `add` and `commit` then stage and commit the temporary project's files on the worktree's branch. The check's own Git, run in-process by the unit tests and by the binary the CLI tests start, read the outer repository too, and its `git diff` rewrote the outer index's stat cache, which Git does even with GIT_OPTIONAL_LOCKS=0. tests/support/git.rs, shared by the unit and CLI tests like temp_dir.rs, removes the variables that point Git at a repository (GIT_DIR, GIT_WORK_TREE, GIT_INDEX_FILE, GIT_COMMON_DIR, GIT_OBJECT_DIRECTORY, GIT_ALTERNATE_OBJECT_DIRECTORIES, GIT_NAMESPACE, GIT_PREFIX) from every test's Git and from the jevgate the CLI tests start. The check now starts Git in one place, revision::git_in, which docs::history::ignored uses too, and removes them there under cfg(test). tests/lint_policy.rs fails when Git is started anywhere else. Measured by running the whole suite with the variables pointing at a throwaway repository that has a linked worktree, in three shapes: GIT_DIR, GIT_WORK_TREE and GIT_INDEX_FILE together; GIT_DIR alone, as `rebase --exec` in a linked worktree exports it; and an absolute GIT_INDEX_FILE alone. Before, 20, 15 and 14 of 572 tests failed, and each run changed the throwaway repository: core.worktree written into the shared configuration; core.bare = true and 7 commits of test files on the worktree's branch; the main index rewritten. After, all 573 tests pass in each shape and all 43 files of the repository (configuration, HEADs, refs and indexes included) keep their sha256. A check still honors the variables: Git sets them for the repository the check runs in, and in a pre-commit hook GIT_INDEX_FILE is the index being committed. git_in's documentation says why.
A check's own Git honors GIT_DIR, GIT_INDEX_FILE and the other variables that point Git at a repository, where the tests' Git now drops them. Git sets them for the hooks, `rebase --exec` and `bisect run` it starts, for the repository the check runs in: in a pre-commit hook GIT_INDEX_FILE is the index being committed (`.git/next-index-*.lock` for a partial commit), and a work tree whose repository is kept elsewhere, as with a bare dotfiles repository, is found only through GIT_DIR and GIT_WORK_TREE. Dropping them would protect nothing: a check only reads, but for `git diff` refreshing the index's stat cache, which leaves what is staged unchanged. The test moves a project's .git out of its work tree and checks that `check --base HEAD` with both variables set judges the changed file. With the variables dropped from the check's Git it fails: "Cannot find revision HEAD ... not a git repository".
The MCP server's instructions and the AGENTS.md snippet on the coding agents page still said "Fix each `review` finding", the policy of 0.25's gate. With 0.26's default gate, reviews of rules still being measured are reported without failing it, and on projects JevGate was never tuned on shared-logic reviews were right 46 of 85 times, file-organization 2 of 5 and sensitive-data 10 of 24: an agent following the old text refactors on findings wrong about half the time. Both now say what 0.27's hook instructions say: fix each finding marked to fail the gate, and weigh the other reviews and considers, fixing one when it is right or saying why the code should stay.
The action example suggested `args: --rule security`. Naming a rule replaces the selection, so the action ran only the five security rules; no level of them is mature, so under 0.26's default gate the check could never fail, and function simplification, the one level that blocks, was off. A dry run on zoxide with `--rule security` shows exactly those rules, `fail_on` ["mature"] and no `fail_on_mature`. Under 0.25.0 the same args still failed on security reviews. The example is now `--rule default --rule security`, as the action's own input description has it, with a sentence on what a selection without a mature level does.
`jevgate init` wrote `maintainability = "review"` and `tests = "review"` from 0.3 to 0.25, the README's first step, so most configurations hold them. Under 0.26 they keep every review of those groups failing the check (shared-logic reviews right 54% of the time on unseen projects) and judge hardcoded values (16%), and from 0.27 the agent hook blocks on them; only the CHANGELOG and the configuration page said so. A command that reads jevgate.toml now says so on stderr while the lines are there with the comment init wrote after them, which lists the group's rules and tells them from a level set by hand: deleting the lines gives the default rules and gate, and deleting only the comments keeps the levels without the notice.
The hero image was JevGate 0.22.0 on zoxide: "gate failed: 3 new review findings", two of them shared-logic reviews, a hardcoded-values consider and eleven notes, each message ending in a probability. 0.26 fails that commit on one function-simplification review, reports the other three reviews and a consider as still being measured, and no longer runs hardcoded values by default. Both images are now 0.26's, replayed from the corpus's answer cache on the same commit (09a18b4) with no request: the default `jevgate check` agent text, and the HTML report with src/util.rs open, its failing review marked. The README and the site's introduction also say what the gate blocks by default and how often it was right (function-simplification reviews, 20 of 23 on unseen projects), and the README what a check costs (the last commits of 118 corpus projects, checked as pull requests with every rule and nothing cached, sent 4.55M first-pass input tokens, $0.19) and that a rerun of unchanged code sends no request. A Vercel AI Gateway key is said to be accepted but untried with a real key: only OpenRouter has answered a live request, and Vercel's `typesafe-ai/jev` shows its cost as unknown.
The tests' temporary directories are canonical, which on Windows is the verbatim form \\?\C:\..., and Git for Windows reads a GIT_DIR in that form as "not a git repository": base_reads_the_repository_git_dir_names failed on the Windows CI job. The test now strips the prefix, as a person would write the path; nothing changes elsewhere.
Contributor
Author
|
Superseded by #43, which carries every roadmap version from 0.26 to 0.30 on one branch, to be released together as 0.30.0. Everything in this pull request is in #43 unchanged (the |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
JevGate 0.26 from the roadmap: a pull request check judges only what the change touches and fails only on rules and levels measured right on projects JevGate was never tuned on, and OpenRouter and Vercel AI Gateway keys work as TypeSafe keys do. The jevgate-action side (a sticky pull request comment and the key kind) is Tech-Byte-Frontier/jevgate-action#2. The version is not bumped; releasing is yours.
A default pull request check (
jevgate check --base, as the action runs it), replayed from the answer cache on the last commit of 118 corpus projects (90 open-source, 28 of the maintainer's own private repositories); nothing was sent:8 of the 9 failing projects and 10 of the 11 failing findings are the maintainer's own repositories; the other is ky's
Kyconstructor, labeled right in this pull request.Block only mature rules by default (approved 2026-09-28;
5a4da12,abe1bc9)mature, is the default: a rule's reviews or considers fail the check when at least 80% of them were right on the 25 projects never used for tuning (11 held out, 14 fresh), over at least 20 findings labeled by hand. The table is insrc/maturity.rs, with how and when it was measured. Every other finding is reported, marked as still being measured, and the check passes. Any explicit level (fail_on,[rules],--fail-on,[[scope]]) replaces it exactly as it says.--rule documentationorall, agent-context considers fail too. Its 22 right findings come from 4 of the maintainer's own repositories.(fails the gate)in the agent text and a line on the reviews still being measured; each finding'sgate(fails,measuring,advisory) andfail_on_maturein the JSON report; GitHub, SARIF and GitLab messages with the rule and level's unseen precision; the HTML report and the MCP tools;jevgate rulescolumns. Wherever a list is capped, failures come first, then reviews.Judge the change (
7ffb23d,87d446e)With
--base, a check asks about and reports only the functions, tests, comments, values and security units on changed lines, copies where either copy changed, a file's outline (or a large document's) only when the change adds a member or heading, and a document the change left alone only in a section naming a path it deleted or renamed. New files are judged whole;--whole-filesasks exactly what 0.25.0 asked. The report'sscopeand the headline say which.On the last commits of the same 118 projects (all rules, tests included):
The 7 not reported on changed lines: 3 outlines of files the commit added no member to (1 labeled wrong, 1 debatable), a comments consider now a note (labeled right), 3 function-simplification considers that became notes in smaller packs (1 wrong). Of 11,693 answers about the same units, 88% were identical to whole-file packs and 36 crossed 0.50 or 0.80. After upgrading, a
--basecheck asks its touched units once more in changed-lines packs (716 new requests and 1.34M tokens against 765 and 1.06M on the corpus cache).Fixed on the way: with
jevgate.tomlbelow the Git top level,--basefound no change and passed.Gateway keys and provider hygiene (
199c9bc..013d30a)jevgate auth loginasks the key's kind (--provider typesafe|openrouter|vercel) and saves it with its provider. A check readsTYPESAFE_API_KEY, then--env-fileor the repository.env(TYPESAFE_API_KEYonly there), then the saved key, thenOPENROUTER_API_KEYorAI_GATEWAY_API_KEY. A key goes only to its provider:jevgate.tomland.envcannot choose a host, a key with another provider's prefix is refused, andJEVGATE_BASE_URLis read only from the environment.jev-1.13.0(TypeSafe, unchanged),typesafe/jev-1.13(OpenRouter),typesafe-ai/jev(Vercel). A name without anx.y.zversion is an alias;/and~are accepted.usagemakes it "cost unknown", never $0. New report fields:provider,paid_models,unmetered_requests,estimated_usd, andrequest_idon judgments and cached answers.retry-after-msand HTTP-dateRetry-After; a 422 shows only field paths and error types; a 402 says credits are exhausted and where to add them; TypeSafe's unknown-model 400 is explained; 20 s per attempt; requests at least 50 ms apart; at most 6 at once.Fixes from the two critiques (19 findings; all checked against the code, all fixed; details in
docs/research/2026-09-28/plans/0.26/integration.md)OPENROUTER_API_KEYoutranked an explicit--env-file, the repository.envand the saved TypeSafe key, moving billing, the data processor and the cache to OpenRouter. The gateways' variables are now read last (705eab6).Failed N: <reason>, such as the 402 message (1b1ba62).3da0e48).Retry-Afterdate with a huge year panicked a request worker (ade1772); a changed-lines check diffed binaries and lockfiles as text on every--watchpoll, andGIT_DIFF_OPTScould widen hunks (d214eb7: a dry run over a changed 150 MB binary is back to 0.10 s from 0.87 s); a body's ownrequest_idreached the report unchecked (004e6e9).concurrencyabove 6 injevgate.tomlhas no effect rather than a notice (4edce15); measuring reviews come before considers in GitHub's 10 warnings (267 of 424 fell past the tenth, 109 now;2f93b13); OpenRouter's datedtypesafe/jev-1.13-20260917is priced (1ec0899); law findings say they were labeled on Bend 2 projects (0e56c5e);--model's short help (4a19246); the action example works with today's@v1(3da0e48).After the stack critique (9 commits on
b061572, pushed; the critique of 0.26 to 0.30 is in the stack workflow's notes)010471d,6e77c1c; fromwip/0.26-gateway-retries,plans/0.26/gateway-retries.md): an answer worth retrying is sent up to 6 times instead of 4, pausing 1, 2, 4, 8 and 8 s, with every provider; a gateway's key sends at most 3 requests at once by default. On 2026-09-28 TypeSafe answered 503 to about two attempts in three for ten minutes, directly and through OpenRouter, and every run then ended incomplete. The CHANGELOG no longer says no gateway was called.83e435b,75deef4): run undergit rebase --execor a hook, whoseGIT_DIRandGIT_INDEX_FILEthey inherited, they had setcore.bare = truein a clone's shared configuration and committed test files into a worktree. The tests' Git drops those variables; a check still honors them. The new GIT_DIR test failed on Windows, where the tests' temporary directories are in the verbatim\\?\form that Git for Windows does not read as a GIT_DIR; it hands Git plain paths now (91180ad).reviewfinding"; they now say to fix what fails the gate and weigh the rest (6651cec).args: --rule security, which replaces the selection with five rules none of which is mature, so the check could never fail; it now says--rule default --rule securityand why (9edaf77).jevgate.tomlthat still holds themaintainability = "review"andtests = "review"linesinitwrote before 0.26, with their comments, gets a notice on stderr naming them (141458e).ee5b336). The images' scripts are inplans/readme-images/.How it was measured (all free except the self-checks)
--cache-only, pinned calibration), joined to the hand labels by fingerprint;docs/research/2026-09-28/scripts/maturity.py, andbands.pyfor the probability bands the doc comment quotes (55%, 46%, 56%, 61%).evaluation/pinned_run.sh budgets-0241 BIN LABELwithJG_FLAGS="--base HEAD~1"for 0.25.0 and 0.26, compared byscripts/scope_compare.py; request counts from--dry-runand--dry-run --refresh.scripts/pr_run.sh(a cache-onlycheck --base HEAD~1per project) andscripts/pr_gate.py. The final binary's replay is identical to the integration's on all 118 projects.--rule all --include-testsrequest bodies of 0.25.0 and this branch are identical on 12 projects (12,145 requests); changed-lines bodies are identical before and after the finish stage's diff change on 16 projects (439 requests).cargo fmt,cargo +1.98.1 clippy --locked --all-targets -- -D warnings,cargo test --locked(530 unit, 37 CLI, 2 lint-policy tests; main has 475, 27, 2),cargo +1.90.0 check --locked;cargo deny checkandcargo packagepass. Atee5b336: 533 unit, 39 CLI and 3 lint-policy tests, the same checks, and 0.25.0's gate (what CI's review runs) on the files the 8 new commits change: passed, no review or consider.check --rule all --include-tests, default gate): gate passed. The three considers in code this version changed are fixed (e0bc674). Left, all hardcoded values (opt-in) on code this version did not touch, whose requests changed with their files: a review ofsrc/units/tests/mod.rs:312, the hardcoded-values tests' fixture text, which failed this pull request's own review (0.25.0 fails on every review) and is accepted in the baseline marked wrong (b061572); and considers onsrc/response.rs:186(the API's"noul") andsrc/units/tests/mod.rs:157(the value 3).Decisions to review
maturity::TABLE.matureis a gate level users can set; each finding recordsgate(fails,measuring,advisory), and the reportfail_on_mature.jevgate initnow writes every group as a commented example; files written before keep theirreviewlevels.--basenow means changed lines.--whole-fileshas nojevgate.tomlkey. The report's newscopefield (changed-lines,whole-files) shares its name with[[scope]]injevgate.toml; rename one before release if that reads badly.--env-fileand the saved key. To know about a key saved before 0.26 (no provider recorded), a check reads the credential store, but only while a gateway's variable is set and nothing before the saved key holds a key.<provider> <key>with aproviderfile beside the credential; 0.25 refuses such a value instead of sending it to TypeSafe.cache_ttl_secs(1 hour); CI runs further apart re-ask their units.jev-1.13; Vercel'stypesafe-ai/jevnames no version and stays "cost unknown".--concurrencyabove 6 is lowered with a notice (a 0.25 script keeps working);concurrencyinjevgate.tomlstays a ceiling, so a 7 or 8 there has no effect.Waiting on you
docs/research/2026-09-28/plans/0.26/provider-probe.sh --canariesfrom your clone (it prints no key; the probes cost well under $0.001 a gateway). Check: whichmodeleach gateway answers with and whether a pinned name works (if one does, make it that provider'sdefault_modelinsrc/provider.rs); whether Vercel returnsusage; which request id each sends; whether/api/v1/keyand/v1/creditsanswer assrc/auth/verify.rsexpects; and that each canary's headline saysvia OpenRouterorvia Vercel AI Gateway.baseinput text still says "only files changed"; chooseapi-key-kindorprovideras the input name.Ky::constructor, TP): review it.Jev spend: after the stack critique $0.013 (0.25.0's and this branch's checks of the new commits); before it: provider $0.029, scope $0.144, integration $0.036, finish $0.094 (JevGate's own check twice), and this pull request's two CI reviews $0.292 (0.25.0 judges the 77 changed files whole, about 1,575 requests a run; the first run failed, so its answers were not cached for the second): about $0.60 of the version's $1.00.