Skip to content

JevGate 0.26: gateway keys, judge the change, block only mature rules - #42

Closed
tauanbinato wants to merge 44 commits into
mainfrom
v0.26
Closed

tauanbinato wants to merge 44 commits into
mainfrom
v0.26

Conversation

@tauanbinato

@tauanbinato tauanbinato commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

JevGate 0.26 from the roadmap: a pull request check judges only what the change touches and fails only on rules and levels measured right on projects JevGate was never tuned on, and OpenRouter and Vercel AI Gateway keys work as TypeSafe keys do. The jevgate-action side (a sticky pull request comment and the key kind) is Tech-Byte-Frontier/jevgate-action#2. The version is not bumped; releasing is yours.

A default pull request check (jevgate check --base, as the action runs it), replayed from the answer cache on the last commit of 118 corpus projects (90 open-source, 28 of the maintainer's own private repositories); nothing was sent:

0.25.0 0.26
Projects the check fails 18 (and one incomplete) 9
Findings that fail it 61 reviews 11 function-simplification reviews
Of those labeled, right / wrong / debatable 32 / 9 / 5 of 46 9 / 0 / 2 of 11
Reviews / considers reported 61 / 140 29 / 70
Projects never used for tuning that fail 4 (10 of 12 labeled right) 2 (3 of 3)

8 of the 9 failing projects and 10 of the 11 failing findings are the maintainer's own repositories; the other is ky's Ky constructor, labeled right in this pull request.

Block only mature rules by default (approved 2026-09-28; 5a4da12, abe1bc9)

  • A new gate level, mature, is the default: a rule's reviews or considers fail the check when at least 80% of them were right on the 25 projects never used for tuning (11 held out, 14 fresh), over at least 20 findings labeled by hand. The table is in src/maturity.rs, with how and when it was measured. Every other finding is reported, marked as still being measured, and the check passes. Any explicit level (fail_on, [rules], --fail-on, [[scope]]) replaces it exactly as it says.
  • Two levels are mature: function-simplification reviews (20 of 23, 87%) and agent-context considers (22 of 24, 92%). Agent context is opt-in (documentation), so default runs fail only on function-simplification reviews; with --rule documentation or all, agent-context considers fail too. Its 22 right findings come from 4 of the maintainer's own repositories.
  • Full runs on the unseen projects: the default gate failed on 122 findings, 57% right (64% leaving debatable ones out), and 17 of 22 projects; now 23 findings, 87% right, and 9 projects. Tuned projects: 468 at 73% to 95 at 83%. 49 right reviews on unseen projects now warn instead of failing.
  • Hardcoded values is opt-in: 6 of its 37 unseen labels were right (16%). A default run asks 44% fewer first-pass requests on 94 corpus projects (22,370 to 12,623).
  • The policy shows everywhere: (fails the gate) in the agent text and a line on the reviews still being measured; each finding's gate (fails, measuring, advisory) and fail_on_mature in the JSON report; GitHub, SARIF and GitLab messages with the rule and level's unseen precision; the HTML report and the MCP tools; jevgate rules columns. Wherever a list is capped, failures come first, then reviews.

Judge the change (7ffb23d, 87d446e)

  • With --base, a check asks about and reports only the functions, tests, comments, values and security units on changed lines, copies where either copy changed, a file's outline (or a large document's) only when the change adds a member or heading, and a document the change left alone only in a section naming a path it deleted or renamed. New files are judged whole; --whole-files asks exactly what 0.25.0 asked. The report's scope and the headline say which.

  • On the last commits of the same 118 projects (all rules, tests included):

    0.25.0 0.26
    Review and consider findings off the changed lines 167 of 311 (54%) 9 of 149 (6%), all by design
    Findings on changed lines 144 140 (137 the same, 3 new)
    Labeled right, on changed lines 73% 75%
    First-pass requests / input tokens, nothing cached 4,828 / 11.05M 2,259 / 4.55M

    The 7 not reported on changed lines: 3 outlines of files the commit added no member to (1 labeled wrong, 1 debatable), a comments consider now a note (labeled right), 3 function-simplification considers that became notes in smaller packs (1 wrong). Of 11,693 answers about the same units, 88% were identical to whole-file packs and 36 crossed 0.50 or 0.80. After upgrading, a --base check asks its touched units once more in changed-lines packs (716 new requests and 1.34M tokens against 765 and 1.06M on the corpus cache).

  • Fixed on the way: with jevgate.toml below the Git top level, --base found no change and passed.

Gateway keys and provider hygiene (199c9bc..013d30a)

  • jevgate auth login asks the key's kind (--provider typesafe|openrouter|vercel) and saves it with its provider. A check reads TYPESAFE_API_KEY, then --env-file or the repository .env (TYPESAFE_API_KEY only there), then the saved key, then OPENROUTER_API_KEY or AI_GATEWAY_API_KEY. A key goes only to its provider: jevgate.toml and .env cannot choose a host, a key with another provider's prefix is refused, and JEVGATE_BASE_URL is read only from the environment.
  • Default models: jev-1.13.0 (TypeSafe, unchanged), typesafe/jev-1.13 (OpenRouter), typesafe-ai/jev (Vercel). A name without an x.y.z version is an alias; / and ~ are accepted.
  • Cost is priced by the model that answered; a response without usage makes it "cost unknown", never $0. New report fields: provider, paid_models, unmetered_requests, estimated_usd, and request_id on judgments and cached answers.
  • Hygiene: request ids in errors; retry-after-ms and HTTP-date Retry-After; a 422 shows only field paths and error types; a 402 says credits are exhausted and where to add them; TypeSafe's unknown-model 400 is explained; 20 s per attempt; requests at least 50 ms apart; at most 6 at once.
  • Tested end to end against a local server in each gateway's shape, and through OpenRouter on 2026-09-28 (key check, answers, usage, price and request ids as expected; plans/0.26/provider-live.md). Vercel AI Gateway has not been tried with a key, and the README says so.

Fixes from the two critiques (19 findings; all checked against the code, all fixed; details in docs/research/2026-09-28/plans/0.26/integration.md)

  • An exported OPENROUTER_API_KEY outranked an explicit --env-file, the repository .env and the saved TypeSafe key, moving billing, the data processor and the cache to OpenRouter. The gateways' variables are now read last (705eab6).
  • The agent text and the MCP tool said "1 files failed." without the reason; they now list Failed N: <reason>, such as the 402 message (1b1ba62).
  • ci.md called the 118 projects open-source; it and the CHANGELOG now name the maintainer's repositories, and 0.25.0 failed 18 projects, not 19 (3da0e48).
  • A Retry-After date with a huge year panicked a request worker (ade1772); a changed-lines check diffed binaries and lockfiles as text on every --watch poll, and GIT_DIFF_OPTS could widen hunks (d214eb7: a dry run over a changed 150 MB binary is back to 0.10 s from 0.87 s); a body's own request_id reached the report unchecked (004e6e9).
  • Docs and messages: concurrency above 6 in jevgate.toml has no effect rather than a notice (4edce15); measuring reviews come before considers in GitHub's 10 warnings (267 of 424 fell past the tenth, 109 now; 2f93b13); OpenRouter's dated typesafe/jev-1.13-20260917 is priced (1ec0899); law findings say they were labeled on Bend 2 projects (0e56c5e); --model's short help (4a19246); the action example works with today's @v1 (3da0e48).

After the stack critique (9 commits on b061572, pushed; the critique of 0.26 to 0.30 is in the stack workflow's notes)

  • Retries and gateway concurrency (010471d, 6e77c1c; from wip/0.26-gateway-retries, plans/0.26/gateway-retries.md): an answer worth retrying is sent up to 6 times instead of 4, pausing 1, 2, 4, 8 and 8 s, with every provider; a gateway's key sends at most 3 requests at once by default. On 2026-09-28 TypeSafe answered 503 to about two attempts in three for ten minutes, directly and through OpenRouter, and every run then ended incomplete. The CHANGELOG no longer says no gateway was called.
  • The tests leave the repository they run in alone (83e435b, 75deef4): run under git rebase --exec or a hook, whose GIT_DIR and GIT_INDEX_FILE they inherited, they had set core.bare = true in a clone's shared configuration and committed test files into a worktree. The tests' Git drops those variables; a check still honors them. The new GIT_DIR test failed on Windows, where the tests' temporary directories are in the verbatim \\?\ form that Git for Windows does not read as a GIT_DIR; it hands Git plain paths now (91180ad).
  • The MCP server's instructions and the coding-agents page's AGENTS.md snippet said "Fix each review finding"; they now say to fix what fails the gate and weigh the rest (6651cec).
  • The CI page suggested args: --rule security, which replaces the selection with five rules none of which is mature, so the check could never fail; it now says --rule default --rule security and why (9edaf77).
  • A jevgate.toml that still holds the maintainability = "review" and tests = "review" lines init wrote before 0.26, with their comments, gets a notice on stderr naming them (141458e).
  • The README's images are 0.26's (the terminal image was 0.22.0's: three failing reviews, a hardcoded-values consider, probabilities), replayed from the answer cache on zoxide at the same commit; the README and the site's introduction say what blocks by default and how often it was right, and the README what a check costs (ee5b336). The images' scripts are in plans/readme-images/.

How it was measured (all free except the self-checks)

  • Maturity table: 0.25.0 replayed from the answer cache over the 94 labeled projects outside Bend 2 (--cache-only, pinned calibration), joined to the hand labels by fingerprint; docs/research/2026-09-28/scripts/maturity.py, and bands.py for the probability bands the doc comment quotes (55%, 46%, 56%, 61%).
  • Scope: evaluation/pinned_run.sh budgets-0241 BIN LABEL with JG_FLAGS="--base HEAD~1" for 0.25.0 and 0.26, compared by scripts/scope_compare.py; request counts from --dry-run and --dry-run --refresh.
  • Pull request replay: scripts/pr_run.sh (a cache-only check --base HEAD~1 per project) and scripts/pr_gate.py. The final binary's replay is identical to the integration's on all 118 projects.
  • Nothing asked again with a TypeSafe key: --rule all --include-tests request bodies of 0.25.0 and this branch are identical on 12 projects (12,145 requests); changed-lines bodies are identical before and after the finish stage's diff change on 16 projects (439 requests).
  • Checks before each commit: cargo fmt, cargo +1.98.1 clippy --locked --all-targets -- -D warnings, cargo test --locked (530 unit, 37 CLI, 2 lint-policy tests; main has 475, 27, 2), cargo +1.90.0 check --locked; cargo deny check and cargo package pass. At ee5b336: 533 unit, 39 CLI and 3 lint-policy tests, the same checks, and 0.25.0's gate (what CI's review runs) on the files the 8 new commits change: passed, no review or consider.
  • JevGate's own check (check --rule all --include-tests, default gate): gate passed. The three considers in code this version changed are fixed (e0bc674). Left, all hardcoded values (opt-in) on code this version did not touch, whose requests changed with their files: a review of src/units/tests/mod.rs:312, the hardcoded-values tests' fixture text, which failed this pull request's own review (0.25.0 fails on every review) and is accepted in the baseline marked wrong (b061572); and considers on src/response.rs:186 (the API's "noul") and src/units/tests/mod.rs:157 (the value 3).
  • CI: every job passes (fmt, clippy, tests on Linux, macOS and Windows, MSRV 1.90, cargo-deny, build, JevGate's review).

Decisions to review

  1. Agent-context considers fail the default gate once the documentation rules are selected: they meet the approved criterion (22 of 24), but the roadmap expected only function-simplification reviews, and 23 of their 24 unseen labels come from the maintainer's repositories. Excluding them is one row of maturity::TABLE.
  2. mature is a gate level users can set; each finding records gate (fails, measuring, advisory), and the report fail_on_mature. jevgate init now writes every group as a commented example; files written before keep their review levels.
  3. --base now means changed lines. --whole-files has no jevgate.toml key. The report's new scope field (changed-lines, whole-files) shares its name with [[scope]] in jevgate.toml; rename one before release if that reads badly.
  4. Key order, changed after the critique: the provider item read the gateways' variables before --env-file and the saved key. To know about a key saved before 0.26 (no provider recorded), a check reads the credential store, but only while a gateway's variable is set and nothing before the saved key holds a key.
  5. Saved gateway keys are stored as <provider> <key> with a provider file beside the credential; 0.25 refuses such a value instead of sending it to TypeSafe.
  6. Gateway models are aliases, whose answers expire after cache_ttl_secs (1 hour); CI runs further apart re-ask their units.
  7. Pricing: OpenRouter's dated endpoint is priced like jev-1.13; Vercel's typesafe-ai/jev names no version and stays "cost unknown".
  8. Concurrency: --concurrency above 6 is lowered with a notice (a 0.25 script keeps working); concurrency in jevgate.toml stays a ceiling, so a 7 or 8 there has no effect.

Waiting on you

  • The gateway done-when: with an OpenRouter and a Vercel AI Gateway key, run docs/research/2026-09-28/plans/0.26/provider-probe.sh --canaries from your clone (it prints no key; the probes cost well under $0.001 a gateway). Check: which model each gateway answers with and whether a pinned name works (if one does, make it that provider's default_model in src/provider.rs); whether Vercel returns usage; which request id each sends; whether /api/v1/key and /v1/credits answer as src/auth/verify.rs expects; and that each canary's headline says via OpenRouter or via Vercel AI Gateway.
  • jevgate-action#2: merge after 0.26.0 ships; its base input text still says "only files changed"; choose api-key-kind or provider as the input name.
  • Release step 1: the two hardcoded-values considers above, to fix or accept as wrong; and the baseline entry added here.
  • The ky label added here (Ky::constructor, TP): review it.
  • The site was not built locally (no mdbook); the Linux Secret Service and Windows auth paths compile only in CI.

Jev spend: after the stack critique $0.013 (0.25.0's and this branch's checks of the new commits); before it: provider $0.029, scope $0.144, integration $0.036, finish $0.094 (JevGate's own check twice), and this pull request's two CI reviews $0.292 (0.25.0 judges the 77 changed files whole, about 1,575 requests a run; the first run failed, so its answers were not cached for the second): about $0.60 of the version's $1.00.

A rule and level is mature when its findings were right at least 80% of
the time on the 25 projects JevGate was never tuned on, over at least 20
hand labels. The default gate level is now `mature`, which fails only on
those; every other finding is reported without failing the check. Any
level set in jevgate.toml or with --fail-on replaces it as it says, and
`mature` can be set by name.

Measured from the corpus labels joined to 0.25.0's findings (replayed
from the answer cache): function-simplification reviews (20 of 23) and
agent-context considers (22 of 24) are mature. On the unseen projects'
full runs the default gate failed on 122 findings, 57% of the labeled
ones right, and on 17 of 22 projects; it now fails on 23, 87% right, and
on 9. The concern probability does not separate right from wrong reviews
there (55%, 46%, 56% and 61% across four bands), so the table in
src/maturity.rs decides.

Each finding records how the gate counted it (`gate`: fails, measuring or
advisory), so the JSON report, annotations, SARIF, GitLab, the HTML report
and the MCP tool agree, and the report names what `mature` stands for in
`fail_on_mature`. The agent text marks findings that fail the gate and
lists them first wherever a list is capped, and says when reviews did not
fail it because their rules are still being measured. `jevgate rules`
shows the levels that fail by default and each rule's precision on unseen
projects. `jevgate init` writes the groups as commented examples, since
`maintainability = "review"` would opt every maintainability review back
in.
On the 25 projects JevGate was never tuned on, 6 of its 37 labeled
reviews and considers were right (16%), against 47 of 85 on the projects
it was tuned on (55%). It stays in the maintainability group and in
`all`; `--rule default --rule hardcoded-values` adds it back.

Without it, a default run plans 44% fewer first-pass requests on the 94
labeled corpus projects (22,370 to 12,623) and uploads 39% fewer bytes;
with `--rule all` every request body is unchanged.
…ateway names

Gateways name Jev `typesafe/jev-1.13`, `~typesafe/jev-latest` and
`typesafe-ai/jev`. JevGate took any name but `jev-latest` and
`jev-preview` for a pinned version, whose cached answers never expire and
whose answer must name exactly that model, and its model check refused the
`/` and `~` in an answering model's name: every gateway answer would have
failed.

A name is pinned only when its last segment ends in an x.y.z version and no
`~` marks it as moving; a pinned request accepts its own version with or
without a gateway's namespace. The default `jev-1.13.0` is unchanged, so
nothing is asked again.
… usage

The cost was priced only when the requested name was exactly
`jev-1.13.0`, so a `jev-latest` run showed none although TypeSafe answers
it with `jev-1.13.0`, and a response without `usage` failed validation.
Gateways need not pass TypeSafe's `usage` through.

Each paid answer is now billed to the model it names; an answer without
usage is accepted, cached without it and counted as unmetered, which makes
the cost unknown ("cost unknown" in the headline, `estimated_usd: null`)
instead of $0, and keeps it out of the bytes-per-token calibration. The
price table moves to `model.rs`, by model line (`jev-1.13` and its
versions under any gateway namespace), checked against TypeSafe's models
page on 2026-09-28.
A pure move: the queue, retry and error tests leave transport.rs (915
lines, half of them tests) for transport/tests.rs, before the provider
exchange tests join them.
…402 and 422

TypeSafe names each request in `x-typesafe-request-id` and its SDKs honor
`retry-after-ms` before `Retry-After`; JevGate kept neither and read
`Retry-After` only as whole seconds. A 402 said "request was not retried"
and nothing about credits; a 422 said nothing about what was invalid.

An error now ends with the request id when the provider sent one (else the
response's own `id`, which OpenRouter sends), and each judgment and cached
answer keeps the id of the request behind it. A retry waits for
`retry-after-ms`, else `Retry-After` in seconds or as an HTTP date, still
capped at 30 s. A 402 says the credits are exhausted and where to add them,
also in the message of every request it stopped; a 422 names each
`detail[].loc` and `type`, never `msg` or `input`, which can echo source; a
404 suggests the model name; a 413 counts as beyond the model's context.
Header values are kept only when they are safe to print.

Messages name the provider through a `Service`, so the gateways can share
them; a mock provider on 127.0.0.1 (tests/support/mock_provider.rs) tests
whole exchanges: the bearer key, the uploaded body without local metadata,
the request id, a 429 with `retry-after-ms`, and a 422.
…pts after 20 s

TypeSafe documents 1,200 requests a minute. On the corpus's largest runs six
workers made 18 to 20 requests a second (0.3 s a request), just under it,
and `--concurrency 8` would make about 27. Requests now start at least
50 ms apart across every worker and retry, and concurrency accepts 1 to 6.

An attempt gave up only after 60 s; TypeSafe's SDKs wait 10 s. It now gives
up after 20 s, and a timed-out request is still sent once more. A mock
provider that answers late shows the retry.
People who already pay for OpenRouter or Vercel AI Gateway can run JevGate
without a TypeSafe account, as they can Abide and Qlty: both gateways serve
TypeSafe's API and the same model at the same price. JevGate read only
TYPESAFE_API_KEY and sent every request to api.typesafe.ai.

The provider follows the key. `jevgate auth login` asks which kind of key
it is (`--provider` for scripts) and saves it with its provider; a check
reads TYPESAFE_API_KEY, OPENROUTER_API_KEY or AI_GATEWAY_API_KEY from the
environment, then `--env-file`, then the saved key. The repository's .env
is read only for TYPESAFE_API_KEY: a gateway's key there is usually the
application's own. `auth status` shows the provider, the endpoint and the
keys set but not used; the headline says `via OpenRouter`.

A key goes only to its own provider. jevgate.toml and .env cannot name a
host; a key with another provider's prefix (`sk-or-`, `vck_`) is refused;
JEVGATE_BASE_URL, from the environment only, takes https or loopback http
for a self-hosted proxy and is announced on stderr. A gateway's saved key
is stored as `<provider> <key>`, which 0.25 refuses as invalid rather than
send to TypeSafe; a plain `provider` file beside it lets a check choose its
default model (`typesafe/jev-1.13`, `typesafe-ai/jev`) before reading the
key, and a mismatch stops the run instead of sending anything.

Tested end to end against a local server in each gateway's shape
(OpenRouter's `id`, `provider` and `usage.cost`; an answer without usage);
no gateway has been called.
A float `sum` over no items is -0.0, so a fully cached run printed
"~$-0.0000" and wrote `"estimated_usd": -0.0`; the canaries replayed from
their caches showed it. The total now folds from 0.0.
TypeSafe answers `"model": "jev-1.13"`, a name its docs use, with HTTP 400
and `{"detail": {"error_type": "api_usage_error", "message": "Unknown
model: jev-1.13"}}` (asked on 2026-09-28; the request id header came with
it). JevGate said only "TypeSafe HTTP 400". It now says "unknown model;
check the model name", reading the message's fixed opening and never
echoing it.

The Models changelog entry said TypeSafe took `jev-1.13`; it does not, so
the entry now says what the alias rule changes: gateways' names and names
other than jev-latest and jev-preview.
…tate

Asked on 2026-09-28: its detail[].input echoes the whole request, which is
why the message reads only loc and type.
…ame two numbers

A key saved by 0.25 was saved as TypeSafe's whatever it was; its prefix is
now checked like the environment's and the files'. A paid answer whose model
name fails validation is billed to "unknown" (no price) rather than to its
raw text. JevGate's own check of this branch noted two numbers a reader
must guess: the provider record's size limit and the days in 400 years.
An `--env-file` is now read for all three providers' variables, and a
template's empty `OPENROUTER_API_KEY=` beside a real TypeSafe key failed
the whole file as an invalid key. An empty value counts as unset, as an
empty environment variable already does.
`--base` compared the files `git diff --name-status` printed from the Git
top level with paths relative to JevGate's root, the nearest directory
holding `jevgate.toml`. In a repository whose `jevgate.toml` is in a
subdirectory, such as one package of a monorepo, no tracked change
matched: the check selected only untracked files, which `ls-files`
already prints from the root, and a pull request of edited files passed
as "no-changed-source". `--relative` makes the diff print paths from the
root and leave out changes outside it.

The unit tests' Git helper moves to `Project::git`, shared by the
documentation tests and the new one.
`--base` selected whole changed files, so a pull request check asked
about, paid for and reported every unit of a touched file: on each
corpus project's last commit, 54% of the review and consider findings in
the changed files sat on lines the commit did not touch.

A check with `--base` now asks about and reports only what the change
touches:

- functions, tests, comments, values and security units whose lines the
  change added, modified or removed lines inside (`git diff -U0`,
  streamed so a large patch never sits in memory, `--text` so a file
  marked `-diff` keeps its lines); a new or untracked file is judged
  whole;
- copy pairs where either copy changed;
- a file's outline, and a large document's, only when the change adds a
  member or heading its base version lacks (read from Git only then);
- a document the change left alone when one of its sections names a
  path the change deleted or renamed, and only that section.

Functions, values, comments, security units, instruction sections and
laws are still packed within runs whose ends are computed over the whole
file, so the changed functions of one run share a pack and the packs of
other runs stay as they were. `--whole-files` keeps the whole-file check
and asks exactly what 0.25.0 asked.

The report says which scope a check used (`scope`: `changed-lines` or
`whole-files`), as do the headline and the HTML report. `baseline
--merge` after a check of changed lines keeps the other entries of the
files it checked, and a finding such a check did not judge is
`non-comparable` rather than `resolved`. The MCP tool takes
`whole_files`.

On the last commits of 118 corpus projects (all rules, tests included),
0.25.0 reported 311 review and consider findings, 167 of them off the
changed lines; this reports 149, 9 off them (5 outlines the change added
members to, 3 copy pairs whose other copy changed, 1 section with lines
removed inside). Of the 144 findings on changed lines, 137 stay: 3
outlines of files the commit added no member to are not asked, a group
of six comments of which the commit touched one is a note, and 3
function-simplification considers asked beside fewer functions are
notes. With nothing cached, the first pass asks 2,259 requests and 4.55M
input tokens instead of 4,828 and 11.05M. The corpus runs cost about
$0.12.
0.26 caps concurrency at 6, and `concurrency = 7` or `8` in jevgate.toml, or
`--concurrency 8`, became an error. 0.25 accepted up to 8, so a CI
configuration valid yesterday would stop every run with exit 2 after an
upgrade, which the stability page's rule for deprecated settings (keep
working, with a warning on stderr) is meant to prevent. Nothing needs more
than 6: requests start at least 50 ms apart whatever the concurrency, and
the transport already clamped the workers to 6.

A value above 6, from the flag or the file, is now lowered to 6 with one
line on stderr; 0 is still an error. The schema keeps 6 as its maximum, so
editors still point at the value to change.
Since 0.26 a run's cost is unknown when a response reports no token usage,
which a gateway need not send, as well as when the model that answered has
no published price. The headline says "cost unknown" for both, but the HTML
report still explained every unknown cost as "No published rate configured
for this model", wrong for the first case, which is the one a Vercel AI
Gateway run is likely to hit.

The page now gets the report's unmetered request count and says "Cost
unknown: 3 responses reported no token usage", or that the model that
answered has no published rate. Rendered in jsdom with both cases and a
known cost.
The three 0.26 changes each updated the pages they touched; read together,
a few places still missed the others:

- ci.md: how the action takes a gateway key (`api-key-kind`, jevgate-action
  1.2), and what the two gate changes add up to for a pull request check:
  replayed on the last commits of 118 corpus projects, the defaults fail 9
  of them, all on function-simplification reviews, where 0.25.0 failed 19.
- output.md: the headline now carries the scope a `--base` check judged,
  the gateway that answered and an unknown cost; say so in one place.
- `jevgate rules`, the rules reference, what-it-finds and coding-agents:
  say plainly that an opt-in rule's mature levels fail the check once it is
  selected, as agent-context considers do with the documentation rules.
- privacy-and-cost: Vercel AI Gateway's public model list marks Jev
  `zdr: none` (checked today), so it cannot apply zero data retention to
  these requests, as the page said it could.
- `check --help`: `--cache-only` never contacts the provider, whichever it
  is.
- "a key that starts as another provider's do" was missing a word.
…ck fails on

The three changes were merged as eleven bullets in the order they
landed. They are now one section: an opening paragraph with what they do
together, then the default gate, `--base`, and gateway keys, each with its
details under it.

The opening gives the number the two gate changes decide together, which
neither change measured: a pull request check with the defaults
(`jevgate check --base`, as the action runs it), replayed from the answer
cache on the last commit of the 118 corpus projects the scope change used.
0.25.0 fails 19 of them on 61 reviews, 31 of the 45 labeled right (9
wrong); 0.26 fails 9 on 11 function-simplification reviews, 8 of the 10
labeled right (none wrong, 2 debatable).

Also: the `--base` table gives the findings lost on changed lines next to
the ones off them, agent-context considers failing the check once the
documentation rules run is said plainly, and the action's `api-key-kind`
is named where gateway keys are.
JevGate's own check of the version (`check --base main --rule all
--include-tests`) left two considers, both in the gateway change's tests:

- `serve` in the mock provider read the request line, the headers and the
  body, logged the request and wrote the reply in one function (0.92).
  Reading the request, its headers and writing the reply are now their own
  functions.
- The rate-limit and timeout tests repeated the same steps: a counter
  that treats the first request differently, then one request through a
  one-worker queue (0.86). `mock_first` and `queued` hold them.

No behavior changes; the same tests pass.
The heading came right after the paragraph above it when the gate and key
sections were merged.
The default gate and the changed-lines scope were each tested alone. A
check with `--base` now runs both: of six functions a scripted provider
calls review-worthy, only the one the change edited is judged, and its
function-simplification review, a mature level, fails the gate on its own;
with `--whole-files` all six do.
OPENROUTER_API_KEY and AI_GATEWAY_API_KEY in the environment came before
--env-file, the repository's .env and the key saved by `jevgate auth
login`. Other tools read both variables, so a user who exported one for
them and upgraded had every check moved to the gateway: billed to that
account, sent through a data processor they never chose for JevGate, and
asked again from scratch, since the model name changed. The critique
reproduced it with `--env-file ts.env` and an exported OpenRouter key:
every request went out with the OpenRouter key.

A check now reads TYPESAFE_API_KEY from the environment, then the
credential file, then the saved key, and only then the gateways'
variables. A key saved before 0.26 has no provider recorded beside it,
so the credential store is read to find it, and only when a gateway's
variable would otherwise be used; `resolve` reads it once. In CI the
system store needs a terminal, so there the gateway's variable is used as
before. `jevgate auth logout` no longer names a gateway's variable as an
override of the saved key, since it is read after it.
A run stopped by exhausted credits printed its headline, "1 files
failed." and exit 2: the reason (`TypeSafe HTTP 402 (credits exhausted;
…)`) was only in the JSON report, with --verbose, or in the GitHub
annotations. The MCP server's `jevgate_check` returns this text, so an
agent could not tell a 402 from a missing key, and 0.26's new messages
for 402, 422, an unknown model or a key issued by another provider went
unseen by default.

The summary now lists each distinct failure reason with how many files
gave it, before the skip reasons it already listed, and counts one file
in the singular. Tested end to end against a mock provider answering 402,
in the agent text and in the MCP tool's result.
`Retry-After: Sun, 06 Nov 500000000000 08:49:37 GMT` panicked in the
request worker ("overflow when adding duration to instant"): the date
parser took any year from 1970 up, and the seconds since 1970 no longer
fit the system time. The request then failed as a worker failure without
its retry, and the run ended incomplete. For larger years the product
wrapped silently in release builds.

The form HTTP requires senders to use has a four-digit year, so any other
is not a date and the usual backoff applies; the sum is also added with
checked_add. 31 Dec 9999 still parses.
`send` put the checked request id (the header, else the body's `id`)
under `request_id`, but when there was none it left a `request_id` the
provider's body carried itself, and the report copied it unchecked: a
mock answering `"request_id": "not an id \u001b[31m<script>…"` put the
escape and newline into files[].judgments[].request_id on a live run,
while the cache dropped it, so a cached replay's report disagreed.

A body's own `request_id` is now replaced by the checked id, or dropped
without one, as the cached answer drops it.
A changed-lines check ran `git diff -U0 --text` over the whole change.
A branch that changes a 150 MB binary made Git write a 301 MB patch: a
dry run took 0.87 s and 460 MB where 0.25.0 took 0.14 s, and `--base
--watch` repeated that diff on every 250 ms poll, keeping Git busy with
nothing edited. A regenerated 10 MB lockfile cost 0.14 s a poll.

The lines are now read after the inputs are known, only for the files
read to be judged, each with its name before a rename so Git still pairs
the two. The paths go to Git in batches of about 16 KiB, well under a
Windows command line, and no path means no diff. On the binary branch a
dry run is back to 0.10 s, as with 0.25.0; on 16 corpus projects the
changed-lines request bodies are identical (439 requests).

Git also lets GIT_DIFF_OPTS outrank -U0: with `--unified=6` set, context
lines counted as changed, and a function next to the edited one was
judged and could fail the gate. The variable is now removed from Git's
environment.
…in --concurrency 0

The CHANGELOG, the configuration page and the schema said a
`concurrency` of 7 or 8 in jevgate.toml is lowered to 6 with a notice.
It never is: the file's value is a ceiling on the flag, whose default is
6, so it stays 6 without a word, as in 0.25.0. Only `--concurrency`
above 6 reaches the cap and its notice. The docs and the schema now say
so, and the unit test lowers `--concurrency 8` through `configure`, so it
fails without the cap; it passed without it before.

`--concurrency 0` answered "0 is not in 1..=4294967295", contradicting
"at most 6"; it now says "Use a whole number from 1 to 6; a higher one is
lowered to 6".
Capped lists put the findings that fail the gate first, then the rest by
rank alone, so a review still being measured competed with higher-ranked
considers: GitHub shows 10 warning annotations a step, and in 0.25 reviews
were errors with a cap of their own. Replaying the order over the
whole-repository runs of 94 corpus projects with the default rules
(results/mat-gate), 267 of 424 such reviews fell past the tenth warning,
in 35 projects. With reviews first, 109 do, all in the 10 projects that
hold more than ten of them.

The same order serves the agent text, the job summary's 50 rows and the
MCP server's 50 findings. A stable sort keeps the rank order within each
part.
OpenRouter's public listing gives `typesafe/jev-1.13`, JevGate's default
model with an OpenRouter key, the canonical slug
`typesafe/jev-1.13-20260917` at $0.000000042 per prompt token (checked
on 2026-09-28, no key needed). The price table knew only `jev-1.13` and
its `.N` versions, so if OpenRouter's answers name the endpoint, every
OpenRouter run would show "cost unknown" and a null `estimated_usd`.

A dated snapshot of a priced line (eight digits after a hyphen) is now
priced like the line. Vercel AI Gateway's `typesafe-ai/jev` stays
unpriced: it names no version, so its price follows whatever it points
to; the privacy page and the CHANGELOG now say so. The maintainer's
probe script shows which name each gateway answers with.
…led yet"

A law finding still being measured said its rule was "still being
measured (none labeled yet on projects JevGate was never tuned on)". The
maturity table leaves Bend 2's labels out, and `tests/laws` judges only
Bend 2 code, so it has no row; but its findings were labeled, on Bend 2
projects (0.22.0 counted 15 right of 23 on projects then unseen), so the
text was wrong for it.

The maturity module now words why a level has no share: law findings
were labeled only on Bend 2 projects, which the table leaves out; any
other level has none labeled yet. GitHub, SARIF and GitLab messages and
the HTML report use that text, `jevgate rules` and the rules reference
say it in their legend, and the CHANGELOG says Bend 2's labels were kept
apart.
`jevgate check -h`, the man page and the completions showed
`--model … [default: jev-1.13.0]` whatever the key, while the default
follows the key's provider (typesafe/jev-1.13 with an OpenRouter key,
typesafe-ai/jev with a Vercel AI Gateway key), as the long help says. A
gateway user reading -h could pass jev-1.13.0, which OpenRouter refuses.
It now reads `[default: the key's provider's model]`, as --rule and
--env-file describe theirs.
…rrect them

ci.md called the 118 corpus projects of the last-commit replays
open-source; 28 of them are the maintainer's private repositories, and 8
of the 9 projects the 0.26 default check fails are among them. ci.md and
the CHANGELOG now say so, and the CHANGELOG adds that 10 of the 11
failing findings are on those repositories and the other, ky's `Ky`
constructor, was labeled right (labeled from the code for this: the
constructor carries `eslint-disable complexity` and folds URL resolution
and a 60-line search-params merge into setup). With that label the 0.26
replay is 9 of 11 right, none wrong, 2 debatable, and 0.25.0's 32 of 46.

0.25.0 failed 18 of those projects, not 19: the 19th's replay lacked one
cached answer and ended incomplete. The losses on changed lines now say
the comments consider was labeled right, as the scope note has it.

The CHANGELOG said twice that nothing is asked again on upgrade; that
holds for the model. A --base check asks its touched units again once, in
changed-lines packs: with a cache holding whole-file answers of the same
commits, the corpus's first pass needed 716 new requests and 1.34M input
tokens against 765 and 1.06M for 0.25.0.

jevgate-action 1.1, today's @v1, has no `api-key-kind` input and passes
`api-key` as TYPESAFE_API_KEY, which 0.26 refuses for an OpenRouter key.
ci.md and the CHANGELOG now give the form that works with both: leave
`api-key` out and set the gateway's variable in the step's `env`.
JevGate's own check (`--rule all --include-tests`) found two considers in
code this version changed:

- `baseline::write` mixed separate jobs (0.94): collecting the findings
  to accept, carrying earlier reasons, and keeping the entries a merge
  keeps, which 0.26 extended for changed-lines checks. Each is now a named
  function; behavior is unchanged.
- The new gateway test's `save_as_0_25` repeated `save_credential` of the
  auth tests (0.99). Both now use one `Project::save_credential(key)`.

`auth::file::save` gets the doc comment its sibling `load` has, saying
the file holds an API key kept as written, since the provider needs it as
is; the check read it as a password kept in plain text (0.83).
…xture

The pull request's own review (JevGate 0.25.0, `--rule all
--include-tests`, where every review fails the gate) failed on
`src/units/tests/mod.rs:312`: "One of this file's constants fixes a value
that differs between deployments (0.98). The constant is `HARDCODED`."
`HARDCODED` is the source text the hardcoded-values tests check,
`const REGION: &str = "eu-west-1"` included, so the finding is wrong
(test data). The code is untouched by this version; its request changed
because the file gained a line above it.

Accepted with `jevgate baseline --merge`, keeping only this entry, and
`jevgate baseline mark wrong src/units/tests/mod.rs:312`, as AGENTS.md
asks for a mistaken finding.
A check with an OpenRouter or Vercel AI Gateway key now sends at most 3
requests at once unless --concurrency or concurrency in jevgate.toml says
otherwise; a TypeSafe key keeps 6. Six workers at JevGate's pacing make up
to TypeSafe's limit of 1,200 requests a minute for one account, and
through a gateway that account is the gateway's, shared with its other
customers.

It is a precaution, not a measured fix. On 2026-09-28 the canaries
through OpenRouter got HTTP 503 on 141 of 224 attempts (63%), but a
self-check with a TypeSafe key at the same time got 22 of 34 (65%), and a
1,083-request TypeSafe run half an hour earlier, 6 at once, had no retry:
the 503s were TypeSafe's. OpenRouter's rounds of a few requests at once
fared only a little better (18 of 36 attempts, against 123 of 188 in
first-pass rounds of up to six; z about 1.7).

concurrency in jevgate.toml now applies without the flag and caps it, as
max_requests does, so a repository can raise a gateway's default up to 6;
above 6 it means 6, silently, as before. The report's concurrency is the
value used.
An answer worth retrying (HTTP 408, 429, 500, 502-504, 520-524, 529) is
now sent up to 6 times instead of 4, with every provider, pausing 1, 2,
4, 8 and 8 s, each up to a quarter longer; the first three pauses are
0.25's, and a longer pause the provider asks for is still honored up to
30 s. A timeout is still sent twice at most, and a connection that fails
before sending 4 times, so an offline run ends no later.

On 2026-09-28 TypeSafe answered 503 to about two attempts in three for at
least ten minutes, directly (a self-check at 10:30: 22 of 34 attempts, 1
of 13 requests gave up) and through OpenRouter (the four canaries, 10:34
to 10:39: 141 of 224 attempts, 16 of 99 requests gave up), each failed
attempt taking about 10 s. Every one of those runs ended incomplete. An
independent failure of 63% per attempt predicts 15.6 of 99 requests
exhausting 4 attempts; 16 did.

A simulation of this queue (shared pause, 50 ms pacing, jitter) puts
numbers on the change. In that brownout, 6 attempts leave 6.3% of
requests unanswered instead of 16%, and a 13-request check completes 42%
of the time instead of 9%. When 1 attempt in 5 fails, a 1,000-request run
completes 95% of the time instead of 21%. A provider failing every
attempt is outlasted for 26 s instead of 8.7. The cost comes only while
attempts fail: a hard outage answered at once ends a 100-request run
incomplete after 8.3 minutes instead of 2.7 at 6 at once (16 instead of
5.3 at a gateway's 3). The 0.27 hook is bounded by its own deadline.

TypeSafe's SDKs retry twice; a gate's run that ends incomplete costs a
rerun of the whole job, where one more pause of 8 s costs little. The
pause stops doubling at 8 s because it holds every worker.
Git exports GIT_DIR to a hook, `git rebase --exec` or `git bisect run`
in a linked worktree, and GIT_INDEX_FILE to a pre-commit hook (absolute
for `commit -a` and a partial commit; probed on Git 2.51.0). The tests
run `git init`, `add` and `commit` in temporary projects, and with those
variables inherited they acted on the outer repository instead. `git
init` with GIT_DIR set to a worktree's administrative directory guesses
a bare repository and writes `core.bare = true` into the configuration
the worktree shares with its main clone; `add` and `commit` then stage
and commit the temporary project's files on the worktree's branch. The
check's own Git, run in-process by the unit tests and by the binary the
CLI tests start, read the outer repository too, and its `git diff`
rewrote the outer index's stat cache, which Git does even with
GIT_OPTIONAL_LOCKS=0.

tests/support/git.rs, shared by the unit and CLI tests like temp_dir.rs,
removes the variables that point Git at a repository (GIT_DIR,
GIT_WORK_TREE, GIT_INDEX_FILE, GIT_COMMON_DIR, GIT_OBJECT_DIRECTORY,
GIT_ALTERNATE_OBJECT_DIRECTORIES, GIT_NAMESPACE, GIT_PREFIX) from every
test's Git and from the jevgate the CLI tests start. The check now starts
Git in one place, revision::git_in, which docs::history::ignored uses
too, and removes them there under cfg(test). tests/lint_policy.rs fails
when Git is started anywhere else.

Measured by running the whole suite with the variables pointing at a
throwaway repository that has a linked worktree, in three shapes:
GIT_DIR, GIT_WORK_TREE and GIT_INDEX_FILE together; GIT_DIR alone, as
`rebase --exec` in a linked worktree exports it; and an absolute
GIT_INDEX_FILE alone. Before, 20, 15 and 14 of 572 tests failed, and
each run changed the throwaway repository: core.worktree written into
the shared configuration; core.bare = true and 7 commits of test files
on the worktree's branch; the main index rewritten. After, all 573 tests
pass in each shape and all 43 files of the repository (configuration,
HEADs, refs and indexes included) keep their sha256.

A check still honors the variables: Git sets them for the repository the
check runs in, and in a pre-commit hook GIT_INDEX_FILE is the index being
committed. git_in's documentation says why.
A check's own Git honors GIT_DIR, GIT_INDEX_FILE and the other variables
that point Git at a repository, where the tests' Git now drops them. Git
sets them for the hooks, `rebase --exec` and `bisect run` it starts, for
the repository the check runs in: in a pre-commit hook GIT_INDEX_FILE is
the index being committed (`.git/next-index-*.lock` for a partial
commit), and a work tree whose repository is kept elsewhere, as with a
bare dotfiles repository, is found only through GIT_DIR and
GIT_WORK_TREE. Dropping them would protect nothing: a check only reads,
but for `git diff` refreshing the index's stat cache, which leaves what
is staged unchanged.

The test moves a project's .git out of its work tree and checks that
`check --base HEAD` with both variables set judges the changed file.
With the variables dropped from the check's Git it fails: "Cannot find
revision HEAD ... not a git repository".
The MCP server's instructions and the AGENTS.md snippet on the coding
agents page still said "Fix each `review` finding", the policy of 0.25's
gate. With 0.26's default gate, reviews of rules still being measured are
reported without failing it, and on projects JevGate was never tuned on
shared-logic reviews were right 46 of 85 times, file-organization 2 of 5
and sensitive-data 10 of 24: an agent following the old text refactors
on findings wrong about half the time.

Both now say what 0.27's hook instructions say: fix each finding marked
to fail the gate, and weigh the other reviews and considers, fixing one
when it is right or saying why the code should stay.
The action example suggested `args: --rule security`. Naming a rule
replaces the selection, so the action ran only the five security rules;
no level of them is mature, so under 0.26's default gate the check could
never fail, and function simplification, the one level that blocks, was
off. A dry run on zoxide with `--rule security` shows exactly those rules,
`fail_on` ["mature"] and no `fail_on_mature`. Under 0.25.0 the same args
still failed on security reviews.

The example is now `--rule default --rule security`, as the action's own
input description has it, with a sentence on what a selection without a
mature level does.
`jevgate init` wrote `maintainability = "review"` and `tests = "review"`
from 0.3 to 0.25, the README's first step, so most configurations hold
them. Under 0.26 they keep every review of those groups failing the check
(shared-logic reviews right 54% of the time on unseen projects) and judge
hardcoded values (16%), and from 0.27 the agent hook blocks on them; only
the CHANGELOG and the configuration page said so.

A command that reads jevgate.toml now says so on stderr while the lines
are there with the comment init wrote after them, which lists the group's
rules and tells them from a level set by hand: deleting the lines gives
the default rules and gate, and deleting only the comments keeps the
levels without the notice.
The hero image was JevGate 0.22.0 on zoxide: "gate failed: 3 new review
findings", two of them shared-logic reviews, a hardcoded-values consider
and eleven notes, each message ending in a probability. 0.26 fails that
commit on one function-simplification review, reports the other three
reviews and a consider as still being measured, and no longer runs
hardcoded values by default. Both images are now 0.26's, replayed from
the corpus's answer cache on the same commit (09a18b4) with no request:
the default `jevgate check` agent text, and the HTML report with
src/util.rs open, its failing review marked.

The README and the site's introduction also say what the gate blocks by
default and how often it was right (function-simplification reviews, 20
of 23 on unseen projects), and the README what a check costs (the last
commits of 118 corpus projects, checked as pull requests with every rule
and nothing cached, sent 4.55M first-pass input tokens, $0.19) and that a
rerun of unchanged code sends no request. A Vercel AI Gateway key is said
to be accepted but untried with a real key: only OpenRouter has answered
a live request, and Vercel's `typesafe-ai/jev` shows its cost as unknown.
The tests' temporary directories are canonical, which on Windows is the
verbatim form \\?\C:\..., and Git for Windows reads a GIT_DIR in that
form as "not a git repository": base_reads_the_repository_git_dir_names
failed on the Windows CI job. The test now strips the prefix, as a person
would write the path; nothing changes elsewhere.
@tauanbinato

Copy link
Copy Markdown
Contributor Author

Superseded by #43, which carries every roadmap version from 0.26 to 0.30 on one branch, to be released together as 0.30.0. Everything in this pull request is in #43 unchanged (the v0.26 branch is its base for the per-version compare views): review 0.26 there through main...v0.26.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant