Skip to content

Count a request whose answer is cached as fitting, whatever the calibration - #38

Merged
tauanbinato merged 2 commits into
split-composefrom
stable-fits
Sep 27, 2026
Merged

tauanbinato merged 2 commits into
split-composefrom
stable-fits

Conversation

@tauanbinato

Copy link
Copy Markdown
Contributor

Stacked on #37 (and so #36): its base is split-compose. The change itself touches none of their code; the base only avoids a conflict on the CHANGELOG's Unreleased section, where #36 adds a line too. Merge #36 and #37 first and retarget this to main, or cherry-pick its first commit alone onto main.

The problem. Two runs of the same release on the same code can give different findings.

  • Whether a request fits the provider's limit is decided from a token estimate. The estimate's calibration (.jevgate/token-budget.json) is replaced after each run by the bytes per token of that run's fresh requests.
  • So a request near the limit, typically a recheck or a kind that sends a whole file, is planned in one run and dropped in the next. The outline then falls back to its first answer, or asks its kind from the outline alone, which is a new question.

I hit this while comparing branches on the corpus (see #36). One pi-fabric outline was a note in one run of 0.23.1 and clear in another, so every comparison needed each clone's calibration restored from a snapshot first. In CI, or in check --watch, the same thing makes a gate flicker without any code change.

The change. Planning counts a request as fitting when the estimate says so, or when the answer cache already holds its answer: the provider took it once.

  • The cache is looked up only for requests the estimate rejects, so planning costs the same.
  • The file-purpose request follows the same rule.
  • --refresh skips the cache, so it plans by the estimate alone.
  • TokenBudget is unchanged. Planning now receives Limits, which pairs the budget with that cache lookup, and FileContext.budget holds it, so no planner changed.

Measured on the corpus, 181 projects, from the cache.

  • 0.23.1 under three calibrations (the last run's, one snapshot's, and a stricter 2.5 bytes per token everywhere) differed in one consider (b2-toon_bend's scripts/claims-audit.py) and five undecided units. Under the stricter one it also asked about 42 thousand new tokens of questions.
  • With this change, pinned to the same snapshot, the only difference from 0.23.1 is that toon_bend consider getting its cached recheck back. Under the stricter calibration, only three units still differ, and nothing new is asked.

Known remainder. Those three units are in b2-bend-collections. Their whole-file rechecks fit a loose estimate, so they're sent, but the provider refuses them as too long. A refused request is not cached, and a unit that has first answers keeps its planned recheck. So the kind fallback isn't asked in that run, and the file stays undecided, while a stricter calibration asks the kind from the outline and decides it. Fixing that means remembering refusals between runs (a request the provider refused fits no more) and letting a unit whose recheck was refused take the outline-only kind in the same run. I left it for a separate change.

Test: an outline whose recheck was answered at 3.0 bytes per token is rechecked from the cache at 2.0 and keeps its status, with no request sent. It fails without the change, because the second run asks the kind from the outline alone.

Tests, clippy (also 1.98), cargo +1.90.0 check --locked and the self-check (no review or consider) pass.

@tauanbinato
tauanbinato merged commit aac50c7 into split-compose Sep 27, 2026
9 checks passed
@tauanbinato
tauanbinato deleted the stable-fits branch September 27, 2026 21:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant