[Feat] OmniJev as a vision decision model for the browser agents - #33
chaimaerachdi wants to merge 3 commits into
Conversation
OmniJev (github.com/tinnel123666888/OmniJev, Apache-2.0) is a System 1 decision model on Qwen3.5 vision-language backbones that answers typed questions over a screenshot. `--model omnijev` runs it in process on the browser agents: OmniJevModel maps the choice and noul questions to its shape, renormalises a choice's probabilities over the offered keys (OmniJev leaves its abstain probability out of them) and requires an image. The model code comes from a local clone named by OMNIJEV_REPO; the omnijev extra brings torch, torchvision, transformers>=5 and peft. The browser front now sends a viewport PNG with every observation to a decision model that reads images, through the probe's run-code executor, and records screenshot_ms per tick. Text-only models see no change. README: a section with four of OmniJev's own v1.1 demo replays, linked from its repository at a pinned commit and credited. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
full (the default) sends the text state, the goal and the rules with every question; short sends the goal and the ask only, the screenshot carrying the page. Measured offline on Jev's 12 recorded Google Flights steps with the 0.8B on CPU: full picks Jev's whole step 3 of 12 times at 178 s a step, short 1 of 12 at 57 s. Full stays the default; short is there for speed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hsliuustc0106
left a comment
There was a problem hiding this comment.
Reviewed commit 1d01e69. Found one actionable P2 issue, detailed inline.
Validation: 207 tests passed, 4 skipped, plus 21 subtests across the OmniJev adapter, screenshot capture, factory, browser setup, browser policy, and shared decision-model tests. A minimal reproduction confirmed that truncation produces identical operation prompts before and after an unsuccessful Search click. No live model or GPU inference performed.
| """The observation's text state, read in front of every question (OmniJev reads text context in the instructions).""" | ||
| state = observation.state | ||
| text = state if isinstance(state, str) else json.dumps(state, ensure_ascii=False, separators=(",", ":")) | ||
| return text[:OMNIJEV_STATE_CHARS] |
There was a problem hiding this comment.
[P2] Preserve action history when truncating context
The browser observation serializes page text before elements and recent_actions. A page with 6,000 characters of text consumes this entire slice, dropping both fields. I reproduced identical operation prompts before and after an unsuccessful Search click, although the browser rules require switching to PRESS_ENTER after that failure. The screenshot cannot supply this history, so the model loses information needed to select the next operation. Limit page text separately and preserve recent actions within the context budget.
There was a problem hiding this comment.
Thanks, good catch. Fixed in e2bf4c5: a browser state now puts its recent actions first and caps the page text at 2,000 characters, so a long page can't push the history out. Test added
The browser state was serialised page text first and cut at 6,000 characters, so a page with a long text dropped the elements and the recent actions: OmniJev saw the same prompt before and after a Search click that did nothing, although the rules then call for PRESS_ENTER. A browser state now puts its recent actions first and caps the page text at 2,000 characters on its own. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Why
Bring OmniJev (Apache-2.0) into system1-agents, integrate
its use cases and show its demos in the README. OmniJev is a System 1 decision model on Qwen3.5 vision-language
backbones (0.8B, 2B, 4B) that answers the same typed questions as Jev over a screenshot. It is also a local model
trained on web tasks (Mind2Web), where Laya, trained for text classification, fails on Google Flights (#9).
The backend, the screenshot and the README are ready. On a GPU, OmniJev-4B completed the Google Flights task in 1 of 7 runs, with the same answer as Jev (easyJet, €72), and reached Search in 2 of 7, at 1.1 to 1.7 s a decision (see Results). Jev completes it every time. The 0.8B on CPU gets lost.
How
s1a/decision_models/omnijev.py:OmniJevModel(supports_images=True, deterministic) over OmniJev'sMSO1.system_one({"images": [...]}, questions). The text state, goal and rules go in front of each question; abrowser element row becomes its label and value. A choice's probabilities sum to
1 - abstainin OmniJev, so theyare renormalised over the offered keys (the key stays OmniJev's own, so a malformed answer still fails validation;
abstainstays inraw). An observation without an image isMODEL_SERVICE_CONFIG_ERROR.s1a/browser/decision_model.py: when the decision model reads images, every tick's observation carries a viewportPNG, taken with
page.screenshot()through the probe's run-code executor (tab brought to the front, 15 s cap);the tick records
screenshot_ms. Jev, Laya and Cua-S1 read text only: no screenshot is taken for them.OMNIJEV_REPOnames a local clone; theomnijevextra brings torch,torchvision (OmniJev's
requirements.txtomits it, but transformers' Qwen3-VL video processor imports it),transformers>=5 and peft.
OMNIJEV_CHECKPOINT/OMNIJEV_REVISION/OMNIJEV_BASEdefault to the 0.8B v1.1.What
--model omnijevon the browser agents (flights,allrecipes). Not ondecide, the MCP server or the toolagents: they send text only.
repository at commit 14dbec4 and credited, labelled as its recorded-trajectory replays, not s1a runs.
docs/decision-models.md,docs/configuration.md,docs/agents.md,docs/architecture.md,.env.example,CHANGELOG.md.Results
OmniJev-4B v1.1 on a GPU (NVIDIA A100 80 GB, 2026-09-29)
Headless Chromium on the server, its own profile,
s1a run flights --model omnijev, default budgets (240 s).Google likely serves headless server traffic differently; a headed browser or a residential connection should
show the results. Not verified yet.
OmniJev-0.8B v1.1, CPU only
name), "not done" (0.23). On the same page Laya answers DONE (0.55).
s1a run flights --model omnijev): the screenshot works (0.7 to 0.8 s per capture). OmniJev's firsttwo decisions are Jev's: "Round trip", then "One way". The third reopened the ticket-type menu.
offline. The OmniJev authors report 216 to 236 ms per request for the 0.8B on an A800 GPU (their earlier release).
the 0.8B does not complete the task. After its first action it clicked the "Frankfurt am Main → San Francisco"
suggestion card under the form, landed on that route's results, and opened the filters panel instead of fixing the
cities. Stopped after four actions (about 13 minutes).
questions OmniJev picks Jev's operation on 9 of 12 steps and the whole step on 3 of 12, at 178 s a step; with
short questions (goal + ask,
OMNIJEV_PROMPT=short) 8 and 1 of 12, at 57 s. Full stays the default._BROWSER_SIMPLE_QUERY_BUDGET_Sintask_tool.py,_BROWSER_SIMPLE_TASK_DEADLINE_Sinruntime.py), two or three decisions at CPU speed. With a GPUit is no issue. A live run also needs a dedicated Chrome profile that nothing else uses: an earlier run on a shared
window followed another tab and does not count.
the 4B on a GPU can complete the task (1 of 7 runs, same flight as Jev) but is not yet reliable.
Open questions
and game ones need environments s1a does not have yet.
Verification
uv run ruff format --check . && uv run ruff check .: 140 files formatted, all checks passed.uv run ty check: 1 diagnostic, ins1a/decision_models/cua.py, the same onmain.uv lock --check: the lock file matches.uv run pytest -q --ignore=tests/test_browser_policy.py: 16 failed, 420 passed on this branch; 16 failed,390 passed on
mainon this Windows machine, the same failing tests by name, none new.tests/test_browser_policy.pydoes not collect onmainhere either (openjiuwen.harness.schema.decision_policymissing in the installed openjiuwen).
tests/test_decision_models_omnijev.py(the shared contract over a fakeMSO1, the mapping, the prompt option, therenormalisation, the errors,
from_env) andtests/test_browser_screenshot.py.scripts/smoke.sh:smoke: ok.🤖 Generated with Claude Code