Skip to content

[Feat] OmniJev as a vision decision model for the browser agents - #33

Open
chaimaerachdi wants to merge 3 commits into
ThinkFlowLab:mainfrom
chaimaerachdi:omnijev
Open

chaimaerachdi wants to merge 3 commits into
ThinkFlowLab:mainfrom
chaimaerachdi:omnijev

Conversation

@chaimaerachdi

@chaimaerachdi chaimaerachdi commented Sep 28, 2026 •

Copy link
Copy Markdown

Why

Bring OmniJev (Apache-2.0) into system1-agents, integrate
its use cases and show its demos in the README. OmniJev is a System 1 decision model on Qwen3.5 vision-language
backbones (0.8B, 2B, 4B) that answers the same typed questions as Jev over a screenshot. It is also a local model
trained on web tasks (Mind2Web), where Laya, trained for text classification, fails on Google Flights (#9).

The backend, the screenshot and the README are ready. On a GPU, OmniJev-4B completed the Google Flights task in 1 of 7 runs, with the same answer as Jev (easyJet, €72), and reached Search in 2 of 7, at 1.1 to 1.7 s a decision (see Results). Jev completes it every time. The 0.8B on CPU gets lost.

How

  • s1a/decision_models/omnijev.py: OmniJevModel (supports_images=True, deterministic) over OmniJev's
    MSO1.system_one({"images": [...]}, questions). The text state, goal and rules go in front of each question; a
    browser element row becomes its label and value. A choice's probabilities sum to 1 - abstain in OmniJev, so they
    are renormalised over the offered keys (the key stays OmniJev's own, so a malformed answer still fails validation;
    abstain stays in raw). An observation without an image is MODEL_SERVICE_CONFIG_ERROR.
  • s1a/browser/decision_model.py: when the decision model reads images, every tick's observation carries a viewport
    PNG, taken with page.screenshot() through the probe's run-code executor (tab brought to the front, 15 s cap);
    the tick records screenshot_ms. Jev, Laya and Cua-S1 read text only: no screenshot is taken for them.
  • OmniJev ships as a repository, not a package: OMNIJEV_REPO names a local clone; the omnijev extra brings torch,
    torchvision (OmniJev's requirements.txt omits it, but transformers' Qwen3-VL video processor imports it),
    transformers>=5 and peft. OMNIJEV_CHECKPOINT / OMNIJEV_REVISION / OMNIJEV_BASE default to the 0.8B v1.1.

What

  • --model omnijev on the browser agents (flights, allrecipes). Not on decide, the MCP server or the tool
    agents: they send text only.
  • README: a section with four of OmniJev's own v1.1 demo replays (web, phone, games, robotics), linked from its
    repository at commit 14dbec4 and credited, labelled as its recorded-trajectory replays, not s1a runs.
  • Docs: docs/decision-models.md, docs/configuration.md, docs/agents.md, docs/architecture.md,
    .env.example, CHANGELOG.md.

Results

OmniJev-4B v1.1 on a GPU (NVIDIA A100 80 GB, 2026-09-29)

Headless Chromium on the server, its own profile, s1a run flights --model omnijev, default budgets (240 s).

Run Decisions Median decision Outcome
1 25 1.10 s Filled the whole form (Zurich, London, one way, November 1, 2026, Done), then pressed Enter in Departure instead of clicking Search, and wandered (Frankfurt card, main menu) until the form budget ran out
2 14 1.13 s Every step, Search included: 13 actions against Jev's 12. Google then answered "No results returned. Oops, something went wrong." on the results page, and the run ended BLOCKED after one WAIT
3 4 1.22 s Stopped early: WAIT with no progress
4 7 1.44 s Typed into "Return" (not needed one way); the value model had nothing to type, BLOCKED
5 25 1.10 s Lost after the form, cut by the budget
6 15 1.74 s DONE: Search clicked, answer "easyJet, €72, 9:30 PM to 10:20 PM, November 1", the same flight Jev found
7 25 1.11 s Lost after the form, cut by the budget
  • The screenshot takes 0.65 to 0.8 s a step, about as long as the decision.
  • Run 2's "No results" is Google's page, not a model choice: the results URL and title ("Zürich to London") are right.
    Google likely serves headless server traffic differently; a headed browser or a residential connection should
    show the results. Not verified yet.
  • 7 runs: 1 complete (14%), 2 reached Search (29%). Jev: every run of the day completed. The failures are after the form is filled (wandering, a wrong field), not in filling it.

OmniJev-0.8B v1.1, CPU only

  • One screenshot of the Google Flights home page: operation "type text" (0.93), element "Where from?" (0.38 by
    name), "not done" (0.23). On the same page Laya answers DONE (0.55).
  • Live run (s1a run flights --model omnijev): the screenshot works (0.7 to 0.8 s per capture). OmniJev's first
    two decisions are Jev's: "Round trip", then "One way". The third reopened the ticket-type menu.
  • Speed: 42 to 122 s per decision on CPU (4 questions with long instructions), against 28 s for 4 short questions
    offline. The OmniJev authors report 216 to 236 ms per request for the 0.8B on an A800 GPU (their earlier release).
  • Longer run, dedicated Chrome profile (2026-09-29, openJiuwen's 240 s browser budget raised for the test only):
    the 0.8B does not complete the task. After its first action it clicked the "Frankfurt am Main → San Francisco"
    suggestion card under the form, landed on that route's results, and opened the filters panel instead of fixing the
    cities. Stopped after four actions (about 13 minutes).
  • Offline, on Jev's 12 recorded steps (screenshot + state + Jev's pick per step, 2026-09-29): with the full
    questions OmniJev picks Jev's operation on 9 of 12 steps and the whole step on 3 of 12, at 178 s a step; with
    short questions (goal + ask, OMNIJEV_PROMPT=short) 8 and 1 of 12, at 57 s. Full stays the default.
  • The 240 s budget: openJiuwen's browser subagent stops a task after 240 s (_BROWSER_SIMPLE_QUERY_BUDGET_S in
    task_tool.py, _BROWSER_SIMPLE_TASK_DEADLINE_S in runtime.py), two or three decisions at CPU speed. With a GPU
    it is no issue. A live run also needs a dedicated Chrome profile that nothing else uses: an earlier run on a shared
    window followed another tab and does not count.
  • Verdict: the integration works end to end. The 0.8B on CPU starts the way Jev does but does not finish;
    the 4B on a GPU can complete the task (1 of 7 runs, same flight as Jev) but is not yet reliable.

Open questions

  • Which of OmniJev's demos should become s1a agents? The web one maps onto the browser agents; the phone, robot
    and game ones need environments s1a does not have yet.
  • Google's "No results" on a headless server browser: run headed, or through another network, to see the results page.

Verification

  • uv run ruff format --check . && uv run ruff check .: 140 files formatted, all checks passed.
  • uv run ty check: 1 diagnostic, in s1a/decision_models/cua.py, the same on main.
  • uv lock --check: the lock file matches.
  • uv run pytest -q --ignore=tests/test_browser_policy.py: 16 failed, 420 passed on this branch; 16 failed,
    390 passed on main on this Windows machine, the same failing tests by name, none new.
    tests/test_browser_policy.py does not collect on main here either (openjiuwen.harness.schema.decision_policy
    missing in the installed openjiuwen).
  • New tests: tests/test_decision_models_omnijev.py (the shared contract over a fake MSO1, the mapping, the prompt option, the
    renormalisation, the errors, from_env) and tests/test_browser_screenshot.py.
  • scripts/smoke.sh: smoke: ok.
  • Live run on Google Flights in a dedicated Chrome profile: the integration runs end to end; the 0.8B does not complete the task (see Results).

🤖 Generated with Claude Code

OmniJev (github.com/tinnel123666888/OmniJev, Apache-2.0) is a System 1 decision
model on Qwen3.5 vision-language backbones that answers typed questions over a
screenshot. `--model omnijev` runs it in process on the browser agents:
OmniJevModel maps the choice and noul questions to its shape, renormalises a
choice's probabilities over the offered keys (OmniJev leaves its abstain
probability out of them) and requires an image. The model code comes from a
local clone named by OMNIJEV_REPO; the omnijev extra brings torch,
torchvision, transformers>=5 and peft.

The browser front now sends a viewport PNG with every observation to a
decision model that reads images, through the probe's run-code executor, and
records screenshot_ms per tick. Text-only models see no change.

README: a section with four of OmniJev's own v1.1 demo replays, linked from
its repository at a pinned commit and credited.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chaimaerachdi
chaimaerachdi marked this pull request as ready for review September 29, 2026 08:50
full (the default) sends the text state, the goal and the rules with every
question; short sends the goal and the ask only, the screenshot carrying the
page. Measured offline on Jev's 12 recorded Google Flights steps with the
0.8B on CPU: full picks Jev's whole step 3 of 12 times at 178 s a step,
short 1 of 12 at 57 s. Full stays the default; short is there for speed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed commit 1d01e69. Found one actionable P2 issue, detailed inline.

Validation: 207 tests passed, 4 skipped, plus 21 subtests across the OmniJev adapter, screenshot capture, factory, browser setup, browser policy, and shared decision-model tests. A minimal reproduction confirmed that truncation produces identical operation prompts before and after an unsuccessful Search click. No live model or GPU inference performed.

"""The observation's text state, read in front of every question (OmniJev reads text context in the instructions)."""
state = observation.state
text = state if isinstance(state, str) else json.dumps(state, ensure_ascii=False, separators=(",", ":"))
return text[:OMNIJEV_STATE_CHARS]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve action history when truncating context

The browser observation serializes page text before elements and recent_actions. A page with 6,000 characters of text consumes this entire slice, dropping both fields. I reproduced identical operation prompts before and after an unsuccessful Search click, although the browser rules require switching to PRESS_ENTER after that failure. The screenshot cannot supply this history, so the model loses information needed to select the next operation. Limit page text separately and preserve recent actions within the context budget.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, good catch. Fixed in e2bf4c5: a browser state now puts its recent actions first and caps the page text at 2,000 characters, so a long page can't push the history out. Test added

The browser state was serialised page text first and cut at 6,000
characters, so a page with a long text dropped the elements and the recent
actions: OmniJev saw the same prompt before and after a Search click that
did nothing, although the rules then call for PRESS_ENTER. A browser state
now puts its recent actions first and caps the page text at 2,000
characters on its own.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants