Skip to content

feat(server): add a reasoning effort picker for Grok - #6887

Closed
torarinvik wants to merge 2 commits into
pingdotgg:mainfrom
torarinvik:feat/grok-reasoning-picker
Closed

feat(server): add a reasoning effort picker for Grok#6887
torarinvik wants to merge 2 commits into
pingdotgg:mainfrom
torarinvik:feat/grok-reasoning-picker

Conversation

@torarinvik

@torarinvik torarinvik commented Aug 15, 2026

Copy link
Copy Markdown

What changed

Grok models now show the same reasoning picker Codex, Claude, and Cursor already have. The levels come from the Grok CLI itself, not from a hardcoded list.

Grok advertises them per model over ACP in ModelInfo._meta.reasoningEfforts, along with the level currently applied in _meta.reasoningEffort. It applies one through session/set_model with _meta: { reasoningEffort: "<level>" }. T3 was ignoring both, so every Grok model shipped with empty capabilities and the composer traits menu showed only the Access section.

File Change
acp/GrokAcpSupport.ts Parse the levels out of the ACP model _meta, read the session's applied effort, and extend applyGrokAcpModelSelection to carry an effort
Layers/GrokProvider.ts Discovered models get a reasoningEffort select descriptor instead of empty capabilities
Layers/GrokAdapter.ts Track the applied effort on the session; apply the selection at session start and per turn
acp/AcpSessionRuntime.ts setSessionModel takes optional request _meta

No client changes were needed — web and mobile both render the picker generically from optionDescriptors, with no per-provider branching.

Before / after

The composer traits menu on Grok 4.6.

Before After

Levels are genuinely per model, so Grok 4.5 gets its own shorter list rather than 4.6's:

Decisions worth reviewing

Which level is marked default. Grok flags more than one level as default (4.6 marks both xhigh and high), so the flags alone are ambiguous. The level actually applied to the model wins, and default is only the fallback.

Effort-only changes resend the model. session/set_model is what carries the effort, so changing effort without changing model re-sends the model Grok already has selected. session/set_mode and model-id suffixes (grok-4.5:low) do not work — I probed both against grok 1.0.4 and only the _meta route applies.

Models that fail discovery get nothing. The grok-build fallback used when ACP discovery fails keeps empty capabilities. The levels are per model, and inventing a list for a model we could not probe risks sending an effort it does not accept.

Changing effort mid-thread is allowed. Only the model is gated behind requiresNewThreadForModelChange, which this does not touch.

Verification

  • checkGrokProviderStatus against the real CLI returns Grok 4.6 with xhigh/high/medium/low (medium default) and Grok 4.5 with high/medium/low (high default).
  • End-to-end in the app: selected a level, sent a turn, and confirmed Grok's own session record (~/.grok/sessions/.../chat_history.jsonl) stamps each assistant message with the effort it ran at — lowmediumhigh across turns, exactly tracking the picker. The persisted thread selection reads {"id":"reasoningEffort","value":"low"} and session/set_model goes out with the _meta field present.
  • New unit tests cover metadata parsing, default resolution, and each set_model application path (model-only, effort-only, both, neither).
  • vp run -r typecheck, vp lint, and the server provider suite (563 tests) all pass.

Note that asking Grok what effort it is running at is not a valid check — it has no tool for that and will answer from ~/.grok/config.toml, which the per-session override deliberately does not rewrite.

Follow-up, not included here

Separately and pre-existing: we create the ACP session on Grok's default model and then call session/set_model, but Grok renders its system prompt at session creation, so it keeps saying "You are Grok 4.6" while routing 4.5. session/new accepts _meta.modelId (verified — the resulting session's stored prompt says "You are Grok 4.5"), which would fix it. Left out to keep this PR to one thing; happy to send it separately if wanted.

Disclosure

This was written by Claude Opus 5 running in Claude Code, at medium reasoning effort with fast mode on, driven and reviewed by me. The protocol findings above (what Grok advertises, what actually applies an effort, what does not) come from probing the installed grok 1.0.4 CLI directly, not from documentation or guesswork. The before/after screenshots are real captures of this branch versus main.


Note

Medium Risk
Changes Grok session model configuration and ACP set_model behavior at start and per turn; logic is covered by tests but incorrect effort handling could affect live Grok runs.

Overview
Grok models can now expose a Reasoning select in the composer traits menu, aligned with other providers. Levels are read from Grok ACP ModelInfo._meta (reasoningEfforts / reasoningEffort) instead of hardcoded empty capabilities.

applyGrokAcpModelSelection now takes current and requested reasoning effort, returns { modelId, reasoningEffort }, and calls session/set_model when either the model or effort changes. Effort is sent as _meta.reasoningEffort; effort-only updates re-send the current model id. On model switches without an explicit effort, prior effort is not carried over so unsupported levels are not applied to the target model.

GrokAdapter tracks currentReasoningEffort on the session and applies user selections at session start and on each turn via provider option reasoningEffort. AcpSessionRuntime.setSessionModel accepts optional request _meta for that payload. Discovered Grok models get capabilities from buildGrokModelCapabilities. Unit tests cover metadata parsing, default resolution, and set_model paths.

Reviewed by Cursor Bugbot for commit 11f2f17. Bugbot is set up for automated code reviews on this repo. Configure here.

Note

Add reasoning effort picker for Grok models

  • Adds a Reasoning select option to Grok model capability descriptors for models that advertise reasoning levels, sourced from model _meta via grokReasoningEffortLevelsFromModelMeta.
  • Extends applyGrokAcpModelSelection to accept and return reasoningEffort; triggers session/set_model for effort-only changes and attaches reasoningEffort in request _meta.
  • Persists currentReasoningEffort in GrokSessionContext so each turn can update effort independently of model ID changes.
  • Extends AcpSessionRuntime.setSessionModel to forward an optional meta object as _meta in the request payload.
  • Behavioral Change: switching models no longer carries over the previous reasoning effort unless explicitly requested.

Macroscope summarized 11f2f17.

Grok advertises per-model reasoning levels over ACP in
`ModelInfo._meta.reasoningEfforts`, and applies one through
`session/set_model` with `_meta.reasoningEffort`. T3 ignored both, so
Grok models shipped with empty capabilities and no picker, while Codex,
Claude, and Cursor all have one.

Read the levels off the discovered models so the existing composer
traits menu renders them, and carry the selection through session start
and each turn. Levels are per model (Grok 4.6 offers xhigh/high/medium/
low, Grok 4.5 offers high/medium/low), so nothing is hardcoded.

Grok flags more than one level as `default`, so the level currently
applied to the model wins and the `default` flags are a fallback.
An effort-only change resends the model Grok already has selected,
since `session/set_model` is what carries the effort.

Models that fail ACP discovery keep empty capabilities rather than
being given an invented level list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 15, 2026

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 0b2c16fc-963c-4794-9b4f-934d0ae1ea6e

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added vouch:unvouched PR author is not yet trusted in the VOUCHED list. size:L 100-499 changed lines (additions + deletions). labels Aug 15, 2026
Comment thread apps/server/src/provider/acp/GrokAcpSupport.ts Outdated

@cursor cursor Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Want fixes drafted automatically? Bugbot Autofix can create code changes for findings. A team admin can enable Autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit c7b7a54. Configure here.

Comment thread apps/server/src/provider/acp/GrokAcpSupport.ts
@macroscopeapp

macroscopeapp Bot commented Aug 15, 2026

Copy link
Copy Markdown
Contributor

Approvability

Verdict: Needs human review

This PR introduces a new user-facing feature (reasoning effort picker for Grok) that modifies core model selection interfaces and propagates new configuration through multiple code paths. New feature enablement with behavioral changes warrants human review.

You can customize Macroscope's approvability policy. Learn more.

Reasoning levels differ per Grok model, so carrying the current effort
onto a different model can send a level the target never advertised.
Probing `grok 1.0.4`, `session/set_model` accepts `grok-4.5` with
`xhigh` without error, applies it, and leaves the session config with no
effort selected at all.

Fall back to the target model's own default when the model changes and
no effort was explicitly requested. An explicit request is still always
sent, including when it matches the effort already applied — deciding
this on `shouldSwitchReasoningEffort` alone would silently drop a
requested level that happened to equal the previous model's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@t3dotgg

t3dotgg commented Aug 23, 2026

Copy link
Copy Markdown
Member

Note

🤖 GPT-5.6 Sol responding on behalf of Theo

Closing this PR after an automated pass over open pull requests. Duplicates trusted Grok reasoning controls in #6386.

@t3dotgg t3dotgg closed this Aug 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L 100-499 changed lines (additions + deletions). vouch:unvouched PR author is not yet trusted in the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants