Skip to content

[RFC]: Reuse shared observations across decision questions — community help wanted #85

Description

@hsliuustc0106

Motivation

Agent workloads often ask several questions about the same observation. Independent full prefills can repeatedly compute an expensive image/video or text prefix.

Valen issue #29 documents a concrete example: its Qwen path shares processor encoding but repeats backbone execution per question and per Score level. Fifteen five-level Score questions require 75 explicit backbone calls. This is a source-level call count, not a measured speedup opportunity.

System1-Omni could make shared-observation serving useful across compatible decision models.

Proposed first step

Produce a focused design and benchmark plan before implementing a general cache or scheduler. Select one runnable model and freeze paired request manifests with one observation and increasing numbers of questions/candidates.

Hypothesis: reusing compatible observation computation reduces total multi-question latency while preserving reference decisions and declared numerical tolerances. Memory cost and cache-hit/miss behavior must be measured alongside latency.

Questions for community feedback

  1. Which reuse boundary is both correct and valuable first: preprocessing, vision features, or the language prefix's execution state?
  2. For Qwen's hybrid attention/recurrent layers, what state must be retained or copied to execute independent branches correctly?
  3. What defines compatibility and invalidation: model/adapter revision, media/preprocessing identity, prompt prefix, positions and precision?
  4. How should memory be bounded and state ownership handled across requests?
  5. Would candidate/question batching be a simpler first improvement for the selected workload?

Preserve independent question semantics. In particular, Valen Score branches see only their own level description; concatenating all levels/questions into a new prompt changes the computation.

RFC deliverables

  • One concrete workload and pinned reference, plus a small reproducible request manifest.
  • Current call counts, a latency breakdown where measurable, and the chosen reuse boundary with alternatives and tradeoffs.
  • A correctness plan covering ordering, different observations/candidates, invalidation and declared numerical tolerances.
  • A comparison plan with frozen controls, a run budget, success/stop criteria, memory measurements and raw results.
  • An explicit API/state-ownership proposal, or a statement that the first slice leaves the API unchanged.
  • Small follow-up implementation tasks after feedback; report negative or inconclusive results as well.

Use the existing benchmark protocol and architecture contracts. Valen measurements depend on its reference-worker integration (#84); an already runnable text model can support an initial study.

Community help wanted

We welcome real multi-question agent traces, workload design, CPU request fixtures, profiling, and review of cache/recurrent-state correctness. Please comment with a workload or a boundary you can investigate.

Feedback is open; no implementation deadline or speedup is committed by this RFC.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

RFCRequest for comments on design changeshelp wantedExtra attention is neededperformancePerformance discussion or regression

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions