Skip to content

Spec: design-token completion, change-count estimates, and shared-logic hardening (web + agent) #341

Description

@elkaix

Spec: design-token completion, change-count estimates, and shared-logic hardening (web + agent)

Parent spec. Status: ready-for-agent. Source: synthesized from grilling round 1 (six confirmed decisions), findings R1–R8, and the repository's own design contract (apps/pythinker-web/DESIGN.md).

Problem Statement

The web app's design-token contract exists on paper and in style.css, but several surfaces still bypass it: the embedded terminal hardcodes two color palettes, the dynamic workflow panel renders statically with raw sizes, and scattered components use pixel font sizes that ignore the user's UI font scale. Separately, when an edit tool touches a large file, transcripts and file summaries show +0/-0, which reads as "nothing changed". Several completed enhancement branches are also not yet integrated, so none of this shipped work is visible on a single branch. Finally, a bounded comparison against a reference implementation suggests small robustness improvements in shared agent logic that have never been audited here.

Solution

Integrate the outstanding enhancement stack into one branch; complete the token alignment (terminal palette from CSS tokens, workflow-panel motion family, typography-scale fixes); make changed-line counts degrade to conservative estimates instead of zero; and land proven correctness fixes in shared logic behind regression tests. All work stays within the existing design vocabulary, passes the full gate matrix, and adds zero brand or non-English content.

User Stories

  1. As a web user, I want the embedded terminal to follow the app's light/dark/system theme from the design tokens, so that the terminal never clashes with the rest of the UI after a theme switch.
  2. As a web user, I want the dynamic workflow panel to animate state changes with the shared motion tokens, so that running work feels alive without being noisy.
  3. As a web user with reduced motion enabled, I want the workflow pulse disabled automatically, so that the interface respects my OS setting.
  4. As a web user, I want every text label in workflow cards, badges, keyboard hints, and dialogs to scale with my chosen UI font size, so that the UI adapts to my eyesight.
  5. As a web user, I want attachment count bubbles to remain legible after scaling, so that scaled text does not overflow its pill.
  6. As a user reviewing an agent session, I want large edits to show estimated changed-line counts, so that a big refactor never reads as "no change".
  7. As a user scanning file summaries, I want those estimates to be clearly derived and stable, so that I trust the counts across sessions.
  8. As a user, I want edit cards to fall back to the same estimates when a full diff is too expensive, so that behavior is consistent between card and summary.
  9. As a developer, I want the terminal palette defined once in CSS and consumed by the terminal, so that theme changes need one edit.
  10. As a developer, I want the workflow rail/motion tokens to be feature-scoped and documented, so that other surfaces do not accidentally consume them.
  11. As a developer, I want a diff-stat helper with prefix/suffix trimming, so that a one-line edit in a huge file is counted cheaply and exactly.
  12. As a developer, I want regression tests over the stat helper's trim, fallback, and newline edge cases, so that future changes cannot silently break counts.
  13. As a maintainer, I want the design audit to end with zero raw pixel font sizes outside the documented exemptions, so that the type contract is enforceable.
  14. As a maintainer, I want shared agent logic fixes to each carry a failing-first regression test, so that every fix proves itself.
  15. As a maintainer, I want every examined subsystem recorded with an adopt-or-not decision and ground, so that future audits do not redo the same comparison.
  16. As a user on the light theme and on the dark theme, I want all token-driven changes to look correct in both, plus the mono accent, so that no theme regresses.
  17. As a reviewer, I want the branch to contain the full enhancement stack with generated bundle output regenerated, so that no stale artifacts or duplicate changelog entries land.

Implementation Decisions

  • Bring the enhancement stack in by merging feat/zcode-wave3 (superset: design-system waves + token alignment) into the run branch; resolve the generated web bundle by regeneration (pnpm run build:web + restage), never by hand-editing bundle output; if the merge re-introduces .changeset/tower-ascii-titles-and-token-usage.md, delete it (that changeset shipped in fix(agent-core-v2): require ASCII tower mission titles and record token usage #329 and was consumed by ci: release packages #330).
  • Terminal: xterm theme values read from new --terminal-* CSS custom properties (12 roles, three theme scopes: light root, [data-color-scheme=dark], system-dark media query); computed styles read at render, theme flips re-trigger via the existing isDark reactivity; keep a hardcoded dark fallback only for the background.
  • Workflow panel: feature-scoped --color-workflow-rule/trace/trace-strong + --wf-t-*/--wf-ease/--wf-beat/--tracking-wf-label tokens; card/status transitions on --wf-t-base; running-dot pulse at --wf-beat; label tracking on card metadata; global reduced-motion rule already disables the pulse.
  • Shape/overlay audit and fix: sweep application components for violations of the design contract's radius/shape hierarchy, the approved 2xl exception list, and overlay roles; every real violation found is fixed onto the contract tokens (documented exceptions stay). Audit is bounded to application surfaces, not the catalog view.
  • Typography: convert raw font-size: Npx declarations to the --text-* scale in DynamicWorkflowPanel, FirstRun, Recovery, Kbd, Badge, AttachmentChip, Composer; icon glyphs and catalog demos stay exempt per the design contract; attachment count pill enlarges (12→16px) to host --text-xs.
  • Diff stats: add computeLineChangeStat(before, after) (logical-line split, common prefix/suffix trim, row-array LCS, 400k-cell conservative fallback) in apps/pythinker-web/src/lib/diffLines.ts; computeEditChangeStat in lib/toolDiff.ts; consumers lib/turnFiles.ts + components/chat/tool-calls/EditTool.vue fall back to estimates when a full diff is too expensive; Write and replace_all stay unestimated (cannot know the before-state) and keep their statsIncomplete flag.
  • Shared-logic sweep: three bounded subsystem comparisons in the reference tree — (a) process/exec lifecycle handling, (b) output bounding/truncation/preview, (c) streaming/tool-call edge handling — against apps/pythinker-code, packages/agent-core-v2, packages/pyaos; fix only defects proven against our code with a failing-first regression test; record every subsystem in the private ledger with adopt/skip ground regardless of outcome.
  • Defensive tiers: presentational changes Tier 1 (reversible, no state); diff-stat helper Tier 1 (pure function, input-bounded by cell budget); sweep fixes Tier 0–1 (localized, guarded by regression tests); no Tier 2/3 work is planned; anything discovered above Tier 1 becomes a reported lead, not a silent expansion.
  • No new screenshot or visual-regression harness; validation is code inspection against the token contract plus the gate matrix.
  • Changesets for new user-visible behavior (estimates, token alignment); none for test-only or merge-only changes.
  • Out of scope, recorded as a lead rather than a ticket: the boundary-planner mapping skill (its graph indexes the private reference checkout's own tree — their paths, symbols, and tooling — so adopting it here would mean authoring a Pythinker feature graph from scratch; ledger row SKIPPED_NOT_APPLICABLE, revisit as its own effort), and the user's untracked architecture-governance work in their checkout (untouched).
  • Deliverable: the goal/continue-complete-porting-to-pythinker run branch only. The run does not push and does not open a PR; the user merges.
  • Changesets: user-visible behavior gets a changeset at the minor bump level by default (patch only when the change is a pure fix of just-shipped behavior); test-only and merge-only changes get none.
  • Sweep ledger: every examined subsystem gets a private ledger row marked adopted or not-applicable with the named ground.

Testing Decisions

  • Seams (existing only, highest available): per-package vitest suites — apps/pythinker-web lib tests (diff stats, toolDiff, turnFiles), package suites for any agent-side fix (agent-core-v2, pyaos, and apps/pythinker-code suites when the sweep touches the CLI); vue-tsc typecheck as the component-level seam; full gate matrix as the outer seam: root pnpm run typecheck, pnpm run lint, pnpm run sherif, pnpm run build, pnpm test, web pnpm run typecheck && pnpm run build; apps/vscode typecheck+test separately whenever that app changes.
  • Known base-state exceptions honored: gateway transcriptContract.e2e S3 case fails at base locally (environment-specific; CI green) — not a regression if unchanged; agent-core-v2 rpc-events double-notify case is quarantined flaky (expires 2026-10-08).
  • Regression tests for every logic fix must fail at the pre-fix commit and pass after (verified by running the focused test against the unpatched file first).

Acceptance Criteria (parent-level)

  • A1: After a theme switch, the terminal palette changes with it (values sourced from CSS custom properties; no duplicated hex palettes in the component beyond the single fallback).
  • A2: Workflow cards animate state transitions; the running indicator pulses; with prefers-reduced-motion, nothing pulses.
  • A3: rg 'font-size:\s*[0-9]' apps/pythinker-web/src --glob '*.vue' (excluding DesignSystemView and documented glyph exemptions) returns no application-text hits.
  • A4: A synthetic large edit (700 changed lines each side) yields non-zero estimated counts in file summaries and the edit card; a small edit in a 1800-line file yields exact {added:1, removed:1}.
  • A5: Every landed shared-logic fix has a test that fails without the fix; every examined subsystem has a ledger row.
  • A6: Full gate matrix green — root typecheck, lint, sherif, build, pnpm test; web typecheck + build; when the sweep touches agent-side packages, their suites plus (if apps/pythinker-code or apps/vscode change) the CLI build and the separate VS Code typecheck/test; leak check sections A–E pass and F/G show no growth; dist-web regenerated whenever web source changed in the same commit.
  • A7: After the audit, rg-checkable radius/shape/2xl/overlay violations on application surfaces are zero outside the documented exceptions, and every token-driven change is correct in all four appearances: light, dark, system-dark, and the mono accent.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    ready-for-agentAgent-ready per triage vocabulary

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions