Skip to content

feat(canvas): packed cell grid for terminals - #459

Closed
phall1 wants to merge 12 commits into
vercel-labs:mainfrom
phall1:upstream/packed-cell-grid
Closed

phall1 wants to merge 12 commits into
vercel-labs:mainfrom
phall1:upstream/packed-cell-grid

Conversation

@phall1

@phall1 phall1 commented Sep 19, 2026

Copy link
Copy Markdown

Terminals currently emit one command per cell and fall off a render cliff. This adds a packed cell-grid command, an AppKit per-row decoder, bold/italic faces, fingerprinting, host draw reuse, and the composite/seams/cadence follow-ups that make it paint correctly.

phall and others added 12 commits September 19, 2026 19:30
Keep a full terminal viewport, complete terminal fonts, safe underlined-run
costing, secondary gesture ownership, and context-menu policy/invariants.
A terminal is the densest widget the toolkit hosts, and the per-view
frame budgets were sized for desktop chrome. A terminal row costs one
display-list command per contiguous same-background run plus one per
contiguous same-foreground run, so a styled 200-column row runs 30-60
commands where a whole three-pane view runs a few hundred. Against the
old 2,048-command budget (1,792 after the widget chrome reserve) a
60-row viewport had ~30 commands per row to spend, and the painter
degraded from the BOTTOM in silence: a realistically colored 200x60
screen (syntax highlighting, htop meters, a colored build log) painted
34 of its 60 rows and the rest showed bare background, with nothing in
any log to say a budget had eaten it.

Measured on a 200x60 grid at the widget tier (terminal_grid_tests.zig
prints both):

  realistic styling (50 fg runs/row):  34/60 rows -> 60/60
  truecolor (distinct fg+bg per cell):  4/60 rows ->  9/60

Changes:

- canvas_limits: max_canvas_commands_per_view 2048 -> 4096 and
  max_canvas_text_bytes_per_view 32 KiB -> 64 KiB (a 300x100 viewport
  is ~30 KB of presented text before chrome, and a split holds back a
  share for its sibling). Both host retained-command caps pin the new
  number (appkit_host.m, gpu_surface_renderer.cpp), and
  canvas.max_display_list_text_bytes / the new
  canvas.max_display_list_commands stay in lockstep by test.
  Measured cost: ~696 B per command slot across the view's retained
  mirrors, so RuntimeView 3.44 -> 4.83 MiB and the 32-slot Runtime
  110.0 -> 154.5 MiB of fixed-capacity address space.

- Truncation is never silent. terminal_grid.paintReport returns the
  rows painted, the rows handed over, and the store that stopped it;
  paint() keeps its exact signature and every paint records budget
  stops on the builder (canvas.DisplayListDegradation), which the
  runtime turns into one teaching log line on the EDGES of a
  degradation rather than once per frame.

- The widget emit path's frame scratch moves off the stack into the
  per-thread pool the frame planner already uses. At the new command
  budget the display list, the chrome copy store and the diff output
  no longer fit a stack frame under the widget emit recursion — a
  measured segfault inside the button emitter.

What this does NOT fix, deliberately: a viewport with a distinct
foreground AND background per cell merges nothing and wants two
commands per cell — ~24,000 for 200x60, ~60,000 for 300x100. At ~696 B
a slot that is 500 MiB and 1.3 GiB across the view slots, and a trial
raise to 8,192 alone (255 MiB) crashed runtime construction. That
density is not a number in canvas_limits; it needs a packed cell-grid
command the host renderers expand themselves. The budgets now carry
realistic styling and say so out loud when they cannot.

Also: widening a clip no longer leaves stale pixels behind. The render
planner erases push_clip/pop_clip into a per-command clip field, so the
retained packet baseline holds no key for a clip and the refined dirty
rect could not name pixels a growing clip revealed over unchanged
content — the frame kept whatever the host last drew there (the stale
columns a terminal pane showed after a split collapsed back to full
width). The baseline now carries its clip rects and the next frame adds
the difference, in both directions. Frames whose clips did not move,
including tweens that move content under a stationary clip, keep their
region-scoped patches.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A terminal is the one surface whose content scales with AREA rather than
with design. Painting it as display-list shapes cost one background
command per contiguous same-colour run plus one text command per
contiguous same-foreground run, which merges into nothing on a styled
screen: a 200x60 truecolor viewport wanted ~24,000 commands against a
per-view budget of 4,096 and painted nine rows of sixty, the rest bare
background. No budget raise reaches that shape — one command slot costs
~700 B across the view's retained mirrors, so 60,000 commands is 40 MiB
per view before a pixel exists.

`CanvasCommand.cell_grid` (src/primitives/canvas/cell_grid.zig) carries
the screen instead: a cols x rows lattice of 20-byte cells, each with its
own background, cluster offset, foreground, and style, which every
renderer expands itself. Geometry is implied by the index, so nothing
per-cell is stored but content.

Measured (terminal_grid_tests.zig prints both):

  200x60 truecolor, distinct fg+bg per cell:
      before   9/60 rows, ~1,600 commands
      after   60/60 rows,      4 commands, 12,000 cells,  36 text bytes
  300x100 truecolor, distinct fg+bg per cell:
      after  100/100 rows,     4 commands, 30,000 cells, 585 KB, 36 text

The command count is now a constant, not a function of styling. Cluster
bytes are INTERNED, so a screen drawn from a 36-character alphabet costs
36 bytes however many cells it has.

Three bug classes close by construction:

- Reflow is safe. A screen is ONE retained key replaced wholesale, so a
  row that loses a run cannot orphan a per-run command keyed by its old
  start column. That was the stale-column bug — a split collapsing back
  to full width left a duplicated prompt line and a truncated hostname
  in the revealed columns. Verified gone by running the consuming app
  against this commit and driving split/close twice.
- Cell geometry is exact. A cell's position is its index, so a combining
  mark or a wide cluster can no longer advance its neighbours out of
  their columns whatever the face does.
- Every SGR attribute has somewhere to live. `TerminalCell` gains bold,
  italic, strikethrough, overline, six underline styles and an underline
  colour; `TerminalCursor` gains `blinking` and `wide`, and
  `TerminalCursorShape` gains `block_hollow` so an emulator-requested
  hollow block stops colliding with the focus-driven outline. All
  additive: the consuming app builds against this commit unmodified.

The reference CPU renderer is the oracle and is complete: two passes
(every background, then every glyph and decoration, so a neighbour's
background cannot erase an overhanging glyph), each cell's ink through
the same `drawGlyphBox` path a text run takes. Decoration geometry lives
in `CellDecoration` so a future host encoder reads the same source.

The GPU packet layer marks `cell_grid` unsupported, which routes terminal
frames to the CPU pixel path — the reference renderer — on every host.
That is correct everywhere from day one and needs no wire-format change;
it is also slower than a native encoder, and closing that is the next
step. Deliberately not shipping an unverified one: automation screenshots
render through the reference path, so a host encoder cannot be validated
here and a wrong one would diverge silently on the user's glass.

Also fixed, all found by the size increase:

- `paintInto` in the terminal tests returned a `Builder` BY VALUE, which
  left every builder-owned slice aimed at a dead stack frame. Text runs
  were small enough to survive it; a 38-cell grid was not.
- `Builder.initAt` and `CanvasDisplayListScratch.reset` replace
  whole-struct assignment on the widget emit path. Both structs carry a
  frame's inline storage, so `x.* = .{}` built a megabyte-scale stack
  temporary and overflowed the thread.
- `runtimeViewInfo` took a multi-megabyte `RuntimeView` by value.

Memory: RuntimeView 4.83 -> 5.45 MiB (the retained cell array), so the
32-slot Runtime goes 154.5 -> 174.5 MiB of fixed-capacity address space.
The command budget could now come back down to 2048 and give ~45 MiB of
that back, since the terminal is what drove it up; that is a separate
change with its own measurement.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The packed cell grid fixed what a terminal could DRAW and broke how it
PRESENTS. `cell_grid` was unsupported by the GPU packet layer, so one
command dropped the whole view to the CPU pixel path: every frame became
a full-surface upload (~2.8 MB at 1x, 11 MB at 2x) and incremental
dirty-region patching turned off entirely.

Measured on the running app at 1100x640, through the automation
snapshot:

                          before        after
  gpu_present_path        pixels    ->  packet
  present_mode            none      ->  patch
  present_patch_bytes     0         ->  789 (1 upsert on a shell prompt)
  present_fallback        unsupported_command -> none (0 frames)
  gpu_input_latency_ns    19,392,000 -> 5,360,000 (budget 16,666,666)
  budget_exceeded         1         ->  0

Three pieces.

ONE GRID COMMAND PER ROW, not per screen. A retained command is the unit
of CHANGE: a screen-wide grid makes a keystroke re-encode and re-upload
every cell, which is the full-surface cost merely moved from the
rasterizer to the wire. A row is the granularity a terminal actually
changes at. Rows are bounded by `max_rows`, so a 300x100 truecolor
screen is 103 commands where per-run painting wanted ~60,000.

Cluster interning moved from per-SCREEN to per-ROW for the same reason,
and this one was measured the hard way: a blob shared across rows put
every row's fingerprint on every other row's characters, so one new
letter re-encoded the screen — 31 upserts and 8,449 bytes per frame.
Per-row it is 1 upsert and ~400-800.

WIRE FORMAT v6 plus the AppKit decoder. `cell_grid` is command kind 14,
its payload implied by the kind (the flag byte is full). Cells encode as
a delta stream — a tag byte per cell whose low bit means "same style as
the previous cell" — so a plain row costs about a byte a column and only
a genuinely per-cell-styled row pays the full 15. Clusters ride inline
per cell, so a row decodes without the rest of the screen.

The host renderer mirrors the reference renderer deliberately, including
its two-pass order (every background, then every glyph and decoration —
one pass lets a neighbour's background erase an overhanging glyph) and
its exact decoration geometry: all six underline styles, underline
colour, strikethrough, overline, wide cells. Stated rather than hidden:
glyph RASTERIZATION differs (CoreText vs the engine's outline filler),
as it already does for every draw_text command; `bold` and `italic` are
carried but not synthesised, matching the reference renderer rather than
getting ahead of it.

Windows moves to v6 and refuses only packets containing a cell grid (an
unknown kind fails validation), so every non-terminal frame keeps the
retained Direct2D path. Not verified on a device — no Windows here.

Also:

- `TerminalCursor.blinking` is real now instead of documented-only. The
  runtime arms the same looping opacity animation a text caret uses,
  keyed on the new `terminal_grid.cursorCommandId`. A painter has no
  clock; blinking is time, so it was never the painter's to stamp.
- The command budget returns to 2048 (the terminal was the only reason
  it went to 4096 and now costs ~100 commands), handing back ~45 MiB of
  the Runtime's fixed address space. Both host retained caps follow.
- Grid command ids move to `0x60_0000 + row`, out of
  `reserved_id_offset` (0x62_0000) where callers layer their focus ring
  — the previous commit put the grid and the ring on the same key.

Suite: 2970 pass / 14 skip / 2984 total.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`bold` and `italic` reached `CellFlags` and every renderer and then
changed no pixels: `drawCellGrid` built its run with the grid's single
`font_id`. So `\x1b[1m` and `\x1b[3m` were carried honestly and drawn
identically to regular.

A cell grid now carries a FONT FAMILY — the regular face plus bold,
italic, and bold-italic companion ids — and both renderers pick per CELL
off the style flags. One row mixes weights freely.

Where the faces come from: the APP registers them
(`Runtime.registerCanvasFont`, ids >= 64) and names them in
`DesignTokens.typography.mono_bold_font_id` /
`mono_italic_font_id` / `mono_bold_italic_font_id`. The SDK bundles only
GeistMono-Regular and does not presume to ship a consumer's type; an app
that supplies nothing still gets visible weight, because:

SYNTHESIS IS THE FALLBACK, and this commit implements it. Stated plainly
rather than dressed up as real faces: a missing companion is faked — bold
by drawing the glyph a second time offset by max(1, size/14) px, italic
by shearing 0.2 about the baseline. Both renderers read the SAME rules
(`cell_grid.CellSynthesis`) because a bold run that renders bold on the
oracle and regular on the host is worse than no bold at all. A half
family is used for the half it covers: a real bold face with no
bold-italic is sheared rather than double-faked.

Measured through the reference renderer, "mono" at one size:
regular 135 inked pixels, bold 189 (+40%), italic 124 (different shape,
not a second pass). Rendered a four-row sheet (regular / bold / italic /
bold-italic) and looked at it: four visibly distinct weights, columns
aligned identically across all four.

CELL GEOMETRY DOES NOT MOVE, which is the invariant the packed cell
rests on and the thing most likely to break here. A bold face has
different advances; in a lattice that must change nothing, because a
cell's position is its INDEX. Pinned by a test that compares every cell
rect, the cell width, the baseline, and the command's raster extent
across regular/bold/italic rows. Faux bold thickens ink inside the cell
and faux italic shears about the baseline — neither touches the pen.

Italic overhang is real ink and is protected by the existing two-pass
order (all backgrounds, then all glyphs). Pinned by rendering an italic
'H' beside a bright background cell and asserting the lean survives.

Wire format v7 carries the three companion ids. Windows follows the
version and still refuses only packets containing a cell grid.

Also: the terminal cursor blink is checked BEFORE the focus-visible
gate. A text caret blinks once focus is VISIBLE (the keyboard ring); a
terminal cursor blinks whenever the terminal holds focus, however it got
it — a click-focused terminal would otherwise sit steady while the
program that asked for `\x1b[1 q` waited.

Suite: 2975 pass / 14 skip / 2989 total.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
canvasGpuCommandFingerprint decides whether a retained canvas command must be
re-encoded and sent to the GPU host. It hashed kind, bounds, opacity,
stroke_width, cap, id, clip, transform, shape, paint, image, text and effect --
and never read command.cells.

The cell_grid arm of gpu.zig puts ALL terminal content into .cells and leaves
shape, paint, text, image and effect empty; CellGrid.bounds() is purely
geometric. So a cell_grid row's packet fingerprint was INVARIANT under every
possible content change.

Downstream, canvas_frame.zig sets upserts[index]=false on a fingerprint match
and then skips re-encoding, and appkit_host.m calls rasterCacheRemoveKey only
for evicted or upserted keys -- so the host re-blits the raster it already
holds. A terminal emits one cell_grid PER ROW under a stable key, so once a
retained baseline existed a row's glyphs could never change on glass again.
Content only appeared when something else forced a full present.

Reported downstream as: terminal output not appearing when it should, an
occasional prompt landing at random, and typed text never showing. Confirmed
fixed on the real app by the reporter.

The cursor is a separate fill_rect whose bounds and paint ARE hashed, so it
kept moving over stale pixels -- which is why this read as "the text is the
same colour as the background" rather than as a stuck frame.

It also defeated every instrument used to chase it: the CPU reference renderer
re-rasterizes from the content-aware display list, so reference screenshots
always looked correct, and a background colour change alters a fill_rect PAINT,
which IS hashed, so that one did reach the glass.

The fix mirrors render_fingerprints.zig's cellGridFingerprint, which already
hashes the grid identity plus std.mem.sliceAsBytes(cells) over a full screen
every frame, so the cost is known-acceptable.

canvas_frame_patch_tests.zig now covers it: the fingerprint must differ for a
cluster change, a colour change, and a flag change. Seen failing before the
fix (exit 1) and passing after.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Carried from #12, which was squash-merged onto the fork's
`main` (34cc9d5) rather than onto the cockpit/v* lineage. This puts it
on the lineage so the next pin and the sdk-head check see the same tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VuMATX7f1VNpDX2y3hwS3D
Upstream v0.10.0 (vercel-labs#398) made appkit_host.m call two functions the Zig
side exports from src/updater/c_api.zig. The cell-grid host test links
the host with clang alone, so it now needs those two symbols. Refuse
every feed and archive; the harness never reaches the updater.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VuMATX7f1VNpDX2y3hwS3D
The composite pass (NATIVE_SDK_GPU_COMPOSITE=1) refuses any command
kind outside its knownKind list, and a refusal reads to the engine as
'this host cannot present packets', so it falls back to the CPU
reference renderer (.missing_service). For an app whose frames are
entirely cell_grid -- a terminal -- every present is refused, so
composite mode replaces the CoreText rasterizer instead of compositing
it, and the NATIVE_SDK_GPU_SHOT_DIR readback never fires.

The kind was drawable all along: both CG raster paths dispatch through
NativeSdkPacketDrawCommandBody, which has drawn cell_grid since the
kind was introduced, and NativeSdkPacketCommandRasterCacheable already
answers YES for it. Only the gate said otherwise.

Verified with scripts/test-appkit-cell-grid-host.sh (real binary
cell-grid decoder, CoreText raster path, retained raster cache): ok.

Co-authored-by: phall <phall@users.noreply.github.com>
Let capture harnesses set NATIVE_SDK_GPU_SHOT_EVERY so settled views can dump their newest composited frame without waiting for the thirtieth content-changing present. Keep 30 as the default, clamp zero to one, and ignore malformed values. Exercise default, configured, clamped, and malformed policies through the real AppKit host harness.

Co-authored-by: phall <phall@noreply.github.com>
Metal Hybrid C signed bump of the native canvas paint ceilings.

- cells: max_canvas_cells_per_view / max_display_list_cells 32768 → 131072 (4x)
- text: max_canvas_text_bytes_per_view / max_display_list_text_bytes 65536 → 131072 (2x)
- glyphs, commands, paths, and atlas_variants_per_glyph unchanged

Cockpit pin follows in no-phux/phux.
@vercel

vercel Bot commented Sep 19, 2026

Copy link
Copy Markdown

@phall1 is attempting to deploy a commit to the Vercel Labs Team on Vercel.

A member of the Team first needs to authorize it.

@phall1

phall1 commented Sep 20, 2026

Copy link
Copy Markdown
Author

This is a whole terminal pipeline (packed grid, host decoder, seams/cadence, ceilings) in one PR — not a useful review unit.

Closing; I'll split it and reopen as focused patches.

@phall1 phall1 closed this Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant