Skip to content

aheui: integrate the generated JIT runtime - #10

Open
youknowone wants to merge 312 commits into
mainfrom
jit-selected-size-getfield
Open

youknowone wants to merge 312 commits into
mainfrom
jit-selected-size-getfield

Conversation

@youknowone

@youknowone youknowone commented Aug 13, 2026 •

Copy link
Copy Markdown
Owner

개요

Aheui의 실행 상태와 opcode dispatch를 aheuinterpreter로 통합합니다.
일반 실행과 JIT는 같은 인터프리터 소스를 사용하며, aheui-jit는 생성물
로더와 시작 시 연결을 담당합니다. raw/tagged 이중모드, banded stack,
moving node GC와 bigint root 전달은 유지합니다.

  • 중복된 일반 인터프리터를 제거하고 canonical mainloop를 공유합니다.
  • aheui-jit/src/lib.rs는 생성물 연결용21줄입니다.
  • Charon LLBC는 runtime의 명시적14개 helper root와 그 callee를 번역합니다.
    생성된 JitCode는134개에서15개로 줄었습니다.
  • Cargo의 majit pin과 CI checkout SHA를 함께 갱신합니다.
  • WASI 실행기의 cold-cache 파일 생성 한도를 capture 한도와 분리합니다.
    stdout/stderr capture 제한8MiB와 메모리 안전장치는 유지합니다.

생성 경계 — 완전한 LLBC 전환은 아님

실행 portal은 aheuinterpreter의 #[jit_interp]가 생성하고 runtime helper는
LLBC pipeline이 생성합니다. 사용되지 않는 별도 mainloop를 closure root로
번역하지 않습니다. 새 명시적 helper-root API는 매크로 포털을 위한 구조적
예외이며, RPython의 일반 configured-driver 경로와 완전히 같다고 주장하지
않습니다. 실제 portal을 그 경로로 생성하는 작업이 남아 있습니다.

검증 기록

이동 작업 당시 ARM Linux의 메모리 제한 VM에서 다음을 통과했습니다.
아래 수치는 rebase 이전 기록이며 새 CI 결과와 구분합니다.

  • Aheui 단위·통합 테스트25개, 별도 macro/helper API 회귀 테스트3개.
  • default62 + jitstress62 검사 및 전후 바이너리124개 카운터/출력 비교.
  • 실제 --no-jit62개 출력/종료 코드 비교.
  • WASI cold-cache 실제 JIT 실행: loop1, abort0, panic0, 제어 실행과 출력 일치.
  • browser WASM과 일반 실행의 num-bigint/rbigint 빌드 검사.
  • 대표6개 프로그램의 교대30회 프로세스 시간: logo 동일, 나머지 약6–26%
    개선. 시작 비용 포함 측정이며 JIT 루프 자체 처리량 측정은 아닙니다.

.jitstats나 .opcensus 기준값은 이 이동/게시 작업에서 변경하지 않았습니다.

게시 직전 rebase된 majit65ee1d86939 기준으로 Aheui 단위·통합25개와
Python 검사11개를 다시 통과했습니다. Cargo/CI pin 일치, formatting과
diff 검사도 통과했습니다. 전체 Pyre 명령 두 개는 실행을 시작했으나,
별도 extract-llbc.py --check가 기존 root LLBC의 source/closure 불일치를
확인하여 빌드 중 중단했습니다. 재추출 전 실행 검증은 하지 않았습니다.

게시된 소스의 release 빌드, default62 + jitstress62 기준선 검사와
기존 바이너리 대비124개 카운터/출력 비교도 다시 통과했습니다.
실제 --no-jit 실행62개의 출력/종료 코드 비교도 통과했습니다.

머지 전 미해결 사항

  • 이전 HEAD bf36b18의 Actions run34290540798는 opcensus.py check에서
    10개 프로그램의 연산 수/출력·종료 코드 검사에 실패했습니다. jitstats
    check/sweep과 별개이며 새 HEAD에서도 반드시 검증해야 합니다.
  • census가 전체 MAJIT_LOG를 모으므로8MiB capture 제한에 걸릴 수 있습니다.
    실행기는 이 경우125를 반환하지만 census는 그 진단을 버립니다.
    제한을 올리지 않고 필요한 이벤트만 수집하고, 실제 backend 연산 수 증가는
    아키텍처별 trace를 비교해 원인을 규명해야 합니다.
  • 전체 Pyre gate는 이전에 LLBC 의미 그래프 생성 중 메모리 제한 내 할당에
    실패했습니다. 인터프리터 이동으로 해결되었다고 주장하지 않습니다.
  • Aheui GC 어댑터의 기존 oversized alloc_zeroed 경로는 node collector의
    회수 대상이 아닙니다. 실제 descriptor 도달 여부를 확인한 뒤 지원 범위를
    명시하거나 회수 owner를 제공해야 합니다. 현재 corpus 누수의 증거는 아닙니다.
  • helper-only 빌드에 남은 빈 virtualizable 설정과 이전 portal의 call-effect
    설정은 실제 사용 여부를 확인해 정리해야 합니다.

고정 C/Rust runtime 본문은 이미 파일 템플릿으로 분리되어 있습니다.
추가 생성기 추상화는 위 실패와 소유권 문제보다 후순위입니다.

Call backend set_new_via_gc(true) after set_gc_allocator so compiled
New() ops route through the nursery-backed GcAllocator instead of
libc::malloc, sharing the interpreter's node pool.

Assisted-by: Claude
Switch the node nursery from free-on-JIT-path to the rpaheui-orthodox
leak+collect model:

- jit_free_node => concrete_only_void: the free runs on the concrete
  path only; compiled traces omit it, so nodes leak on the JIT path.
- Nursery gains `chunks: Vec<*mut Node>` (grow() records each base) and
  `collect(keep)`: a non-moving mark-sweep that marks every node
  reachable from the 28 `pools[*].head` chains, `port.head`, and the
  in-flight `keep` node, then sweeps unmarked chunk slots onto the free
  list. It runs from alloc() when the free list is empty and the bump
  region is exhausted, before growing.
- set_gc_roots(storage): mainloop registers the storage as the root set
  after refreshing pools.
- jit_alloc_node stays `residual_ref`, which forces pending head-stores
  to memory before the call, so the roots are current at collect time.
- NurseryGcAllocator: the 256MB cap now applies only to the oversized
  `alloc_zeroed` fallback; the node path self-bounds via the chunk cap
  and the collector.

Verification (dynasm): all 6 samples byte-identical to naive (logo
996310). collect fires ~1000x on logo, reclaiming ~262k nodes per
cycle; the nursery stays at 1 chunk. With the collector disabled, logo
leaks to the 64-chunk cap and exits 99 at 40960 bytes, so the collector
is load-bearing.

Caveats: on dynasm logo compiles one loop and enters compiled code, but
every entry deopts on a guard failure (a guard-saturated trace), so most
of the run executes on the deopt/interpreter path; the compiled
jit_alloc_node+collect path is therefore exercised only briefly.
cranelift logo hits a pre-existing recompile storm on the same
guard-saturated loop (present without this change), so the compiled path
is not yet validated end-to-end.

Assisted-by: Claude
Declared `int(i32)`, the `stacksize` red carried an `as i64` sign-extend
(IntLshift/IntRshift by 32) that the macro inserts on every per-op
`stackok` comparison and delta write. `int(i64)` matches the IR's native
Int width, so those boundary casts become identity: the logo loop trace
drops 5684 shift ops (pre-opt 50673 -> 44989), output byte-identical
(996310 bytes / exit 42), all samples unchanged.

Assisted-by: Claude
Restore stackok to the greens list (rpaheui aheui.py:29 parity), removing
the green->red demotion. With stackok green the per-op
jit_effective_stacksize_delta folds to a constant: on logo.aheui the 5080
residual delta-calls drop to 0, the post-opt trace shrinks 42733 -> 26144
ops (-38.8%), and logo runs ~2.2x faster. Output is byte-identical (logo
996310/exit42 and the hello/hello-world/fibonacci/factorial/99bottles
samples). The demotion's green-key-explosion concern does not reproduce on
the current base: logo compiles a single bounded loop.

Assisted-by: Claude
Expose the nursery bump-pointer slot addresses via
storage::nursery_bump_addrs() and wire them into the GcAllocator's
nursery_free_addr / nursery_top_addr (previously stubbed to 0). Switch the
jit_alloc_node call policy from residual_ref to nursery_alloc_ref, tagging the
node allocator's CallR descr with PyreHelperKind::NurseryAlloc so the dynasm
backend emits an inline headerless nursery bump (bump free by the 16-byte node
size, compare against end, cond-call jit_alloc_node only on exhaustion) instead
of a full residual call.

Assisted-by: Claude
AHEUI_COPY_COLLECT=1 switches Nursery::collect from the non-moving
mark-sweep to a semispace copying collection: live chains are evacuated
Cheney-style into a to-space chunk (forwarding pointer in the old node's
next field + side bitmap), the root fields (pool heads, queue tail, port
head, the allocating call's next argument) are forwarded in place, and
the from-space chunks are recycled as the next to-space. Default off;
the sweep path is unchanged.

collect now takes the allocation's next pointer by &mut so the caller
uses the forwarded address. Queue::push re-reads self.tail after
alloc_node since a copying collect can move the sentinel.

NODE_ROOT_WALK_HOOK is an inert hook for a later slice to register a
jitframe root walker; without it the copying mode is only sound for the
interpreter path (set_gc_roots is only called under the jit today, so
the gate stays off by default).

Adds a unit test that drives real copying collections through Storage
(stack LIFO + queue FIFO integrity across forced collects); the test
fails against a stale-scan-limit Cheney loop and passes with the
current one.

Assisted-by: Claude
Register the libc-jitframe tracer and install NODE_ROOT_WALK_HOOK at
mainloop setup. The walker visits every GcRef root slot the majit-gc
minor collection enumerates — shadow-stack refs, jitframe roots plus
jf_gcmap-marked slots of libc jitframes, blackhole ref register banks,
resume ref slices, and extra/trace-constant roots — so a copying node
collection under AHEUI_COPY_COLLECT=1 forwards jit-held node pointers
in place. Slots holding non-node refs are ignored by the collector's
chunk range check.

Assisted-by: Claude
Remove the AHEUI_COPY_COLLECT gate and the non-moving mark-sweep
(collect_sweep/mark_chain). Nursery::collect now always evacuates live
chains Cheney-style (trace_and_drag_out_of_nursery shape) and resets
the bump region, so the inline nursery-bump fast path stays hot after
the first chunk exhaustion instead of routing every allocation through
the slow-path free-list pop.

Queue::push passes the old tail through allocation as an explicit keep
root (via the new sentinel's next field) instead of re-reading
self.tail after alloc_node: the collector updates the tail root through
the GC_ROOTS pointer while push holds &mut self, and release codegen is
free to reuse the stale pre-collection read under noalias — observed as
a ~46K-node queue-chain truncation in the multi-chunk survivor test.

logo (aarch64, interleaved x9): jit min 2.10s vs naive min 2.66s
(0.79x); rpaheui snippet corpus byte-identical between jit and naive
except pi.jinseo (pre-existing, aheui-rust#6) and capacity-exit
artifacts.

Assisted-by: Claude
The stack OP_CMP arm called jit_val_ge, which lowered to a residual
CALL_PURE in the trace (~70ms self-time on logo). Replace it with an
operator-only fast path: when both operands carry the small tag
(bit 0 set), compare the tagged words directly — the tag is
order-preserving — and tag the 0/1 result; fall back to jit_val_ge
otherwise. The #[jit_interp] macro lowers this to IntAnd + GuardTrue +
IntGe, with the residual jit_val_ge as the guard-failure path.

Add raw tagged-word BitAnd and PartialOrd impls to Val so the
operator-only source compiles on the concrete path; both operate on
the packed i64 and are documented as JIT-fast-path-only.

logo byte-identical (996310/42), jit==naive across the corpus.

Assisted-by: Claude
The stack OP_ADD/OP_SUB arms called val_add/val_sub, which lowered to
residual CALL_PURE ops in the trace. Replace them with operator-only
fast paths: when both operands carry the small tag, untag by an
arithmetic shift, add/subtract the untagged words (their sum/difference
cannot overflow i64 since each is bounded by 2^62), verify the result
round-trips through the small tag with ((r << 1) >> 1) == r, and retag;
fall back to val_add/val_sub otherwise. The macro lowers this to
IntRshift/IntAdd/IntAnd/IntEq + GuardTrue with the residual call as the
guard-failure path.

Add Val Shr (raw tagged-word arithmetic shift = untag) and a full-width
val_retag_small helper ((v<<1)|1, no i32 truncation, unlike jit_tag_val).

logo byte-identical (996310/42), jit==naive across the corpus.

Assisted-by: Claude
Register jit_retag_small in native_tag_small so the OP_ADD/OP_SUB/OP_CMP
smallint fast-path retag lowers to native (x<<1)|1 instead of a CallI.
Route OP_CMP's tag through jit_retag_small (identical to jit_tag_val on
0/1) so the hot compare tag is native as well.

Assisted-by: Claude
Untag both operands, sign-extend to 32 bits so the i64 product cannot
overflow, multiply, and re-tag via native_tag_small; fall back to val_mul
when an operand is a bigint or exceeds the ±2^31 fast range. This nets a
win now that the retag lowers to native IR — previously the added retag
CallI made this path slower than the val_mul call.

Assisted-by: Claude
…ge_push

Declare the pools element type in pool_arrays (`storage_ref => jit_sel_get_ref ->
Stack`) and rewrite the OP_MOV push as a green branch on the operand: VAL_QUEUE /
VAL_PORT keep the jit_storage_push residual; any other stack pushes inline through
jit_sel_get_ref(storage_ref, target) (getarrayitem + node alloc + setfield
head/size), mirroring the OP_PUSH selected_ref inline.

Assisted-by: Claude
OP_PUSH, OP_DUP, and the OP_MOV push lower node allocation to NodeJit struct
literals via `struct_allocs = { NodeJit => jit_alloc_node }` and
`headerless_structs = { NodeJit }`: the trace emits a headerless New plus
value@0/next@8 field stores that the optimizer folds when the node does not
escape, while concrete execution calls jit_alloc_node. Stack.head is left
deferrable at guards, so a virtual head surviving to a side-exit is
reconstructed on the deopt path rather than force-materialized.

A debug-only canary in refresh_state_from_storage walks every ordinary stack
and asserts its node-chain length equals its size field.

Assisted-by: Claude
Val::from_big allocated big integers with Box::into_raw and nothing freed
them, so every value above 2^62 leaked its Box<BigInt>.

Stand up majit-gc's MiniMarkGC in the JIT build and register big integers as
a non-moving old-gen GC type with a drop_in_place destructor and an
external-size accounting for the malachite limbs. Route Val::from_big through
an allocation hook that allocates into the old generation; a big-integer root
walk enumerates every Node.value across the interpreter storage chains, the
JIT node roots, port.last_push, and transient in-flight values.
maybe_collect_bigints runs a non-moving old-gen sweep once 8 MiB of big
integers has accumulated, wired into the interpreter push / _put_value
safepoints.

aheui-runtime stays majit-gc-free: allocation, collection, and the root walk
are reached through function-pointer hooks installed by the JIT build; the
naive feature keeps no-op stubs.

Under --no-jit dead big integers are reclaimed (fibonacci.tokigun peak RSS
grows sub-linearly instead of ~0.5 MiB per MiB of output). Under --jit
collection does not yet fire: node virtualization lets compiled code allocate
big integers without reaching a GC safepoint, and invoking the sweep from the
allocation site hangs because the mid-residual shadow stack is not walkable.
The JIT path awaits a loop-back-edge safepoint poll.

Assisted-by: Claude
Assisted-by: Claude

Assisted-by: Codex
Trigger maybe_collect_bigints() from alloc_bigint_oldgen so bignum
collection fires in compiled node-virtualized code, where the interpreter
push/pop hooks and the loop merge point are bypassed once a trace is
compiled and collection would otherwise never run under --jit. The
freshly-built AheuiBigInt value is still Rust-owned (not yet a GC object),
so it cannot be swept, and the just-consumed operands are reclaimable.

Walk bignum roots from Storage only in walk_bigint_root_values; the untyped
JIT node shadow-stack hook could follow .next from non-node slots and walk
arbitrary memory.

Assisted-by: Claude
majit #573 replaced PipelineConfig.portal: Option<PortalSpec> with
jit_drivers: Vec<JitDriverSpec>, whose portal field is a CallPath rather
than a name string, and dropped ProgramPipelineResult.opcode_dispatch.
Point the mainloop driver at CallPath::from_segments(["mainloop"]) and
report jitcodes.len() in the build diagnostic.

Assisted-by: Claude
Rewrite linkedlist_jit::stack_pop as a #[jit_inline] helper taking
`stack: ref(Stack)` with Stack::head=>Node / Node::next=>Node field
typing, so the OP_POP Stack branch traces as native getfield/setfield
plus nursery free instead of a residual int call. Register it inline_int
and replace the hand-inlined OP_POP arm with lj::stack_pop(selected_ref).

Feature-gate the majit deps that #[jit_inline] pulls in behind a new
aheui-runtime `jit` feature (majit-ir/majit-macros/majit-metainterp made
optional; linkedlist_jit module cfg-gated), enabled by aheui-jit. Keeps
the wasm32 and interpreter-only builds free of the JIT crates.

Assisted-by: Claude
Rewrite stack_swap's body to explicit head/next/value field operations, annotate
it with #[jit_inline] typed ref-param and ref-field metadata, register it as
inline_void, and call it from the OP_SWAP arm instead of hand-inlining the node
value exchange in the mainloop.

Assisted-by: Claude
Rewrite stack_add/stack_sub/stack_mul/stack_cmp with explicit head/next/value
field operations and their native smallint fast paths, annotate each with
#[jit_inline] ref-param/ref-field metadata, native_tag_small = { val_retag_small },
and the elidable val_*/free_node_jit call policies, register them as inline_void,
and call them from the OP_ADD/OP_SUB/OP_MUL/OP_CMP arms instead of hand-inlining
the pop/arith/free sequence in the mainloop.

Assisted-by: Claude
majit_translate::JitDriverSpec gained an autoreds field. The aheui driver
declares explicit reds (stacksize/storage/selected), so set autoreds:
false.

Assisted-by: Claude
Rewrite stack_push/stack_dup with explicit head/value/next field
operations and a super::linkedlist::Node struct-literal init, annotate
each with #[jit_inline] ref-param/ref-field metadata plus
struct_allocs = { Node => alloc_node_jit } and headerless_structs =
{ Node }, register them as inline_void, and call them from the
OP_PUSH/OP_DUP arms instead of hand-inlining the alloc/link sequence in
the mainloop.

Assisted-by: Claude
Rewrite queue_pop/queue_swap with explicit head/value/next field
operations mirroring the Stack helpers (Queue.pop/swap are the default
LinkedList impls; head/size share Stack's offsets, tail is untouched),
annotate each with #[jit_inline] ref(Queue) metadata, and register
queue_pop as inline_int and queue_swap as inline_void instead of residual.
The OP_POP/OP_POPNUM/OP_POPCHAR/OP_SWAP arms already call these helpers in
the is_queue branch.

Add examples/queue_hot_gate: a pure-queue program (single SEL 21; a POPNUM
loop over K pushed values) run through mainloop at a low threshold (the
loop compiles) and a huge threshold (interpreted). Byte-identical stdout
verifies the compiled inlined queue_pop matches interpretation. Pure-queue
plus a mid-program HALT keep the trace clear of two pre-existing
low-threshold crashes (mixed-storage FieldDescr, program-boundary).

Assisted-by: Claude
The inline_int/inline_void storage helpers (stack_push/pop/add/sub/mul/dup/
swap/cmp and queue_pop/swap) splice their self.size setfield into the
trace, so the heapcache already invalidates the len(selected) getfield on
that store; they no longer need a residual_writes size-write declaration.
Remove them, leaving only the still-residual mutators, and refresh the
Phase D-1 registration and residual_writes comments that described every
storage op as residual.

Assisted-by: Claude
…aborting

OP_POP called lj::stack_pop / lj::queue_pop (registered inline_int) in
discarded statement position. inline_int lowers only in value position, so
the arm classified Unsupported and emitted a BC_ABORT stub, which kept the
loop from closing and made the single-walker re-execute the traced iteration
on abort (duplicated output at low threshold). Bind the popped value
(discarded), mirroring OP_POPNUM's pop shape, so the arm is traceable and the
loop compiles.

Assisted-by: Claude
queue_dup was the only queue mutator still registered `residual_void`,
while stack_dup and queue_swap use inline_void. A residual queue_dup
does not invalidate the optimizer's cached head/size for the queue, so
a pop after the DUP+POP idiom (pc91/92 in the output loop) read the
pre-dup head/size and set size to size-1 instead of size, dropping an
extra queue element each iteration.

Give queue_dup a #[jit_inline] body mirroring stack_dup (head-only, the
queue's tail is untouched) so the head/size mutation is emitted as
tracked setfields and the following pop reads the post-dup values.

99dan no longer hangs at low MAJIT_THRESHOLD; logo/99bottles/99dan are
byte-identical to the interpreter at THRESHOLD 5/20/40/default.

Assisted-by: Claude
…terialization

Virtualized headerless NodeJit allocations that escape into a stack head field are recorded as TAGVIRTUAL pendingfields on merge-point guards. On guard-fail deopt the blackhole resume path materializes them via the registered BlackholeAllocator bh_new/bh_setfield_gc_* methods. aheui registered none, so deopt fell back to NullAllocator (bh_new -> 0, setfield no-op), leaving the pushed node unlinked while its size increment was committed; the resulting size=chain+1 desync derefs a NULL head at low MAJIT_THRESHOLD.

Assisted-by: Claude
…Port

Storage::pools type-puns its *mut Stack slots over Stack, Queue, and Port,
which share head@0/size@8 but mint distinct JIT field descriptors. A
residual head-mutating call (queue/stack div/mod, pop-based ops) only
invalidated the Stack-typed getfield, so a trace specialized on the Queue
or Port layout kept a stale cached head after the call and reused a
popped/freed node, deref'ing head.next -> SIGSEGV (fibonacci.puzzlet
runs the queue layout).

Declare Queue::head and Port::head as ref fields and add residual_writes
aliases for head and size across all three layouts so force_from_effectinfo
invalidates whichever descriptor the specialized trace cached.

Assisted-by: Claude
OP_BRZ compares the popped word to the mode's zero directly.
Depth analysis uses an integer lattice and an IndexMap worklist.
opcensus records MAJIT_LOG_OPS events and rejects incomplete
measurements.

Assisted-by: Claude
ENABLE_OPTS drops unroll so a 27k-op recording is optimized
once. floor_div_i64 / floor_mod_i64 carry int.py_div / int.py_mod
so rewrite can fold x//2**k. The jitstress opcensus baselines
are that compiled shape. Pin majit to d0c2819206f.

Assisted-by: Claude
@youknowone

Copy link
Copy Markdown
Owner Author

Pushed d916090 on top of the generated-JIT integration.

  • ENABLE_OPTS is again ALL_OPTS minus unroll, so a ~27k-op logo recording is optimized once (depends on pyre #1782: unread simple-loop JUMP slots stay general).
  • floor_div_i64 / floor_mod_i64 carry int.py_div / int.py_mod, so rewrite.optimize_call_int_py_div folds x // 2**k.
  • jitstress opcensus baselines are that compiled shape (vable static-field GETFIELDs + the fold). opcensus.py check is 10/10; jitstats.py check --no-log is 62+62.
  • logo@50 on the recording Mac: min 107.7 ms / p50 109 ms.
  • majit pin and the snippet-matrix checkout are d0c2819206f.

— commented by Claude

bands and cap are no-arg elidable constants for the run. Promoting
them at every merge only re-guards values that cannot change.
In-loop MAJIT_HANDOFF / SPDIAG bodies move behind elidable flags
and residual helpers.

Assisted-by: Claude
Dropping stackok stops jit_effective_stacksize_delta from folding
and leaves logo at ~21k residual ops. Dropping is_queue does not
change aheui.aheui's ~11k guard-failure count.

Assisted-by: Claude
RPython's 200 leaves tens of thousands of cold guards without a
bridge. 20 still gives logo one bridge and cuts aheui.aheui
99bottles/quine.40col roughly in half. jitstats/opcensus match
that compile shape.

Assisted-by: Claude
The value::floor_div_i64 wrappers are oopspec CALL_PURE, so they
never appear in the symbolic fnaddr table. Pin majit to f6a44feddea.

Assisted-by: Claude
The helper pipeline now names the oopspec wrappers, not the
ahsembler consts. An unbound wrapper hash aborted logo traces.

Assisted-by: Claude
enable_opts defaults to "all"; recordings longer than 6000 ops take
the simple-loop path.

Assisted-by: Claude Fable 5.1
examples/allocsites.rs runs a program through the JIT under a counting
System allocator and reports total allocations and sampled call sites,
by stack and by innermost majit/aheui frame. AHEUI_ALLOCSITES* set the
capture window.

Assisted-by: Grok 4.6
Assisted-by: Claude Fable 5.1
The majit attribute macros expand to ::linkme::distributed_slice.

Assisted-by: Claude Fable 5.1
The majit attribute macros expand to ::ctor::ctor on wasm targets.

Assisted-by: Claude Fable 5.1
Stack and Queue carry a value buffer, its capacity and (for the queue) a
front index after the shared ListBase prefix; the jit storage helpers
index that buffer instead of walking Node chains. AheuiState keeps a
virtualizable value ring sized by jit_cap() and a band count. The JIT
mirror structs, the declared field keys and the struct-shape test follow
the new layout. main prints the jitprof summary under MAJIT_STATS.

Assisted-by: Grok 4.6
Assisted-by: Claude Fable 5.1
The baselines predate the unroll enable and the ring-buffer storage:
peeled loops emit a preamble and a body (Label 1 -> 2), and the storage
ops are inlined as indexed loads and stores.

Assisted-by: Claude Fable 5.1
kind (stack / queue / port / band) is written at init and in OP_SEL from
the pc-derived operand, and replaces is_queue in the merge-point greens.
The dispatch arms test kind instead of comparing state.selected.

aheui.aheui + 99bottles: recorded guards 16301 -> 12294, recorded ops
34538 -> 30140; loops, bridges and guard failures unchanged.

Assisted-by: Grok 4.6
Assisted-by: Claude Fable 5.1
Each trace guards kind once with GuardValue where it compared
state.selected before.

Assisted-by: Claude Fable 5.1
band_base (selected * cap) is a state field written in OP_SEL beside
kind. A band arm promotes state.sp & cap_mask and band_base once and
derives its slot indices from the promoted values.

aheui.aheui + 99bottles: recorded ops 30140 -> 24640, guards
12294 -> 12160; loops, bridges and guard failures unchanged.

Assisted-by: Grok 4.6
Assisted-by: Claude Fable 5.1
Eight programs emit fewer ops; huntcook emits one more IntAdd and
quine.puzzlet.40col one more GuardValue.

Assisted-by: Claude Fable 5.1
cap is 0 before the buffer exists and STACK_CAP after, and size is 0
while cap is 0, so size == cap is the first push (allocate the buffer)
or a full buffer (overflow). stack_push and stack_dup make that one
comparison in place of cap == 0 and size == 65536.

aheui.aheui + 99bottles: recorded guards 12160 -> 11513.

Assisted-by: Grok 4.6
Assisted-by: Claude Opus 5
2e65-print emits one more IntEq; pi.jinseo and quine.puzzlet.40col
emit fewer ops.

Assisted-by: Claude Opus 5
Band pool `i` is now `kind = KIND_BAND + i`, and band tests use
`kind >= KIND_BAND`. Band arms compute their ring base as
`(kind - KIND_BAND) * cap` from the green instead of promoting the
`band_base` state field, which is removed. OP_SEL indexes `depths` and
OP_MOV identifies the selected band through `kind` instead of the red
`selected`.

aheui.aheui+99bottles: recorded guards 11513 -> 11039; cycles (min of 5)
707.3M -> 687.1M. +quine40: guards 13945 -> 13386; cycles 1490.5M ->
1420.1M. logo counters unchanged. Output, loop, bridge and guard-failure
counts are unchanged.

Assisted-by: Grok 4.7
Assisted-by: Claude Opus 5
`sp` was only ever `stacksize as usize`, written after the per-op delta and
at the end of OP_SEL. Band arms now bind `let sp = state.stacksize as usize;`
next to their use, and the `sp` state field, its `state_fields` entry and
its two writes are removed.

On aheui.aheui+99bottles the loop-header `same_as` ops on duplicated boxes
drop from 605 to 20 (MAJIT_LOG dump); recorded guards, output, loop, bridge
and guard-failure counts are unchanged.

Assisted-by: Grok 4.7
Assisted-by: Claude Opus 5
The previous pin, c30e57188e3, no longer built this tree: aheui-jit's
build.rs passes `PipelineConfig.helper_graphs` and six arguments to
`analyze_helper_pipeline_with_modules`. The new pin also changes
`register_extra_root_walker` to take a label, so the bigint root walker
registers as "aheui_bigint_roots".

Baselines re-recorded against the new pin:
- jitstress jitstats:
  - pi/pi.jinseo: loops_compiled 8 -> 5
  - quine/quine.puzzlet.40col: loops_compiled 11 -> 8
  - Both move with majit's bound_reached counter decay and its typed
    bound_reached cell. Outputs are unchanged.
- opcensus:
  - logo/logo: 1249 -> 1377 ops. IntGe 16 -> 274, IntMul 57 -> 187,
    IntSub 143 -> 13, UintLt 130 -> 0. This is 130 range tests of the form
    `(x >= lo) * (hi >= x)`. They used to fuse into
    `uint_lt(int_sub(x, lo), span)`, a pyre-only intbounds rewrite that
    majit #1899 removed. The count is the same when this tree's baseline
    commit 089f6d7 is built against the new pin.
  - literary/huntcook: +1 op.
  - quine/quine.puzzlet: +3 ops.

All 62 default and 62 jitstress jitstats programs pass, and so does
opcensus. Workspace tests other than compaheuiler pass.

Assisted-by: Claude Opus 5.5
`TRACE_LIMIT = 70000` is removed. The driver now sets `trace_limit` only
from the `MAJIT_TRACE_LIMIT` override, as rpaheui does with
`RPAHEUI_TRACE_LIMIT`.

The limit was raised because logo's whole-program recording exceeds the
default. At the default it aborted, and the segmented loop never closed
its bridges. With majit 262b993771e the bound_reached cell carries the
driver's greens, so the segmented loop and its bridges share one cell.

logo now aborts once as too long and compiles a segmented loop plus
2 bridges:
- default and jitstress jitstats: loops_aborted 0 -> 2, guard_failures
  21 -> 41
- opcensus: 1377 -> 1413 ops in 3 traces instead of 2 (IntAddOvf +20,
  GuardNoOverflow +21)
- instructions retired: +1.9%; max RSS 24.2 -> 16.4 MiB

Output is unchanged. aheui.aheui with 99bottles and with quine40 records
the same loops and bridges as before, and instructions retired are
unchanged.

Assisted-by: Claude Opus 5.5
The tip is rebased onto upstream/main 33a06674e0c. It includes the
jit_interp fix that merges a green reassignment inside an inline
dispatch arm.

Assisted-by: Claude Opus 5.5
The interpreter loop's MAJIT_HANDOFF and MAJIT_SPDIAG blocks now compile
only with the new `diag` feature (aheuinterpreter, forwarded by aheui-jit
and the root package). The diagnostic functions and their `calls`
registrations stay.

The traced jitcode is unchanged (99bottles walker steps and residual
calls are identical). Whole-run instructions on aheui.aheui 99bottles
drop by about 7-10M, all of it in the interpreter loop.

Assisted-by: Claude Opus 5.5
…gram

`Program` gains `cap` and `bands`, declared immutable through a
hand-written `__MAJIT_IMMUTABLE_FIELDS` const; ahsembler takes no majit
dependency. `mainloop` runs on a copy of the program carrying the resolved
cap and band count, and the loop head reads `program.cap` /
`program.bands` instead of calling `jit_cap()` / `jit_band_count()`, which
are removed together with their call-policy and wasm word-ABI entries.
`CAP_SELECTED` / `BAND_COUNT` stay for the root walk.

On aheui.aheui + 99bottles the traced mainloop makes one residual R_I call
per merge point instead of three (optrace R_I 15087 -> 4045 counted
against the S25 build; recorded ops and compiled traces are unchanged
against a build of the same majit without this change). Needs the majit-macros change
that types the portal env binding.

Assisted-by: Claude Opus 5.5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant