Skip to content

ci(pack): run shared runtime ecosystem suite on release - #3397

Closed
fireairforce wants to merge 1 commit into
nextfrom
zoomdong/release-shared-runtime-ecosystem
Closed

fireairforce wants to merge 1 commit into
nextfrom
zoomdong/release-shared-runtime-ecosystem

Conversation

@fireairforce

Copy link
Copy Markdown
Member

Summary

Merge order

Merge utooland/utoo-ecosystem-ci#5 first, then this PR, before cutting the next Utoopack tag. The reusable workflow is loaded from @main, so suite: release must exist there when the release runs.

Validation

  • actionlint -shellcheck= .github/workflows/pack-release.yml passed. The full actionlint command reports two pre-existing ShellCheck infos on unrelated lines.
  • The repository pre-push checks passed: cargo fmt --check, tombi format --check, biome ci, and typos.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T09:47:13.990248Z 86441fc PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@github-actions

Copy link
Copy Markdown

📊 Performance Benchmark Report (with-antd)

Utoopack Performance Report

Report ID: utoopack_performance_report_20260929_100313
Generated: 2026-09-29 10:03:13
Trace File: trace_antd.json (0.3GB, 0.82M spans)
Test Project: examples/with-antd


Executive Summary

Metric Value Assessment
Total Wall Time 6,390.9 ms Baseline
Total Thread Work (de-duped) 19,190.0 ms Non-overlapping busy time
Effective Parallelism 3.0x thread_work / wall_time
Working Threads 10 Threads with actual spans
Thread Utilization 30.0% ⚠️ Suboptimal
Total Spans 824,788 All B/E + X events
Meaningful Spans (>= 10us) 207,783 (25.2% of total)
Tracing Noise (< 10us) 617,005 (74.8% of total)

Build Phase Timeline

Shows when each build phase is active and how much CPU it consumes.
Self-Time is the time spent exclusively in that phase (excluding children).

Phase Spans Inclusive (ms) Self-Time (ms) Wall Range (ms)
Resolve 47,165 6,315.6 1,734.4 2,989.6
Parse 8,175 1,388.1 1,040.7 5,632.6
Analyze 134,538 41,668.8 8,406.9 5,539.0
Chunk 5,190 5,232.4 781.3 2,265.4
Codegen 10,106 2,105.1 1,371.9 1,938.7
Emit 31 42.7 21.4 10.6
Other 2,578 7,641.1 4,273.2 6,390.9

Workload Distribution by Diagnostic Tier

Category Spans Inclusive (ms) % Work Self-Time (ms) % Self
P0: Scheduling & Resolution 181,992 48,578.7 253.1% 10,374.1 54.1%
P1: I/O & Heavy Tasks 2,863 114.1 0.6% 92.8 0.5%
P2: Architecture (Locks/Memory) 0 0.0 0.0% 0.0 0.0%
P3: Asset Pipeline 22,015 8,774.2 45.7% 3,222.7 16.8%
P4: Bridge/Interop 0 0.0 0.0% 0.0 0.0%
Other 913 6,926.8 36.1% 3,940.1 20.5%

Top 20 Tasks by Self-Time

Self-time is the exclusive duration: time spent in the task itself, not in sub-tasks.
This is the most accurate indicator of where CPU cycles are actually spent.

Self (ms) Inclusive (ms) Count Avg Self (us) P95 Self (ms) Max Self (ms) % Work Task Name Top Caller
4,119.2 24,205.7 90,094 45.7 0.1 8.8 21.5% module module (60%)
2,475.4 2,912.0 20 123769.4 401.3 481.9 12.9% save snapshot persist (5%)
2,048.5 2,179.0 2,377 861.8 2.9 223.9 10.7% analyze ecmascript module module (73%)
1,379.5 14,220.0 32,731 42.1 0.1 5.8 7.2% process module process module (82%)
1,163.9 3,123.0 25,612 45.4 0.1 26.2 6.1% internal resolving internal resolving (76%)
977.4 1,324.9 6,007 162.7 0.6 68.7 5.1% parse ecmascript parse ecmascript (65%)
758.7 831.4 7,843 96.7 0.4 5.7 4.0% precompute code generation generate merged code (44%)
717.3 812.3 6,744 106.4 0.4 125.2 3.7% compute async module info compute async module info (57%)
642.2 2,053.7 649 989.6 2.2 262.2 3.3% generate merged code chunking (65%)
605.9 4,773.1 3,940 153.8 0.2 49.3 3.2% chunking chunking (62%)
562.2 3,184.3 20,889 26.9 0.0 4.8 2.9% resolving module (57%)
437.9 437.9 9 48651.5 213.5 229.8 2.3% blocking save snapshot (67%)
432.1 432.1 329 1313.5 1.1 257.8 2.3% generate source map code generation (83%)
359.8 767.5 134 2684.9 4.6 213.1 1.9% emit code generate merged code (41%)
311.9 578.7 1,309 238.3 0.1 176.1 1.6% write all entrypoints to disk write all entrypoints to disk (15%)
181.1 841.5 1,934 93.6 0.2 44.5 0.9% code generation code generation (84%)
171.5 455.1 1,187 144.5 0.2 33.1 0.9% compute async chunks compute async chunks (46%)
92.0 111.8 712 129.2 0.1 30.4 0.5% compute binding usage info compute binding usage info (57%)
63.2 63.2 2,166 29.2 0.0 1.9 0.3% read file parse ecmascript (91%)
37.5 67.1 1,863 20.1 0.0 13.1 0.2% collect mergeable modules collect mergeable modules (99%)

Critical Path Analysis

The longest sequential dependency chains that determine wall-clock time.
Focus on reducing the depth of these chains to improve parallelism.

Rank Self-Time (ms) Depth Path
1 711.8 3 persist → save snapshot → blocking
2 476.3 6 chunking → generate merged code → emit code → emit code → emit code → read file
3 401.8 2 save snapshot → blocking
4 298.8 4 chunking → generate merged code → emit code → generate source map
5 227.1 9 module → module → process module → ... → process module → analyze ecmascript module → analyze ecmascript module

Batching Candidates

High-volume tasks dominated by a single parent. If the parent can batch them,
it drastically reduces scheduler overhead.

Task Name Count Top Caller (Attribution) Avg Self P95 Self Total Self
process module 32,731 process module (82%) 42.1 us 0.06 ms 1,379.5 ms
internal resolving 25,612 internal resolving (76%) 45.4 us 0.08 ms 1,163.9 ms

Duration Distribution

Range Count Percentage
<10us 617,005 74.8%
10us-100us 133,964 16.2%
100us-1ms 63,319 7.7%
1ms-10ms 10,279 1.2%
10ms-100ms 182 0.0%
>100ms 39 0.0%

Action Items

  1. [P0] Focus on tasks with the highest Self-Time — these are where CPU cycles are actually spent.
  2. [P0] Use Batching Candidates to identify callers that should use try_join or reduce #[turbo_tasks::function] granularity.
  3. [P1] Check Build Phase Timeline for phases with disproportionate wall range vs. self-time (= serialization).
  4. [P1] Inspect P95 Self (ms) for heavy monolith tasks. Focus on long-tail outliers, not averages.
  5. [P1] Review Critical Paths — reducing the longest chain depth directly improves wall-clock time.
  6. [P2] If Thread Utilization < 60%, investigate scheduling gaps (lock contention or deep dependency chains).

Report generated by Utoopack Performance Analysis Agent

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant