Skip to content

[ot] Keep mark-specific property checks out of line - #503

Closed
behdad wants to merge 1 commit into
perf/share-contextual-matchingfrom
perf/outline-mark-check
Closed

behdad wants to merge 1 commit into
perf/share-contextual-matchingfrom
perf/outline-mark-check

Conversation

@behdad

@behdad behdad commented Oct 6, 2026

Copy link
Copy Markdown
Member

Stacked on #502, which is stacked on the extended-layout work in #495.
This PR contains only the two-line mark-check change.

Keep match_properties_mark out of line so its filtering-set and attachment-class
checks are shared across the generic matching loops. The common non-mark path
stays inline; matching behavior is unchanged.

Size and performance

Measured on Linux x86-64 with Rust 1.89, release optimization, fat LTO, and one
codegen unit, against the base including #502:

  • 12,836 bytes saved in ELF text+data, in both hr-shape and the C API library.
  • 12 KiB smaller stripped files, in both artifacts.
  • Nastaliq timing is essentially unchanged. Amiri samples are about 1–2.4% slower;
    the other measured samples are roughly within ±1%. This is a deliberate size
    tradeoff, not a claim of universal performance neutrality.

Timing checks used CPU-pinned alternating before/after runs across six short
samples, Arabic text with diacritics, the existing Arabic/Indic/Latin benchmark
corpora, and a variable Source Serif sample. Glyphs, positions, and flags match
the baseline in all four directions for those inputs.

Forcing all apply methods or generic matching functions out of line grew the
binary; outlining the shared mutation/application helpers barely helped. Only
the mark-specific helper is kept here.

Validation

  • Workspace all-feature tests, including 6073 shaping cases.
  • Strict workspace/all-target/all-feature Clippy and rustfmt.
  • Rust 1.85 no-std build for thumbv7em-none-eabihf with libm.

Inlining this shared helper duplicates mark-filtering and attachment
class checks across generic matching loops. Keep one out-of-line body
while leaving the non-mark fast path inline.

In the measured release/LTO build this saves about 12.5 KiB of ELF
text+data (12 KiB in stripped CLI and C API library file sizes). The
Nastaliq sample is unchanged; Amiri samples slow by up to about 2.4%.

Workspace all-feature tests, including 6073 shaping cases, strict
Clippy, formatting, and Rust 1.85 no-std ARM build pass.

Assisted-by: OpenAI Codex
@behdad

behdad commented Oct 6, 2026

Copy link
Copy Markdown
Member Author

Too dependent on codegen compiler decisions.

@behdad behdad closed this Oct 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant