Skip to content

Runtime performance: ~700× slower than native on a geo benchmark; analysis and a 2× prototype fix #66

Description

@mipastgt

After a long time of silence I evaluated rustc_codegen_jvm (5ed31a3) again on a real-world crate: geo 0.33 buffering a line string and doing point-in-polygon tests. The results are identical to native in every run, including the i_overlay buffer construction. Very impressive!

Performance is about 700× native. It's not the JVM: a line-by-line Java port of the kernel runs at 1.5× native, and the same program compiled to Wasm and run through Chicory on the same JVM reaches about 3× native. Profiles put ~90% of the time in org.rustlang.runtime, mainly memory views (field reads through &T into Vec/Box storage decode and register whole structs), projectStructField's by-name lookups and Range::next allocations.

The attached patch (~220 lines) lowers primitive field reads through shared references to Freeze types to a single scalar load at a constant offset. It makes buffer and classify 2× faster, and python3 Tester.py --release passes.

The attached zip has the full write-up (geo_buffer_bench/docs/FINDINGS.md), the patch, the benchmark and the probes, with commands to reproduce everything.

This analysis was conducted with the help of Claude Opus 5.5.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions