Export JVM metrics and capture a heap dump on OOM - #129
Merged
Merged
Conversation
The server exported its own cache-size gauges but nothing about the JVM, and an OOM left nothing behind. Memory problems were invisible until the process died and unexplainable afterwards. Metrics.init now calls DefaultExports.initialize(), so heap, GC, thread and class-loading collectors are scrapeable. The signal worth alerting on is heap that fails to fall back after a GC, rather than peak heap. simpleclient_hotspot was already on the runtime classpath, but only transitively through prometheus-proxy, so importing from it directly would have rested on someone else's dependency graph. It is now declared explicitly in the catalog under the existing prometheus version ref. Heap dumps are enabled for local runs via applicationDefaultJvmArgs and the uber target, and ReadingBatServer reports at startup whether the JVM will write one, warning when it will not. The flags are set by whoever launches the JVM rather than by this code, so a startup report is the difference between setting an env var and knowing it took effect. Production is configured in readingbat-site. CLAUDE.md gains a warning against measuring memory inside a test task. Kover's coverage agent holds a strong reference to every classloader it sees, so script evaluation appears to leak about 1 MB per Kotlin eval under ./gradlew test and does not leak at all in production -- a phantom that survived an entire investigation (#128) before a GC-root trace found the agent holding it. Verified the MXBean read reports enabled=false with no flags and enabled=true with them, and that removing DefaultExports.initialize() fails the new test with all four metric families missing. applicationDefaultJvmArgs uses listOf rather than a collection literal: -Xcollection-literals applies to project sources via configureKotlin(), not to Gradle's own Kotlin DSL compilation. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
3 tasks
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
"The flags cost nothing until an OOM" appeared twice, once in the opening paragraph and again in the paragraph explaining why the function logs rather than enforces. The second paragraph now makes only its own point. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
3 tasks
pambrose
added a commit
that referenced
this pull request
Sep 27, 2026
…nd cut 3.5.0 (#130) common-utils 5 and prometheus-proxy 4.2 moved to the Prometheus Java client 1.x, and common-utils' MetricsService now serves PrometheusRegistry.defaultRegistry. A collector left in the old 0.x CollectorRegistry still compiles and runs but is never scraped -- which is what would have happened to the JVM collectors added in #129 had DefaultExports.initialize() been kept. They are now registered with JvmMetrics.builder().register(), the 30 recording sites move from labels(...) to labelValues(...), and server_start_time_seconds is set explicitly because 1.x dropped setToCurrentTime(). JvmMetricsTest now asserts against the registry that is actually served, using the 1.x metric names (jvm_memory_used_bytes, jvm_classes_currently_loaded). The public Metrics properties are now 1.x types, which is source-breaking for anything that records to them from outside the library. Since common-utils 4.x, FileSystemSource.file(path) resolves against pathPrefix itself, so Challenge's call -- which prepended pathPrefix as well -- doubled it on the upgrade. Invisible for a "./" root, fatal for "../", where the Python test content lives. Paths are now relative to the source root. The docs site did not build from the project's own uv environment: zensical.toml named the emoji extension by its Material for MkDocs path, and zensical remaps material.extensions only in YAML configs. CI masked it by also installing mkdocs-material; the config now names zensical.extensions.emoji and the workflow installs zensical alone. Also fixes an unresolvable KDoc link on Endpoints.STATIC_PATH; bumps Gradle 9.8.0, Kotlin 2.4.20, Flyway 13.8.0, kotlinter 5.7.0, buildconfig 6.1.2 and the website's Python deps; and documents the release (3.5.0, 2026-09-27) across CHANGELOG, RELEASE_NOTES, README, llms.txt, CLAUDE.md and the docs site. 381 tests, 0 failures, 6 skipped; lint and detekt clean. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The server exported its own cache-size gauges but nothing about the JVM, and an OOM left nothing behind — memory problems were invisible until the process died and unexplainable afterwards. This fixes both halves.
Metrics.initnow callsDefaultExports.initialize(), so heap, GC, thread and class-loading collectors are scrapeable. The signal worth alerting on is heap that fails to fall back after a GC, not peak heap.simpleclient_hotspotwas already on the runtime classpath only transitively throughprometheus-proxy; it is now declared explicitly in the catalog under the existingprometheusversion ref.applicationDefaultJvmArgsand theubertarget. Production is configured separately in readingbat-site.ReadingBatServerlogs whether the JVM will write a dump on OOM, and warns when it will not. The flags are set by whoever launches the JVM rather than by this code, so this is the difference between setting an env var and knowing it took effect../gradlew testand does not leak at all in production. That phantom survived an entire investigation (Kotlin script evaluation retains ~1.3 MB of heap per eval, permanently #128) before a GC-root trace found the agent holding it.applicationDefaultJvmArgsuseslistOfrather than a collection literal:-Xcollection-literalsapplies to project sources viaconfigureKotlin(), not to Gradle's own Kotlin DSL compilation.Test plan
JvmMetricsTestassertsjvm_memory_bytes_used,jvm_gc_collection_seconds,jvm_threads_currentandjvm_classes_loadedare registered afterMetrics.init. Verified it bites: removingDefaultExports.initialize()fails it with all four missing.HotSpotDiagnosticMXBeanread behind the startup report returnsenabled=falsewith no flags andenabled=true path=/tmp/dumpswith them.-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=build.make lintclean.🤖 Generated with Claude Code