Diagnostics

JFR: The Events You Actually Need

Sixty seconds of default settings produced 4.3 MB, three quarters of it one event type, and none of it answered whether the application was lock-contended.

JFR: The Events You Actually Need — Diagnostics article cover
On this page

Sixty seconds of the default configuration produced 4.3 MB. Three quarters of it was one event type.

jdk.GCPhaseParallel fired 116,632 times in that minute — 73% of the events and 73% of the bytes. Meanwhile the workload contended on a single monitor in every request, and the recording contained zero events saying so.

Both facts come from the same file. Together they are the argument for treating the event catalogue as a triage rather than a menu.

The workload below is eight threads, each allocating request-sized buffers, entering one shared monitor, throwing one exception in fifty and appending to a file. JDK 26.0.2.1, macOS 26.6.2, Apple M4 Pro.

The catalogue is 196 events and seven knobs

JDK 26 defines 196 event types. default.jfc enables 88 of them; profile.jfc enables 93. The five extra events are not what separates the two configurations — 31 events carry different settings, and most of that is retuning: jdk.ExecutionSample moves from a 20 ms to a 10 ms period, jdk.JavaMonitorEnter from a 20 ms to a 10 ms threshold, jdk.ObjectAllocationSample from 150 to 300 samples a second, jdk.Deoptimization gains a stack trace.

A .jfc file is not 196 switches either. It is a small set of controls the events subscribe to:

gc                  off | normal | detailed | high | all
allocation-profiling off | low | medium | high
exceptions          off | throttled | all
method-profiling, compiler, thread-dump, memory-leaks, locking-threshold, class-loading

Those names go straight on the command line, which is the part most people miss:

java -XX:StartFlightRecording:settings=default,gc=off,exceptions=all,filename=r.jfr

Where the bytes actually go

jfr summary sorts a recording by what it cost to write. On the same sixty seconds:

Configuration Bytes Events
settings=default 4,319,142 158,929
settings=profile 7,868,815 359,185
default, jdk.GCPhaseParallel#enabled=false 1,208,185 43,204

The third row is the useful one. Turning off a single event removed 72% of the file, and jdk.GarbageCollection, jdk.GCPhasePause, jdk.YoungGarbageCollection and jdk.GCHeapSummary all survived — every pause number a GC log would have given you is still there. What went was the per-phase parallel-worker detail, which is a collector-engineering event that reached a general-purpose default profile.

The profile configuration has its own version of this: jdk.PromoteObjectOutsidePLAB, one of the seven events it adds, is 2.8 MB of its 7.9 MB. Two GC-internal event types are 76% of a profiling recording.

Throughput is not the axis

Oracle’s documentation puts default at “typically, less than 1% overhead” and says profile “records more events”. JMH on this machine, @Fork(2), five warmup and five measurement iterations:

Benchmark No recording default profile
allocate 9,740 ± 357 ops/ms 9,433 ± 75 9,443 ± 63
throwAndCatch 1,978 ± 32 ops/ms 1,887 ± 63 1,868 ± 86

Both recorded configurations land 3–5% below the baseline, and — the point — they land in the same place. default and profile are inside each other’s error bars on both benchmarks while differing 1.8x in bytes written. These are microbenchmarks whose entire body is the thing being sampled, so this is the worst case rather than a contradiction of the 1% figure; what it shows is that choosing between the two configurations is a decision about disk and retention window, not about speed.

Cost against answer, on one question

Take a single question — how many exceptions is this application throwing — and look at the two events that answer it.

jdk.ExceptionStatistics       59 events        662 bytes
jdk.JavaExceptionThrow     5,773 events    158,432 bytes

jdk.ExceptionStatistics is a periodic counter, once a second, and it is exact: 79,666 throws in the minute. jdk.JavaExceptionThrow costs 239 times more, and because it is throttled to 100 a second it captured 7.2% of them. It is still the event you want, because it carries the stack trace and therefore answers where. But it does not answer how many, and a count taken by summing those events is wrong by a factor of nearly fourteen.

That pairing repeats across the catalogue: a cheap periodic or statistics event that counts, and an expensive sampled event that attributes. Knowing which of the two you are reading is most of the skill.

What the defaults cannot answer

Eight threads entered the same monitor on every request. The default recording:

jdk.JavaMonitorEnter        0 events        0 bytes

Not a small number — zero. jdk.JavaMonitorEnter has a 20 ms threshold, and no individual acquisition came close, even though the aggregate contention was total. Lowering the threshold, same workload, 30 seconds each:

Threshold Events Bytes
20 ms (default) 0 0
5 ms 0 0
1 ms 170 3,587

The default configuration answers “is any single lock acquisition catastrophic”, not “is this application contended”. Those are different questions, and the second one costs 3.6 KB per thirty seconds to start answering.

Getting the file out of a running process

None of this requires a restart. A process started with no JFR flag at all:

$ jcmd 66435 JFR.start name=triage settings=default maxsize=100m maxage=10m
Started recording 1.

$ jcmd 66435 JFR.check
Recording 1: name=triage maxsize=100.0MB maxage=10m (running)

$ jcmd 66435 JFR.dump name=triage filename=live.jfr
Dumped recording "triage", 1.3 MB written to:

maxage is what makes a continuous recording useful: the last ten minutes are always in the buffer, and JFR.dump copies them out without stopping anything. That is the difference from a thread dump, which answers an instantaneous question and requires you to be holding the camera when the incident happens. JFR is for everything that is not instantaneous.

To write a configuration rather than pass flags, jfr configure produces a real .jfc from a base plus deltas:

jfr configure --input default --output triage.jfc \
    jdk.GCPhaseParallel#enabled=false \
    jdk.JavaMonitorEnter#threshold=1ms \
    jdk.ObjectAllocationSample#throttle=300/s

A working default, and three deltas

Start from default, then apply what the question needs:

  1. Always. jdk.GCPhaseParallel#enabled=false. It costs nothing to lose and roughly three and a half times the wall-clock window a fixed maxsize buffer holds.
  2. Hunting contention. locking-threshold=1ms — one control, and it moves jdk.JavaMonitorEnter and jdk.ThreadPark together. Cheap, and the defaults are blind here.
  3. Hunting allocation. allocation-profiling=high, and read jfr view allocation-by-site rather than the raw events.
  4. Hunting a stall. jdk.SafepointLatency#enabled=true, disabled in both shipped profiles, which reports per-thread time to safepoint with a stack.

Then stop reading event names. JDK 26 ships 82 built-in views — contention-by-site, allocation-by-site, safepoints, hot-methods — and jfr view allocation-by-site recording.jfr answers in one line what assembling the same picture out of jdk.ObjectAllocationSample takes an afternoon to do badly.

Frequently asked

Is the profile configuration much more expensive than default?
Not in throughput. On two JMH benchmarks the two were within each other's error bars, both around 3 to 5 percent below an unrecorded baseline. They differ in what they write - 7.9 MB against 4.3 MB for the same minute - so the decision is about disk and retention window, not about speed.
Why does my recording contain no lock contention events?
Because jdk.JavaMonitorEnter has a 20 ms threshold by default and most contention is far below it. A workload that contended on every request recorded zero events at 20 ms, zero at 5 ms, and 170 at 1 ms.

Search the site

Arrow keys to move, Enter to open.