Runtime
The Cost of Optional
Identical source, five call sites, one JDK 26 run - nothing allocated at the first and twelve bytes at three others. The wrapper is not what decides.

On this page
The same source allocated nothing at one call site and twelve bytes at three others.
One JDK 26 run, one wrapper class, five places it was used. Nothing about Optional differed between them, and no change was made to the calling code — what differed was what C2 was able to prove at each site, and that turned out to be the whole of the cost.
All measurements below are from JDK 26.0.2.1 (Homebrew build), macOS 26.6.2, Apple M4 Pro, twelve cores, under JMH 1.37 with three forks and five warmup plus five measurement iterations.
Five call sites, one wrapper
The lookup is the same in every case: an array of sixteen strings, every fourth entry null, wrapped with Optional.ofNullable and consumed with map(String::length).orElse(0).
interface Store { Optional<String> get(int key); }
// monomorphic: the field is declared as the concrete class
return mono.get(i++).map(String::length).orElse(0);
// megamorphic: four implementations rotate through one call site
return poly[n & 3].get(n).map(String::length).orElse(0);
-prof gc reports bytes allocated per operation, which is a far steadier number than a timing figure and the one that answers the question directly:
| Call site | ns/op | B/op |
|---|---|---|
nullCheck — no wrapper at all |
0.825 ± 0.005 | ≈ 0 |
monomorphic |
1.451 ± 0.025 | ≈ 0 |
stored — wrapper written to a field |
2.314 ± 0.059 | 12 |
acrossBoundary — returned through a method C2 may not inline |
2.357 ± 0.016 | 12 |
megamorphic — four receiver types |
3.865 ± 0.480 | 12 |
The first two rows are the reason the folklore exists, and the last three are the reason it keeps being contradicted by allocation profiles. Both are correct about a case and wrong about the general claim.
Twelve bytes is three quarters of an allocation
An Optional is not twelve bytes. A probe that allocates exactly one escaping Optional per operation and does nothing else reports 16.000 ± 0.001 B/op. That is what the layout predicts: on this build UseCompressedOops and UseCompressedClassPointers are on and UseCompactObjectHeaders is off, so the instance is an eight-byte mark word, a four-byte compressed class pointer and a four-byte reference to the value, already aligned to the eight-byte boundary.
Twelve is what sixteen becomes when a quarter of the calls allocate nothing. Every fourth entry in the array is null, Optional.ofNullable(null) returns the shared EMPTY instance rather than constructing one, and the same probe pointed at ofNullable(null) reports nothing at all. Three sixteens across four operations is twelve, which is what the profiler reported to three decimal places.
The arithmetic is worth doing rather than assuming, because it also settles what is not in the number. map(String::length) produces a second Optional and boxes an int, and neither appears in the total: the boxed lengths are 2 and 3, which come from the Integer cache, and the second wrapper is scalar-replaced even in the runs where the first one is not. Only the wrapper that crossed the failed proof survives.
What the proof needs
The arithmetic explains the number; it does not explain why one call site produced zero and three produced sixteen bytes per present value.
Scalar replacement requires C2 to see the allocation and every use of it in one compiled unit, and each of the three failing rows removes that in a different way. The stored case is the simplest: a field write makes the object reachable after the method returns, so the escape is real rather than a limitation. acrossBoundary returns the wrapper from a method marked DONT_INLINE; the allocation is still local in principle, but the compiler is not permitted to look at the code that would prove it.
megamorphic is the one that reaches production code without anyone deciding to write it. TypeProfileWidth is 2 on this build: the profile records at most two receiver types, and C2’s guarded inlining stops at the same number — a monomorphic or bimorphic call site gets an inlined body behind a type check, and a third type leaves nothing to guard. The callee is then not inlined, the allocation inside it is never visible to the caller, and the wrapper materialises — for a reason that lives in a completely different file from the one being optimised.
Reading the decision
None of this has to be inferred. -XX:+UnlockDiagnosticVMOptions -XX:+PrintInlining prints the decision that produced each row:
megamorphic: @ 21 OptionalCost$Store::get (0 bytes) failed to inline: virtual call
monomorphic: @ 15 OptionalCost$ArrayStore::get (18 bytes) inline (hot)
Two lines, same source, and the (0 bytes) on the first is the tell: C2 never resolved a method body to measure, because there was no single body to resolve.
That pair is the whole diagnostic. -prof gc says whether the wrapper survived; the inlining log says which proof failed. Together they replace the question “is Optional slow” with one that has an answer.
They also correct the optimistic reading. monomorphic allocates nothing and still costs 0.63 ns more than the null check it replaced — the null test, the branch and the orElse are real work that escape analysis was never going to remove. Scalar replacement makes the wrapper free of allocation, which is not the same claim as free.
The wrapper was never the variable
Three failures, three mechanisms, and the same twelve bytes. The wrapper does not record which proof failed, which is why an allocation profile can show Optional at the top of the list and say nothing about what to change.
The test is cheap and specific: run the hot path under -prof gc and compare gc.alloc.rate.norm against zero. If it is not zero, -XX:+PrintInlining names the call site, and the fix is at that call site — one fewer implementation on the hot interface, a boundary removed, a field that did not need to hold the wrapper — rather than in a style rule about Optional.
This is the first of five articles measuring the same claim against different abstractions, and it is the claim the rest of the series rests on: the cost of an abstraction is a property of the call site, not of the abstraction. Streams are next, and they fail in the same place for the same reason.
The mechanism underneath all of it — what escape analysis proves, and what it does with the proof — is the allocation that never happened.
Frequently asked
- Does Optional allocate?
- That is a question about a call site, not about Optional. On JDK 26 the same wrapper allocated nothing where C2 could prove it did not escape, and 16 bytes per present value where it could not - at a megamorphic call site, behind a boundary C2 was not allowed to inline through, and when the wrapper was stored in a field.
- Is it enough to keep Optional out of fields and parameters?
- No. Storing the wrapper is one of three failures measured here, and the other two happen in code that follows the API note exactly. A method that returns Optional through an interface whose call site sees three or more receiver types allocates whatever the style guide says.


