Runtime

Streams Versus Loops: What the JIT Sees

The same pipeline cost 7.05 times the loop over eight elements and 1.31 times over a million. The ratio everyone quotes is a measurement at a trip count.

Streams Versus Loops: What the JIT Sees — Runtime article cover
On this page

Over eight elements the pipeline cost 7.05 times the loop. Over a million it cost 1.31 times.

Same source, same JDK 26 run, same JMH configuration. Nothing about the pipeline changed between those two figures except how many elements went through it, which is why the published ratios for “streams versus loops” disagree with each other and with the reader’s own profile: each of them is a measurement taken at one trip count.

All measurements are from JDK 26.0.2.1 (Homebrew build), macOS 26.6.2, Apple M4 Pro, twelve cores, under JMH 1.37 with three forks and five warmup plus five measurement iterations.

The ratio is a function of trip count

The work is a filter and a sum, written three ways: a for loop over int[], Arrays.stream(int[]).filter(...).asLongStream().sum(), and the same over Integer[].

Elements loop (ns/op) IntStream boxed Stream IntStream ÷ loop
8 3.470 ± 0.338 24.470 ± 0.242 32.944 ± 1.597 7.05
1,024 151.956 ± 8.838 215.149 ± 11.310 1127.680 ± 14.890 1.42
1,048,576 152802.743 ± 2384.841 200759.049 ± 37470.557 1136572.577 ± 23700.540 1.31

The boxed column is the one worth pausing on. It sits at 7.42 and 7.44 times the loop at the two larger sizes — close enough to the 4.5-to-6.5 range the canonical 2015 benchmark reported for simple operations that the ratio circulating as the cost of streams is often the cost of Integer. The primitive pipeline converges somewhere near 1.3, and it converges downwards, which is the opposite of what a per-element overhead would do.

A fixed cost, not a per-element one

-prof gc says why. The IntStream pipeline allocates 336.000 B/op over eight elements and 336.002 B/op over 1,024 — the same number, to three decimal places, for work that is 128 times larger. At 1,048,576 elements the figure reads 320.9 ± 27.4 B/op, the same constant measured against a denominator large enough that the profiler’s sampling shows through.

That constant is the pipeline: the spliterator, the sink chain, the stream object. It is built once per call and then, as the package documentation puts it, “filtering, mapping, and summing can be fused into a single pass on the data, with minimal intermediate state”. The fusion is real, the per-element cost after it is close to the loop’s, and the fixed cost is paid whether the call is going to process eight elements or a million.

Which makes the ratio arithmetic rather than mystery. A fixed cost of roughly 20 ns divided over eight elements is 2.6 ns each; divided over a million it disappears. The loop’s own allocation stays at or below 1.1 B/op at every size, a factor of roughly 300 under the pipeline.

What the first call costs

There is a second fixed cost, and it does not appear in any steady-state benchmark because JMH’s warmup exists to remove it.

A lambda’s invokedynamic call site is not linked until it first executes, and linkage, in the words of LambdaMetafactory, “may involve dynamically loading a new class that implements the target interface”; the stream machinery loads on first use as well. Measured as a single shot with no warmup, twenty forks, one measurement each:

First call µs/op
coldLoop 9.352 ± 0.183
coldIntStream 1228.769 ± 58.742

A factor of 131, or about 1.2 ms of wall clock. That is nothing in a process that runs for a week and a real number in one that starts a thousand times a day — and it is paid again on every restart, which is the connection to the startup work later in this plan rather than to anything about streams as a style.

Pollution needs a call to survive

The interesting failure mode is a shared pipeline. Four different lambdas reaching one helper method should make the predicate call site megamorphic, and what happens then was the whole subject of part one of this series.

The obvious experiment measures nothing. Four predicate classes through one shared filteredSum helper produced no slowdown at all — the “polluted” arrangement came out 6.8% faster than four separate local pipelines, well outside the error bars. C2 had inlined the helper at each of its four call sites, and each inlined copy saw exactly one predicate type. There was no shared call site left to pollute.

Pollution requires the call to survive. With the helper held back from inlining — the state a real helper reaches by being large, or by being called from thirty places rather than four — the same body behaves differently:

1,048,576 elements ns/op
one predicate class at the call site 772,742 ± 121,313
four predicate classes at the call site 8,018,375 ± 30,815

10.4 times, for identical work. TypeProfileWidth is 2 on this build, so a third receiver type leaves C2’s guarded inlining nothing to guard: the predicate is no longer inlined into the loop, the fusion the package documentation describes stops happening, and each element pays a virtual call.

That is the same mechanism as part one, on a different abstraction, and it is why the failed experiment is worth publishing. The question is never how many lambda classes exist. It is how many of them reach one surviving call site.

Three questions instead of one ratio

The honest answer to “are streams slower” is that the question is missing three of its terms.

How many elements per invocation — because the pipeline’s cost is fixed and the loop’s is not, and the ratio is that arithmetic. Whether the values are boxed — because that is where the large published ratios actually come from. And how many lambda classes reach the pipeline’s call site, which decides whether the per-element cost is a fused loop or a virtual call.

None of that requires a decision about style. It requires the trip count, which the code already knows, and a look at the compilation decisions that follow from it.

Frequently asked

Are streams slower than loops?
At eight elements per call, on JDK 26, by a factor of 7.05. At a million elements per call, by a factor of 1.31. The published figures disagree because each was taken at one trip count, and the pipeline's cost is almost entirely fixed rather than per-element.
Does boxing matter more than the pipeline?
Yes, and by a wide margin. A stream of Integer ran 7.4 times the loop at every size at or above 1024 elements, while the primitive pipeline converged on 1.31. The large ratios in circulation are usually measuring boxing.

Search the site

Arrow keys to move, Enter to open.