You have an ArrayList of orders and you need a Set of SKUs, a Map from id to order, or an unmodifiable list of the active ones. The pipeline in the middle is Streams basics and Collectors in the advanced post. This post is the collection-side of that trip: how a collection opens a stream, what its Spliterator advertises, and how you land in the right collection — not another pass through map and filter.
Leave with stream(). Land with a collector that names the type. Parallel is a split, not a lock.
Same checkout records as the rest of the series:
public record LineItem(String sku, int quantity, BigDecimal unitPrice) {}
public record Order(
String id,
String customerEmail,
List<LineItem> items,
BigDecimal total,
boolean active) {}
stream() and parallelStream() leave the collection
Every Collection can open a stream. The list (or set, or queue) still holds the data. The stream is a pipeline over it.
List<Order> orders = new ArrayList<>();
orders.add(new Order("o-100", "ada@example.com", List.of(), new BigDecimal("12.00"), true));
orders.add(new Order("o-101", "linus@example.com", List.of(), new BigDecimal("4.50"), false));
Stream<Order> sequential = orders.stream();
Stream<Order> parallel = orders.parallelStream();
stream() is sequential: one encounter order, one thread (unless you later call parallel()). parallelStream() asks the JDK to split the source and run partitions on the common ForkJoin pool. Under the hood both are StreamSupport.stream(spliterator(), parallel) — the collection’s spliterator is the source. Both are views of the source at runtime, not copies. Do not add or remove from orders while either pipeline runs — fail-fast collections throw ConcurrentModificationException; concurrent collections may show some later updates. That iterator story is on the hub.
parallelStream() is not a free speedup and it is not thread-safety for a shared ArrayList you happen to add to inside a lambda. Small lists and I/O-bound work get slower. CPU-bound work on a large, split-friendly source is the case that wins. Measure.
A Map is not a Collection. You do not call byId.stream(). You stream a view:
Map<String, Order> byId = new HashMap<>();
byId.put("o-100", orders.get(0));
Stream<String> ids = byId.keySet().stream();
Stream<Order> values = byId.values().stream();
Stream<Map.Entry<String, Order>> rows = byId.entrySet().stream();
Those views carry the map’s characteristics: a ConcurrentHashMap key set streams as CONCURRENT and not ORDERED; a LinkedHashMap key set streams in encounter order. Pick the view that matches the element you want, then land with a collector as below.
The pipeline itself — lazy intermediates, single-use streams, when a for-loop is clearer — stays in Streams basics. Production pitfalls and groupingBy stay in Streams advanced. Custom intermediate ops that remember earlier elements are Gatherers, not a collector.
What a Spliterator advertises
Collection.spliterator() is how the stream knows whether it can split, whether order matters, and whether the source is sized. You rarely call trySplit yourself. You do read the characteristics when an interview (or a surprise parallel result) asks why ArrayList and ConcurrentHashMap behave differently under parallelStream().
| Characteristic | Means |
|---|---|
SIZED | estimateSize is an exact count |
SUBSIZED | splits are also SIZED — even partitions |
ORDERED | encounter order is defined and must be preserved unless you say otherwise |
DISTINCT | no duplicate elements (Set) |
SORTED | follows a sort; getComparator() may be non-null |
NONNULL | no null elements |
IMMUTABLE | the source will not be structurally modified |
CONCURRENT | the source may be modified concurrently; traversal is weakly consistent |
trySplit() conceptually carves off a prefix (or a bin range) and returns a second spliterator, or null if this source cannot usefully split.
Spliterator<Order> remaining = orders.spliterator();
Spliterator<Order> prefix = remaining.trySplit();
// ArrayList: prefix is roughly half the index range; remaining keeps the rest
// prefix == null means "cannot split usefully from here"
An ArrayList splits an index range in half — cheap, SIZED and SUBSIZED, ORDERED. A ConcurrentHashMap splits table bins — CONCURRENT, DISTINCT, NONNULL, not ORDERED. A default spliterator-from-iterator (some custom collections) splits poorly; parallelStream() then spends its time coordinating, not computing.
Spliterator<Order> split = orders.spliterator();
System.out.println(split.hasCharacteristics(Spliterator.ORDERED)); // true for ArrayList
System.out.println(split.hasCharacteristics(Spliterator.SIZED)); // true
System.out.println(split.hasCharacteristics(Spliterator.DISTINCT)); // false — List allows duplicates
Spliterator<String> keys = new ConcurrentHashMap<String, Order>().keySet().spliterator();
System.out.println(keys.hasCharacteristics(Spliterator.CONCURRENT)); // true
System.out.println(keys.hasCharacteristics(Spliterator.ORDERED)); // false
System.out.println(keys.hasCharacteristics(Spliterator.DISTINCT)); // true
ArrayList advertises ordered, sized splits. That is why parallelStream() on a large list can actually divide work. ConcurrentHashMap advertises CONCURRENT: another thread may put while you stream; you will not get fail-fast CME, and you will not see a defined encounter order.
List.of results typically advertise IMMUTABLE (and ORDERED, SIZED). HashSet advertises DISTINCT but not ORDERED — a parallel stream over a hash set does not owe you insert order, because the set never did.
You do not need to bitwise-decode every collection. Remember the two that interviews ask: list vs concurrent map, and ORDERED vs DISTINCT.
Land in a collection: collectors that name the type
Finishing a pipeline with collect is how a stream becomes a collection again. The advanced Streams post owns groupingBy, partitioningBy, joining, and merge functions in depth. Here we only pick the destination type.
import static java.util.stream.Collectors.collectingAndThen;
import static java.util.stream.Collectors.toCollection;
import static java.util.stream.Collectors.toList;
import static java.util.stream.Collectors.toMap;
import static java.util.stream.Collectors.toSet;
import static java.util.stream.Collectors.toUnmodifiableList;
import static java.util.stream.Collectors.toUnmodifiableMap;
import static java.util.stream.Collectors.toUnmodifiableSet;
List<LineItem> items = List.of(
new LineItem("SKU-MUG", 2, new BigDecimal("12.00")),
new LineItem("SKU-TEA", 1, new BigDecimal("4.50")),
new LineItem("SKU-MUG", 1, new BigDecimal("12.00")));
toList / toSet / toCollection. toList() has no guaranteed concrete class or mutability (in practice an ArrayList). toSet() is typically a HashSet — uniqueness, no encounter order. toCollection(ArrayList::new) (or TreeSet::new, LinkedList::new) is how you name the class.
List<String> skusMutable = items.stream()
.map(LineItem::sku)
.collect(toList());
skusMutable.add("SKU-BIN"); // typically fine — but the spec does not promise ArrayList
Set<String> uniqueSkus = items.stream()
.map(LineItem::sku)
.collect(toSet());
// SKU-MUG, SKU-TEA — HashSet-like, order unspecified
ArrayList<String> named = items.stream()
.map(LineItem::sku)
.collect(toCollection(ArrayList::new));
named.add("SKU-BIN"); // this is an ArrayList, on purpose
toUnmodifiableList / toUnmodifiableSet / toUnmodifiableMap (Java 10). The result rejects mutation and rejects nulls (NPE). Same family as List.of.
List<String> frozen = items.stream()
.map(LineItem::sku)
.collect(toUnmodifiableList());
frozen.add("SKU-BIN"); // UnsupportedOperationException
Stream.toList() (Java 16). Unmodifiable, preserves encounter order, allows nulls — unlike List.of and unlike toUnmodifiableList(). Prefer it when the pipeline is already a stream and the caller must not add. Prefer toCollection(ArrayList::new) when they will.
List<String> java16 = items.stream()
.map(LineItem::sku)
.toList();
java16.add("SKU-BIN"); // UnsupportedOperationException
| Finish | Typical result | Nulls | Mutate? |
|---|---|---|---|
collect(toList()) | unspecified (often ArrayList) | allowed | typically yes; not guaranteed |
collect(toCollection(ArrayList::new)) | ArrayList | allowed | yes |
collect(toUnmodifiableList()) | unmodifiable list | no | no |
stream.toList() (16+) | unmodifiable list | yes | no |
collect(toSet()) | unspecified Set | allowed | typically yes |
collect(toUnmodifiableSet()) | unmodifiable set | no | no |
toMap. Keys from one function, values from another. Duplicate keys throw unless you pass a merge function — that detail is in Streams advanced. Collection-side: you get a Map, typically HashMap. toUnmodifiableMap freezes it and rejects nulls. toMap(..., ..., ..., LinkedHashMap::new) names the map class the same way toCollection names the list.
Map<String, Order> byId = orders.stream()
.collect(toMap(Order::id, o -> o));
Map<String, Order> frozenById = orders.stream()
.collect(toUnmodifiableMap(Order::id, o -> o));
A checkout helper that must name both uniqueness and encounter order does not hope toSet() is a LinkedHashSet. It says so:
LinkedHashSet<String> skusInPackOrder = items.stream()
.map(LineItem::sku)
.collect(toCollection(LinkedHashSet::new));
First occurrence of SKU-MUG wins; later duplicates drop; iteration follows first-seen order. toSet() would drop duplicates too and forget that order. toUnmodifiableSet() would freeze a hash set. The supplier is the type decision.
collectingAndThen. Collect, then run a finishing function on the result — wrap, copy, or freeze without exposing the mutable intermediate.
List<String> snapshot = items.stream()
.map(LineItem::sku)
.collect(collectingAndThen(toCollection(ArrayList::new), List::copyOf));
That is “mutable while the collector fills, unmodifiable when the caller receives it.” Collections::unmodifiableList here would wrap the ArrayList the collector just built; nobody else holds that list, so the view is effectively a freeze. List::copyOf matches the factory contract, including the null rejection.
Parallel does not make your accumulator safe
The collector protocol is: a supplier per partition, an accumulator that accepts one element, a combiner that merges two partial results. Built-in toList() works under parallelStream() because the framework combines those partial lists. Your outer ArrayList is not in that protocol.
List<String> ids = new ArrayList<>();
orders.parallelStream()
.map(Order::id)
.forEach(ids::add); // race — ArrayList.add is not thread-safe
That is the bug. parallelStream() runs the forEach on several threads. They all call add on the same list. You can lose elements, blow the array, or both. Wrapping ids in Collections.synchronizedList still makes forEach + add a smell — you wanted collect.
List<String> ids = orders.parallelStream()
.map(Order::id)
.collect(toCollection(ArrayList::new));
parallelStream() does not make an arbitrary collector thread-safe. A custom collector that closes over a shared ArrayList and adds in the accumulator is still a race. Concurrent collectors (groupingByConcurrent, toConcurrentMap) exist when the result map should be filled concurrently. The built-in toList() is safe under parallel because of supplier + combiner, not because parallel “upgrades” whatever you pass.
forEach on a parallel stream also does not promise encounter order. If you needed order, you needed an ordered collect, not a shared sink.
Interview lens
Interviewers want stream vs parallelStream, what a spliterator advertises, and how you finish.
| Question | Honest answer |
|---|---|
stream() vs parallelStream()? | Sequential vs split-and-run on the common pool. Parallel is not automatically faster and not a lock. |
ORDERED vs DISTINCT on a spliterator? | ORDERED: encounter order is a contract (ArrayList). DISTINCT: no duplicates (Set, ConcurrentHashMap keys). A list is ordered and not distinct. A hash set is distinct and not ordered. |
Collectors.toList() vs Stream.toList()? | toList() collector: unspecified, typically mutable. Stream.toList() (16+): unmodifiable, nulls allowed. toUnmodifiableList(): unmodifiable, nulls forbidden. |
Why is forEach + add on an outer ArrayList a race under parallel? | Several threads call ArrayList.add. Use collect. Parallel does not wrap your sink. |
ConcurrentHashMap spliterator? | CONCURRENT (and DISTINCT, NONNULL). Weakly consistent. Not ORDERED. |
Wrong answer: “parallelStream() makes any collector thread-safe.” It splits the source. Collectors that follow supplier / accumulator / combiner combine partitions. A lambda that mutates a list you created outside the pipeline is still a data race.
What to draw: collection → spliterator (characteristics) → stream → collector → collection. Point at CONCURRENT on the map and ORDERED|SIZED|SUBSIZED on the list.
Cheat sheet
Leave collection.stream() sequential
collection.parallelStream() split; not a lock; not a free speedup
map.keySet()/values()/entrySet().stream() Map is not a Collection
Spliterator
ORDERED DISTINCT SORTED SIZED SUBSIZED NONNULL IMMUTABLE CONCURRENT
trySplit carve a partition or return null
ArrayList ORDERED, SIZED, SUBSIZED
ConcurrentHashMap CONCURRENT, DISTINCT, NONNULL — not ORDERED
HashSet DISTINCT — not ORDERED
List.of IMMUTABLE (typically), ORDERED, SIZED
Land
toCollection(ArrayList::new) name the class, mutable
toList() / toSet() unspecified implementation
toUnmodifiableList/Set/Map Java 10, no nulls, no mutate
Stream.toList() Java 16, unmodifiable, nulls allowed
toMap(key, value) duplicate keys throw (merge in advanced post)
collectingAndThen(c, finisher) collect, then freeze/copy/wrap
Don't parallel forEach + outer ArrayList.add
Don't re-teach map/filter here — basics + advanced posts
Gatherers = custom intermediate ops, not a collector
/posts/java/java-24-stream-gatherers/
Do:
- Open with
stream(); finish withtoCollection/toUnmodifiable*/Stream.toList()depending on mutability and nulls. - Read spliterator characteristics when parallel results look “out of order” or a concurrent map does not throw CME.
- Collect under parallel; do not share an
ArrayListaccumulator.
Don’t:
- Treat
parallelStream()as a thread-safety switch for collectors or forforEachsinks. - Use
collect(toList())when you needed a guaranteedArrayList— name it withtoCollection. - Confuse
Stream.toList()(nulls OK) withList.of/toUnmodifiableList()(nulls forbidden). - Invent peek-buffers for windows; that is a Gatherer.
Wrap-up
A collection opens a stream with stream() or parallelStream(). The Spliterator is what it advertises to that stream: ordered sized splits on ArrayList, concurrent unordered distinct keys on ConcurrentHashMap. You put the result back with a collector that names the type — toCollection, toUnmodifiable*, Stream.toList(), toMap — not with forEach into an outer list. Parallel splits work; parallel does not lock your sink. The pipeline in between is already documented; custom intermediate state is a Gatherer.
First, last, and reversed on the collection you landed in are a type in Java 21, not another stream trick.