Skip to main content

Command Palette

Search for a command to run...

Java Streams: Simplifying Data Processing

Updated
•8 min read•View as Markdown

firstly let’s discover some data structures.
Arrays : Fixed-size container for elements of the same type.

int[] poolers= {1, 2, 3, 4}; Integer[] talents= {10, 20, 30};

Memory:

  • Stack: reference variable (numbers, integers)

  • Heap: actual array object storing the elements

  • String Pool (if array contains Strings): string literals

Important Methods / Properties:

  • array.length → size of the array

  • Access elements via array[index]

  • Arrays are fixed-size → cannot add or remove elements

Lists (implements Collection) : Ordered collection, allows duplicates, dynamic size.

Common Implementations:

  • ArrayList → backed by a resizable array

  • LinkedList → backed by nodes (doubly-linked) // you can dive more into it i dont have much about it.

Memory:

  • Heap: List object + elements

  • Stack: Reference variable

Important Methods:

  • add(element) → add to list

  • get(index) → get element

  • remove(index) → remove element

  • size() → number of elements

  • contains(element) → check if exists

  • isEmpty() → check if list is empty

Sets (implements Collection) : Unordered collection, no duplicates.

Common Implementations:

  • HashSet → fastest, unordered

  • LinkedHashSet → preserves insertion order // you can dive more.

  • TreeSet → sorted order // you can dive more too haha

Memory:

  • Heap: Set object + elements

  • Stack: Reference variable

Important Methods:

  • add(element) → add to set

  • remove(element) → remove element

  • contains(element) → check existence

  • size() → number of elements

  • isEmpty() → check if set is empty

Maps (key-value pairs, not Collection) : Stores key-value pairs, keys are unique.

Common Implementations:

  • HashMap → unordered, fastest

  • LinkedHashMap → preserves insertion order // you know

  • TreeMap → sorted by key // too

Memory:

  • Heap: Map object + entries (key and value objects)

  • Stack: Reference variable

Important Methods:

  • put(key, value) → add/update entry

  • get(key) → retrieve value

  • remove(key) → delete entry

  • containsKey(key) / containsValue(value)

  • size() → number of entries

  • keySet() / values() / entrySet()

Collections Framework Overview :

Collection (interface) :

┌───────────────┐
│ List          │ -> ArrayList, LinkedList
│ Set           │ -> HashSet, LinkedHashSet, TreeSet
│ Queue / Deque │ -> PriorityQueue, ArrayDeque
└───────────────┘

Map<K,V> : separate interface, not part of Collection

Java Collections (with an s) :

  • Collections is a final class in java.util package.

  • It cannot be instantiated (private constructor).

  • It provides static utility methods to perform common operations on Collection objects (List, Set, etc.), like sorting, searching, reversing, shuffling, synchronizing, etc.

Important Methods

MethodDescriptionExample
sort(List<T> list)Sorts a list in natural orderCollections.sort(list);
sort(List<T> list, Comparator<T> c)Sorts a list using a custom comparatorCollections.sort(list, (a,b)->b-a);
reverse(List<?> list)Reverses order of listCollections.reverse(list);
shuffle(List<?> list)Randomly shuffles elementsCollections.shuffle(list);
min(Collection<? extends T> c)Returns minimum elementCollections.min(list);
max(Collection<? extends T> c)Returns maximum elementCollections.max(list);

When to Use Collections

When you want to manipulate existing collections without writing loops manually.

as we said , we were talking about data structure that holds data.
now we’ll jump into another different concept , which is STREAMS (dakshi dial twitch w kda hh , i kandhk rwah tchouf twitch kidayra)

Stream : not a data structure. It’s a pipeline for processing data from a source (arrays, lists, sets, etc.).

  • Streams support functional-style operations like filter, map, reduce, collect.

  • They don’t store data themselves — they just process it.

List<String> names = List.of("alouhab", "oqritel", "abdeladim","balhbib");

names.stream() // Source: list in heap
    .filter(n -> n.startsWith("a")) // Intermediate: lazy
    .map(String::toUpperCase) // Intermediate: lazy
    .forEach(System.out::println); // Terminal: triggers execution

Memory View

ComponentMemory LocationNotes
Source (names list)HeapOriginal list object
Stream objectHeapReference stored in stack variable if assigned
Lambda objects (n -> n.startsWith("A"))HeapParameters n live temporarily on stack during execution
Terminal operation result (forEach output)DependsMay produce new objects in heap if collect used

Important: Streams are lazy, so nothing happens until a terminal operation is called.

Pipeline Explained :

A pipeline = Source → Intermediate Operations → Terminal Operation

  1. Source: Where the stream comes from (array, list, set).

  2. Intermediate Operations: Transform or filter data.

    • Lazy, returns a new stream, doesn’t execute immediately.

    • Examples: filter, map, sorted, distinct, limit.

  3. Terminal Operation: Triggers execution and produces a result.

    • Examples: collect, reduce, forEach, count.

Diagram:

List<String> names
       |
     Stream
       |
  filter(n -> n.startsWith("a"))
       |
  map(String::toUpperCase)
       |
  collect(Collectors.toList())  // terminal operation triggers execution

Intermediate vs Terminal Operations

TypeDescriptionExamples
IntermediateLazy (mashi lazybob hh), return a new Stream, can chainfilter(), map(), sorted(), distinct(), limit()
Terminalproduce result, ends the pipelineforEach(), collect(), reduce(), count(), anyMatch()

Example:

List<Integer> numbers = List.of(1, 2, 3, 4, 5);

// Intermediate operations
Stream<Integer> stream = numbers.stream()
                                .filter(n -> n % 2 == 0) // lazy
                                .map(n -> n * 2);        // lazy

// Terminal operation triggers execution
int sum = stream.reduce(0, Integer::sum);
System.out.println(sum); // 12

Lambdas in Streams :

  • Lambdas are anonymous functions used as arguments in streams.

  • Stored in heap, reference parameters are on stack.

Example:

numbers.stream()
       .filter(n -> n > 2)      // lambda stored in heap
       .map(n -> n * 10)        // lambda stored in heap
       .forEach(System.out::println);
  • Execution happens element by element; parameters (n) are stack variables during execution.

Important Stream Methods

Intermediate

  • filter(Predicate<T>) → keep elements that satisfy condition

  • map(Function<T,R>) → transform elements

  • distinct() → remove duplicates

  • sorted() → natural or custom sort

  • limit(n) → take first n elements

  • skip(n) → skip first n elements

Terminal

  • forEach(Consumer<T>) → iterate and perform action

  • collect(Collectors.toList()) → collect into a list

  • reduce(BinaryOperator<T>) → combine elements into single result

  • count() → number of elements

  • anyMatch(Predicate<T>) → check if any element matches condition

  • allMatch(Predicate<T>) → check all elements

  • findFirst() / findAny() → get first/any element

Streams vs Collections – Key Differences

FeatureCollectionStream
Stores dataYesNo
ReusableYesNo (one-time use)
IterationExternal (for-loop, iterator)Internal (lambda, functional)
Lazy/EagerEager(does it immediatly)Lazy until terminal op
OperationsCRUD, add/removeTransform/filter/reduce

Quick Example

List<String> words = List.of("apple", "banana", "avocado", "pear");

List<String> result = words.stream()
                           .filter(w -> w.startsWith("a"))
                           .map(String::toUpperCase)
                           .sorted()
                           .toList();

System.out.println(result); // [APPLE, AVOCADO]
  • Memory:

    • words list → heap

    • Stream & lambdas → heap

    • Temporary stack variables during execution → stack

    • Result list → heap

If you stop at an intermediate operation like filter() without a terminal operation, nothing actually happens.
Why?

  • Intermediate operations in streams are lazy.

  • They just describe what to do, they don’t process the data yet.

  • A terminal operation is required to trigger execution.


Example

List<Integer> numbers = List.of(1, 2, 3, 4, 5);

Stream<Integer> stream = numbers.stream()
                                .filter(n -> n % 2 == 0); // just described the filter
  • At this point:

    • No filtering happened yet

    • Stream pipeline exists in memory (heap), but no elements have been processed

    • Nothing is printed, nothing is returned

// Now we add terminal operation
int sum = stream.reduce(0, Integer::sum); // triggers execution
System.out.println(sum); // 6
  • Only when reduce(), collect(), or forEach() is called, the filter is actually applied and results are computed.

Simple analogy:

  • Intermediate operations = instructions on what you want to do

  • Terminal operation = actually run the instructions

do let’s say it again :
A Stream pipeline is the chain of operations (intermediate + terminal) you define on a stream.

Think of it as a recipe or instructions for processing data.

  • Components of a Stream Pipeline:

    1. Source → Where data comes from (array, list, set, map)

    2. Intermediate operations → Transform/filter data (filter, map, distinct, etc.) lazy

    3. Terminal operation → Executes the pipeline (collect, forEach, reduce, etc.) eager


Memory perspective

List<Integer> numbers = List.of(1, 2, 3, 4, 5);

Stream<Integer> stream = numbers.stream()
                                .filter(n -> n % 2 == 0)  // intermediate
                                .map(n -> n * 10);         // intermediate
  • Heap:

    • The stream object exists in memory (pipeline description)

    • Lambdas for filter and map exist in heap

  • Stack:

    • Reference variable stream points to the pipeline
  • Important:

    • No elements have been processed yet → lazy evaluation

    • Actual processing happens only when a terminal operation is called


Adding terminal operation triggers execution

List<Integer> result = stream.collect(Collectors.toList());  // terminal op
  • Now the pipeline is executed:

    1. filter applied to each element

    2. map applied to filtered elements

    3. Result collected into a new list


Simple analogy:

  • Pipeline = the instructions or plan

  • Terminal operation = the “go” button that actually runs it

Parallel Streams in Java

  • parallelStream() is a special kind of stream that splits data processing across multiple CPU cores automatically.

  • It’s part of the Streams API and works like a regular stream, but operations run in parallel.


When to Use parallelStream()

Good for:

  • Large datasets (thousands/millions of elements)

  • CPU-intensive operations (complex calculations, transformations)

  • When order of results doesn’t matter (unless using forEachOrdered)

❌ Avoid for:

  • Small collections → overhead of parallelism may slow it down

  • Operations with side effects (modifying shared variables) → can cause concurrency issues


Example

List<Integer> numbers = List.of(1, 2, 3, 4, 5, 6, 7, 8);

// Parallel processing
int sum = numbers.parallelStream()
                 .filter(n -> n % 2 == 0)
                 .map(n -> n * 2)
                 .reduce(0, Integer::sum);

System.out.println(sum); // 40
  • The stream splits the elements into chunks, processes them on multiple threads, then combines results.

💡 Tip:

  • Always test performance! Sometimes parallelStream() doesn’t improve speed for small collections.

THANK YOU.
@louhabali.