Parallel Programming

Reviewed & published by Brayan K

By the end of this lesson you'll take a CPU-bound loop that crawls on one core and spread it across all your cores with Parallel.For, Parallel.ForEach, and PLINQ's .AsParallel() — and you'll do it safely, because you'll know exactly how to stop two threads corrupting the same variable. This is how you turn an 8-core machine into an 8x speed-up instead of a pile of subtle bugs.

Part of the free C# course at LearnCodingFast — hands-on lessons with worked examples and the output they print, plus practice exercises and a quick quiz.

What You'll Learn

💡 Real-World Analogy

Picture a supermarket with one long queue of shoppers. Sequential code is a single checkout till: every shopper waits for the one in front, and the queue only moves as fast as that one cashier. Parallel code is opening more tills — eight cashiers serving the same queue at once, so the line clears roughly eight times faster. Parallel.For and PLINQ are the store manager shouting "open more tills!" — they automatically split the shoppers across however many cashiers (CPU cores) you have. But watch the shared till roll: if two cashiers reach for the same cash drawer at the same moment, the total goes wrong. That collision is a data race, and most of this lesson is about giving each cashier the right kind of drawer.

The mental model: parallel = many workers, async = one worker not waiting

These two ideas get confused constantly, so pin them down now. They solve different problems and you pick by asking one question: is my program busy, or is it waiting?

The golden rule: parallelise CPU work, await I/O work. Running Parallel.ForEach over a list of HTTP calls is a classic mistake — it ties up thread-pool threads doing nothing but waiting, which is exactly what async was built to avoid.

📊 Parallelism Toolbox (and async, for contrast)

ToolUse it forExample
Parallel.ForCPU loop over a number rangeParallel.For(0, n, i => ...)
Parallel.ForEachCPU loop over a collectionParallel.ForEach(items, x => ...)
Parallel.InvokeRun a few different methods at onceParallel.Invoke(A, B, C)
.AsParallel()Make a LINQ query parallel (PLINQ)q.AsParallel().Where(...)
.AsOrdered()Keep original order in PLINQ output.AsParallel().AsOrdered()
InterlockedAtomic counter / sum, lock-freeInterlocked.Add(ref t, x)
lockGuard a multi-step shared updatelock (gate) { ... }
ConcurrentDictionaryThread-safe map, no locks neededdict.AddOrUpdate(k, 1, ...)
async / awaitI/O waiting (the other lesson)await client.GetAsync(url)

Reach for the Parallel/PLINQ rows when the CPU is busy; reach for the last row when the program is just waiting.

1. Parallel.For, Parallel.ForEach & Parallel.Invoke

Parallel.For takes a start and an exclusive end index and runs the loop body across thread-pool threads — each iteration may land on a different core, in any order. Parallel.ForEach does the same over a collection, and Parallel.Invoke fires off several different methods at once. Because iterations run simultaneously, you must never assume an order: in the worked example below every iteration writes to its own array slot (safe), and the printed lines come out scrambled. Read it, run it twice, and notice the order changes but the work always completes.

using System;
using System.Diagnostics;
using System.Threading;
using System.Threading.Tasks;

class Program
{
    // A deliberately heavy, CPU-bound calculation so the cores have real work.
    static double HeavyComputation(int n)
    {
        double result = 0;
        for (int i = 0; i < 200_000; i++)
            result += Math.Sqrt(i * (n + 1));
        return result;
    }

    static void Main()
    {
        const int items = 100;

        // SEQUENTIAL: one core, one item at a time. The times add up.
        var sw = Stopwatch.StartNew();
        double[] seqResults = new double[items];
        for (int i = 0; i < items; i++)
            seqResults[i] = HeavyComputation(i);
        sw.Stop();
        Console.WriteLine($"Sequential for-loop: {sw.ElapsedMilliseconds}ms");

        // Parallel.For — same loop, spread across ALL cores automatically.
        // Each iteration 'i' may run on a different thread, in ANY order.
        // It is SAFE here because every iteration writes to its OWN slot.
        sw.Restart();
        double[] parResults = new double[items];
        Parallel.For(0, items, i =>
        {
            parResults[i] = HeavyComputation(i);   // no shared state -> no race
        });
        sw.Stop();
        Console.WriteLine($"Parallel.For:        {sw.ElapsedMilliseconds}ms (uses all cores)");

        // Parallel.ForEach — same idea but over a collection instead of a range.
        // MaxDegreeOfParallelism caps how many threads run at once (here: 3).
        Console.WriteLine("\n=== Parallel.ForEach (order is NOT deterministic) ===");
        string[] pages = { "page1", "page2", "page3", "page4", "page5" };
        Parallel.ForEach(pages, new ParallelOptions { MaxDegreeOfParallelism = 3 }, page =>
        {
            Console.WriteLine($"  processing {page} on thread {Thread.CurrentThread.ManagedThreadId}");
        });

        // Parallel.Invoke — run several DIFFERENT methods at the same time.
        Console.WriteLine("\n=== Parallel.Invoke (three jobs at once) ===");
        Parallel.Invoke(
            () => Console.WriteLine("  job A done"),
            () => Console.WriteLine("  job B done"),
            () => Console.WriteLine("  job C done")
        );
        Console.WriteLine("All work finished.");
    }
}

Your turn. The program below ships and invoices six orders in parallel. Fill in the three ___ blanks to call Parallel.ForEach and Parallel.For correctly. The printed order will be different every run — that's expected — so you check the count at the end, not the order.

using System;
using System.Threading;
using System.Threading.Tasks;

class Program
{
    static void Main()
    {
        // 🎯 YOUR TURN — fill in the blanks marked with ___
        // Goal: process 6 orders in parallel. Order of printing is RANDOM,
        // so we print a COUNT at the end, not order-sensitive output.

        string[] orders = { "A1", "B2", "C3", "D4", "E5", "F6" };

        // 1) Use Parallel.ForEach to walk the array across threads.
        Parallel.___(orders, order =>          // 👉 method name:  ForEach
        {
            Console.WriteLine($"  shipped {order} on thread {Thread.CurrentThread.ManagedThreadId}");
        });

        // 2) Now use Parallel.For to run iterations 0..(count-1).
        //    Parallel.For's end index is EXCLUSIVE, so pass orders.Length.
        Parallel.For(0, orders.___, i =>       // 👉 the array's item count:  Length
        {
            // each iteration touches its own index 'i' — safe, no shared writes
            Console.WriteLine($"  invoiced item #{i}");
        });

        // Order above will vary every run — but the COUNT never does.
        Console.WriteLine($"\nProcessed {orders.Length} orders in parallel.");

        // ✅ Expected output (order of the 'shipped'/'invoiced' lines varies):
        //    ...6 'shipped' lines in some order...
        //    ...6 'invoiced' lines in some order...
        //
        //    Processed 6 orders in parallel.
    }
}

2. PLINQ — Parallel LINQ

Already know LINQ? Then you already know PLINQ. Drop .AsParallel() into any query and the engine partitions the data across cores and runs your .Where/.Select/.Sum on each chunk. Order-independent reductions like .Count() and .Sum() are the sweet spot — the answer is identical to the sequential version, just faster. The catch: by default the order of streamed results is arbitrary, so add .AsOrdered() when sequence matters.

using System;
using System.Diagnostics;
using System.Linq;

class Program
{
    static bool IsPrime(int n)
    {
        if (n < 2) return false;
        for (int i = 2; i <= Math.Sqrt(n); i++)
            if (n % i == 0) return false;
        return true;
    }

    static void Main()
    {
        int range = 1_000_000;

        // Sequential LINQ — single-threaded.
        var sw = Stopwatch.StartNew();
        int seqCount = Enumerable.Range(2, range).Where(IsPrime).Count();
        sw.Stop();
        Console.WriteLine($"Sequential LINQ: {seqCount} primes in {sw.ElapsedMilliseconds}ms");

        // PLINQ — literally just add .AsParallel() and the query splits
        // the data across cores. .Count() is order-independent, so this is
        // a perfect fit and the result is IDENTICAL to the sequential one.
        sw.Restart();
        int parCount = Enumerable.Range(2, range).AsParallel().Where(IsPrime).Count();
        sw.Stop();
        Console.WriteLine($"PLINQ:           {parCount} primes in {sw.ElapsedMilliseconds}ms");

        // By default PLINQ returns results in ARBITRARY order. If you need the
        // original order back, add .AsOrdered() — small cost, correct sequence.
        Console.WriteLine("\n=== First 10 primes, order preserved with .AsOrdered() ===");
        var firstPrimes = Enumerable.Range(2, 1000)
            .AsParallel()
            .AsOrdered()                 // keep the original ordering
            .Where(IsPrime)
            .Take(10);
        Console.WriteLine(string.Join(", ", firstPrimes));

        // PLINQ also does .Select / .Sum etc. Here we square-and-sum in parallel.
        long sumOfSquares = Enumerable.Range(1, 1000)
            .AsParallel()
            .Select(n => (long)n * n)    // transform on many threads
            .Sum();                      // PLINQ does the safe reduction for you
        Console.WriteLine($"\nSum of squares 1..1000 (PLINQ): {sumOfSquares}");
    }
}

Now you write a PLINQ pipeline. Take 1 to 100, keep the even numbers, square them, and sum the result — all in parallel. Because Sum doesn't care about order, the total is the same on every run. Fill in the four ___ blanks:

using System;
using System.Linq;

class Program
{
    static void Main()
    {
        // 🎯 YOUR TURN — fill in the blanks marked with ___
        // Goal: take 1..100, keep the EVEN numbers, square them, and SUM them
        //       — all in PARALLEL. Sum is order-independent, so PLINQ is ideal
        //       and the TOTAL is the same on every run.

        int[] numbers = Enumerable.Range(1, 100).ToArray();

        long total = numbers
            // 1) Turn the query parallel.
            .___()                       // 👉 method name:  AsParallel
            // 2) Keep only even numbers.
            .Where(n => n % 2 ___ 0)     // 👉 the 'equals zero' operator:  ==
            // 3) Square each remaining number (use long to avoid overflow).
            .Select(n => (long)n ___ n)  // 👉 the multiply operator:  *
            // 4) Add them all up.
            .___();                      // 👉 method name:  Sum

        Console.WriteLine($"Sum of squares of evens 1..100 = {total}");

        // ✅ Expected output:
        //    Sum of squares of evens 1..100 = 171700
    }
}

3. Thread Safety: Races, Interlocked, lock & Concurrent Collections

Here's where parallel code bites. The instant two threads write to the same variable, you have a data race. The classic example is count++: it's secretly three steps — read, add one, write back — so two threads can read the same value and both write back the same result, silently losing increments. The worked example below shows the broken version (the total comes out too low), then three fixes: Interlocked for atomic counters, lock for multi-step updates, and ConcurrentDictionary for a thread-safe map.

using System;
using System.Collections.Concurrent;
using System.Linq;
using System.Threading;
using System.Threading.Tasks;

class Program
{
    static void Main()
    {
        // ❌ WRONG — many threads do unsafeCount++ on the SAME variable.
        // ++ is read-modify-write: two threads can read the same value and
        // both write back the same +1, so increments are LOST (a data race).
        int unsafeCount = 0;
        Parallel.For(0, 1_000_000, _ => { unsafeCount++; });
        Console.WriteLine($"Unsafe ++ (should be 1000000): {unsafeCount}  <- usually too low!");

        // ✅ Fix 1: Interlocked — a single ATOMIC, lock-free increment.
        int safeCount = 0;
        Parallel.For(0, 1_000_000, _ => { Interlocked.Increment(ref safeCount); });
        Console.WriteLine($"Interlocked.Increment:        {safeCount}");

        // ✅ Fix 2: lock — only one thread inside the block at a time.
        long lockTotal = 0;
        object gate = new object();
        Parallel.For(0, 1_000_000, _ =>
        {
            lock (gate) { lockTotal++; }   // correct, but slower than Interlocked
        });
        Console.WriteLine($"lock { } around ++:            {lockTotal}");

        // ✅ Fix 3: ConcurrentDictionary — a thread-safe map. AddOrUpdate
        // tallies hits per key without you writing any locks at all.
        var hits = new ConcurrentDictionary<string, int>();
        Parallel.For(0, 1000, i =>
        {
            string key = $"bucket{i % 4}";
            hits.AddOrUpdate(key, 1, (k, v) => v + 1);   // atomic add-or-update
        });
        Console.WriteLine("\n=== ConcurrentDictionary tally ===");
        foreach (var kvp in hits.OrderBy(k => k.Key))
            Console.WriteLine($"  {kvp.Key}: {kvp.Value} hits");
        Console.WriteLine($"  total: {hits.Values.Sum()}");   // always 1000
    }
}

4. Fast Parallel Sums with Thread-Local State

Locking on every single iteration is correct but slow — the threads spend their time queuing for the lock instead of computing. The professional pattern uses the four-lambda Parallel.For overload: each thread keeps a private running subtotal (zero contention), and only when a thread finishes does it merge its subtotal into the shared grand total once, atomically. Notice the key property — the threads finish in random order, but the final total is exactly the same every run. Determinism of the result, non-determinism of the schedule.

using System;
using System.Threading;
using System.Threading.Tasks;

class Program
{
    static void Main()
    {
        // The 4-lambda Parallel.For overload is the FAST way to total things.
        // Each thread keeps a private 'subtotal' (no contention), then merges
        // its subtotal into the shared grand total ONCE, under Interlocked.
        long grandTotal = 0;

        Parallel.For(
            1, 1_000_001,                 // sum 1..1_000_000 (end is exclusive)
            () => 0L,                     // localInit: each thread starts at 0
            (i, state, subtotal) =>       // body: add i to THIS thread's subtotal
            {
                return subtotal + i;      // no shared writes here -> no contention
            },
            subtotal =>                   // localFinally: merge once per thread
            {
                Interlocked.Add(ref grandTotal, subtotal);
            });

        // The TOTAL is deterministic even though threads finished in random order.
        Console.WriteLine($"Sum 1..1000000 = {grandTotal}");   // 500000500000
        Console.WriteLine($"Formula check  = {1_000_000L * 1_000_001L / 2}");
    }
}

5. Exceptions: AggregateException

When a parallel body throws, the failure can't just bubble up one stack — several threads might fail at once. So the Parallel methods and PLINQ collect every error into a single AggregateException. You catch that one type, then loop over its .InnerExceptions to see each underlying problem. (Remaining iterations may still run, so don't assume the loop stops dead at the first throw.)

using System;
using System.Threading.Tasks;

class Program
{
    static void Main()
    {
        // When a parallel body throws, the failures from all threads are
        // bundled into ONE AggregateException. Catch it and inspect
        // .InnerExceptions to see each underlying error.
        try
        {
            Parallel.For(0, 10, i =>
            {
                if (i == 3 || i == 7)
                    throw new InvalidOperationException($"item {i} failed");
            });
        }
        catch (AggregateException ex)
        {
            Console.WriteLine($"Caught {ex.InnerExceptions.Count} inner exception(s):");
            foreach (var inner in ex.InnerExceptions)
                Console.WriteLine($"  - {inner.Message}");
        }

        Console.WriteLine("Program continues after handling the failures.");
        // ✅ Expected: two inner exceptions (items 3 and 7), in either order.
    }
}

🔎 Deep Dive: when NOT to parallelise

Parallelism is not free. Splitting the work, scheduling threads, and merging results all cost time and memory. If the per-item work is tiny or the collection is small, that overhead can make the parallel version slower than the simple loop. Don't guess — measure with a Stopwatch before and after.

// CPU-bound + independent + big   -> Parallel / PLINQ
Parallel.For(0, 1_000_000, i => Crunch(i));

// I/O-bound (waiting on the network) -> async, NOT Parallel
var pages = await Task.WhenAll(urls.Select(u => http.GetStringAsync(u)));

Putting It Together: a Batch Image Processor

Here's a realistic CPU-bound job: process 200 "images", each needing heavy per-item work. It uses Parallel.ForEach capped to Environment.ProcessorCount cores, collects results in a thread-safe ConcurrentBag, and reports order-independent stats. You understand every line now — the parallelism, the core cap, the thread-safe collection, and why the totals are deterministic.

using System;
using System.Collections.Concurrent;
using System.Linq;
using System.Threading.Tasks;

class Program
{
    // Pretend each "image" needs heavy CPU work (resize/filter).
    static long ProcessImage(int id)
    {
        long checksum = 0;
        for (int i = 0; i < 50_000; i++)
            checksum += (long)Math.Sqrt(i * (id + 1));
        return checksum;
    }

    static void Main()
    {
        int[] imageIds = Enumerable.Range(1, 200).ToArray();

        // Collect per-image results in a THREAD-SAFE collection.
        var results = new ConcurrentBag<long>();

        // Cap parallelism to the core count — more threads than cores just
        // adds context-switching overhead for pure CPU work.
        var options = new ParallelOptions
        {
            MaxDegreeOfParallelism = Environment.ProcessorCount
        };

        Parallel.ForEach(imageIds, options, id =>
        {
            results.Add(ProcessImage(id));   // ConcurrentBag.Add is thread-safe
        });

        // These reductions are order-independent, so the numbers are
        // deterministic even though the images finished in random order.
        Console.WriteLine($"Processed {results.Count} images on up to " +
                          $"{options.MaxDegreeOfParallelism} cores.");
        Console.WriteLine($"Total checksum: {results.Sum()}");
        Console.WriteLine($"Max checksum:   {results.Max()}");
    }
}

The images finish in a random order, but Count, Sum and Max are order-independent — so the reported numbers are identical on every run.

Pro Tips

Common Errors (and the fix)

📋 Quick Reference

TaskCodeNotes
Parallel range loopParallel.For(0, n, i => ...)end is exclusive
Parallel collection loopParallel.ForEach(items, x => ...)order varies
Several methods at onceParallel.Invoke(A, B, C)fire-and-wait
Limit threadsnew ParallelOptions { MaxDegreeOfParallelism = n }cap to cores
Parallel LINQq.AsParallel().Where(...).Sum()PLINQ
Keep order in PLINQ.AsParallel().AsOrdered()small cost
Atomic counterInterlocked.Increment(ref n)lock-free
Atomic addInterlocked.Add(ref total, x)lock-free
Guard a blocklock (gate) { ... }one thread inside
Thread-safe mapdict.AddOrUpdate(k, 1, (k,v)=>v+1)ConcurrentDictionary
Catch parallel errorscatch (AggregateException ex)see InnerExceptions

Frequently Asked Questions

Q: What's the difference between parallel and async?

Parallel uses many threads to do CPU work faster (more cooks). Async uses one thread efficiently while it waits on I/O (one cook not standing idle). Parallelise computation; await network/disk/database. Mixing them up — e.g. Parallel.ForEach over HTTP calls — wastes threads.

Q: Why is my output in a different order every time I run it?

Because parallel iterations run on different threads with no guaranteed schedule. That's normal and expected. If order matters, use PLINQ's .AsOrdered(), or collect into a list and sort it afterwards. For totals and counts, order doesn't matter at all.

Q: My parallel sum gives a different (too-low) total each run — why?

That's a data race. A plain total += x across threads loses updates because += isn't atomic. Use Interlocked.Add(ref total, x), or the thread-local Parallel.For overload that sums private subtotals and merges once. The fixed version is deterministic — same total every run.

Q: Is parallel always faster?

No. Splitting work and scheduling threads has overhead, so for small or cheap workloads the simple loop wins. Parallelism pays off when the data is large and each item does real CPU work. Always measure both with a Stopwatch before committing.

Mini-Challenge: Thread-Safe Parallel Sum

No blanks this time — just a brief and an outline. Sum every number from 1 to 1,000,000 using Parallel.For, but do it safely: a plain total += i is a data race, so make the add atomic with Interlocked.Add (or use the thread-local subtotal overload for extra credit). The whole point: even though the threads run in a random order, the total is identical on every run. Run it and check against the expected line in the comments.

using System;
using System.Threading;
using System.Threading.Tasks;

class Program
{
    static void Main()
    {
        // 🎯 MINI-CHALLENGE: Parallel sum of a range, the THREAD-SAFE way
        // 1. You want the sum of all numbers from 1 to 1_000_000.
        // 2. Declare a shared 'long total = 0;'.
        // 3. Use Parallel.For(1, 1_000_001, i => { ... }) to add each i to
        //    'total'. A plain 'total += i;' is a DATA RACE — instead make the
        //    add atomic with Interlocked.Add(ref total, i).
        //    (Bonus: the faster way is the 4-lambda overload — each thread sums
        //     a local subtotal, then Interlocked.Add merges it once at the end.)
        // 4. Print the total. Even though the threads run in a random ORDER,
        //    the TOTAL is exactly the same every single run — that's the point.
        //
        // ✅ Expected output:
        //    Sum 1..1000000 = 500000500000

        // your code here
    }
}

🎉 Lesson Complete

Practice quiz

When should you reach for Parallel.For / PLINQ rather than async/await?

  • For I/O-bound waiting
  • For network calls
  • For CPU-bound work that keeps cores busy
  • For disk reads

Answer: For CPU-bound work that keeps cores busy. Parallelise CPU-bound work to keep every core busy; use async/await for I/O-bound waiting.

Is the end index of Parallel.For(0, n, ...) inclusive or exclusive?

  • Exclusive
  • Inclusive
  • It depends on the body
  • It loops forever

Answer: Exclusive. Parallel.For's end index is exclusive, just like a standard for loop using i < n.

What turns an ordinary LINQ query into a parallel (PLINQ) one?

  • .ToList()
  • .Where()
  • .Select()
  • .AsParallel()

Answer: .AsParallel(). Dropping .AsParallel() into a query partitions the data across cores.

By default, what is the order of streamed PLINQ results?

  • Always the original order
  • Arbitrary unless you add .AsOrdered()
  • Reverse order
  • Sorted ascending

Answer: Arbitrary unless you add .AsOrdered(). PLINQ returns results in arbitrary order by default; add .AsOrdered() to restore the original sequence.

Why does 'count++' across threads silently lose increments?

  • ++ is a read-modify-write, so two threads can read the same value and both write back the same result
  • ++ is atomic
  • The compiler removes it
  • It overflows

Answer: ++ is a read-modify-write, so two threads can read the same value and both write back the same result. ++ is read-modify-write and not atomic, so concurrent threads can clobber each other's updates — a data race.

Which is the lock-free way to safely increment a shared counter?

  • count++
  • lock (count)
  • Interlocked.Increment(ref count)
  • count = count + 1

Answer: Interlocked.Increment(ref count). Interlocked.Increment performs a single atomic, lock-free increment.

When a body inside Parallel.For throws, how do the failures surface?

  • As the original exception type
  • Bundled into one AggregateException with .InnerExceptions
  • They are swallowed
  • The program crashes instantly

Answer: Bundled into one AggregateException with .InnerExceptions. Failures from all threads are bundled into a single AggregateException; inspect .InnerExceptions for each error.

What does MaxDegreeOfParallelism control?

  • The loop's end index
  • The result order
  • The exception type
  • How many threads run at once

Answer: How many threads run at once. MaxDegreeOfParallelism caps how many threads run concurrently — match it to Environment.ProcessorCount for CPU work.

Why is the four-lambda Parallel.For overload faster for summing?

  • It skips iterations
  • Each thread keeps a private subtotal and merges once, avoiding per-iteration contention
  • It uses no threads
  • It runs sequentially

Answer: Each thread keeps a private subtotal and merges once, avoiding per-iteration contention. Thread-local subtotals have no contention; each thread merges its subtotal into the grand total just once, atomically.

Running Parallel.ForEach over a list of HTTP calls is a mistake because...

  • HTTP is too fast
  • It always throws
  • It ties up thread-pool threads doing nothing but waiting — use async instead
  • It needs a lock

Answer: It ties up thread-pool threads doing nothing but waiting — use async instead. I/O-bound work should use async/await + Task.WhenAll; Parallel blocks thread-pool threads that are merely waiting.

Continue this course

Related lessons