Concurrency in C++

Reviewed & published by Brayan K

By the end of this lesson you'll be able to run work on several threads at once with std::thread, pass data in safely, protect shared state from data races with a std::mutex, and collect results back with std::async and std::future — the toolkit behind every fast, responsive C++ program.

Part of the free C++ course at LearnCodingFast — hands-on lessons with worked examples and the output they print, plus practice exercises and a quick quiz.

What You'll Learn

💡 Real-World Analogy

Think of your program as a kitchen. A single-threaded program is one chef doing every step in order. Threads are extra chefs working at the same time — the meal comes out faster. But if two chefs reach for the same chopping board at once, they collide: that's a data race. The fix is a mutex — a single "talking stick" only one chef may hold while using the board. And std::async is like handing a chef a ticket: you walk away, and later redeem the ticket for the finished dish (a std::future).

1. Creating Threads with std::thread

A thread is a second line of execution that runs at the same time as your main(). You create one by handing std::thread a function plus any arguments — it starts running immediately. Two rules matter: every thread must be join()ed (wait for it to finish) or detach()ed (let it run on its own), and arguments are copied by default. To let a thread modify one of your variables, wrap it in std::ref so it's passed by reference. Read this worked example and run it.

#include <iostream>
#include <thread>   // std::thread lives here
#include <string>
using namespace std;

// A plain function each thread will run.
void greet(const string& name, int times) {
    for (int i = 0; i < times; i++) {
        cout << "Hello from " << name << " #" << i << "\n";
    }
}

// A thread that MODIFIES a variable owned by main().
// It takes int& so it can write back to the caller's variable.
void addTo(int& total, int amount) {
    total += amount;   // changes main()'s 'total'
}

int main() {
    // 1) Create a thread. It starts running greet(...) IMMEDIATELY,
    //    in parallel with main(). Extra arguments are passed after the function.
    thread worker(greet, "Worker", 3);

    // 2) join() = "wait here until that thread finishes".
    //    Skipping this is a crash (see Common Errors below).
    worker.join();   // main() pauses until 'worker' is done

    // 3) By default, std::thread COPIES its arguments. To let a thread
    //    change one of YOUR variables, wrap it in std::ref (a reference).
    int total = 0;
    thread adder(addTo, ref(total), 10);  // ref(total) -> pass by reference
    adder.join();                         // wait, THEN read the result
    cout << "total after thread = " << total << "\n";  // 10

    // hardware_concurrency() hints how many threads run truly in parallel.
    cout << "cores available: " << thread::hardware_concurrency() << "\n";

    return 0;
}
// ✅ Output (the greet lines may appear in any order relative to nothing,
//    but here only one thread prints, so):
//    Hello from Worker #0
//    Hello from Worker #1
//    Hello from Worker #2
//    total after thread = 10
//    cores available: 8   (varies by machine)

// ⚠️ No expected-output panel for this one, on purpose.
// Threads interleave differently on every run and on every machine, so
// there is no output to promise.
//
// One real run on the machine that builds this site printed:
//    Hello from Worker #0
//    Hello from Worker #1
//    Hello from Worker #2
//    total after thread = 10
//    cores available: 4

Your turn. The program below starts one thread but is missing two things: how to pass the variable by reference, and how to wait for the thread. Fill in the blanks marked ___, then run it.

#include <iostream>
#include <thread>
using namespace std;

void doubleIt(int& n) {
    n = n * 2;   // writes back through the reference
}

int main() {
    // 🎯 YOUR TURN — fill in the blanks marked with ___

    int value = 21;

    // 1) Start a thread that runs doubleIt and passes 'value' BY REFERENCE
    //    so the thread can change it. Reference wrapper = ref(...).
    thread t(doubleIt, ___);   // 👉 ref(value)

    // 2) Wait for the thread to finish before you read 'value'.
    t.___;                     // 👉 join()

    cout << "value = " << value << "\n";

    // ✅ Expected output:  value = 42
}

2. Data Races & std::mutex

When two threads touch the same variable and at least one writes, you have a data race. The trap: counter++ looks like one step but is really three — read, add one, write back. Two threads can interleave those steps and lose updates, giving a wrong, random answer. The C++ standard calls a data race undefined behaviour, which means anything can happen. The fix is a std::mutex (short for "mutual exclusion"): a lock only one thread can hold at a time. Wrap it in a std::lock_guard, which locks on creation and automatically unlocks when it goes out of scope.

#include <iostream>
#include <thread>
#include <mutex>    // std::mutex, std::lock_guard
using namespace std;

int counter = 0;       // shared mutable state — DANGEROUS without a lock
mutex mtx;             // the "talking stick": only one thread holds it at a time

// UNSAFE: many threads do counter++ at once. counter++ is really
// read -> add 1 -> write, and two threads can interleave those steps,
// losing updates. This is a DATA RACE (undefined behaviour).
void raceIncrement() {
    for (int i = 0; i < 100000; i++) counter++;   // ❌ no protection
}

// SAFE: lock_guard locks 'mtx' on construction and unlocks when it goes
// out of scope (even if an exception is thrown). Only one thread is
// inside the critical section at a time, so no updates are lost.
void safeIncrement() {
    for (int i = 0; i < 100000; i++) {
        lock_guard<mutex> guard(mtx);   // lock now, auto-unlock at }
        counter++;                      // protected: one thread at a time
    }
}

int main() {
    // Run the SAFE version on two threads.
    counter = 0;
    thread a(safeIncrement);
    thread b(safeIncrement);
    a.join();
    b.join();
    cout << "safe counter = " << counter << "\n";   // always 200000

    // The unsafe version would print a wrong, RANDOM number < 200000
    // (different every run) because updates collide. Try it to see!
    return 0;
}
// ✅ Output:  safe counter = 200000

// ⚠️ No expected-output panel for this one, on purpose.
// Threads interleave differently on every run and on every machine, so
// there is no output to promise.
//
// One real run on the machine that builds this site printed:
//    safe counter = 200000

3. Getting Results Back: std::async & std::future

A raw std::thread runs a void function — it can't easily hand you a return value. std::async solves this. You give it a function and arguments; it runs them on another thread and returns a std::future, a "ticket" you redeem later by calling .get(). .get() blocks until the result is ready, then returns it (and rethrows any exception the task threw). For sharing a single value without a mutex, std::atomic makes operations like ++ indivisible at the hardware level — perfect for a shared counter or flag.

#include <iostream>
#include <future>   // std::async, std::future
#include <atomic>   // std::atomic
using namespace std;

// A function that returns a value (unlike a void thread function).
long long sumTo(long long n) {
    long long total = 0;
    for (long long i = 1; i <= n; i++) total += i;
    return total;
}

int main() {
    // std::async launches sumTo on another thread and immediately hands you
    // a std::future — a "ticket" you redeem later for the return value.
    future<long long> ticket = async(launch::async, sumTo, 1000000);

    // main() is free to do other work while sumTo runs in parallel.
    cout << "Working on something else...\n";

    // .get() blocks until the result is ready, then returns it.
    // (You can only call get() once per future.)
    long long answer = ticket.get();
    cout << "sum 1..1000000 = " << answer << "\n";   // 500000500000

    // Bonus: std::atomic is a lock-free way to share a SINGLE value safely.
    // No mutex needed — the ++ is guaranteed indivisible by the hardware.
    atomic<int> hits{0};
    auto f1 = async(launch::async, [&hits]() { for (int i = 0; i < 1000; i++) hits++; });
    auto f2 = async(launch::async, [&hits]() { for (int i = 0; i < 1000; i++) hits++; });
    f1.get(); f2.get();             // wait for both
    cout << "atomic hits = " << hits << "\n";   // always 2000

    return 0;
}
// ✅ Output:
//    Working on something else...
//    sum 1..1000000 = 500000500000
//    atomic hits = 2000

// ⚠️ No expected-output panel for this one, on purpose.
// Threads interleave differently on every run and on every machine, so
// there is no output to promise.
//
// One real run on the machine that builds this site printed:
//    Working on something else...
//    sum 1..1000000 = 500000500000
//    atomic hits = 2000

Now you try. Launch a function on another thread with std::async and pull the answer back out of the future. Fill in the two blanks:

#include <iostream>
#include <future>
using namespace std;

int square(int x) {
    return x * x;
}

int main() {
    // 🎯 YOUR TURN — fill in the blanks marked with ___

    // 1) Launch square(9) on another thread with std::async.
    //    Store the ticket in a future<int>.
    future<int> result = async(launch::async, ___, 9);  // 👉 square

    // 2) Redeem the ticket: block until the answer is ready and read it.
    int answer = result.___;   // 👉 get()

    cout << "9 squared = " << answer << "\n";

    // ✅ Expected output:  9 squared = 81
}

🔎 Deep Dive: which tool do I reach for?

std::thread — a manual background worker. You control its life and must join() or detach() it. Best for long-running tasks that don't return a value.

std::async + std::future — fire off a task that returns something (or might throw) and collect it later with .get(). Less bookkeeping than a raw thread.

std::mutex + std::lock_guard — protect a block of code that touches shared data so only one thread runs it at a time.

std::atomic — a single shared value (a counter, a flag) updated safely without a lock.

Pro Tips

Common Errors (and the fix)

📋 Quick Reference

TaskCode
Create a threadstd::thread t(func, args...);
Wait for itt.join();
Run it independentlyt.detach();
Pass by referencestd::thread t(f, std::ref(x));
Protect shared datastd::lock_guard<std::mutex> g(mtx);
Run task, get resultauto f = std::async(launch::async, fn);
Read the resultauto v = f.get();
Lock-free counterstd::atomic<int> n{0};

Mini-Challenge: Safe Shared Total

No blanks this time — just a brief and an outline to keep you on track. Two threads add to the same total, so the std::mutex is what keeps the answer correct. Build it, run it, and check your output against the expected line in the comments.

#include <iostream>
#include <thread>
#include <mutex>
using namespace std;

int main() {
    // 🎯 MINI-CHALLENGE: Safe shared total
    // 1. Declare an int "sum" set to 0 and a std::mutex called "mtx".
    // 2. Write a lambda that loops 50000 times and adds i to "sum",
    //    protecting each add with a lock_guard<mutex>(mtx).
    // 3. Run that lambda on TWO std::threads (t1 and t2).
    // 4. join() BOTH threads, then print "sum = " << sum.
    //
    // Why the lock? Without it, t1 and t2 race on "sum" and you get a
    // wrong, random total every run.
    //
    // ✅ Expected output:  sum = 2499950000

    // your code here
}

🎉 Lesson Complete

Practice quiz

What does calling join() on a std::thread do?

  • Starts the thread
  • Kills the thread immediately
  • Makes the calling thread wait until that thread finishes
  • Detaches the thread

Answer: Makes the calling thread wait until that thread finishes. join() blocks until the thread completes, then cleans it up.

What happens if a joinable std::thread is destroyed without join() or detach()?

  • The program calls std::terminate
  • Nothing, it cleans up
  • It silently leaks
  • It auto-joins

Answer: The program calls std::terminate. Every thread must be joined or detached before destruction, or std::terminate is called.

By default, how are arguments passed to a std::thread's function?

  • By reference
  • By pointer
  • By move only
  • They are copied

Answer: They are copied. Arguments are copied; to let a thread modify your variable, wrap it in std::ref so it's passed by reference.

What is a data race?

  • Two threads finishing at the same time
  • Two+ threads accessing the same memory with at least one writing, and no synchronization
  • A thread that runs too fast
  • A loop that never ends

Answer: Two+ threads accessing the same memory with at least one writing, and no synchronization. Unsynchronized concurrent access with at least one writer is a data race — undefined behaviour in C++.

Why does std::lock_guard make protecting shared data safer than locking by hand?

  • It locks on construction and automatically unlocks when it goes out of scope, even on exceptions
  • It is faster
  • It never needs a mutex
  • It allows many threads in at once

Answer: It locks on construction and automatically unlocks when it goes out of scope, even on exceptions. RAII: lock_guard unlocks automatically at the end of the scope, so you can't forget to unlock.

What does std::async return so you can collect a task's result later?

  • A std::thread
  • A std::mutex
  • A std::future you redeem with .get()
  • void

Answer: A std::future you redeem with .get(). async returns a std::future; calling .get() blocks until the result is ready and returns it (rethrowing any exception).

How many times can you call .get() on a single std::future?

  • Unlimited
  • Exactly once
  • Twice
  • Once per thread

Answer: Exactly once. A future delivers its value once; calling get() again is invalid.

When is std::atomic enough instead of a std::mutex?

  • Always
  • When you must update several values together
  • Never
  • For a single shared value like a counter or flag

Answer: For a single shared value like a counter or flag. atomic makes individual operations on one value indivisible; use a mutex when several related values must stay consistent.

Why does multi-threaded cout output come out in a different order each run?

  • A compiler bug
  • The OS schedules threads independently, so the order they reach cout is non-deterministic
  • cout is broken
  • Threads always reverse order

Answer: The OS schedules threads independently, so the order they reach cout is non-deterministic. Thread scheduling is non-deterministic; never rely on thread output ordering.

Two threads each hold one mutex and wait for the other's. What is this called, and what prevents it?

  • A data race; use atomic
  • A spin; use detach()
  • A deadlock; always lock multiple mutexes in the same order (or use scoped_lock)
  • A leak; use join()

Answer: A deadlock; always lock multiple mutexes in the same order (or use scoped_lock). Circular waiting is a deadlock; consistent lock ordering or std::scoped_lock taking them together avoids it.

Continue this course

Frequently asked questions

What is the difference between join() and detach()?

join() makes the calling thread wait until the other thread finishes, then cleans it up. detach() lets the thread run independently in the background and severs the handle. Every std::thread must be either joined or detached before it is destroyed, or the program calls std::terminate.

What is a data race and why is it dangerous?

A data race happens when two or more threads access the same memory at the same time and at least one is writing, with no synchronisation. Operations like counter++ are not indivisible, so threads interleave and lose updates. The C++ standard calls this undefined behaviour: the program may produce wrong results, crash, or appear to work and fail later. Protect shared data with a std::mutex or use std::atomic.

When should I use std::async instead of std::thread?

Use std::async when your task RETURNS a value or might throw — async hands you a std::future that delivers the result (or rethrows the exception) when you call get(), and it manages the thread for you. Use std::thread when you want full manual control over a long-running background worker.

Do I still need a mutex if I use std::atomic?

Not for a single value. std::atomic<int> makes individual reads, writes, and operations like ++ indivisible without a lock, so it is perfect for a shared counter or flag. Reach for a std::mutex when you must keep several related values consistent together, or guard a larger critical section.

Why does the output from my threads come out in a different order each run?

Because the operating system schedules threads independently, the order in which they reach a cout statement is non-deterministic — it can change from run to run and machine to machine. Never rely on thread output ordering. If you need ordering, synchronise explicitly (for example with join() or a condition variable).

Related lessons