Atomic Operations
Reviewed & published by Brayan K
By the end of this lesson you'll be able to build a thread-safe counter with no mutex, update shared values lock-free with compare-and-swap, write a spinlock from an atomic_flag, and choose the right memory ordering for the job — the toolkit behind high-performance concurrent C++.
Part of the free C++ course at LearnCodingFast — hands-on lessons with worked examples and the output they print, plus practice exercises and a quick quiz.
What You'll Learn
- Why a shared int++ across threads is a data race — and how atomic<int> fixes it
- Use fetch_add, load, store, and exchange for indivisible updates
- Apply compare_exchange_weak/strong (CAS) for lock-free read-modify-write
- Build a spinlock from std::atomic_flag with test_and_set / clear
- Tell relaxed, acquire/release, and seq_cst ordering apart
- Know when to pick atomics over a mutex (and when not to)
💡 Real-World Analogy
A normal counter++ is three steps: read the number, add one, write it back. Picture two people sharing one whiteboard tally. Both read "5", both write "6" — and one count vanished. That's a data race. An atomic operation is like a turnstile counter: each click is one sealed, indivisible action that the hardware guarantees can't overlap another. No locking, no waiting — just a count that's always exactly right. A mutex, by contrast, is the deli ticket queue: safe, but you wait your turn. Atomics skip the queue for the small jobs.
1. Atomic Counters vs Data Races
An atomic<int> guarantees every read and write is indivisible — no other thread can ever catch it half-finished. A plain int shared across threads has no such promise: counter++ compiles to read-add-write, and two threads can interleave those steps and lose updates. The worked example below races eight threads at both kinds of counter so you can see the difference. Read every comment, then run it.
#include <iostream>
#include <atomic> // std::atomic, std::memory_order
#include <thread> // std::thread
#include <vector>
using namespace std;
// An atomic<int> can be read/modified by many threads at once
// WITHOUT a data race. Each operation is indivisible: no thread can
// ever see it "half done".
atomic<int> safeCounter{0}; // {0} sets the starting value to 0
int plainCounter = 0; // a normal int — NOT safe to share
// Each thread runs this. It bumps both counters n times.
void work(int n) {
for (int i = 0; i < n; i++) {
safeCounter.fetch_add(1); // atomic +1 (same as safeCounter++)
plainCounter++; // DATA RACE: read, +1, write — not atomic
}
}
int main() {
const int THREADS = 8;
const int PER_THREAD = 100000;
// Launch 8 threads, all hammering the same two counters.
vector<thread> pool;
for (int t = 0; t < THREADS; t++)
pool.emplace_back(work, PER_THREAD);
for (auto& th : pool) th.join(); // wait for all to finish
int expected = THREADS * PER_THREAD; // 800000
// safeCounter is ALWAYS exact — atomic adds can't collide.
cout << "Atomic counter: " << safeCounter.load()
<< " (expected " << expected << ")\n";
// plainCounter is almost always TOO LOW — increments were lost
// when two threads read the same old value before writing back.
cout << "Plain counter: " << plainCounter
<< " (expected " << expected << ", usually wrong)\n";
return 0;
}
// Expected:
// Atomic counter: 800000 (expected 800000)
// Plain counter: <some number < 800000> (expected 800000, usually wrong)
//
// Run-note: fetch_add() above uses the DEFAULT ordering (seq_cst).
// For a pure counter, memory_order_relaxed is also correct and faster,
// because you only need the count to be atomic, not ordered against
// other variables. We cover ordering in section 4 — when in doubt,
// keep the default.
// ⚠️ No expected-output panel for this one, on purpose.
// The plain counter is a data race: each run loses a different number of
// increments. The atomic one is always 800000, and the gap between the two
// is the whole point.
//
// One real run on the machine that builds this site printed:
// Atomic counter: 800000 (expected 800000)
// Plain counter: 248006 (expected 800000, usually wrong)Your turn. The program below counts hits across four threads — fill in the two blanks marked ___ to make the counter atomic and read its final value.
#include <iostream>
#include <atomic>
#include <thread>
#include <vector>
using namespace std;
// 🎯 YOUR TURN — replace each ___ then press "Try it Yourself".
atomic<int> hits{0}; // a shared, thread-safe counter
void countHits(int n) {
for (int i = 0; i < n; i++) {
// 1) Atomically add 1 to 'hits'
___; // 👉 hits.fetch_add(1); (or hits++;)
}
}
int main() {
vector<thread> pool;
for (int t = 0; t < 4; t++)
pool.emplace_back(countHits, 250);
for (auto& th : pool) th.join();
// 2) Atomically read the final value of 'hits'
cout << "Total hits: " << ___ << "\n"; // 👉 hits.load()
// ✅ Expected output:
// Total hits: 1000
return 0;
}2. load, store, exchange & fetch_add
You don't read an atomic with plain = — you call methods so the intent (and the ordering) is explicit. load() reads, store(x) writes, and exchange(x) writes a new value while handing you the old one, all atomically. fetch_add(n) and fetch_sub(n) add or subtract and return the value before the change. These five cover almost everything you'll do with a single atomic.
#include <iostream>
#include <atomic>
using namespace std;
int main() {
atomic<int> val{10};
// load() reads the current value atomically.
cout << "load(): " << val.load() << "\n"; // 10
// store() writes a new value atomically.
val.store(42);
cout << "after store: " << val.load() << "\n"; // 42
// exchange() writes a NEW value and returns the OLD one,
// both in a single atomic step.
int previous = val.exchange(99);
cout << "exchange: old=" << previous
<< " new=" << val.load() << "\n"; // old=42 new=99
// fetch_add / fetch_sub return the value BEFORE the change.
int before = val.fetch_add(1); // before=99
cout << "fetch_add: before=" << before
<< " now=" << val.load() << "\n"; // now=100
return 0;
}
// ✅ Expected output:
// load(): 10
// after store: 42
// exchange: old=42 new=99
// fetch_add: before=99 now=1003. Compare-and-Swap (CAS) for Lock-Free Updates
fetch_add only handles simple arithmetic. For anything else — "double it", "set it only if it's still what I last saw" — you need compare-and-swap. compare_exchange_strong(expected, desired) swaps in desired only if the value still equals expected; otherwise it loads the real current value back into expected and returns false. That refresh-on-failure is why CAS lives in a loop: read, compute, try to swap, and retry if someone beat you to it. Use compare_exchange_weak inside that loop (it can fail spuriously but is faster); use strong when you aren't looping.
#include <iostream>
#include <atomic>
using namespace std;
int main() {
// compare_exchange does: "IF the value equals 'expected',
// replace it with 'desired' and return true. OTHERWISE write the
// ACTUAL current value back into 'expected' and return false."
// This is the building block of every lock-free algorithm.
atomic<int> val{100};
int expected = 100;
bool ok = val.compare_exchange_strong(expected, 200);
cout << "CAS 100->200: " << (ok ? "ok" : "fail")
<< ", val=" << val.load() << "\n"; // ok, val=200
// Now 'expected' is stale (we think it's still 100, but it's 200).
expected = 100;
ok = val.compare_exchange_strong(expected, 300);
cout << "CAS 100->300: " << (ok ? "ok" : "fail")
<< ", expected now=" << expected << "\n"; // fail, expected now=200
// The CAS LOOP: how you do a lock-free read-modify-write.
// Here we atomically double the value, retrying on collision.
int cur = val.load();
while (!val.compare_exchange_weak(cur, cur * 2)) {
// compare_exchange_weak may fail spuriously, so we loop.
// On failure 'cur' is refreshed to the real value for us.
}
cout << "doubled: " << val.load() << "\n"; // 400
return 0;
}
// ✅ Expected output:
// CAS 100->200: ok, val=200
// CAS 100->300: fail, expected now=200
// doubled: 400Now you try. Fill in the three blanks to swap in a value with exchange, then bump it once with a CAS.
#include <iostream>
#include <atomic>
using namespace std;
int main() {
// 🎯 YOUR TURN — replace each ___ then press "Try it Yourself".
atomic<int> level{1};
// 1) Swap in the value 5 and capture what was there before.
int old = ___; // 👉 level.exchange(5)
cout << "old level was " << old << "\n"; // old level was 1
// 2) Try to bump 5 -> 6 with CAS. Fill the expected value (5)
// and the desired value (6).
int expected = ___; // 👉 5
level.compare_exchange_strong(expected, ___); // 👉 6
cout << "new level is " << level.load() << "\n"; // new level is 6
// ✅ Expected output:
// old level was 1
// new level is 6
return 0;
}4. Spinlocks & Memory Ordering
A std::atomic_flag is the simplest atomic — just set or clear, and the only type guaranteed lock-free on every platform. With its test_and_set / clear pair you can build a spinlock: a lock that busy-loops instead of sleeping, ideal for critical sections so short that sleeping would cost more than spinning.
Every atomic operation also takes a memory ordering that controls how it's ordered against other memory accesses across threads. The three you'll meet: memory_order_relaxed (atomic, but no cross-thread ordering — fine for a lone counter), acquire/release (a paired handshake: a release store publishes everything written before it to whoever does an acquire load — this is what the spinlock uses), and seq_cst (the default: one global order, easiest to reason about, slightly slower). Default to seq_cst and only relax when profiling demands it.
#include <iostream>
#include <atomic>
#include <thread>
#include <vector>
using namespace std;
// A spinlock is the simplest lock: "spin" (loop) until you grab it.
// atomic_flag is the only type guaranteed lock-free everywhere.
class SpinLock {
atomic_flag flag = ATOMIC_FLAG_INIT; // starts "clear" (unlocked)
public:
void lock() {
// test_and_set returns the OLD value and sets the flag.
// While it returns true, someone else holds the lock — keep spinning.
// acquire ordering: nothing after lock() can move before it.
while (flag.test_and_set(memory_order_acquire)) {
// busy-wait — cheap for VERY short critical sections only
}
}
void unlock() {
// release ordering: everything before unlock() is visible to
// the next thread that acquires the lock.
flag.clear(memory_order_release);
}
};
SpinLock spin;
int shared = 0; // protected by 'spin'
void addMany() {
for (int i = 0; i < 10000; i++) {
spin.lock();
shared++; // safe: only the lock holder runs this
spin.unlock();
}
}
int main() {
vector<thread> pool;
for (int t = 0; t < 4; t++) pool.emplace_back(addMany);
for (auto& th : pool) th.join();
cout << "shared = " << shared << " (expected 40000)\n";
return 0;
}
// Expected:
// shared = 40000 (expected 40000)
//
// Run-note: the acquire on lock() pairs with the release on unlock().
// That pairing is what publishes 'shared' safely between threads —
// using memory_order_relaxed here would compile but be WRONG.
// ⚠️ No expected-output panel for this one, on purpose.
// Threads interleave differently on every run and on every machine, so
// there is no output to promise.
//
// One real run on the machine that builds this site printed:
// shared = 40000 (expected 40000)Pro Tips
- 💡 Default to seq_cst: the standard ordering is the safest. Only drop to acquire/release or relaxed when a profiler shows a real bottleneck.
- 💡 Atomics are per-operation: they make one access indivisible. To update two values together as a unit, you still need a mutex.
- 💡 Check is_lock_free(): if it returns false, the library is using a hidden mutex behind your atomic — you lost the speed win.
- 💡 Prefer compare_exchange_weak in loops: it's faster and you're retrying anyway; use strong only when you're not looping.
Common Errors (and the fix)
- Non-atomic shared counter (the data race): a plain int counter; incremented from several threads loses updates and is undefined behaviour. Make it atomic<int> and use fetch_add(1) (or counter++).
- Misusing memory_order_relaxed: relaxed makes the op atomic but gives no cross-thread ordering, so using it to publish other data (set data, then flip a relaxed ready flag) lets the reader see ready before data. Use release on the writer and acquire on the reader.
- The ABA problem: CAS only checks the value equals expected — if it went A → B → A in between, the CAS still succeeds, which can corrupt lock-free structures that reuse memory. Guard with a version/tag counter that only ever increments.
- Compound atomicity assumed: atomic<int> a, b; if (a.load() > b.load()) reads two atomics in two steps — they can change between them. Wrap the whole check in a mutex if it must be consistent.
- Reading an atomic with = expecting a copy: atomic<int> b = a; won't compile — atomics aren't copyable. Use atomic<int> b{a.load()};.
📋 Quick Reference — Memory Orders
| Ordering | Guarantee | Use it for |
|---|---|---|
| relaxed | Atomic only; no cross-thread ordering | Standalone counters, statistics |
| acquire | No later read/write moves before this load | The reader/consumer side |
| release | No earlier read/write moves after this store | The writer/producer side |
| acq_rel | Acquire + release on one read-modify-write | CAS in a lock-free loop |
| seq_cst | Single global order across all threads (default) | Everything, until proven a bottleneck |
acquire on a load pairs with release on the matching store — that pairing is what safely publishes data from one thread to another. Mismatch them and the compiler/CPU is free to reorder your code.
Mini-Challenge: a Shared Stop Flag
No blanks this time — just a brief and an outline. Use an atomic<bool> to tell a worker thread when to stop and an atomic<int> to count how far it got. This is the everyday pattern for cleanly shutting a thread down. Build it, run it, and check the shape of your output against the example.
#include <iostream>
#include <atomic>
#include <thread>
#include <vector>
using namespace std;
// 🎯 MINI-CHALLENGE: a shared "stop" flag
// 1. Declare an atomic<bool> called "stop" initialised to false.
// 2. Declare an atomic<int> called "ticks" initialised to 0.
// 3. Start ONE worker thread that loops: while (!stop) ticks++;
// (load stop atomically each time; fetch_add or ++ on ticks).
// 4. In main, sleep briefly, set stop = true, join the worker.
// 5. Print: "Worker ran <ticks> times before stopping".
//
// ✅ Expected (example): Worker ran 5821394 times before stopping
// (the exact number varies — any large count is correct)
int main() {
// your code here
return 0;
}🎉 Lesson Complete
- ✅ A shared plain int incremented by many threads is a data race; atomic<int> makes each update indivisible
- ✅ load, store, exchange, and fetch_add/fetch_sub cover single-atomic work
- ✅ compare_exchange_weak/strong (CAS) power lock-free read-modify-write loops
- ✅ std::atomic_flag with test_and_set/clear builds a spinlock
- ✅ Memory order: relaxed (no ordering), acquire/release (paired handshake), seq_cst (the safe default)
- ✅ Atomics are per-operation; reach for a mutex when several values must change together, and watch for the ABA problem
Practice quiz
Why is a plain 'int' shared across threads and incremented with counter++ unsafe?
- int is too small
- int cannot be shared at all
- counter++ is read-add-write, so threads can interleave and lose updates
- It is only unsafe on one core
Answer: counter++ is read-add-write, so threads can interleave and lose updates. The non-atomic read-modify-write can interleave between threads, losing increments — a data race and undefined behaviour.
What guarantee does an atomic<int> operation provide?
- Each operation is indivisible; no thread sees it half-done
- It is faster than a normal int always
- It locks all other threads out of the program
- It orders all memory in the program
Answer: Each operation is indivisible; no thread sees it half-done. Atomic operations are indivisible, so concurrent reads/writes never collide mid-operation.
What does exchange(x) on an atomic do?
- Reads x without changing it
- Adds x to the value
- Compares and swaps
- Writes the new value x and returns the OLD value, atomically
Answer: Writes the new value x and returns the OLD value, atomically. exchange writes the new value and hands back the previous one in a single atomic step.
fetch_add(1) on an atomic returns:
- The new value after adding
- The value BEFORE the addition
- Always 1
- void
Answer: The value BEFORE the addition. fetch_add (and fetch_sub) return the value as it was before the change.
compare_exchange_strong(expected, desired) when the value does NOT equal expected:
- Writes the actual current value into 'expected' and returns false
- Throws an exception
- Sets the value to desired anyway
- Blocks until they match
Answer: Writes the actual current value into 'expected' and returns false. On failure it refreshes 'expected' with the real current value and returns false — which is why CAS lives in a loop.
Why prefer compare_exchange_weak inside a retry loop?
- It never fails
- It is the only one that compiles in a loop
- It may fail spuriously but is faster, and you are looping anyway
- It locks the atomic
Answer: It may fail spuriously but is faster, and you are looping anyway. weak can fail for no real reason on some hardware, but is cheaper; in a loop the spurious failure just retries.
What is the ABA problem?
- Two atomics with the same name
- A value goes A→B→A, so a CAS succeeds even though state changed and changed back
- Adding before subtracting
- A deadlock between atomics
Answer: A value goes A→B→A, so a CAS succeeds even though state changed and changed back. CAS only checks the value still equals 'expected', not that it never changed; a version counter is the classic fix.
What does memory_order_relaxed guarantee?
- A single global order across threads
- Nothing at all
- Acquire-release semantics
- The operation is atomic, but gives no cross-thread ordering against other variables
Answer: The operation is atomic, but gives no cross-thread ordering against other variables. Relaxed makes the op itself atomic but adds no ordering — fine for a standalone counter, wrong for publishing other data.
In a spinlock, lock() uses memory_order_acquire and unlock() uses release. Why?
- For speed only
- The acquire/release pair publishes the protected data safely between threads
- To make the lock recursive
- Because relaxed would not compile
Answer: The acquire/release pair publishes the protected data safely between threads. The release on unlock pairs with the acquire on the next lock, so writes inside the critical section become visible.
Which type is guaranteed lock-free on every platform and is used to build a spinlock?
- std::atomic<int>
- std::mutex
- std::atomic_flag
- std::atomic<bool>
Answer: std::atomic_flag. std::atomic_flag with test_and_set / clear is the only atomic type guaranteed lock-free everywhere.
Continue this course
- Previous: Mutexes, Locks, Deadlock Avoidance & Thread-safe Patterns
- Next: Memory Allocation Internals: new/delete, allocators & custom allocators — How new/delete work and how to write custom allocators for performance
- Quick reference: C++ cheat sheet
Frequently asked questions
When should I use std::atomic instead of a mutex?
Reach for atomics when you are protecting a single, simple value such as a counter, a flag, or a pointer that swaps as a unit. They are lock-free and far faster than a mutex for that case. The moment you need to update two or more values together as one consistent step, use a mutex — atomics only make each individual operation indivisible, not a group of them.
What is the difference between compare_exchange_weak and compare_exchange_strong?
Both do the compare-and-swap: replace the value only if it still equals what you expected. compare_exchange_strong fails only when the values genuinely differ. compare_exchange_weak may also fail 'spuriously' (for no real reason) on some hardware, so you use it inside a retry loop where you are looping anyway. Use weak in a loop for best performance, and strong when you are not looping.
What is the ABA problem?
A CAS only checks whether the value still equals what you expected — not whether it changed and changed back. If a value goes A -> B -> A between your read and your CAS, the CAS succeeds as if nothing happened, even though state you cared about (like a freed and reused node) is different. The classic fix is to attach a version counter that always increments, so A-with-version-1 never equals A-with-version-2.
What does memory_order_relaxed actually relax?
Relaxed guarantees the operation itself is atomic, but gives NO ordering guarantee relative to other variables across threads. It is correct for a standalone counter where you only care about the final count. It is wrong when one atomic is meant to 'publish' other data — there you need release on the writer and acquire on the reader so the data is visible.
Which memory ordering should I use by default?
Use the default, memory_order_seq_cst, until profiling proves it is a bottleneck. It gives a single total order that is the easiest to reason about and the hardest to get wrong. Drop to acquire/release only for proven hot paths, and to relaxed only for independent counters or statistics where ordering genuinely does not matter.