Performance Profiling & Benchmarking

Reviewed & published by Brayan K

By the end of this lesson you'll be able to measure how fast your code runs instead of guessing — timing a block with Stopwatch, comparing two approaches fairly, writing rigorous benchmarks with BenchmarkDotNet, watching memory allocations, and knowing when to reach for a full profiler. The golden rule of all performance work: measure first, optimise second.

Part of the free C# course at LearnCodingFast — hands-on lessons with worked examples and the output they print, plus practice exercises and a quick quiz.

What You'll Learn

⏱️ Real-World Analogy

Think of two tools an athlete uses. A Stopwatch is the handheld stopwatch a coach clicks at the start and end of a sprint — instant, simple, and perfect for a quick "how long did that take?". BenchmarkDotNet is the full fitness tracker: it makes you do warmup laps, runs the test many times, throws out the flukes, reports the average and the variation, and even logs how much energy (memory) you burned. The stopwatch tells you a rough number in seconds; the fitness tracker gives you trustworthy, repeatable data you can stake a decision on. The rule both share: never judge fitness by feel — measure it. Optimising code you haven't measured is like training harder at the wrong thing.

Three tools, three jobs

ToolWhat it's forPrecisionReach for it when…
StopwatchQuick, ad-hoc timing of a blockRough (ms)You want a fast "how long?" answer right now
BenchmarkDotNetRigorous A/B benchmarks with stats + memoryHigh (ns, statistical)You're comparing approaches and the result matters
Profiler (dotTrace / PerfView)Finding where time/allocations go in a whole appWhole-programSomething's slow but you don't know which method

A benchmark answers "is A faster than B?". A profiler answers "which of my 500 methods is the slow one?". You usually profile to find the hotspot, then benchmark to fix it.

1. Measure, Don't Guess: Stopwatch

The fastest way to answer "how long does this take?" is System.Diagnostics.Stopwatch. Stopwatch.StartNew() creates and starts a high-resolution timer in one line; you do your work, call .Stop(), then read .ElapsedMilliseconds. Crucially, do not use DateTime.Now for this — it can jump backwards (clock changes, NTP syncs) and has poor resolution. Stopwatch is monotonic and built for exactly this. Read this worked example, run it, then you'll time your own block.

using System;
using System.Diagnostics;

class Program
{
    static void Main()
    {
        // Stopwatch is a high-resolution timer — far more accurate than
        // DateTime.Now for measuring elapsed time. StartNew() creates it
        // AND starts it in one line.
        var sw = Stopwatch.StartNew();

        // === The block of work you want to time ===
        long total = 0;
        for (int i = 0; i < 10_000_000; i++)
            total += i;                 // some real work to measure
        // ==========================================

        sw.Stop();                      // stop the clock

        // ElapsedMilliseconds = whole milliseconds (a long).
        // sw.Elapsed gives a TimeSpan if you want finer detail.
        Console.WriteLine($"Sum     = {total}");
        Console.WriteLine($"Elapsed = {sw.ElapsedMilliseconds} ms");

        // ⚠️ No expected-output panel here, on purpose: "Elapsed" is a
        // stopwatch reading and changes with the machine and its load. One
        // real run where this site is built printed:
        //
        //    Sum     = 49999995000000
        //    Elapsed = 24 ms
        //
        // The sum is fixed; the millisecond count is not, and a page that
        // promised you a number here would be wrong on your machine.
    }
}

Your turn. The program below times a counting loop — it just needs the three Stopwatch calls filled in. Replace each ___ using the hints, then run it.

using System;
using System.Diagnostics;

class Program
{
    static void Main()
    {
        // 🎯 YOUR TURN — fill in the blanks marked with ___, then run it.

        // 1) Create AND start a Stopwatch in one line.
        var sw = ___;               // 👉 Stopwatch.StartNew()

        // The work we want to time: count up to a big number.
        long count = 0;
        for (int i = 0; i < 5_000_000; i++)
            count++;

        // 2) Stop the clock.
        sw.___;                     // 👉 Stop()

        // 3) Print how many whole milliseconds it took.
        Console.WriteLine($"Counted {count} in {sw.___} ms"); // 👉 ElapsedMilliseconds

        // ✅ Expected output (the ms will vary by machine):
        //    Counted 5000000 in 9 ms
    }
}

2. Comparing Two Approaches

A single timing in isolation rarely tells you much — what you really want is "is A faster than B?". The classic example: building a big string. Using += in a loop secretly creates a brand-new string on every pass and copies everything so far, so the cost grows with the square of the size. A StringBuilder keeps one growable buffer and appends into it. Time both side by side and the gap is dramatic — and now you can prove it rather than assert it.

using System;
using System.Diagnostics;
using System.Text;

class Program
{
    static void Main()
    {
        const int N = 50_000;

        // --- Approach A: build a string with += in a loop ---
        // Each += makes a BRAND-NEW string and copies everything so far,
        // so the work grows with the square of N. This is the slow trap.
        var swA = Stopwatch.StartNew();
        string a = "";
        for (int i = 0; i < N; i++)
            a += "x";               // allocates a new string EVERY time
        swA.Stop();

        // --- Approach B: build it with a StringBuilder ---
        // StringBuilder keeps ONE growable buffer and appends into it.
        var swB = Stopwatch.StartNew();
        var sb = new StringBuilder();
        for (int i = 0; i < N; i++)
            sb.Append("x");         // no new string each time
        string b = sb.ToString();
        swB.Stop();

        Console.WriteLine($"+= concat     : {swA.ElapsedMilliseconds} ms");
        Console.WriteLine($"StringBuilder : {swB.ElapsedMilliseconds} ms");
        Console.WriteLine($"Same result?  : {a.Length == b.Length}");

        // ⚠️ No expected-output panel here, on purpose: two of these three
        // lines are stopwatch readings. One real run printed:
        //
        //    += concat     : 538 ms
        //    StringBuilder : 0 ms
        //    Same result?  : True
        //
        // The exact milliseconds will differ on your machine. The ORDER OF
        // MAGNITUDE will not: += rebuilds the whole string every iteration,
        // so it is quadratic, and StringBuilder is linear.
    }
}

Now you try. Time both approaches yourself and print them so you can compare. Fill in the three ___ blanks — two stopwatch calls and one result read:

using System;
using System.Diagnostics;
using System.Text;

class Program
{
    static void Main()
    {
        const int N = 30_000;

        // 🎯 YOUR TURN — time BOTH approaches and compare them.

        // --- Approach A: += in a loop (the slow way) ---
        var swA = Stopwatch.StartNew();
        string a = "";
        for (int i = 0; i < N; i++)
            a += "*";
        // 1) Stop stopwatch A.
        swA.___;                    // 👉 Stop()

        // --- Approach B: StringBuilder (the fast way) ---
        // 2) Create AND start a second stopwatch.
        var swB = ___;              // 👉 Stopwatch.StartNew()
        var sb = new StringBuilder();
        for (int i = 0; i < N; i++)
            sb.Append("*");
        swB.Stop();

        // 3) Print both timings so you can compare them.
        Console.WriteLine($"+= concat     : {swA.ElapsedMilliseconds} ms");
        Console.WriteLine($"StringBuilder : {swB.___} ms");  // 👉 ElapsedMilliseconds

        // ✅ Expected output (the StringBuilder line should be far smaller):
        //    += concat     : 250 ms
        //    StringBuilder : 0 ms
    }
}

🔎 Deep Dive: why your Stopwatch number can lie

The JIT warmup trap. .NET compiles your method to machine code the first time it runs (Just-In-Time compilation). So the very first call includes the cost of compiling — it can be 10–100x slower than steady state. If you time only one run, you're often timing the compiler, not your code. The fix: do one or two throwaway runs first, then start the stopwatch.

Debug vs Release. In a Debug build the JIT turns off optimisations so debugging is easier. Always measure performance in Release (dotnet run -c Release); Debug numbers are not representative.

Tiny operations. A Stopwatch resolves to fractions of a millisecond at best. Trying to time something that takes nanoseconds (a single method call) by running it once is hopeless — measurement noise swamps the signal. To time tiny things you must run them millions of times in a loop and divide... which is exactly the fiddly, error-prone work that BenchmarkDotNet does correctly for you.

work();                       // 1) warmup — let the JIT compile it
var sw = Stopwatch.StartNew();
for (int i = 0; i < 1000; i++) work();   // 2) many runs, not one
sw.Stop();
double perRun = (double)sw.ElapsedMilliseconds / 1000;  // 3) average

3. The Right Tool: BenchmarkDotNet

BenchmarkDotNet is the industry-standard .NET benchmarking library, and it handles every pitfall above for you. You tag methods with [Benchmark], mark one as the Baseline, and call BenchmarkRunner.Run<YourClass>(). It warms up the JIT, runs each method many times, discards outliers, computes the mean and standard deviation, and prints a clean comparison table. You read it, not write the timing scaffolding yourself. (BenchmarkDotNet needs a real .NET project, so study this worked example here, then run it in your own project with dotnet run -c Release.)

using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;
using System.Text;

// Run benchmarks with:  dotnet run -c Release
// NEVER benchmark in Debug mode — the JIT skips optimisations there,
// so the numbers are meaningless.

[MemoryDiagnoser]              // also report allocations + GC, not just time
public class StringBenchmarks
{
    private readonly string[] _words = System.Linq.Enumerable
        .Range(0, 1000)
        .Select(i => $"word{i}")
        .ToArray();

    // Baseline = true marks the row everything else is compared against.
    [Benchmark(Baseline = true)]
    public string Concatenation()
    {
        var result = "";
        foreach (var word in _words)
            result += word + " ";     // O(n^2) — re-copies the whole string
        return result;
    }

    [Benchmark]
    public string StringBuilder_()
    {
        var sb = new StringBuilder();
        foreach (var word in _words)
            sb.Append(word).Append(' ');  // O(n) — one growable buffer
        return sb.ToString();
    }

    [Benchmark]
    public string StringJoin() => string.Join(" ", _words);  // most idiomatic
}

class Program
{
    // BenchmarkRunner does the warmup, runs many iterations, and does
    // the statistics for you — then prints the results table below.
    static void Main() => BenchmarkRunner.Run<StringBenchmarks>();
}

// ✅ Expected output (a sample BenchmarkDotNet results table):
//
// |          Method |       Mean | Ratio |   Gen0 | Allocated |
// |---------------- |-----------:|------:|-------:|----------:|
// |   Concatenation | 1,250.3 us |  1.00 | 450.00 |   2.10 MB |
// | StringBuilder_  |    12.5 us |  0.01 |   2.00 |   9.80 KB |
// |      StringJoin |     8.2 us |  0.01 |   1.25 |   6.10 KB |
//
// Read it like this: StringJoin is ~150x faster (Ratio 0.01) AND
// allocates ~350x less memory (6 KB vs 2.1 MB) than the += baseline.

Notice what the table gives you that a bare Stopwatch can't: a Mean (in microseconds, far finer than ms), a Ratio against the baseline so the relative speed is obvious at a glance, and — because we added [MemoryDiagnoser] — the memory each method allocated. That last column is the one beginners forget.

4. Allocations & GC Pressure

Speed isn't the only cost. Every object you create on the heap is future work for the garbage collector (GC) — the part of .NET that reclaims memory you're no longer using. Allocate a lot and the GC runs more often, and a GC pause briefly freezes your program. Under load that's frequently what makes a "fast" method slow in production. Add [MemoryDiagnoser] to a benchmark and watch the Allocated and Gen0 columns: a method that's twice as fast but allocates ten times more memory may lose once real traffic hits it.

using BenchmarkDotNet.Attributes;

// The Allocated and Gen0 columns from [MemoryDiagnoser] often matter
// MORE than Mean. Every allocation is future work for the garbage
// collector — and under load, GC pauses are what slow real apps down.

[MemoryDiagnoser]
public class AllocationBenchmarks
{
    private readonly int[] _data = System.Linq.Enumerable
        .Range(0, 1_000_000).ToArray();

    // Allocates an iterator + delegates behind the scenes.
    [Benchmark(Baseline = true)]
    public int Linq_Sum() => System.Linq.Enumerable.Sum(_data);

    // A plain loop allocates NOTHING on the heap — zero GC pressure.
    [Benchmark]
    public int ForLoop_Sum()
    {
        int sum = 0;
        for (int i = 0; i < _data.Length; i++)
            sum += _data[i];
        return sum;
    }
}

// ✅ Expected output (sample — note the Allocated column):
//
// |      Method |      Mean | Ratio |   Gen0 | Allocated |
// |------------ |----------:|------:|-------:|----------:|
// |    Linq_Sum | 1,180.0 us|  1.00 |      - |      40 B |
// | ForLoop_Sum |   430.0 us|  0.36 |      - |       - B |
//
// The loop is ~2.7x faster AND allocates 0 bytes. "Allocated = -"
// is the goal on a hot path: no allocations means no GC pauses.

🔎 Deep Dive: profilers — finding where the time goes

A benchmark compares two known options. But when a whole app is slow and you don't know which method is to blame, you need a profiler — a tool that watches your program run and reports where the CPU time and the allocations actually went.

The healthy workflow: profile to find the one method eating 80% of the time (it's almost never the one you'd guess), then benchmark two fixes for that method to prove which is actually better.

Common hotspots to look for

When you profile real C# code, the same culprits show up again and again:

Putting It Together: a fair timing helper

Here's a small, reusable helper that puts the whole lesson into practice: it does a warmup run (to dodge the JIT trap), runs the work many times, and reports both total and per-run time — so you can compare two approaches fairly. It's a poor man's BenchmarkDotNet, but it's honest about warmup and repeats, which is more than a naive single timing ever is.

using System;
using System.Diagnostics;

class Program
{
    // A reusable helper: time any block of code, print the result.
    // Run the block ONCE to warm up (let the JIT compile it), then time it.
    static void Time(string label, int repeats, Action work)
    {
        work();                          // warmup run — JIT-compile the method
        var sw = Stopwatch.StartNew();
        for (int i = 0; i < repeats; i++)
            work();                      // the timed runs
        sw.Stop();

        double perRun = (double)sw.ElapsedMilliseconds / repeats;
        Console.WriteLine($"{label,-14}: {sw.ElapsedMilliseconds} ms total, {perRun:F3} ms/run");
    }

    static void Main()
    {
        int[] data = new int[1_000_000];
        for (int i = 0; i < data.Length; i++) data[i] = i;

        // Compare two ways to find the max value, fairly (same warmup + repeats).
        Time("Manual loop", 50, () =>
        {
            int max = int.MinValue;
            foreach (int n in data) if (n > max) max = n;
        });

        Time("LINQ Max()", 50, () =>
        {
            int max = System.Linq.Enumerable.Max(data);
        });

        // ⚠️ No expected-output panel here, on purpose: these are stopwatch
        // readings, and this site only prints output that is the same every
        // time. Five runs on the machine that builds it gave, consistently:
        //
        //    Manual loop   : 141 ms total, 2.820 ms/run
        //    LINQ Max()    : 16 ms total, 0.320 ms/run
        //
        // Note WHICH WAY ROUND that is. The received wisdom is that a
        // hand-written loop beats LINQ, and on .NET Framework it did. On
        // .NET 8, Enumerable.Max over an int[] takes a vectorised (SIMD) path
        // and reads several values per instruction, so it wins by about 8x —
        // every run, not marginally. Which is the lesson: measure the runtime
        // you actually ship on, and let the measurement overrule the folklore.
    }
}

Notice the helper runs work() once before starting the clock — that warmup run gets the JIT compilation out of the way so it doesn't pollute the measurement. For anything you'll publish a number about, graduate to BenchmarkDotNet.

Pro Tips

Common Errors (and the fix)

📋 Quick Reference

TaskCodeNotes
Start a timervar sw = Stopwatch.StartNew();Creates + starts in one line
Stop itsw.Stop();Freezes the elapsed time
Read elapsed (ms)sw.ElapsedMillisecondsA long, whole ms
Read elapsed (fine)sw.ElapsedA TimeSpan
Mark a benchmark[Benchmark]On a method to time
Set the baseline[Benchmark(Baseline = true)]Others compared to it
Track memory[MemoryDiagnoser]Adds Allocated/Gen0 columns
Run benchmarksBenchmarkRunner.Run<T>();Use dotnet run -c Release

Frequently Asked Questions

Q: When is Stopwatch enough and when do I need BenchmarkDotNet?

Stopwatch is great for a quick, rough "how long did that take?" on a chunk of work that takes milliseconds or more. The moment you're comparing two approaches and the answer matters — or the operation is tiny (microseconds) — use BenchmarkDotNet. It handles warmup, many iterations, statistics, and memory for you, so the number you report is trustworthy.

Q: Why is my first timed run so much slower than the rest?

That's JIT compilation. .NET compiles each method to machine code the first time it runs, and that one-off cost lands on your first measurement. Do a throwaway warmup run before you start the Stopwatch (BenchmarkDotNet does this automatically).

Q: Should I worry about memory or just speed?

Both. Memory allocations create work for the garbage collector, and GC pauses are a leading cause of latency spikes under load. A method that's faster but allocates far more can perform worse in production. Use [MemoryDiagnoser] and check the Allocated column, not just Mean.

Q: What's the difference between a benchmark and a profiler?

A benchmark answers "is approach A faster than B?" for code you already suspect. A profiler answers "which method in my whole app is the slow one?" — it watches a running program and shows where time and allocations actually go. Profile to find the hotspot, then benchmark to fix it.

Q: Isn't optimising early just good engineering?

No — "premature optimisation" makes code harder to read for a gain the program may never need. Write it correctly and clearly first, measure, and then optimise only the spots a profiler proves are hot. Most code isn't on a hot path at all.

Mini-Challenge: Benchmark Two Ways to Sum an Array

No blanks this time — just a brief and an outline to keep you on track. Build a large array, sum it two ways (a for loop and LINQ's .Sum()), time each with a Stopwatch after a warmup run, and report which is faster. Run it and check your output against the example in the comments — this is the core skill of the whole lesson: let the measurement decide.

using System;
using System.Diagnostics;

class Program
{
    static void Main()
    {
        // 🎯 MINI-CHALLENGE: which way of summing an array is faster?
        // 1. Build an int[] of 5,000,000 numbers (fill it in a loop).
        // 2. Approach A — sum it with a for loop. Time it with a Stopwatch.
        // 3. Approach B — sum it with LINQ's .Sum(). Time that too.
        //    (Add 'using System.Linq;' at the top, then call data.Sum().)
        // 4. Print BOTH timings AND which approach won.
        // 5. TIP: do one throwaway warmup run of each first, so the JIT
        //    has compiled the code before you start the clock.
        //
        // ✅ Example output (numbers vary, the for loop usually wins):
        //    for loop : 6 ms
        //    LINQ Sum : 22 ms
        //    Winner   : for loop

        // your code here
    }
}

🎉 Lesson Complete

Practice quiz

What is the golden rule of performance work?

  • Optimise everything first
  • Never optimise
  • Measure first, optimise second
  • Guess the hotspot

Answer: Measure first, optimise second. Always measure (profile or benchmark) before optimising — optimising a guess wastes effort.

What does Stopwatch.StartNew() do?

  • Creates and starts a high-resolution timer in one line
  • Stops a timer
  • Returns the current date
  • Resets to zero only

Answer: Creates and starts a high-resolution timer in one line. StartNew() creates a Stopwatch and starts it immediately in one call.

Why should you use Stopwatch instead of DateTime.Now for timing?

  • DateTime.Now is faster
  • They are the same
  • DateTime.Now can't measure time
  • Stopwatch is monotonic and high-resolution; DateTime.Now can jump backwards and has poor resolution

Answer: Stopwatch is monotonic and high-resolution; DateTime.Now can jump backwards and has poor resolution. DateTime.Now isn't monotonic (clock changes can make elapsed time negative); Stopwatch is built for elapsed timing.

What is the 'JIT warmup trap' when timing code?

  • The CPU overheats
  • The first run includes the cost of compiling the method, so it can be far slower
  • The timer is too slow
  • The loop never ends

Answer: The first run includes the cost of compiling the method, so it can be far slower. .NET JIT-compiles a method the first time it runs, so timing only one run often measures the compiler. Do a throwaway warmup first.

Why must you benchmark in Release, not Debug?

  • In Debug the JIT skips optimisations, so the numbers are meaningless
  • Debug is slower to start
  • Release has no logging
  • Debug can't run benchmarks

Answer: In Debug the JIT skips optimisations, so the numbers are meaningless. Debug builds disable JIT optimisations; always measure with dotnet run -c Release.

Why is building a big string with += in a loop slow?

  • It uses too many threads
  • It allocates nothing
  • Each += creates a brand-new string and copies everything, so cost grows with the square of the size
  • The compiler forbids it

Answer: Each += creates a brand-new string and copies everything, so cost grows with the square of the size. += in a loop re-allocates and copies the whole string each pass (O(n^2)); StringBuilder uses one growable buffer.

What does the [MemoryDiagnoser] attribute add to a BenchmarkDotNet report?

  • Only the Mean column
  • Allocation and GC columns (Allocated, Gen0)
  • Thread counts
  • The source code

Answer: Allocation and GC columns (Allocated, Gen0). [MemoryDiagnoser] reports allocations and GC pressure, not just time — often the cost that matters most under load.

What does BenchmarkDotNet do for you that a bare Stopwatch can't?

  • Nothing
  • It runs in Debug
  • It deletes slow methods
  • Warmup, many iterations, outlier removal, statistics, and a comparison table

Answer: Warmup, many iterations, outlier removal, statistics, and a comparison table. BenchmarkDotNet handles warmup, repeated runs, statistics and memory, so reported numbers are trustworthy.

What question does a profiler (dotTrace, PerfView) answer that a benchmark doesn't?

  • Is A faster than B?
  • Which method in the whole app is the slow one?
  • What is the syntax?
  • How to compile?

Answer: Which method in the whole app is the slow one?. A profiler finds where time/allocations go across a whole app; profile to find the hotspot, then benchmark to fix it.

Why can a method that's faster but allocates much more memory lose in production?

  • It uses more disk
  • It can't compile
  • More allocations create work for the GC, and GC pauses cause latency spikes under load
  • It is always slower

Answer: More allocations create work for the GC, and GC pauses cause latency spikes under load. Every allocation is future GC work; under load, GC pauses freeze the program, so a higher-allocating method can perform worse.

Continue this course