Computer Vision Basics

Reviewed & published by Brayan K

Teach a computer to "see" — turn images into numbers, sharpen and blur them by hand, then understand how convolutional neural networks (CNNs) recognise objects.

Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.

What You'll Learn in This Lesson

👁️ How Your Eyes and Brain "See"

When you look at a dog, your eyes do not send "DOG" to your brain. They send millions of tiny light measurements. Your visual system then builds meaning in layers: first it spots edges (where light meets dark), then it groups edges into shapes (an ear, a snout), then it recognises textures (fur), and only at the end does it conclude "that's a dog."

A computer starts in exactly the same place: an image arrives as a grid of brightness numbers. A convolutional neural network (CNN) then rebuilds the same ladder — early layers find edges, middle layers find shapes, deep layers find whole objects. The whole lesson is about understanding that ladder, one rung at a time, starting from raw numbers.

1 An Image Is Just a Grid of Numbers

A pixel ("picture element") is one dot of an image. In a grayscale image each pixel is a single number from 0 (black) to 255 (white), with greys in between. You store the whole image as a list of rows — a nested list — and read any pixel with image[row][col].

A colour image adds a third dimension: every pixel becomes three numbers — Red, Green and Blue (RGB). That is why a 224×224 colour photo is 224 × 224 × 3 = 150,528 numbers. Run the worked example below to read pixels and measure an image's size.

# An image is just a grid of numbers. No libraries needed.
# Each number is a "pixel": 0 = black, 255 = white, in between = grey.

# A tiny 3x3 grayscale image stored as a nested list (a list of rows)
image = [
    [  0, 128, 255],   # row 0: black, grey, white
    [128, 255, 128],   # row 1
    [255, 128,   0],   # row 2
]

# Print it like a picture so you can "see" the numbers
print("=== Pixel values ===")
for row in image:
    print(row)

# Read one pixel: image[row][col]
print()
print("Top-left pixel  image[0][0]:", image[0][0])   # 0   (black)
print("Centre pixel    image[1][1]:", image[1][1])   # 255 (white)

# Size of the image
height = len(image)        # number of rows
width  = len(image[0])     # number of columns in the first row
print()
print("Height x Width:", height, "x", width)         # 3 x 3

# Expected output:
# === Pixel values ===
# [0, 128, 255]
# [128, 255, 128]
# [255, 128, 0]
#
# Top-left pixel  image[0][0]: 0
# Centre pixel    image[1][1]: 255
#
# Height x Width: 3 x 3

2 Basic Operations: Brightness and Threshold

Once an image is numbers, editing it is just arithmetic. Brightening adds a fixed amount to every pixel — but you must clip the result to the 0–255 range, because a pixel can never be darker than black or brighter than white. Thresholding turns the image pure black-and-white: any pixel at or above a cutoff becomes 255, everything else becomes 0. That is the simplest way to separate a bright object from a dark background.

# Two of the most common image operations, written by hand.

image = [
    [  0, 128, 255],
    [128, 255, 128],
    [255, 128,   0],
]

# 1) BRIGHTEN: add a value to every pixel, then "clip" to the 0-255 range.
#    Clipping matters: pixels can never go below 0 or above 255.
def brighten(img, amount):
    out = []
    for row in img:
        new_row = []
        for pixel in row:
            value = pixel + amount
            value = max(0, min(255, value))   # clip into 0..255
            new_row.append(value)
        out.append(new_row)
    return out

# 2) THRESHOLD: turn the image black & white. Pixel >= cutoff -> 255, else 0.
#    This is how you separate a bright object from a dark background.
def threshold(img, cutoff):
    return [[255 if p >= cutoff else 0 for p in row] for row in img]

print("=== Brightened by 50 ===")
for row in brighten(image, 50):
    print(row)

print()
print("=== Threshold at 128 ===")
for row in threshold(image, 128):
    print(row)

# Expected output:
# === Brightened by 50 ===
# [50, 178, 255]
# [178, 255, 178]
# [255, 178, 50]
#
# === Threshold at 128 ===
# [0, 255, 255]
# [255, 255, 255]
# [255, 255, 0]

3 Convolution and Filters (Blur and Edges)

Convolution is the single most important idea in computer vision, and it is much simpler than it sounds: slide a small grid of numbers (a "filter" or "kernel") across the image and combine the pixels underneath it. A 2×2 average filter replaces each region with the average of its four pixels — that softens the image (a blur). The result is slightly smaller than the input, because the window cannot hang off the edge.

# Convolution sounds scary; it is just "slide a small grid over the image
# and combine the numbers underneath." Here is a 2x2 AVERAGE (blur) filter.

image = [
    [ 10,  20,  30],
    [ 40,  50,  60],
    [ 70,  80,  90],
]

# Slide a 2x2 window across the image. For a 3x3 image, the window fits in
# 2 positions across and 2 down -> the result is 2x2 (it shrinks at the edges).
def average_2x2(img):
    out = []
    for i in range(len(img) - 1):          # rows 0,1
        new_row = []
        for j in range(len(img[0]) - 1):   # cols 0,1
            # the four pixels under the 2x2 window
            block = [img[i][j],   img[i][j+1],
                     img[i+1][j], img[i+1][j+1]]
            new_row.append(sum(block) // 4)   # integer average
        out.append(new_row)
    return out

print("=== Original 3x3 ===")
for row in image:
    print(row)

print()
print("=== After 2x2 average filter (blurred, 2x2) ===")
for row in average_2x2(image):
    print(row)

# The top-left output = average of 10,20,40,50 = 120 // 4 = 30
# Expected output:
# === Original 3x3 ===
# [10, 20, 30]
# [40, 50, 60]
# [70, 80, 90]
#
# === After 2x2 average filter (blurred, 2x2) ===
# [30, 40]
# [60, 70]

Swap the filter's numbers and the same sliding machinery does something else entirely. An edge-detection kernel gives a strong response wherever brightness changes sharply, and near zero across flat regions. That is the intuition behind "finding edges." The numpy example below applies a real edge kernel to a 5×5 image.

import numpy as np   # the real world uses numpy, not nested lists

# A 5x5 image with a bright square in the middle
image = np.array([
    [0, 0, 0, 0, 0],
    [0, 1, 1, 1, 0],
    [0, 1, 1, 1, 0],
    [0, 1, 1, 1, 0],
    [0, 0, 0, 0, 0],
], dtype=float)

# A 3x3 edge-detection kernel: it fires where the centre differs from neighbours
edge_kernel = np.array([
    [-1, -1, -1],
    [-1,  8, -1],
    [-1, -1, -1],
], dtype=float)

def convolve(img, kernel):
    kh, kw = kernel.shape
    out = np.zeros((img.shape[0] - kh + 1, img.shape[1] - kw + 1))
    for i in range(out.shape[0]):
        for j in range(out.shape[1]):
            out[i, j] = np.sum(img[i:i+kh, j:j+kw] * kernel)
    return out

print(convolve(image, edge_kernel))

# Expected output:
# [[ 5.  3.  5.]
#  [ 3.  0.  3.]
#  [ 5.  3.  5.]]
# Notice: the centre is 0 (flat, no edge) and the corners are high (edges!).
# 🎯 YOUR TURN — fill in the blanks marked with ___

image = [
    [  0, 128, 255],
    [255,  64,   0],
]

# Invert means: a black pixel becomes white and vice-versa.
# The rule for an 8-bit pixel is:  new_value = 255 - old_value
def invert(img):
    out = []
    for row in img:
        new_row = []
        for pixel in row:
            new_row.append(___)   # 👉 replace ___ with the invert formula
        out.append(new_row)
    return out

for row in invert(image):
    print(row)

# ✅ Expected output:
# [255, 127, 0]
# [0, 191, 255]
# 🎯 YOUR TURN — fill in the blanks marked with ___

image = [
    [ 4,  8, 12],
    [16, 20, 24],
    [28, 32, 36],
]

# Compute just the TOP-LEFT value of a 2x2 average filter.
# The 2x2 window covers image[0][0], image[0][1], image[1][0], image[1][1].
top_left = image[0][0] + image[0][1] + ___ + ___   # 👉 add the two bottom pixels
average  = top_left ___ 4                            # 👉 use integer division //

print("Window sum:", top_left)
print("Average:   ", average)

# ✅ Expected output:
# Window sum: 48
# Average:    12

4 From Filters to CNNs (Conv, Pool, Feature Maps)

A CNN stacks the idea you just built. A convolution layer applies many filters at once — but instead of you choosing the numbers, the network learns them during training. Each filter produces a feature map: a grid showing where that pattern was found. Early layers learn edge filters, deeper layers learn shape and object filters — exactly the eye/brain ladder from the analogy.

A pooling layer (usually MaxPool 2×2) then shrinks each feature map by keeping only the strongest value in each 2×2 block. This throws away precise positions but keeps "was the feature here, roughly?", which makes the network smaller and more robust. Conv → Pool → Conv → Pool repeats until a Flatten turns the maps into a vector for a final Dense classifier.

In real frameworks you describe this stack in a few lines. The example below sketches an OpenCV pre-processing step (note the BGR→RGB fix) and a small Keras CNN. The shapes shrink layer by layer — read the # Expected output to see how.

# ── OpenCV reads images as BGR, not RGB! A classic bug. ──
import cv2                       # pip install opencv-python
img = cv2.imread("cat.jpg")      # shape (H, W, 3) in B, G, R order
img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)   # convert before showing/feeding
img = cv2.resize(img, (224, 224))            # CNNs need a fixed input size
img = img / 255.0                            # normalise pixels to 0..1

# ── A small image classifier in Keras (TensorFlow) ──
from tensorflow import keras
from tensorflow.keras import layers

model = keras.Sequential([
    layers.Input((28, 28, 1)),               # 28x28 grayscale (like MNIST)
    layers.Conv2D(32, 3, activation="relu"), # 32 filters detect simple edges
    layers.MaxPooling2D(2),                  # shrink + keep the strongest signal
    layers.Conv2D(64, 3, activation="relu"), # deeper filters detect shapes
    layers.MaxPooling2D(2),
    layers.Flatten(),                        # 2D feature maps -> 1D vector
    layers.Dense(128, activation="relu"),
    layers.Dense(10, activation="softmax"),  # 10 classes, probabilities sum to 1
])
model.summary()

# Expected output (shapes shrink layer by layer):
# Conv2D       -> (None, 26, 26, 32)
# MaxPooling2D -> (None, 13, 13, 32)
# Conv2D       -> (None, 11, 11, 64)
# MaxPooling2D -> (None,  5,  5, 64)
# Flatten      -> (None, 1600)
# Dense        -> (None, 128)
# Dense        -> (None, 10)
# Total params: ~225,000 — tiny, because Conv layers SHARE weights.

🗂️ The Three Core Vision Tasks

Almost every computer-vision product is one of these three jobs, in increasing difficulty:

5 Common Errors (And How to Fix Them)

These four mistakes trip up nearly every beginner. Spotting them saves hours.

❌ Forgetting to normalise pixels

Feeding raw 0–255 pixels into a network. The large values make training unstable and slow.

✅ Fix: scale to 0–1 before training:

image = image / 255.0   # now every pixel is between 0.0 and 1.0

❌ Wrong channel order (BGR vs RGB)

OpenCV's cv2.imread returns pixels in BGR order, but most models and display libraries expect RGB. Colours come out swapped (blue skies look orange).

✅ Fix: convert right after loading:

img = cv2.cvtColor(img, cv2.COLOR_BGR2RGB)

❌ Not resizing to a fixed input size

CNNs expect every input to be the same shape (e.g. 224×224). Passing mixed sizes raises a shape error like expected (224,224,3), got (480,640,3).

✅ Fix: resize every image first:

img = cv2.resize(img, (224, 224))

❌ Giant dense layers instead of convolution

Flattening a raw image straight into a Dense layer creates millions of weights and overfits instantly.

✅ Fix: use Conv2D + MaxPooling2D first to shrink and share weights, and flatten only near the end.

📋 Quick Reference

TermWhat It MeansExample / Note
PixelOne dot of an image0=black … 255=white
GrayscaleOne value per pixelshape H×W×1
RGBRed, Green, Blue per pixelshape H×W×3
NormaliseScale pixels to 0–1img / 255.0
ThresholdMake black & white255 if p >= c else 0
ConvolutionSlide a filter over the imageblur, sharpen, edges
Conv2DLearns filters → feature mapsshares weights
MaxPoolShrinks, keeps strongest value2×2 halves H and W
Flatten2D maps → 1D vectorfeeds Dense layer
CV tasksThree core jobsclassify · detect · segment

🎯 Mini Challenge: Posterise an Image

Time to fly solo. Snap every pixel to the nearest of three levels (0, 128, 255) — a classic poster-art effect. The starter below gives only the brief and the data; you write the logic. Check yourself against the expected output in the comments.

# 🎯 MINI-CHALLENGE: posterise an image
# A 3x3 grayscale image is given. Write a "posterise" function that snaps
# every pixel to the NEAREST of three levels: 0, 128, or 255.
#
# Rule:
#   pixel < 64           -> 0
#   64 <= pixel < 192    -> 128
#   pixel >= 192         -> 255
#
# Steps:
# 1. Loop over every row, then every pixel.
# 2. Apply the rule above to choose 0, 128 or 255.
# 3. Build and print the new image (a list of rows).
#
# ✅ Expected (for the image below):
# [0, 128, 255]
# [128, 128, 0]
# [255, 0, 128]

image = [
    [ 30, 100, 240],
    [130, 150,  10],
    [200,  20, 120],
]

# your code here

Lesson complete — you can make a computer see!

You now know that images are grids of numbers, you can brighten, threshold, blur and find edges by hand, and you understand how convolution, pooling and feature maps stack into a CNN that classifies, detects, or segments. The "magic" of computer vision is just arithmetic on pixels, repeated in layers.

Practice quiz

In a grayscale image, what does a pixel value of 0 represent?

  • White
  • Red
  • Black
  • Transparent

Answer: Black. Grayscale pixels run from 0 (black) to 255 (white), with greys in between.

How many numbers represent each pixel in an RGB image?

  • 3
  • 1
  • 2
  • 4

Answer: 3. RGB stores three values per pixel — Red, Green, and Blue — so it has three channels.

What does convolution do?

  • Sorts pixels by brightness
  • Deletes the image edges
  • Converts colour to grayscale
  • Slides a small filter across the image and combines the pixels underneath

Answer: Slides a small filter across the image and combines the pixels underneath. Convolution slides a small kernel over the image, combining the pixels it covers to produce each output value.

An averaging (mean) filter applied to an image produces:

  • Sharper edges
  • A blur
  • A rotation
  • A brighter image only

Answer: A blur. A 2x2 average filter replaces each region with the mean of its pixels, softening (blurring) the image.

What does a max-pooling layer do in a CNN?

  • Keeps the strongest value in each block, shrinking the feature map
  • Adds more filters
  • Normalises pixels to 0..1
  • Flattens the image into a vector

Answer: Keeps the strongest value in each block, shrinking the feature map. MaxPool 2x2 keeps the largest value in each 2x2 block, shrinking the map while keeping the strongest signal.

Why do CNNs use far fewer parameters than a dense layer over a raw image?

  • They use smaller images
  • They skip the output layer
  • A convolution shares one small filter across the whole image
  • They ignore colour channels

Answer: A convolution shares one small filter across the whole image. Convolution shares a small filter across all positions, so it learns with thousands of weights, not millions.

What does an edge-detection kernel respond strongly to?

  • Flat, uniform regions
  • Places where brightness changes sharply
  • The image corners only
  • The darkest pixel

Answer: Places where brightness changes sharply. An edge kernel gives a strong response where brightness changes sharply and near zero across flat regions.

What does the Flatten layer do in a CNN?

  • Blurs the feature maps
  • Adds padding to the edges
  • Converts RGB to grayscale
  • Turns 2D feature maps into a 1D vector for the dense layer

Answer: Turns 2D feature maps into a 1D vector for the dense layer. Flatten reshapes the 2D feature maps into a 1D vector so a final Dense classifier can process them.

Which computer-vision task labels every individual pixel?

  • Classification
  • Segmentation
  • Object detection
  • Normalisation

Answer: Segmentation. Segmentation labels each pixel, giving an exact outline; classification gives one label, detection draws boxes.

Why normalise pixels to 0..1 before training a CNN?

  • It makes images larger
  • It converts colour to grayscale
  • Large 0..255 values make training unstable and slow
  • It is required to read the file

Answer: Large 0..255 values make training unstable and slow. Dividing by 255 scales pixels to 0..1, keeping gradients well-behaved so training is stable and faster.

Continue this course

Frequently asked questions

Why is an image just a grid of numbers?

A camera sensor measures light at thousands of tiny points. Each measurement becomes a number (a pixel). Grayscale uses one number per pixel (0=black to 255=white); colour uses three numbers per pixel for Red, Green and Blue.

What is the difference between grayscale and RGB?

Grayscale stores one brightness value per pixel, so a 28x28 image is 28x28x1 numbers. RGB stores three values per pixel (red, green, blue), so a 28x28 colour image is 28x28x3 numbers. RGB has three 'channels'; grayscale has one.

What does convolution actually do?

It slides a small grid of numbers (a filter or kernel) across the image and combines the pixels underneath to produce a new value. Different filters highlight different things: averaging blurs, while an edge kernel lights up wherever brightness changes sharply.

Why use a CNN instead of a regular neural network for images?

A 224x224x3 image is over 150,000 numbers. A plain dense layer over that needs millions of weights and ignores the 2D structure. CNNs share a small filter across the whole image, so they learn with far fewer parameters and respect spatial layout.

What are classification, detection and segmentation?

Classification answers 'what is in this image?' with one label. Object detection draws boxes around each object and labels them. Segmentation labels every individual pixel, giving the exact outline of each object.

Related lessons