Neural Networks Introduction
Reviewed & published by Brayan K
By the end of this lesson you'll be able to compute a single neuron's output by hand, write ReLU and sigmoid in plain Python, and explain how layers learn by adjusting weights.
Part of the free AI & Machine Learning course at LearnCodingFast — hands-on lessons with examples you run in your browser, plus practice exercises and a quick quiz.
What You'll Learn in This Lesson
- You'll be able to describe a neuron as weighted sum + bias + activation
- You'll be able to compute a neuron's output in plain Python with lists
- You'll be able to write the ReLU and sigmoid activation functions yourself
- You'll be able to explain what a layer is and how layers stack
- You'll be able to trace a forward pass from inputs to prediction
- You'll be able to explain how learning nudges weights to cut error
🧠 Real-World Analogy: Brain Neurons
Your brain has billions of neurons. Each one receives little electrical signals from its neighbours through connections called synapses. Some connections are strong and some are weak — they decide how much each incoming signal counts. When the combined signal crosses a threshold, the neuron fires and passes a signal on to the next neurons.
An artificial neuron copies this idea with arithmetic. The "synapse strengths" become numbers called weights, the firing threshold becomes a bias, and the "fire or not" decision becomes an activation function. Learning is just gradually turning the strength of each connection up or down until the whole network responds the way you want.
1 The Neuron — Weighted Sum + Bias + Activation
A neuron (also called a perceptron when it's on its own) is the smallest building block of a neural network. It takes some inputs and produces a single number. It does this in three tiny steps:
- Weighted sum — multiply each input by its weight and add the results together.
- Add a bias — add one extra number that shifts the total up or down.
- Activation — pass the total through a function that decides the output.
z = (x1·w1 + x2·w2 + … + xn·wn) + bias
The weights and bias are the neuron's knowledge. Everything a network learns ends up stored as weights and biases. Here is a single neuron written in plain Python — read each comment, then run it.
# A single neuron (a "perceptron") — plain Python, no libraries
# A neuron does 3 tiny steps:
# 1. weighted sum: multiply each input by its weight and add them up
# 2. add a bias: a number that shifts the result up or down
# 3. activation: squash the result through a function (here: a step)
# The neuron's "knowledge" lives in these numbers.
inputs = [1.0, 0.0, 1.0] # 3 features going in (e.g. yes/no signals)
weights = [0.7, 0.3, 0.9] # how much each input matters
bias = -1.0 # a threshold-shifter
# Step 1 + 2: weighted sum + bias -> this number is called "z"
z = bias
for x, w in zip(inputs, weights):
z += x * w # add input * weight, one pair at a time
print("inputs =", inputs)
print("weights =", weights)
print("bias =", bias)
print("z (weighted sum + bias) =", round(z, 2)) # 0.6
# Step 3: activation — a "step" turns z into a clean 0 or 1 decision
output = 1 if z > 0 else 0
print("neuron output =", output) # 1 (z was positive)
# ✅ Expected output:
# inputs = [1.0, 0.0, 1.0]
# weights = [0.7, 0.3, 0.9]
# bias = -1.0
# z (weighted sum + bias) = 0.6
# neuron output = 12 Activation Functions — ReLU and Sigmoid
The activation function is what makes a network powerful. It introduces non-linearity — a fancy way of saying "the output can bend and curve instead of being one straight line." Two activations cover almost everything a beginner needs:
- ReLU (Rectified Linear Unit): returns the input if it's positive, otherwise 0. It's fast and is the default choice for hidden layers.
- Sigmoid: squashes any number into the range 0 to 1 using an S-shaped curve. Perfect for an output that represents a probability ("how likely is this a cat?").
Both are just a couple of lines of plain Python — the only thing you need from the standard library is math.exp for sigmoid.
import math # only the standard library — no numpy
# Activation functions decide what a neuron "fires".
# Without them, stacking neurons would just be one big straight line.
def relu(x):
# ReLU = "Rectified Linear Unit": keep positives, zero out negatives
return x if x > 0 else 0.0
def sigmoid(x):
# Sigmoid squashes ANY number into the range 0..1 (an S-curve)
return 1 / (1 + math.exp(-x))
# Try them on a few numbers so you can see the shape
for x in [-2.0, -0.5, 0.0, 0.5, 2.0]:
print(f"x={x:>4} relu={relu(x):>4} sigmoid={sigmoid(x):.3f}")
# ✅ Expected output:
# x=-2.0 relu= 0.0 sigmoid=0.119
# x=-0.5 relu= 0.0 sigmoid=0.378
# x= 0.0 relu= 0.0 sigmoid=0.500
# x= 0.5 relu= 0.5 sigmoid=0.622
# x= 2.0 relu= 2.0 sigmoid=0.881# 🎯 YOUR TURN — finish the neuron's forward pass
# Fill in each ___ . Expected output is at the bottom so you can self-check.
inputs = [2.0, 3.0]
weights = [0.5, -1.0]
bias = 1.0
# 1) Start z at the bias value
z = ___ # 👉 replace ___ with the bias variable
# 2) Add input * weight for each pair
for x, w in zip(inputs, weights):
z = z + ___ # 👉 replace ___ with x * w
# 3) Apply a step activation: 1 if z is positive, else 0
output = 1 if z ___ 0 else 0 # 👉 replace ___ with the > sign
print("z =", z)
print("output =", output)
# ✅ Expected output:
# z = -1.0
# output = 0import math
# 🎯 YOUR TURN — implement the two activations from scratch.
# Fill in the ___ blanks, then run to compare against the expected output.
def relu(x):
# Return x when it is positive, otherwise return 0.0
return ___ if x > 0 else 0.0 # 👉 replace ___ with x
def sigmoid(x):
# The S-curve: 1 / (1 + e^(-x)). math.exp(n) computes e^n.
return 1 / (1 + math.exp(___)) # 👉 replace ___ with -x
print("relu(-3) =", relu(-3))
print("relu(4) =", relu(4))
print("sigmoid(0) =", sigmoid(0))
# ✅ Expected output:
# relu(-3) = 0.0
# relu(4) = 4
# sigmoid(0) = 0.53 Layers and the Forward Pass
One neuron can only draw a single straight boundary, which is too weak for most real problems. The fix is to use many neurons arranged in layers:
- Input layer — your raw features (e.g. pixel values, sensor readings).
- Hidden layer(s) — neurons that each look for a different pattern in the inputs.
- Output layer — produces the final prediction.
The forward pass is simply running data through the network from left to right: every neuron in a layer computes weighted sum + bias + activation, and its outputs become the inputs to the next layer. Stack enough layers and the network can approximate almost any function — that's the whole magic.
inputs → [hidden layer: many neurons] → [output layer] → prediction
4 How Learning Adjusts Weights
A fresh network starts with random weights, so its first predictions are basically guesses. Training fixes that with a repeating loop:
- Forward pass: run an example through the network to get a prediction.
- Measure error: compare the prediction to the correct answer (the "loss").
- Backpropagation: work out how much each weight contributed to the error.
- Update: nudge every weight and bias a little in the direction that lowers the error.
That nudge size is controlled by the learning rate. Repeat this loop over thousands of examples and the weights slowly settle into values that make good predictions. You don't have to compute the gradients by hand — libraries like TensorFlow and PyTorch do it for you. Here's the same neuron idea written the professional way, plus a tiny Keras network, shown as a read-only reference.
# The SAME idea, written the way professionals do it.
# numpy does the weighted sum for a whole layer in one line.
import numpy as np
inputs = np.array([1.0, 0.0, 1.0])
weights = np.array([0.7, 0.3, 0.9])
bias = -1.0
z = np.dot(inputs, weights) + bias # weighted sum + bias, vectorised
output = 1 / (1 + np.exp(-z)) # sigmoid activation
print(round(float(output), 3)) # 0.646
# Expected output:
# 0.646
# A whole network in a few lines with Keras (TensorFlow).
# This builds 2 inputs -> 4 hidden (ReLU) -> 1 output (sigmoid):
from tensorflow import keras
from tensorflow.keras import layers
model = keras.Sequential([
layers.Dense(4, activation="relu", input_shape=(2,)), # hidden layer
layers.Dense(1, activation="sigmoid"), # output layer
])
model.compile(optimizer="adam", loss="binary_crossentropy")
# model.fit(X, y, epochs=200) # training adjusts every weight automatically
# Expected output (model summary, abridged):
# Total params: 17 (12 in hidden layer + 5 in output layer)5 Common Errors (And How to Fix Them)
These four mistakes trip up nearly everyone who builds their first network.
❌ No non-linearity = a glorified linear model
If every layer uses no activation (or only a linear one), stacking layers collapses into a single straight line. The network can't learn curves and will fail on problems like XOR.
✅ Fix: put a non-linear activation (ReLU) on every hidden layer:
# ❌ hidden = weighted_sum # linear — no power
# ✅ hidden = relu(weighted_sum) # adds the curve the network needsSigmoid and tanh flatten out for large inputs, so their slope (gradient) becomes almost 0. In deep networks the update signal shrinks to nothing and early layers stop learning.
✅ Fix: use ReLU in hidden layers; keep sigmoid for the final output only.
# hidden layers -> relu(z) # gradient stays healthy
# output layer -> sigmoid(z) # 0..1 probability is fine hereFeeding raw values on wildly different scales (e.g. age 0–100 next to salary 0–100000) makes training unstable — the big numbers dominate the weighted sum.
✅ Fix: scale features to a similar range before training.
# scale each feature to roughly 0..1
normalized = [(v - low) / (high - low) for v in feature]❌ Learning rate too big
If each weight update is too large, the network overshoots the good values and the loss bounces around or explodes to nan instead of going down.
✅ Fix: start small (e.g. 0.01) and only increase if learning is too slow.
# learning_rate = 10.0 # ❌ loss jumps around / becomes nan
learning_rate = 0.01 # ✅ steady, reliable improvement📋 Quick Reference
| Term | What it is | In code / formula |
|---|---|---|
| Weight | How much an input matters | x * w |
| Bias | Shifts the sum up/down | z = ... + bias |
| Weighted sum (z) | Inputs·weights + bias | sum(x*w) + bias |
| ReLU | Keep positives, zero negatives | x if x > 0 else 0 |
| Sigmoid | Squash into 0..1 | 1/(1+math.exp(-x)) |
| Layer | A group of neurons | input / hidden / output |
| Forward pass | Inputs → prediction | layer by layer |
| Learning rate | Size of each weight nudge | w += lr * ... |
🎯 Mini-Challenge: A Neuron with a Real Activation
Time to fly with less support. Build a 2-input neuron that ends with a sigmoid activation. Only a comment outline is given — fill in the logic yourself, then check against the expected output in the comments.
import math
# 🎯 MINI-CHALLENGE: a 2-input neuron with a real activation
# Brief:
# 1. Define sigmoid(x) -> 1 / (1 + math.exp(-x))
# 2. inputs = [1.0, 1.0] weights = [2.0, 2.0] bias = -3.0
# 3. Compute z = bias + sum of input*weight (use a loop or zip)
# 4. Pass z through sigmoid to get the output (a probability 0..1)
# 5. print("z =", z) and print("output =", round(output, 3))
#
# ✅ Expected (inputs 1,1):
# z = 1.0
# output = 0.731
# your code hereLesson complete — you understand how neurons think!
You can now compute a neuron as weighted sum + bias + activation, write ReLU and sigmoid in plain Python, describe how layers stack into a forward pass, and explain how training nudges weights to shrink error. These are the exact foundations every deep learning model is built on.
Practice quiz
What three steps does a single neuron perform?
- Sort, filter, and average the inputs
- Tokenize, embed, and classify
- Weighted sum of inputs, add a bias, then apply an activation function
- Split, fit, and score
Answer: Weighted sum of inputs, add a bias, then apply an activation function. A neuron computes weighted sum + bias, then passes the result through an activation function.
What is the role of the bias in a neuron?
- It shifts the weighted sum up or down so the neuron can fire at a different threshold
- It multiplies the inputs together
- It removes negative inputs
- It normalizes the data
Answer: It shifts the weighted sum up or down so the neuron can fire at a different threshold. The bias shifts z before activation; without it every neuron would be forced through zero.
Why do neural networks need non-linear activation functions?
- To make training slower
- To remove the bias term
- To shrink the dataset
- Without them, stacking layers collapses into a single linear model
Answer: Without them, stacking layers collapses into a single linear model. Non-linearity lets the network bend and curve; otherwise stacked layers are just one linear function.
What does the ReLU activation do?
- Squashes any number into 0 to 1
- Returns the input if positive, otherwise 0
- Always returns 1
- Returns the negative of the input
Answer: Returns the input if positive, otherwise 0. ReLU keeps positives and zeroes out negatives — fast and the default for hidden layers.
What range does the sigmoid activation output?
- 0 to 1
- -1 to 1
- Any real number
- 0 to 100
Answer: 0 to 1. Sigmoid squashes any number into the range 0 to 1, ideal for representing a probability.
What is a forward pass?
- Updating every weight to reduce error
- Splitting data into train and test
- Running inputs through the network layer by layer to produce a prediction
- Removing the output layer
Answer: Running inputs through the network layer by layer to produce a prediction. The forward pass feeds inputs through each layer's weighted sum + activation to get a prediction.
Why is sigmoid usually avoided in deep hidden layers?
- It is too fast
- It causes vanishing gradients because its slope flattens for large inputs
- It cannot output probabilities
- It only works on images
Answer: It causes vanishing gradients because its slope flattens for large inputs. Sigmoid saturates, so gradients shrink toward zero and early layers stop learning.
In what order does training adjust a network?
- Update weights, then forward pass
- Backpropagation only
- Measure error, then stop
- Forward pass, measure error, backpropagation, update weights
Answer: Forward pass, measure error, backpropagation, update weights. Each step: forward pass to predict, measure loss, backpropagate, then nudge weights.
What does the learning rate control?
- The number of layers
- The size of each weight update during training
- How many inputs a neuron has
- The activation function used
Answer: The size of each weight update during training. The learning rate sets how big each nudge is; too large overshoots, too small is slow.
Why can't a single neuron solve the XOR problem?
- XOR has too many inputs
- XOR requires sigmoid
- One neuron can only draw a single straight boundary; XOR needs a hidden layer
- XOR has no reward signal
Answer: One neuron can only draw a single straight boundary; XOR needs a hidden layer. A single neuron is linear; XOR is not linearly separable, so it needs at least one hidden layer.
Continue this course
- Previous: Decision Trees & Random Forests
- Next: Deep Learning Fundamentals — Train deep neural networks with backpropagation and activation functions
- Quick reference: AI & Machine Learning cheat sheet › Common Algorithms
- From the blog: Neural Networks: An Introduction · Reinforcement Learning: Q-Learning Explained
Frequently asked questions
What is a neuron (perceptron) in a neural network?
It is a tiny function that multiplies each input by a weight, adds those products together, adds a bias number, and passes the result through an activation function. The output is the neuron's decision.
What is the bias for?
The bias shifts the weighted sum up or down before activation, so the neuron can fire at a different threshold. Without it, every neuron would be forced to pass through zero, which limits what the network can learn.
Why do neural networks need activation functions?
Activation functions add non-linearity. If you remove them, stacking layers collapses into a single straight-line (linear) model that can only solve linearly separable problems. ReLU and sigmoid let the network bend and curve to fit complex data.
What is the difference between ReLU and sigmoid?
ReLU returns the input if it is positive and 0 otherwise — fast and the default for hidden layers. Sigmoid squashes any number into 0..1, which is ideal for an output that represents a probability, but it causes vanishing gradients in deep hidden layers.
What is a forward pass?
The forward pass is running inputs through the network layer by layer to produce a prediction: for each neuron you compute weighted sum + bias, apply the activation, then feed the results into the next layer.
How does a neural network learn?
It compares its prediction to the correct answer to measure error, then nudges every weight and bias a little in the direction that reduces that error. Repeating this over many examples (using backpropagation and gradient descent) is what we call training.