Skip to main content
Interactive Deep Learning Laboratory

Neural Network Visualization

Master multi-layer perceptron (MLP) architectures, non-linear activation dynamics, and backpropagation calculus. Build custom layer topologies, train on non-linear datasets, and step through forward and backward passes in real time.

Technical Overview: Artificial Neural Networks, Forward Propagation, and Backpropagation Calculus

1. What is an Artificial Neural Network (ANN)?

An artificial neural network is a mathematical computational architecture consisting of layers of interconnected artificial neurons inspired by biological brains. An input feature vector x ∈ ℝn is transformed through hidden layers via parameterized linear transformations z[l] = W[l] a[l-1] + b[l] and non-linear activation functions a[l] = σ(z[l]) to produce predictive outputs ŷ.

2. Mathematical Formulations of Deep Learning

The training process of a feedforward neural network is governed by four foundational mathematical equations:

  • Forward Linear Transformation: z[l] = W[l] a[l-1] + b[l]
  • Non-linear Activation: a[l] = σ(z[l])
  • Output Layer Error Delta: δ[L] = ∇a L ⊙ σ'(z[L])
  • Hidden Layer Error Recurrence (Chain Rule): δ[l] = ((W[l+1])T δ[l+1]) ⊙ σ'(z[l])
  • Weight & Bias Gradients: ∂L / ∂W[l] = δ[l] (a[l-1])T and ∂L / ∂b[l] = δ[l]
  • Gradient Descent Update: W[l] := W[l] - η (∂L / ∂W[l]) and b[l] := b[l] - η (∂L / ∂b[l])

3. Activation Function Comparison

Activation Function Mathematical Formula Derivative σ'(z) Output Range Key Characteristics
ReLU (Rectified Linear Unit) f(z) = max(0, z) 1 if z > 0 else 0 [0, ∞) Prevents vanishing gradients, highly computationally efficient, industry standard
GELU (Gaussian Error Linear Unit) f(z) = z Φ(z) ≈ 0.5z(1 + tanh(√(2/π)(z + 0.044715z3))) Smooth continuous derivative (-0.17, ∞) Standard in modern Transformer architectures (GPT, BERT, Llama)
Sigmoid (Logistic) σ(z) = 1 / (1 + e-z) σ(z)(1 - σ(z)) (0, 1) Outputs valid probabilities; prone to vanishing gradient in deep layers
Tanh (Hyperbolic Tangent) tanh(z) = (ez - e-z) / (ez + e-z) 1 - tanh2(z) (-1, 1) Zero-centered output; stronger gradient flow than Sigmoid

4. Loss Functions in Deep Learning

Loss Function Equation Primary Application
Mean Squared Error (MSE) L = (1/2) ∑ (yi - ŷi)2 Continuous regression, function approximation
Binary Cross-Entropy (BCE) L = - [y log(ŷ) + (1 - y) log(1 - ŷ)] Binary classification with Sigmoid output
Categorical Cross-Entropy (CCE) L = - ∑ yk log(ŷk) Multi-class classification with Softmax output