Overview
This tutorial builds a scalar-valued automatic differentiation engine from scratch to expose what neural-network training actually does beneath a modern framework. It begins with numerical derivatives, showing that a gradient measures how a small change in an input affects an output. A Value object is then developed to store scalar data, operation history, graph dependencies, and gradients. Local derivatives for addition, multiplication, tanh, exponentiation, and powers are combined through the chain rule, while topological sorting ensures that reverse-mode differentiation processes the computation graph in the correct order. The implementation also reveals a crucial detail: gradients must accumulate when a value contributes through multiple graph branches. These primitives are assembled into neurons, layers, and a multilayer perceptron, then trained on a small binary-classification dataset using a loss function and gradient descent. The training loop repeatedly performs a forward pass, clears old gradients, runs backpropagation, and updates parameters in the loss-reducing direction. A comparison with PyTorch confirms matching scalar results while explaining why production systems use tensors for parallel efficiency. The central conclusion is that large neural networks differ enormously in scale and engineering complexity, but their core training mechanism remains the same compact sequence of computation, differentiation, and optimization.
Sections
Core Concepts
The principal terms needed to understand the implementation and training process.
- Automatic differentiation, or autograd, is machinery that evaluates derivatives by recording a computation graph and applying local derivative rules backward through it.
- Backpropagation is reverse-mode differentiation through a computation graph, implemented as recursive applications of the chain rule.
- A derivative or gradient is a local sensitivity measure describing how an output changes when an input is changed slightly.
- A computation graph is a directed acyclic graph whose nodes hold values and whose edges record the operations and dependencies that produced them.
- A local derivative describes an operation's output sensitivity to one of its immediate inputs, independent of the larger graph containing that operation.
- A loss function reduces prediction performance to a scalar arranged so that lower values represent behavior closer to the desired targets.
- Gradient descent repeatedly updates parameters in the direction opposite the loss gradient to reduce the loss locally.
- A multilayer perceptron is a sequence of neuron layers, where each neuron computes a weighted sum plus bias and passes it through a nonlinearity.
- A tensor is a multidimensional array of scalars that allows many equivalent scalar operations to execute efficiently in parallel.
- Learning-rate decay progressively reduces the update step size so training can move quickly early and refine parameters more carefully later.
Implementation Details
Specific structures, derivative rules, and training mechanics used to build Micrograd.
- Value stores data, grad, the set of predecessor nodes, the operation label, and a backward closure for propagating gradients.
- For z = x + y, the local derivatives with respect to both inputs are 1, so the incoming gradient is routed unchanged to each input.
- For z = x × y, the gradient contribution to x is y × z.grad, while the contribution to y is x × z.grad.
- For z = tanh(x), the local derivative is 1 − z², so x.grad accumulates (1 − z²) × z.grad.
- For z = exp(x), the local derivative is exp(x), which is already available as z.data.
- For z = xⁿ with constant n, the local derivative is n × xⁿ⁻¹; division can then be expressed as multiplication by a power of −1.
- Backward first builds a topological ordering from the output, initializes the output gradient to 1, and invokes each node's local backward function in reverse order.
- Gradient propagation uses += rather than assignment because multiple graph branches may contribute derivatives to the same value.
- The example MLP has three inputs, two hidden layers of four neurons, one output neuron, and 41 trainable scalar parameters.
- The demonstration loss sums squared differences between predictions and targets; its minimum is zero when every prediction equals its target.
- PyTorch leaf tensors must explicitly request gradients in this example, and double precision is selected to match Python's scalar floating-point behavior.
Execution Checklist
Concrete actions for implementing and training a minimal neural network correctly.
- Implement a Value object that records scalar data, gradient state, parent nodes, producing operation, and a local backward function.
- Add arithmetic primitives one at a time and validate every local derivative with finite-difference gradient checks.
- Accumulate gradient contributions with += so reused values and branching graphs differentiate correctly.
- Build a topological order before the backward pass and traverse it in reverse from the scalar output.
- Reset every parameter gradient to zero before each new backward pass.
- Construct neuron, layer, and MLP abstractions that expose all trainable weights and biases through a parameters method.
- Train by repeatedly calculating predictions and loss, clearing gradients, backpropagating, and updating parameters opposite the gradients.
- Tune the learning rate conservatively and reduce it when large updates cause overshooting or unstable loss.
- Compare forward values and parameter gradients against PyTorch on identical expressions to verify the implementation.
Higher-Level Insights
Broader implications derived from the implementation and debugging process.
- The conceptual complexity of neural-network training is much smaller than the engineering complexity of production frameworks. Large libraries primarily add tensor kernels, device support, data types, numerical safeguards, and performance machinery around a compact mathematical core.
- Successful loss reduction is insufficient evidence that an implementation is correct. The missing zero-grad bug still trained the toy problem, demonstrating why gradient checks and framework comparisons are necessary even when outputs look plausible.
- Autograd is modular because differentiation knowledge belongs to operations rather than entire models. Once each operation defines its forward computation and vector of local derivatives, arbitrary larger graphs become differentiable automatically.
- Scalar engines and tensor frameworks form a clean pedagogical progression: scalar execution exposes every dependency, while tensor execution preserves the same semantics and changes the computational scale.
- The same core training pattern scales from a 41-parameter classifier to language models with hundreds of billions of parameters; the major changes are architecture, loss choice, batching, optimizer design, and systems efficiency.