Deep Learning
Learned hierarchical representations, trained end to end by gradient descent: the mechanics, the architectures, and what actually goes wrong.
Deep learning is machine learning with models built as deep stacks of differentiable operations, trained end to end by gradient descent. The defining idea is that you do not hand-design the features: the network learns a hierarchy of representations, each layer transforming the previous one, and the useful abstractions emerge as a side effect of minimizing the loss. Everything else — architectures, optimizers, normalization tricks — exists to make that optimization stable at depth and scale.
The mechanics underneath
A deep network is a composition of parameterized functions applied to tensors. Training runs a loop: take a mini-batch, compute the forward pass to get predictions, compute a scalar loss, use reverse-mode automatic differentiation (backpropagation) to get the gradient of that loss with respect to every parameter, and let an optimizer take a step. Repeat until the held-out metric stops improving.
Three consequences of that loop are worth internalising. First, the loss must be differentiable with respect to the parameters, which is why so much design effort goes into smooth relaxations of discrete decisions. Second, the gradient is computed on a sample, so training is stochastic and batch size interacts with learning rate. Third, the backward pass needs the intermediate activations from the forward pass, so memory during training scales with depth and batch size, not just with parameter count — usually the binding constraint long before FLOPs are.
Architecture families and what each assumes
An architecture is an inductive bias made concrete — an assumption about the structure of the data, baked into the wiring.
- Fully connected (MLP) layers assume nothing about structure. Flexible, parameter-hungry, the default for generic vector inputs and for the projection blocks inside larger models.
- Convolutional networks assume locality and translation invariance: the same filter is useful everywhere in the image. Weight sharing makes them parameter-efficient on grid-shaped data.
- Recurrent networks assume sequential dependence and carry a hidden state along the sequence. They process in order, which limits parallelism and makes long-range credit assignment hard.
- Transformers replace recurrence with attention: every position can attend to every other in one step, and position is supplied explicitly rather than implied by processing order. The cost is that attention work grows quadratically with sequence length.
- Graph neural networks assume an explicit relational structure and pass messages along edges.
- Generative families — autoregressive models, diffusion models, autoencoders — differ in how they factorize and sample from a distribution rather than in how they stack layers.
Matching the bias to the data is worth more than depth. A convolution applied to data with no spatial structure buys nothing but constraints.
What makes depth trainable
Naively stacked deep networks are hard to optimize: gradients shrink or explode as they propagate through many layers. A small set of components fixed this and now appear almost everywhere.
Residual (skip) connections give the gradient a short path back to early layers, so adding depth degrades optimization far less. Normalization layers rescale activations to keep them in a well-conditioned range, which makes larger learning rates usable. Careful initialization sets the initial scale so signal neither vanishes nor blows up on the first forward pass. Non-saturating activations avoid regions where the derivative is essentially zero. Gradient clipping caps the update when a rare batch produces an enormous gradient.
Optimization in practice
The learning rate is the hyperparameter that matters most; almost everything else is second order. Too high and the loss diverges or oscillates; too low and training stalls in a way that is easy to mistake for a modeling problem. A warmup period followed by a decaying schedule is the common pattern, because early steps are taken from a poorly conditioned starting point.
Adaptive optimizers maintain per-parameter step sizes and are the usual default. Weight decay regularizes by shrinking parameters and, in adaptive optimizers, behaves differently from an L2 penalty added to the loss — a distinction worth knowing before copying a configuration between frameworks. Mixed-precision training keeps most of the computation in a reduced float format while holding a higher-precision copy of the weights and, where the format requires it, scaling the loss to keep small gradients representable.
Generalization and regularization
Deep networks have enough capacity to memorize their training set outright, so the interesting question is why they generalize at all. In practice the levers are: more and more varied data, data augmentation that encodes invariances you actually believe in, dropout and weight decay, early stopping, and transfer from a model pretrained on a larger corpus. Transfer learning is often the highest-leverage of these — starting from learned representations and adapting them needs far less labeled data than training from scratch.
Failure modes
- The loss goes to NaN. Usually a learning rate that is too high, an unstable numerical operation (a logarithm of zero, a division by a near-zero denominator), or an underflow in reduced precision.
- Training loss falls, validation loss does not. Overfitting, or a leak that makes the training split easier than reality.
- Train/eval mismatch. Dropout and normalization layers behave differently in the two modes. Forgetting to switch modes produces metrics that are wrong in a way that looks plausible.
- Silent data problems. Mislabeled examples, duplicated records spanning splits, or an augmentation that destroys the label all cap achievable quality with no error raised.
- Non-reproducibility. Unseeded shuffling, non-deterministic GPU kernels and unpinned library versions make a result impossible to recover later.
- Distribution shift after deployment. A network is confident by construction and will extrapolate off-distribution inputs without signaling that it is doing so.
When not to use it
Deep learning earns its cost when the input is unstructured, the dataset is large, and a pretrained starting point exists. It is usually the wrong first choice for small tabular datasets, where gradient-boosted trees are stronger and cheaper; where a decision must be explained line by line to a regulator; or where inference must run under a tight latency or power budget without the engineering to make that work. If you do need a large model on a constrained target, that engineering is its own discipline — see model compression — and running it reliably is covered under MLOps.