ESC
Type to search guides, tutorials, and reference documentation.
← AI and Machine Learning
🤖

Machine Learning

Fitting a function to data so it generalizes to data it has never seen: the workflow, the metrics, the leaks, and when a rule beats a model.

Machine learning is the practice of fitting a function to data so that it makes useful predictions on data it has never seen. Instead of an engineer writing the rules, an optimization procedure searches for parameters that minimize a loss on a training set, and the whole exercise is judged on how well those parameters transfer to new inputs. Everything difficult about ML follows from that one sentence: the objective you can measure (training loss) is not the objective you care about (behavior in production), and the gap between them is where projects fail.

The shapes of learning

Most practical work falls into a few families. Supervised learning maps inputs to labels you already have — classification when the label is a category, regression when it is a number. Unsupervised learning finds structure without labels: clustering, dimensionality reduction, density estimation. Self-supervised learning manufactures labels from the data itself (predict the masked token, predict the next frame), which is how most modern representation learning is bootstrapped. Reinforcement learning optimizes a policy against a reward signal collected by acting, and is the right frame only when you genuinely have a sequential decision problem with feedback.

Choosing the family is usually forced by the data you can actually obtain. If nobody can tell you what a correct answer looks like, you do not have a supervised problem no matter how much you want one.

The workflow that actually matters

The modeling step is the smallest part of the job. In order of how much they determine the outcome:

  • Problem framing. What decision does the prediction change? If the answer is "none", stop here.
  • Label definition. What exactly counts as a positive? Ambiguous label definitions produce noisy ceilings that no model can exceed.
  • Splitting. How train, validation and test are separated determines whether your numbers mean anything.
  • A baseline. Majority class, last observed value, or a hand-written rule. A model that cannot beat the baseline is not a model.
  • Features and representation. Still the dominant lever on tabular data.
  • Model and tuning. Usually the easiest part to get adequately right.
  • Evaluation, deployment, monitoring. Where the value is realized or lost.

Splits and leakage

Data leakage — information from the target or from the future reaching the model at training time — is the single most common reason a model looks excellent offline and fails in production. It arrives quietly: a feature computed over the full dataset before splitting, an identifier that correlates with the label, a "days since resolution" column that only exists once the outcome is known, a duplicated row landing in both train and test.

Splitting strategy is the main defense. Random splits are appropriate only when rows are independent. For anything with a time dimension, split temporally: train on the past, validate on the future, because that is the only arrangement that mimics deployment. When multiple rows belong to the same entity — a user, a device, a patient — split by group, or the model memorizes the entity rather than the pattern. Any preprocessing that learns from data (scaling, imputation, target encoding, feature selection) must be fitted on the training fold alone and applied to the others.

Choosing a metric that matches the decision

Accuracy is a poor default. On an imbalanced problem, a model that always predicts the majority class can score well while being useless. Precision and recall separate the two ways of being wrong, and which one you weight is a business decision, not a modeling one — a false positive in fraud review costs an analyst's time, a false negative costs the fraud.

ROC-AUC measures ranking quality across all thresholds and is insensitive to class balance, which makes it flattering on rare-event problems; precision-recall AUC is usually the more honest summary there. If your system consumes a probability rather than a decision — expected-value calculations, risk pricing, triage ordering — you also need calibration: a score of 0.3 should mean the event happens about three times in ten. A well-ranked but badly calibrated model breaks any downstream arithmetic.

Finally, the decision threshold is a separate, tunable artifact. Train the model to rank; pick the threshold against the cost of each error type; revisit it when those costs change.

Generalization, bias and variance

A model that is too simple to represent the pattern underfits — high bias, poor on both training and held-out data. A model with enough capacity to memorize the training set overfits — low training error, poor held-out error. Regularization, more data, and simpler hypothesis classes push in one direction; more capacity and richer features push in the other. Cross-validation gives a more stable estimate than a single split when data is scarce, at the cost of compute.

There is a subtler trap: repeatedly tuning against a validation set eventually overfits it too. Keep a test set that you look at rarely, ideally once, and treat any number you have optimized against as optimistic.

Common failure modes

  • Train/serve skew. The feature computed in the training pipeline is not quite the feature computed at inference — different defaults, different time windows, different null handling. It degrades quality silently.
  • Proxy labels. You could not measure "customer was satisfied", so you trained on "did not open a ticket". The model learns the proxy, including its biases.
  • Selection bias in the training data. If historical decisions filtered who appears in your data, the model learns the filter. Models trained on approved-loan outcomes never see the rejected applicants.
  • Distribution shift. Inputs move away from the training distribution; performance decays without any error being raised.
  • Feedback loops. The model's predictions change the behavior that generates its next training set.

When ML is the wrong tool

Prefer explicit rules when the logic is known, stable, and must be auditable — a deterministic rule is cheaper to build, easier to explain, and does not drift. Prefer statistics or simple heuristics when you have very few examples. Avoid ML where an individual wrong answer is catastrophic and no human review sits in the loop. And be honest about feedback: without a mechanism to observe outcomes, you cannot evaluate or retrain, and the system will degrade with nobody noticing.

For tabular problems, gradient-boosted decision trees remain a strong, fast default and are what a deep network usually has to beat. Reach for deep learning when the input is unstructured — images, audio, text, graphs — or when representation learning is the point.

RELATED GUIDES