This introduction has two parts. The first explains what artificial intelligence, machine learning and deep learning are, and how they fit together. The second is a refresher of the maths notation used in machine learning, one symbol at a time, each with a small graphic.
Introduction to Machine Learning
Artificial intelligence and machine learning are often used as if they meant the same thing. They are related but distinct, and deep learning is a third, narrower term: each one sits inside the previous one.
The labels inside each field are examples of well-known techniques; you do not need to know them to follow this introduction.
Artificial Intelligence
Artificial intelligence (AI) is the broad field of building systems that handle tasks we would call intelligent if a person did them, such as holding a conversation, recognising a face or planning a route. Its main areas:
- Natural language processing (NLP): working with human language, as in translation, summaries or chat assistants.
- Computer vision: making sense of images and videos, from reading a scanned document to spotting a defect on a production line.
- Robotics: machines that sense their surroundings and move in the physical world, such as warehouse robots or robot vacuum cleaners.
- Expert systems: programs that encode the knowledge of human experts as rules and apply them to new cases.
AI is mostly meant to support people rather than replace them: it sifts through more data than anyone could read, suggests decisions and takes over repetitive work. It is already used in:
- Healthcare: spotting signs of disease in medical images, speeding up the search for new drugs.
- Finance: flagging fraudulent card payments, assessing risk, managing investments.
- Cybersecurity: detecting intrusions, phishing and malware, and helping analysts respond to them.
Machine Learning
Machine learning (ML) is the part of AI where a system learns from data instead of following rules that a programmer wrote. The algorithm looks at many examples, finds what they have in common, then uses it to judge new examples it has never seen.
Take phishing detection. Instead of writing rules by hand, we give the algorithm thousands of emails labelled "phishing" or "legitimate". It works out on its own which signals matter (number of links, age of the sender's domain, pressure words such as "urgent") and how much. These measurable signals are called features, and the importance the model gives each one is its weight. Shown a new email, it gives a score between 0 and 1: the closer to 1, the more likely it is phishing.
ML comes in three main types:
- Supervised learning: every example comes with the right answer, its label, and the algorithm learns to give that answer itself. Examples: reading handwritten digits, estimating the price of a house, flagging phishing emails.
- Unsupervised learning: the examples have no labels, so the algorithm looks for groups and patterns on its own. Examples: sorting news articles by topic, noticing an unusual login on a network, summarising many measurements with a few numbers.
- Reinforcement learning: an agent (the program that acts) learns by trial and error, collecting rewards for good moves and penalties for bad ones. Examples: a program that learns to play chess or Go, a robot arm that learns to pick up objects, traffic lights that learn when to change.
ML is used in almost every industry:
| Industry | Applications |
|---|---|
| Healthcare | Diagnosis, drug discovery, personalised treatment |
| Finance | Fraud detection, credit risk, algorithmic trading |
| Marketing | Customer segments, targeted ads, recommendations |
| Cybersecurity | Threat detection, intrusion prevention, malware analysis |
| Transportation | Traffic forecasts, autonomous vehicles, route planning |
Deep Learning
Deep learning (DL) is the part of ML that uses neural networks with many layers (hence "deep") to learn directly from complex data. It works best on raw, messy data such as photos, sound recordings and text, where each example is made of thousands of values.
Three things set deep learning apart:
- Hierarchical features: each layer builds on the one before, from simple to abstract. On a screenshot of a web page, the first layers react to edges, the next to text blocks and input fields, the last to a logo above a password field.
- End-to-end learning: the network goes straight from raw input to answer, without a person deciding which features to measure.
- Scalability: it keeps improving as data and computing power grow, where classic algorithms level off.
The most common kinds of networks:
- Convolutional neural networks (CNNs): built for images and video. Small filters slide over the image to detect local patterns, and stacked layers combine them into larger shapes.
- Recurrent neural networks (RNNs): built for sequences such as text and speech. A loop carries information from one step to the next, as a short-term memory.
- Transformers: the most recent of the three, now the standard for language. Self-attention lets each word look at every word of the text, itself included, however far apart, and decide how much each one matters. The large language models behind chat assistants are transformers.
Deep learning is behind most recent progress in AI. It finds and outlines objects in photos. It translates, summarises and writes text. It turns speech into text and text into speech. Paired with reinforcement learning, it has beaten the best human players at Go.
The Relationship Between AI, ML and DL
ML and DL are the parts of AI that learn from data. ML lets a system learn from examples instead of fixed rules. DL is one family of ML methods: it also learns which features to look at, straight from raw data such as pixels or text.
They usually work together:
- A photo app that recognises your friends uses a deep network (DL) trained on labelled faces (supervised learning).
- An email service can pair a simple ML model that scores spam with a transformer that suggests replies.
- A warehouse robot arm learns by trial and error (reinforcement learning) how to grasp new objects, with a deep network reading its camera.
The same task can be solved at each level. A phishing filter can be a set of rules written by an analyst (AI without learning). It can be a model that learns how much each chosen signal matters (ML). Or it can be a network that reads the raw email and finds its own signals (DL):
Every deep learning model is a machine learning model, and every machine learning model is an AI system. The reverse is not true: a rule-based expert system is AI without any learning.
Mathematics Refresher
ML papers and documentation use the same few dozen symbols again and again. This section explains each one with an example and a small graphic. You do not need to master it all: use it as a reference when a formula gets in the way.
A vector is a list of numbers, such as ; a matrix is a grid of numbers in rows and columns, and a size such as 2×3 means 2 rows and 3 columns. Lowercase letters usually stand for numbers or vectors (, ), uppercase letters for matrices, sets and random variables (, , ). A dot () means multiplication and is often left out between a matrix and a vector, so means . Most entries can be read on their own; the set entries share the sets defined at the start of their group.
Algebraic Notation
Subscript:
A subscript usually indexes a variable: is the value of at position of a sequence, often a time step. Take five daily temperatures: , , , , . The value on day 3 is , and always means the value just before .
In AI: time series, the words of a sentence, the steps of an agent.
Superscript:
A superscript usually means a power: multiplies copies of together. So , the area of a 3 by 3 square, and . A few superscripts are labels instead, such as in and in below.
In AI: squared errors , where is the true value and ("y hat") the prediction; polynomials; exponential growth.
Norm:
A norm, written , measures the length of a vector. The usual one is the Euclidean norm , the straight-line length. Two others are common: adds up the sizes of the components, ignoring their signs, and keeps only the largest one:
For : , (moving along the grid) and (the largest component).
In AI: distances between data points, keeping a model's weights small (regularisation), scaling vectors to length 1.
Summation:
The summation sign adds up a series of terms: . With , .
In AI: averages; the loss, a single number that adds up a model's errors over every example; probabilities that add up to 1.
Logarithms and Exponentials
Logarithm Base 2:
answers "how many times must you double 1 to reach ?", or in other words "to what power must 2 be raised to give ?". Four doublings take 1 to 16 (1 → 2 → 4 → 8 → 16), so , since .
In AI: measuring information in bits. A fair coin toss carries bit; picking one of 16 equally likely options carries bits.
Natural Logarithm:
is the logarithm in base , Euler's number: the power to which must be raised to give . So and . It grows quickly at first, then more and more slowly.
In AI: turns a product into a sum, , so the many small probabilities a model multiplies can be added instead. Common loss functions for classifiers, such as cross-entropy, are built on it.
Exponential Function:
raises Euler's number to the power : and . It is always positive and grows faster and faster.
In AI: turning scores into probabilities (softmax, sigmoid), modelling growth and decay, the normal distribution.
Exponential Base 2:
doubles every time grows by 1: , , …, .
In AI: binary representations ( bits encode values) and information measured in bits.
Matrix and Vector Operations
Matrix-Vector Multiplication:
Each entry of is one row of the matrix times the vector: multiply element by element, then add. Row 1 gives , row 2 gives :
In AI: a fully connected layer of a neural network multiplies its input vector by a matrix of weights (then adds a bias).
Matrix-Matrix Multiplication:
Entry of is row of times column of :
The order matters: in general .
In AI: chaining transformations, processing a whole batch of examples at once.
Transpose:
The transpose swaps rows and columns: row of becomes column of , so a 2×3 matrix becomes a 3×2 matrix.
In AI: multiplying two vectors entry by entry and adding the results, written ; reshaping data so that matrix sizes match.
Inverse:
The inverse undoes a matrix: , the identity matrix that changes nothing. If stretches the x-axis by 2, shrinks it back by half:
In AI: solving systems of equations, for example to find in one step the straight line that best fits some data (linear regression).
Determinant:
The determinant of a square matrix (as many rows as columns) is a single number whose size tells by how much the matrix scales areas; a negative sign means the shape is also flipped. For a 2×2 matrix it is . With , : every area doubles. A determinant of 0 means the matrix squashes space flat and has no inverse.
In AI: checking whether a matrix can be inverted (its determinant is not 0) before using it to solve equations.
Trace:
The trace of a square matrix is the sum of its diagonal:
In AI: a quick summary of a matrix; for example, adding up the variances of all the features of a dataset at once.
Set Theory
A set is a collection of distinct elements, written between braces. The examples below take the universe , all the elements we are considering, and two sets inside it: , the even numbers, and , the multiples of 3.
Cardinality:
The cardinality of a set is its number of elements. For , the set above, .
In AI: counting, as in probabilities computed as "favourable cases over all cases".
Union:
The union holds every element that is in , in , or in both: .
In AI: merging datasets or lists of results.
Intersection:
The intersection holds the elements that are in both and : .
In AI: finding what two sets have in common, as in the overlap between a predicted region and the true one.
Complement:
The complement holds every element of the universe that is not in : .
In AI: probabilities of "not A", since .
Comparison Operators
Comparisons check how two values relate and answer true or false. With and :
| Maths | Code | Meaning | Result |
|---|---|---|---|
b >= a |
Greater than or equal to | true | |
a <= b |
Less than or equal to | true | |
a == b |
Equal to | false | |
a != b |
Not equal to | true |
In AI: thresholds ("flag the email if its score "), stopping conditions, filters on data.
Eigenvalues and Scalars
Lambda:
The Greek letter λ (lambda) usually stands for a scalar, a plain number that multiplies a vector, or for an eigenvalue (below). Multiplying by keeps the vector on its line and scales its length: doubles it, halves it.
In AI: λ also often names a setting chosen by hand, such as how strongly to keep a model's weights small (regularisation).
Eigenvector:
An eigenvector of a matrix is a non-zero vector that the matrix only stretches, without changing its line: , where is its eigenvalue. The matrix triples (eigenvalue 3) and leaves as it is (eigenvalue 1), but turns into , which points elsewhere.
In AI: principal component analysis (PCA) uses eigenvectors to find the directions in which the data spreads out most, so that many features can be summarised by a few.
Functions and Operators
Maximum:
returns the largest value: .
In AI: the ReLU function inside neural networks, which keeps positive values and turns negative ones into 0; its cousin returns where the maximum is, as in picking the class with the highest score.
Minimum:
returns the smallest value: .
In AI: training is a minimisation, the search for the parameters with the smallest error.
Reciprocal:
The reciprocal of is , the number that gives 1 when multiplied by : and . The bigger , the smaller .
In AI: averages ( times a sum), rates and normalisations.
Ellipsis:
The ellipsis stands for every term in between in an obvious pattern: is the sum of all the terms from to .
In AI: writing sums, sequences and vectors of any length.
Functions and Probability
Function Notation:
is the function applied to the input . With , the input 3 gives .
In AI: a trained model is a function from inputs to predictions.
Conditional Probability:
is the probability of given that is true: only the cases where holds still count. Over 10 days, it rained 4 times, so . But 5 of the days were cloudy, and it rained on 3 of them:
In AI: a phishing filter estimates , the probability that an email is phishing given what it contains.
Expectation:
A random variable is a quantity whose value depends on chance, such as the score of a dice roll. Its expected value , also called its mean, is its average outcome, each value weighted by its probability: . If takes the values 1, 2, 3, 4 with probabilities 0.1, 0.2, 0.3, 0.4:
In AI: average losses and rewards, decisions under uncertainty.
Variance:
The variance measures how spread out values are around their mean: the average of the squared distances to the mean, . The values 4, 5, 6 and 1, 5, 9 both have mean 5, but variances of and .
In AI: describing how spread out a feature is, and how uncertain a prediction is.
Standard Deviation:
The standard deviation is the square root of the variance, , so it is in the same unit as the data: about and for the two sets above. For bell-shaped (normal) data, about 68% of the values lie within one of the mean .
In AI: putting features on a common scale (mean 0, standard deviation 1) with .
Covariance:
The covariance tells whether two variables move together: . It is positive when they tend to rise together, negative when one rises as the other falls, and close to 0 when there is no linear link.
Take three students: hours of revision and marks . The means are 2 and 60, so the distances to the mean are for and for . Multiplied pair by pair, they give 10, 0 and 10:
It is positive: more revision goes with higher marks.
In AI: finding features that move together; the table of covariances between every pair of features, the covariance matrix, is the starting point of PCA.
Correlation:
The correlation is the covariance scaled to lie between −1 and 1:
Close to 1, the variables rise together along a line; close to −1, one falls as the other rises; close to 0, there is no linear relationship. For the three students above, and , so : their points lie exactly on a rising line.
In AI: spotting redundant features and the features most related to the value to predict. Correlation does not prove that one variable causes the other.