< BACK TO TERMINAL

AI Fundamentals

ID: 0x00c // DATE: 11-10-2026 // 23 MIN READ

This introduction has two parts. The first explains what artificial intelligence, machine learning and deep learning are, and how they fit together. The second is a refresher of the maths notation used in machine learning, one symbol at a time, each with a small graphic.

Introduction to Machine Learning

Artificial intelligence and machine learning are often used as if they meant the same thing. They are related but distinct, and deep learning is a third, narrower term: each one sits inside the previous one.

The labels inside each field are examples of well-known techniques; you do not need to know them to follow this introduction.

Artificial Intelligence

Artificial intelligence (AI) is the broad field of building systems that handle tasks we would call intelligent if a person did them, such as holding a conversation, recognising a face or planning a route. Its main areas:

  • Natural language processing (NLP): working with human language, as in translation, summaries or chat assistants.
  • Computer vision: making sense of images and videos, from reading a scanned document to spotting a defect on a production line.
  • Robotics: machines that sense their surroundings and move in the physical world, such as warehouse robots or robot vacuum cleaners.
  • Expert systems: programs that encode the knowledge of human experts as rules and apply them to new cases.

AI is mostly meant to support people rather than replace them: it sifts through more data than anyone could read, suggests decisions and takes over repetitive work. It is already used in:

  • Healthcare: spotting signs of disease in medical images, speeding up the search for new drugs.
  • Finance: flagging fraudulent card payments, assessing risk, managing investments.
  • Cybersecurity: detecting intrusions, phishing and malware, and helping analysts respond to them.

Machine Learning

Machine learning (ML) is the part of AI where a system learns from data instead of following rules that a programmer wrote. The algorithm looks at many examples, finds what they have in common, then uses it to judge new examples it has never seen.

Take phishing detection. Instead of writing rules by hand, we give the algorithm thousands of emails labelled "phishing" or "legitimate". It works out on its own which signals matter (number of links, age of the sender's domain, pressure words such as "urgent") and how much. These measurable signals are called features, and the importance the model gives each one is its weight. Shown a new email, it gives a score between 0 and 1: the closer to 1, the more likely it is phishing.

ML comes in three main types:

  • Supervised learning: every example comes with the right answer, its label, and the algorithm learns to give that answer itself. Examples: reading handwritten digits, estimating the price of a house, flagging phishing emails.
  • Unsupervised learning: the examples have no labels, so the algorithm looks for groups and patterns on its own. Examples: sorting news articles by topic, noticing an unusual login on a network, summarising many measurements with a few numbers.
  • Reinforcement learning: an agent (the program that acts) learns by trial and error, collecting rewards for good moves and penalties for bad ones. Examples: a program that learns to play chess or Go, a robot arm that learns to pick up objects, traffic lights that learn when to change.

ML is used in almost every industry:

Industry Applications
Healthcare Diagnosis, drug discovery, personalised treatment
Finance Fraud detection, credit risk, algorithmic trading
Marketing Customer segments, targeted ads, recommendations
Cybersecurity Threat detection, intrusion prevention, malware analysis
Transportation Traffic forecasts, autonomous vehicles, route planning

Deep Learning

Deep learning (DL) is the part of ML that uses neural networks with many layers (hence "deep") to learn directly from complex data. It works best on raw, messy data such as photos, sound recordings and text, where each example is made of thousands of values.

Three things set deep learning apart:

  • Hierarchical features: each layer builds on the one before, from simple to abstract. On a screenshot of a web page, the first layers react to edges, the next to text blocks and input fields, the last to a logo above a password field.
  • End-to-end learning: the network goes straight from raw input to answer, without a person deciding which features to measure.
  • Scalability: it keeps improving as data and computing power grow, where classic algorithms level off.

The most common kinds of networks:

  • Convolutional neural networks (CNNs): built for images and video. Small filters slide over the image to detect local patterns, and stacked layers combine them into larger shapes.
  • Recurrent neural networks (RNNs): built for sequences such as text and speech. A loop carries information from one step to the next, as a short-term memory.
  • Transformers: the most recent of the three, now the standard for language. Self-attention lets each word look at every word of the text, itself included, however far apart, and decide how much each one matters. The large language models behind chat assistants are transformers.

Deep learning is behind most recent progress in AI. It finds and outlines objects in photos. It translates, summarises and writes text. It turns speech into text and text into speech. Paired with reinforcement learning, it has beaten the best human players at Go.

The Relationship Between AI, ML and DL

ML and DL are the parts of AI that learn from data. ML lets a system learn from examples instead of fixed rules. DL is one family of ML methods: it also learns which features to look at, straight from raw data such as pixels or text.

They usually work together:

  • A photo app that recognises your friends uses a deep network (DL) trained on labelled faces (supervised learning).
  • An email service can pair a simple ML model that scores spam with a transformer that suggests replies.
  • A warehouse robot arm learns by trial and error (reinforcement learning) how to grasp new objects, with a deep network reading its camera.

The same task can be solved at each level. A phishing filter can be a set of rules written by an analyst (AI without learning). It can be a model that learns how much each chosen signal matters (ML). Or it can be a network that reads the raw email and finds its own signals (DL):

Every deep learning model is a machine learning model, and every machine learning model is an AI system. The reverse is not true: a rule-based expert system is AI without any learning.

Mathematics Refresher

ML papers and documentation use the same few dozen symbols again and again. This section explains each one with an example and a small graphic. You do not need to master it all: use it as a reference when a formula gets in the way.

A vector is a list of numbers, such as v=(3,4)v = (3, 4); a matrix is a grid of numbers in rows and columns, and a size such as 2×3 means 2 rows and 3 columns. Lowercase letters usually stand for numbers or vectors (xx, vv), uppercase letters for matrices, sets and random variables (AA, SS, XX). A dot (⋅\cdot) means multiplication and is often left out between a matrix and a vector, so AvAv means A⋅vA \cdot v. Most entries can be read on their own; the set entries share the sets defined at the start of their group.

Algebraic Notation

Subscript: xtx_t

A subscript usually indexes a variable: xtx_t is the value of xx at position tt of a sequence, often a time step. Take five daily temperatures: x1=12x_1 = 12, x2=14x_2 = 14, x3=13x_3 = 13, x4=15x_4 = 15, x5=16x_5 = 16. The value on day 3 is x3=13x_3 = 13, and xt−1x_{t-1} always means the value just before xtx_t.

In AI: time series, the words of a sentence, the steps of an agent.

Superscript: xnx^n

A superscript usually means a power: xnx^n multiplies nn copies of xx together. So 32=3⋅3=93^2 = 3 \cdot 3 = 9, the area of a 3 by 3 square, and 25=322^5 = 32. A few superscripts are labels instead, such as TT in ATA^T and cc in AcA^c below.

In AI: squared errors (y−y^)2{(y - \hat{y})^2}, where yy is the true value and y^\hat{y} ("y hat") the prediction; polynomials; exponential growth.

Norm: ∥v∥\lVert v \rVert

A norm, written ∥v∥\lVert v \rVert, measures the length of a vector. The usual one is the Euclidean norm ∥v∥2\lVert v \rVert_2, the straight-line length. Two others are common: ∥v∥1\lVert v \rVert_1 adds up the sizes of the components, ignoring their signs, and ∥v∥∞\lVert v \rVert_\infty keeps only the largest one:

∥v∥2=v12+⋯+vn2\lVert v \rVert_2 = \sqrt{v_1^2 + \dots + v_n^2} ∥v∥1=∣v1∣+⋯+∣vn∣\lVert v \rVert_1 = |v_1| + \dots + |v_n| ∥v∥∞=max⁡i∣vi∣\lVert v \rVert_\infty = \max_i |v_i|

For v=(3,4)v = (3, 4): ∥v∥2=9+16=5{\lVert v \rVert_2 = \sqrt{9 + 16} = 5}, ∥v∥1=3+4=7{\lVert v \rVert_1 = 3 + 4 = 7} (moving along the grid) and ∥v∥∞=4{\lVert v \rVert_\infty = 4} (the largest component).

In AI: distances between data points, keeping a model's weights small (regularisation), scaling vectors to length 1.

Summation: Σ\Sigma

The summation sign adds up a series of terms: ∑i=1nai=a1+a2+⋯+an\sum_{i=1}^{n} a_i = a_1 + a_2 + \dots + a_n. With a=(2,4,1,3)a = (2, 4, 1, 3), ∑i=14ai=10\sum_{i=1}^{4} a_i = 10.

In AI: averages; the loss, a single number that adds up a model's errors over every example; probabilities that add up to 1.

Logarithms and Exponentials

Logarithm Base 2: log⁡2(x)\log_2(x)

log⁡2(x)\log_2(x) answers "how many times must you double 1 to reach xx?", or in other words "to what power must 2 be raised to give xx?". Four doublings take 1 to 16 (1 → 2 → 4 → 8 → 16), so log⁡2(16)=4\log_2(16) = 4, since 24=162^4 = 16.

In AI: measuring information in bits. A fair coin toss carries log⁡2(2)=1\log_2(2) = 1 bit; picking one of 16 equally likely options carries log⁡2(16)=4\log_2(16) = 4 bits.

Natural Logarithm: ln⁡(x)\ln(x)

ln⁡(x)\ln(x) is the logarithm in base e≈2.718e \approx 2.718, Euler's number: the power to which ee must be raised to give xx. So ln⁡(1)=0\ln(1) = 0 and ln⁡(e)=1\ln(e) = 1. It grows quickly at first, then more and more slowly.

In AI: ln⁡\ln turns a product into a sum, ln⁡(a⋅b)=ln⁡(a)+ln⁡(b){\ln(a \cdot b) = \ln(a) + \ln(b)}, so the many small probabilities a model multiplies can be added instead. Common loss functions for classifiers, such as cross-entropy, are built on it.

Exponential Function: exe^x

exe^x raises Euler's number to the power xx: e0=1{e^0 = 1} and e1≈2.718{e^1 \approx 2.718}. It is always positive and grows faster and faster.

In AI: turning scores into probabilities (softmax, sigmoid), modelling growth and decay, the normal distribution.

Exponential Base 2: 2x2^x

2x2^x doubles every time xx grows by 1: 20=12^0 = 1, 21=22^1 = 2, …, 25=322^5 = 32.

In AI: binary representations (nn bits encode 2n2^n values) and information measured in bits.

Matrix and Vector Operations

Matrix-Vector Multiplication: AvAv

Each entry of AvAv is one row of the matrix times the vector: multiply element by element, then add. Row 1 gives 2⋅1+1⋅2=4{2 \cdot 1} + {1 \cdot 2} = 4, row 2 gives 0⋅1+3⋅2=6{0 \cdot 1} + {3 \cdot 2} = 6:

[2103][12]=[46]\begin{bmatrix} 2 & 1 \\ 0 & 3 \end{bmatrix} \begin{bmatrix} 1 \\ 2 \end{bmatrix} = \begin{bmatrix} 4 \\ 6 \end{bmatrix}

In AI: a fully connected layer of a neural network multiplies its input vector by a matrix of weights (then adds a bias).

Matrix-Matrix Multiplication: ABAB

Entry (i,j)(i, j) of ABAB is row ii of AA times column jj of BB:

[2103][1021]=[4163]\begin{bmatrix} 2 & 1 \\ 0 & 3 \end{bmatrix} \begin{bmatrix} 1 & 0 \\ 2 & 1 \end{bmatrix} = \begin{bmatrix} 4 & 1 \\ 6 & 3 \end{bmatrix}

The order matters: in general AB≠BAAB \neq BA.

In AI: chaining transformations, processing a whole batch of examples at once.

Transpose: ATA^T

The transpose swaps rows and columns: row ii of AA becomes column ii of ATA^T, so a 2×3 matrix becomes a 3×2 matrix.

[123456]T=[142536]\begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \end{bmatrix}^T = \begin{bmatrix} 1 & 4 \\ 2 & 5 \\ 3 & 6 \end{bmatrix}

In AI: multiplying two vectors entry by entry and adding the results, written uTv=u1v1+⋯+unvnu^T v = u_1 v_1 + \dots + u_n v_n; reshaping data so that matrix sizes match.

Inverse: A−1A^{-1}

The inverse undoes a matrix: A⋅A−1=IA \cdot A^{-1} = I, the identity matrix that changes nothing. If AA stretches the x-axis by 2, A−1A^{-1} shrinks it back by half:

A=[2001]A−1=[0.5001]A = \begin{bmatrix} 2 & 0 \\ 0 & 1 \end{bmatrix} \quad A^{-1} = \begin{bmatrix} 0.5 & 0 \\ 0 & 1 \end{bmatrix}

In AI: solving systems of equations, for example to find in one step the straight line that best fits some data (linear regression).

Determinant: det⁡(A)\det(A)

The determinant of a square matrix (as many rows as columns) is a single number whose size tells by how much the matrix scales areas; a negative sign means the shape is also flipped. For a 2×2 matrix [abcd]\begin{bmatrix} a & b \\ c & d \end{bmatrix} it is ad−bcad - bc. With A=[2101]A = \begin{bmatrix} 2 & 1 \\ 0 & 1 \end{bmatrix}, det⁡(A)=2⋅1−1⋅0=2\det(A) = 2 \cdot 1 - 1 \cdot 0 = 2: every area doubles. A determinant of 0 means the matrix squashes space flat and has no inverse.

In AI: checking whether a matrix can be inverted (its determinant is not 0) before using it to solve equations.

Trace: tr⁡(A)\operatorname{tr}(A)

The trace of a square matrix is the sum of its diagonal:

tr⁡[410253106]=4+5+6=15\operatorname{tr}\begin{bmatrix} 4 & 1 & 0 \\ 2 & 5 & 3 \\ 1 & 0 & 6 \end{bmatrix} = 4 + 5 + 6 = 15

In AI: a quick summary of a matrix; for example, adding up the variances of all the features of a dataset at once.

Set Theory

A set is a collection of distinct elements, written between braces. The examples below take the universe U={1,2,3,4,5,6}U = \{1, 2, 3, 4, 5, 6\}, all the elements we are considering, and two sets inside it: A={2,4,6}A = \{2, 4, 6\}, the even numbers, and B={3,6}B = \{3, 6\}, the multiples of 3.

Cardinality: ∣S∣|S|

The cardinality of a set SS is its number of elements. For S={2,4,6}S = \{2, 4, 6\}, the set AA above, ∣S∣=3|S| = 3.

In AI: counting, as in probabilities computed as "favourable cases over all cases".

Union: A∪BA \cup B

The union holds every element that is in AA, in BB, or in both: A∪B={2,3,4,6}A \cup B = \{2, 3, 4, 6\}.

In AI: merging datasets or lists of results.

Intersection: A∩BA \cap B

The intersection holds the elements that are in both AA and BB: A∩B={6}A \cap B = \{6\}.

In AI: finding what two sets have in common, as in the overlap between a predicted region and the true one.

Complement: AcA^c

The complement holds every element of the universe that is not in AA: Ac={1,3,5}A^c = \{1, 3, 5\}.

In AI: probabilities of "not A", since P(Ac)=1−P(A)P(A^c) = 1 - P(A).

Comparison Operators

Comparisons check how two values relate and answer true or false. With a=3a = 3 and b=5b = 5:

Maths Code Meaning Result
b≥ab \geq a b >= a Greater than or equal to true
a≤ba \leq b a <= b Less than or equal to true
a=ba = b a == b Equal to false
a≠ba \neq b a != b Not equal to true

In AI: thresholds ("flag the email if its score ≥0.5\geq 0.5"), stopping conditions, filters on data.

Eigenvalues and Scalars

Lambda: λ\lambda

The Greek letter λ (lambda) usually stands for a scalar, a plain number that multiplies a vector, or for an eigenvalue (below). Multiplying by λ\lambda keeps the vector on its line and scales its length: λ=2{\lambda = 2} doubles it, λ=0.5{\lambda = 0.5} halves it.

In AI: λ also often names a setting chosen by hand, such as how strongly to keep a model's weights small (regularisation).

Eigenvector: Av=λvAv = \lambda v

An eigenvector of a matrix is a non-zero vector that the matrix only stretches, without changing its line: A⋅v=λ⋅v{A \cdot v = \lambda \cdot v}, where λ\lambda is its eigenvalue. The matrix A=[3001]A = \begin{bmatrix} 3 & 0 \\ 0 & 1 \end{bmatrix} triples (1,0)(1, 0) (eigenvalue 3) and leaves (0,1)(0, 1) as it is (eigenvalue 1), but turns (1,1)(1, 1) into (3,1)(3, 1), which points elsewhere.

In AI: principal component analysis (PCA) uses eigenvectors to find the directions in which the data spreads out most, so that many features can be summarised by a few.

Functions and Operators

Maximum: max⁡(… )\max(\dots)

max⁡\max returns the largest value: max⁡(4,9,2,6)=9\max(4, 9, 2, 6) = 9.

In AI: the ReLU function max⁡(0,x){\max(0, x)} inside neural networks, which keeps positive values and turns negative ones into 0; its cousin arg⁡max⁡\arg\max returns where the maximum is, as in picking the class with the highest score.

Minimum: min⁡(… )\min(\dots)

min⁡\min returns the smallest value: min⁡(4,9,2,6)=2\min(4, 9, 2, 6) = 2.

In AI: training is a minimisation, the search for the parameters with the smallest error.

Reciprocal: 1/x1/x

The reciprocal of xx is 1/x1/x, the number that gives 1 when multiplied by xx: 1/2=0.5{1/2 = 0.5} and 1/4=0.25{1/4 = 0.25}. The bigger xx, the smaller 1/x1/x.

In AI: averages (1n\frac{1}{n} times a sum), rates and normalisations.

Ellipsis: …\dots

The ellipsis stands for every term in between in an obvious pattern: a1+a2+⋯+an{a_1 + a_2 + \dots + a_n} is the sum of all the terms from a1a_1 to ana_n.

In AI: writing sums, sequences and vectors of any length.

Functions and Probability

Function Notation: f(x)f(x)

f(x)f(x) is the function ff applied to the input xx. With f(x)=2x+1f(x) = 2x + 1, the input 3 gives f(3)=7f(3) = 7.

In AI: a trained model is a function from inputs to predictions.

Conditional Probability: P(x∣y){P(x \mid y)}

P(x∣y){P(x \mid y)} is the probability of xx given that yy is true: only the cases where yy holds still count. Over 10 days, it rained 4 times, so P(rain)=0.4P(\text{rain}) = 0.4. But 5 of the days were cloudy, and it rained on 3 of them:

P(rain∣cloudy)=35=0.6P(\text{rain} \mid \text{cloudy}) = \frac{3}{5} = 0.6

In AI: a phishing filter estimates P(phishing∣email){P(\text{phishing} \mid \text{email})}, the probability that an email is phishing given what it contains.

Expectation: E[X]E[X]

A random variable XX is a quantity whose value depends on chance, such as the score of a dice roll. Its expected value E[X]E[X], also called its mean, is its average outcome, each value weighted by its probability: E[X]=∑ixi P(xi)E[X] = \sum_i x_i \, P(x_i). If XX takes the values 1, 2, 3, 4 with probabilities 0.1, 0.2, 0.3, 0.4:

E[X]=1⋅0.1+2⋅0.2+3⋅0.3+4⋅0.4=3\begin{aligned} E[X] &= 1 \cdot 0.1 + 2 \cdot 0.2 \\ &\quad + 3 \cdot 0.3 + 4 \cdot 0.4 = 3 \end{aligned}

In AI: average losses and rewards, decisions under uncertainty.

Variance: Var⁡(X)\operatorname{Var}(X)

The variance measures how spread out values are around their mean: the average of the squared distances to the mean, Var⁡(X)=E[(X−E[X])2]\operatorname{Var}(X) = {E[(X - E[X])^2]}. The values 4, 5, 6 and 1, 5, 9 both have mean 5, but variances of 0.670.67 and 10.6710.67.

In AI: describing how spread out a feature is, and how uncertain a prediction is.

Standard Deviation: σ(X)\sigma(X)

The standard deviation is the square root of the variance, σ(X)=Var⁡(X){\sigma(X) = \sqrt{\operatorname{Var}(X)}}, so it is in the same unit as the data: about 0.820.82 and 3.273.27 for the two sets above. For bell-shaped (normal) data, about 68% of the values lie within one σ\sigma of the mean μ\mu.

In AI: putting features on a common scale (mean 0, standard deviation 1) with z=(x−μ)/σz = {(x - \mu)} / \sigma.

Covariance: Cov⁡(X,Y)\operatorname{Cov}(X, Y)

The covariance tells whether two variables move together: Cov⁡(X,Y)=E[(X−E[X])(Y−E[Y])]\operatorname{Cov}(X, Y) = {E[(X - E[X])(Y - E[Y])]}. It is positive when they tend to rise together, negative when one rises as the other falls, and close to 0 when there is no linear link.

Take three students: hours of revision X=1,2,3X = 1, 2, 3 and marks Y=50,60,70Y = 50, 60, 70. The means are 2 and 60, so the distances to the mean are −1,0,1-1, 0, 1 for XX and −10,0,10-10, 0, 10 for YY. Multiplied pair by pair, they give 10, 0 and 10:

Cov⁡(X,Y)=10+0+103≈6.67\operatorname{Cov}(X, Y) = \frac{10 + 0 + 10}{3} \approx 6.67

It is positive: more revision goes with higher marks.

In AI: finding features that move together; the table of covariances between every pair of features, the covariance matrix, is the starting point of PCA.

Correlation: ρ(X,Y)\rho(X, Y)

The correlation is the covariance scaled to lie between −1 and 1:

ρ(X,Y)=Cov⁡(X,Y)σ(X) σ(Y)\rho(X, Y) = \frac{\operatorname{Cov}(X, Y)}{\sigma(X)\,\sigma(Y)}

Close to 1, the variables rise together along a line; close to −1, one falls as the other rises; close to 0, there is no linear relationship. For the three students above, σ(X)≈0.82\sigma(X) \approx 0.82 and σ(Y)≈8.16\sigma(Y) \approx 8.16, so ρ(X,Y)=1\rho(X, Y) = 1: their points lie exactly on a rising line.

In AI: spotting redundant features and the features most related to the value to predict. Correlation does not prove that one variable causes the other.

← Previous article