Machine Learning & Artificial Intelligence

Supervised and unsupervised learning, model evaluation, optimization, neural networks, deep learning, computer vision, NLP and the foundations behind modern AI systems.

Machine learning is one of the areas where I found it particularly useful to combine practical experimentation with a return to mathematical foundations.

It is relatively easy to train a model using a modern library. The more interesting questions begin when we try to understand why a model works, why it fails, how it generalizes to unseen data and what actually happens during training.

What is a model actually learning? Why does gradient descent improve it? How do we know whether it will generalize? When should we use classification, regression or clustering? Why do neural networks need nonlinear activation functions? What makes deep learning different from traditional machine learning?

This section combines two complementary perspectives: a practical path through modern machine-learning workflows and a deeper theoretical view of the mathematics and principles behind deep learning.


Topics in This Section

Supervised Learning · Unsupervised Learning · Classification · Regression · Clustering · Semi-Supervised Learning · Model Evaluation · Cross-Validation · Bias & Variance · Feature Engineering · Gradient Descent · Logistic Regression · Support Vector Machines · Decision Trees · Random Forests · Ensemble Learning · Dimensionality Reduction · PCA · Neural Networks · Backpropagation · Optimization · Regularization · CNNs · Sequence Models · Attention · NLP · Autoencoders · GANs · Diffusion Models · Reinforcement Learning


Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow

Aurélien Géron

3rd Edition — O’Reilly Media

Level
Foundation → Advanced

Best for
Learning machine learning through practical examples while progressively understanding the concepts behind the algorithms.

Hands-On Machine Learning is the book I use when I want to move quickly between a machine-learning concept and an actual implementation.

It starts from the fundamentals of supervised and unsupervised learning, moves through classical algorithms and model evaluation, and then builds toward neural networks, computer vision, natural-language processing, generative models and reinforcement learning.

What I particularly like is that the algorithms are rarely presented in isolation. The book repeatedly connects model selection, data preparation, evaluation and optimization into a complete machine-learning workflow.

What I Use It For

  • understanding the main categories of machine learning;
  • supervised versus unsupervised learning;
  • classification and regression;
  • training, validation and test sets;
  • cross-validation;
  • precision, recall, F1 and ROC analysis;
  • gradient descent;
  • linear and logistic regression;
  • Support Vector Machines;
  • decision trees and random forests;
  • ensemble learning and boosting;
  • dimensionality reduction and PCA;
  • clustering and anomaly detection;
  • semi-supervised learning;
  • neural networks and backpropagation;
  • deep-learning optimization;
  • CNNs;
  • sequence models;
  • NLP and attention;
  • autoencoders and generative models;
  • reinforcement learning;
  • training and deploying models at scale.

Chapters Worth Reading

Machine-Learning Foundations

Chapter 1 — The Machine Learning Landscape
An introduction to what machine learning is, the main categories of learning systems and the challenges involved in building models that generalize.

  • supervised and unsupervised learning;
  • batch and online learning;
  • instance-based and model-based learning;
  • overfitting and underfitting;
  • training and test data;
  • hyperparameter tuning;
  • model selection.
End-to-End Machine-Learning Workflow

Chapter 2 — End-to-End Machine Learning Project
A complete workflow from defining the problem to preparing data, selecting models and evaluating the final system.

  • problem framing;
  • data exploration;
  • data cleaning;
  • feature transformation;
  • training pipelines;
  • cross-validation;
  • hyperparameter tuning;
  • model evaluation.
Classification

Chapter 3 — Classification
How classification models are trained and evaluated when the goal is to predict discrete categories.

  • binary classification;
  • multiclass classification;
  • confusion matrices;
  • precision and recall;
  • F1 score;
  • ROC curves;
  • error analysis.
Regression & Gradient Descent

Chapter 4 — Training Models
The mathematical mechanisms behind several fundamental machine-learning algorithms.

  • linear regression;
  • Normal Equation;
  • batch gradient descent;
  • stochastic gradient descent;
  • mini-batch gradient descent;
  • polynomial regression;
  • learning curves;
  • regularization;
  • logistic regression;
  • softmax regression.
Support Vector Machines & Decision Trees

Chapter 5 — Support Vector Machines
Linear and nonlinear classification, margins, kernels and SVM regression.

Chapter 6 — Decision Trees
Tree-based classification and regression, decision boundaries and the strengths and limitations of tree models.

Ensemble Learning

Chapter 7 — Ensemble Learning and Random Forests
Combining multiple models to create a stronger predictive system.

  • voting classifiers;
  • bagging;
  • Random Forests;
  • Extra-Trees;
  • AdaBoost;
  • Gradient Boosting;
  • stacking.
Dimensionality Reduction

Chapter 8 — Dimensionality Reduction
Reducing the number of features while attempting to preserve the information that matters.

  • curse of dimensionality;
  • projection;
  • Principal Component Analysis;
  • explained variance;
  • Random Projection;
  • Locally Linear Embedding.
Clustering, Semi-Supervised Learning & Anomaly Detection

Chapter 9 — Unsupervised Learning Techniques
Learning structure from data without relying entirely on labelled examples.

  • k-means;
  • DBSCAN;
  • Gaussian Mixture Models;
  • clustering;
  • semi-supervised learning;
  • anomaly detection;
  • novelty detection.
Neural Networks & Backpropagation

Chapter 10 — Introduction to Artificial Neural Networks with Keras
The transition from classical machine learning to neural networks.

  • artificial neurons;
  • Perceptrons;
  • multilayer perceptrons;
  • backpropagation;
  • activation functions;
  • classification and regression networks;
  • network architecture;
  • hyperparameter tuning.
Training Deep Neural Networks

Chapter 11 — Training Deep Neural Networks
The practical problems involved in training deeper models efficiently and reliably.

  • vanishing and exploding gradients;
  • weight initialization;
  • normalization;
  • gradient clipping;
  • transfer learning;
  • optimizers;
  • learning-rate scheduling;
  • regularization.
Computer Vision

Chapter 14 — Deep Computer Vision Using Convolutional Neural Networks
Convolutional architectures and the mechanisms that made deep learning particularly effective for visual information.

  • convolutional layers;
  • pooling;
  • CNN architectures;
  • ResNet;
  • transfer learning;
  • object detection;
  • semantic segmentation.
Sequences, NLP & Attention

Chapter 15 — Processing Sequences Using RNNs and CNNs
Models designed for sequential and time-dependent information.

Chapter 16 — Natural Language Processing with RNNs and Attention
Text processing, language models and attention-based approaches.

Generative AI Foundations

Chapter 17 — Autoencoders, GANs, and Diffusion Models
Several major approaches for learning representations and generating new data.

  • autoencoders;
  • variational autoencoders;
  • GANs;
  • generative modelling;
  • diffusion models.
Reinforcement Learning

Chapter 18 — Reinforcement Learning
Learning through interaction, rewards and sequential decision-making.

  • policies;
  • rewards;
  • Markov Decision Processes;
  • Q-Learning;
  • Deep Q-Learning;
  • policy gradients.
Training & Deployment at Scale

Chapter 19 — Training and Deploying TensorFlow Models at Scale
Moving from experimentation toward model serving and distributed training.

  • model serving;
  • cloud deployment;
  • GPU acceleration;
  • distributed training;
  • data parallelism;
  • hyperparameter tuning at scale.

My Suggested Learning Path

Machine-Learning Foundations
Chapters 1–3

How Models Learn
Chapter 4

Classical Algorithms
Chapters 5–7

Unsupervised Learning
Chapters 8–9

Neural Networks
Chapters 10–11

Computer Vision & NLP
Chapters 14–16

Generative Models
Chapter 17

Reinforcement Learning & Deployment
Chapters 18–19


Deep Learning

Ian Goodfellow, Yoshua Bengio & Aaron Courville

MIT Press

Level
Intermediate → Advanced

Best for
Understanding the mathematical and conceptual foundations behind neural networks and deep-learning algorithms.

Deep Learning is the book I use when I want to move below the framework APIs and understand the mathematics and principles behind neural networks.

It begins by rebuilding the mathematical foundations needed for machine learning, then develops the theory behind feedforward networks, regularization, optimization, convolutional networks and sequence models.

For me, its role in this library is different from Géron’s book: it is less about getting a model running quickly and more about understanding why the underlying methods behave as they do.

What I Use It For

  • linear algebra for machine learning;
  • probability and information theory;
  • numerical computation;
  • maximum likelihood and estimation;
  • bias and variance;
  • gradient-based learning;
  • deep feedforward networks;
  • backpropagation;
  • regularization;
  • optimization of neural networks;
  • convolutional networks;
  • sequence modelling;
  • practical deep-learning methodology;
  • representation learning.

Chapters Worth Reading

Mathematical Foundations

Chapter 2 — Linear Algebra
Vectors, matrices, tensors, norms, eigendecomposition, Singular Value Decomposition and the mathematical language used throughout machine learning.

Chapter 3 — Probability and Information Theory
Random variables, probability distributions, expectation, Bayes’ rule and information-theoretic concepts.

Chapter 4 — Numerical Computation
Numerical stability, gradient-based optimization and the computational foundations behind model training.

Machine-Learning Foundations

Chapter 5 — Machine Learning Basics
The theoretical foundations needed to reason about learning algorithms and generalization.

  • learning algorithms;
  • capacity;
  • overfitting and underfitting;
  • hyperparameters;
  • estimators;
  • bias and variance;
  • maximum likelihood;
  • supervised learning;
  • unsupervised learning;
  • stochastic gradient descent.
Deep Feedforward Networks

Chapter 6 — Deep Feedforward Networks
The central architecture behind multilayer neural networks.

  • feedforward networks;
  • hidden units;
  • activation functions;
  • output units;
  • cost functions;
  • backpropagation;
  • gradient computation.
Regularization

Chapter 7 — Regularization for Deep Learning
Techniques designed to improve generalization and reduce overfitting.

  • parameter regularization;
  • data augmentation;
  • robustness to noise;
  • early stopping;
  • parameter sharing;
  • dropout;
  • ensemble methods.
Optimization

Chapter 8 — Optimization for Training Deep Models
Why optimization becomes difficult in deep networks and the techniques used to make training practical.

  • gradient-based optimization;
  • stochastic gradient descent;
  • momentum;
  • parameter initialization;
  • adaptive learning rates;
  • optimization strategies.
Convolutional Networks

Chapter 9 — Convolutional Networks
The principles behind neural architectures designed to exploit spatial structure.

  • convolution;
  • pooling;
  • structured outputs;
  • efficient convolution;
  • representation learning for images.
Sequence Modelling

Chapter 10 — Sequence Modeling: Recurrent and Recursive Nets
Neural architectures designed to model sequences and dependencies over time.

  • recurrent neural networks;
  • bidirectional RNNs;
  • encoder-decoder architectures;
  • long-term dependencies;
  • LSTM and gated units.
Practical Methodology

Chapter 11 — Practical Methodology
A systematic approach to diagnosing and improving machine-learning systems.

  • performance metrics;
  • baseline models;
  • data collection;
  • hyperparameter selection;
  • debugging strategies.

My Suggested Learning Path

Mathematical Foundations
Chapters 2–4

Machine-Learning Theory
Chapter 5

Neural-Network Foundations
Chapter 6

Generalization & Optimization
Chapters 7–8

Vision & Sequences
Chapters 9–10

Practical Methodology
Chapter 11


Why I Keep Both Books

Hands-On Machine Learning

Practice first.

  • complete ML workflow;
  • classification and regression;
  • model evaluation;
  • classical algorithms;
  • clustering;
  • neural networks;
  • computer vision;
  • NLP;
  • generative models;
  • reinforcement learning.

Deep Learning

Foundations first.

  • linear algebra;
  • probability;
  • numerical optimization;
  • learning theory;
  • backpropagation;
  • regularization;
  • deep-network optimization;
  • convolution;
  • sequence modelling.

One book helps me build and experiment with models. The other helps me understand the mathematics and principles that make those models possible.


Topic → Book Map

Machine-Learning Foundations

Hands-On Machine Learning: Chapters 1–2
Deep Learning: Chapter 5

Classification

Hands-On Machine Learning: Chapters 3–7
Deep Learning: Chapters 5–6 for theoretical foundations

Regression

Hands-On Machine Learning: Chapters 4–6
Deep Learning: Chapters 5–6

Gradient Descent & Optimization

Hands-On Machine Learning: Chapters 4 and 11
Deep Learning: Chapters 4 and 8

Clustering & Unsupervised Learning

Hands-On Machine Learning: Chapters 8–9
Deep Learning: Chapters 5 and 15 for broader foundations

Neural Networks & Backpropagation

Hands-On Machine Learning: Chapters 10–11
Deep Learning: Chapters 6–8

Computer Vision

Hands-On Machine Learning: Chapter 14
Deep Learning: Chapter 9

Sequence Models

Hands-On Machine Learning: Chapters 15–16
Deep Learning: Chapter 10

Regularization & Generalization

Hands-On Machine Learning: Chapters 4 and 11
Deep Learning: Chapters 5 and 7


How I Use These Books

I find machine learning easier to understand when I begin with the behaviour of a model and then trace that behaviour back to the data, objective function and optimization process behind it.

“The model performs extremely well on the training data but badly on new examples.”

That leads to overfitting, model capacity, regularization, validation strategies and the bias-variance trade-off.

“Fraudulent transactions represent only a tiny fraction of the dataset. Is accuracy still a useful metric?”

That points toward class imbalance, confusion matrices, precision, recall, F1 score and threshold selection.

“How does changing the model parameters actually reduce the prediction error?”

That leads to loss functions, gradients, backpropagation and gradient-based optimization.

“I have thousands of unlabelled observations. Can they still reveal useful structure?”

That opens the door to clustering, dimensionality reduction, anomaly detection and semi-supervised approaches.

Start from the model behaviour, identify the mathematical mechanism behind it, then return to the implementation with a clearer mental model.


Related Areas

Machine learning connects naturally with several other areas in this library.

Algorithms & Data Structures

Optimization, search, computational complexity and the algorithms underlying model training.

Databases & Data Management

Data preparation, storage, feature extraction, analytical workloads and large-scale datasets.

Distributed Systems

Distributed training, data parallelism, model serving and large-scale processing.

Containers, Cloud & DevOps

Model deployment, GPU workloads, reproducible environments and scalable inference infrastructure.

Cybersecurity & Cryptography

Adversarial attacks, privacy, secure training and protection of sensitive data.

Privacy-Enhancing Technologies

Differential privacy, federated learning and privacy-preserving machine learning.


← Back to My Technical Library