Start with the structure of the subject

Machine learning is easier to study when you treat it as a set of connected decisions rather than a catalogue of algorithms. For almost every problem, you need to answer the same questions:

  1. What is the prediction or decision target?
  2. What data is available, and what could make it misleading?
  3. Which model is suitable for the data and objective?
  4. How will performance be measured without leakage?
  5. How will the model be improved, interpreted and deployed safely?

This structure gives you a way to organise lectures, textbook chapters, coding exercises and past questions. It also prevents a common mistake: memorising that an algorithm exists without knowing when to use it or how to judge its output.

There is no single international machine learning exam structure. A university assessment may combine mathematical derivations, programming, short-answer questions, model interpretation or a project. Check your course outline, assessment rubric and official syllabus for the exact weighting and permitted materials. The method below works across those formats because it separates knowledge, calculation, implementation and explanation.

Build one board for each connected topic

Do not begin by making a flashcard for every sentence in your notes. First make a board that shows the relationships between ideas. Group material into sections such as problem formulation, data preparation, models, evaluation and limitations.

Use one example dataset throughout a topic where possible. For instance, a binary classifier predicting whether a transaction is fraudulent can connect class imbalance, logistic regression, threshold selection, precision, recall, calibration and data leakage. One consistent example is easier to reason about than a collection of unrelated definitions.

A board on this topic ends up looking like this:

Machine Learning Supervised learning: classificationStudy
Core methods and evaluationSupervised learning: classification5 sections · 3 columns
Problem formulation4 due
  • Target — predict fraud y ∈ {0,1} from transaction features x
  • Split — train, validation and test data; fit preprocessing on training data only
  • Baseline — majority-class classifier and a simple logistic model
Regularisation3 due
  • L2 — adds λ||w||²; shrinks coefficients towards zero
  • L1 — adds λ||w||₁; can produce exact zero coefficients
  • Increase λ to reduce variance, but excessive regularisation can underfit
Logistic regression

p(y=1|x) = σ(wᵀx + b), where σ(z) = 1/(1+e⁻ᶻ). The threshold changes predicted labels but not the fitted probabilities.

Failure modes
  • Leakage — a feature contains information unavailable when the prediction is made
  • Imbalance — accuracy can be misleading when fraud is rare
  • Calibration — predicted probabilities should reflect observed frequencies
Metrics
MetricFormula / meaning
PrecisionTP/(TP+FP): of flagged transactions, how many are fraud?
RecallTP/(TP+FN): of fraud cases, how many were found?
F12PR/(P+R), the harmonic mean of precision and recall
ROC-AUCRanking performance across classification thresholds

The board is not a summary to reread passively. It is the source for questions, calculations and explanations. Add a short worked example to each difficult section: calculate precision from a confusion matrix, show one gradient update, or explain why a random split is unsafe for time-ordered data.

Reduce the board to a must-know core

After organising a topic, reduce it to the claims you must be able to use without searching. A useful core contains relationships and conditions, not just vocabulary. “Random forest” is a label; “many decorrelated decision trees reduce variance by averaging, but can still inherit biased training data” is usable knowledge.

Keep formulas with their assumptions. For example, know what each metric measures and when it can mislead. Know that cross-validation helps estimate generalisation, but it does not repair leakage if preprocessing or duplicate observations cross the folds. Know the difference between a model parameter, such as a learned weight, and a hyperparameter, such as the regularisation strength.

Here is the type of core checklist to produce after studying a classification block:

A concise core checklist for this topic looks like this:

Must not miss coreSupervised learning: classification
Separate training, validation and test roles; fit scaling, imputation and feature selection inside the training process to avoid leakage
For binary logistic regression, p(y=1|x) = σ(wᵀx+b); changing the decision threshold changes precision and recall, not the fitted probabilities
Accuracy = (TP+TN)/(TP+TN+FP+FN); precision = TP/(TP+FP); recall = TP/(TP+FN); F1 = 2PR/(P+R)
L1 regularisation can set coefficients to zero; L2 usually shrinks them without producing exact zeros
Use stratified splitting for class imbalance when observations are independent; use time-based or grouped splitting when the data-generating structure requires it

Test the core in both directions. Given a formula, explain its practical meaning. Given a practical problem, select a metric or validation method and justify it. This two-way practice matters because assessments often present the situation first and ask for the method second.

Include maths only at the level you need to use

Machine learning study becomes inefficient when you either avoid all mathematics or copy derivations without understanding them. For each formula, record four things:

  • what each symbol means;
  • what the formula calculates;
  • what assumptions or conditions apply;
  • how the result changes when one input changes.

For gradient descent, for example, understand that parameters are updated in the direction that reduces the loss, that the learning rate controls the step size, and that a poor learning rate can make training unstable or painfully slow. For a decision tree, be able to explain how a split reduces impurity and why an unrestricted tree can overfit.

Then calculate small examples by hand before relying on code. A two-feature linear model, a four-cell confusion matrix or one step of gradient descent is enough to expose misunderstandings quickly.

Turn each section into retrieval practice

Once the board and core are stable, use short questions instead of repeated highlighting. A good card has one answer and enough context to prevent guessing from the wording. Separate definition cards from comparison cards and calculation cards.

For machine learning, useful cards should ask things such as “What does recall measure?”, “Why does standardising before a random split cause leakage?” or “What changes when the classification threshold is lowered?” They should not ask for a whole lecture in one response.

A flashcard drill from the classification board looks like this:

Cards — Supervised learning: classification8 due

What is the role of a validation set?

To choose models or hyperparameters during development without using the final test estimate.

All 8 cards
What is the role of a validation set?To choose models or hyperparameters during development without using the final test estimate.
Write the probability model for binary logistic regression.p(y=1|x) = σ(wᵀx+b), where σ(z)=1/(1+e⁻ᶻ).
What does precision measure?TP/(TP+FP): among the observations predicted positive, the proportion that are truly positive.
What does recall measure?TP/(TP+FN): among the truly positive observations, the proportion detected by the model.
Why can accuracy be poor for a rare-event problem?A classifier can achieve high accuracy by predicting the majority class while missing most positive cases.
What is the usual effect of increasing L2 regularisation strength?It penalises large weights more strongly and generally reduces model complexity; excessive strength can cause underfitting.
Why must feature scaling be fitted only on training data?Using validation or test observations to calculate scaling parameters lets information from those sets influence training, causing leakage.
When is a time-based split preferable to a random split?When the model will predict future observations and the data distribution or available information changes over time.

That is a MySummaries deck, filled with machine learning material. Yours is written from your own notes. Start free

Work through the cards actively: read the question, answer aloud or in writing, reveal the answer, then grade the response honestly. “Hard” should mean the fact was absent or materially wrong, not merely that the answer was phrased differently. Mark a card down if you named a metric but could not give its formula or interpretation.

Use spaced repetition, but do not let the schedule replace problem-solving. A card can confirm that you know the formula for F1; only a question can show whether you can choose F1 appropriately when false negatives and false positives have different costs.

Practise the complete workflow with questions

At least once per study session, take a problem from data description to model decision. Use this order:

  1. Define the target and the unit of observation.
  2. Identify leakage, imbalance, missingness and possible distribution shift.
  3. Choose a baseline and a candidate model.
  4. Select a validation strategy that matches how predictions will be used.
  5. Choose metrics linked to the practical cost of errors.
  6. Interpret the result, including uncertainty and limitations.

For coding work, write the pipeline in this order as well. Separate preprocessing, fitting and evaluation. Inspect the labels and splits before tuning the model. Keep a record of the baseline, hyperparameters, validation result and final test result so you can explain what changed and why.

Compare similar algorithms instead of memorising isolated lists

Make short comparison tables from your own notes. For example:

  • Linear regression versus logistic regression: continuous target versus binary probability model.
  • Decision tree versus random forest: one interpretable, high-variance tree versus an ensemble that usually reduces variance through averaging.
  • Bagging versus boosting: parallel variance reduction through aggregation versus sequential fitting that concentrates on previous errors.
  • Generative versus discriminative approaches: modelling a data-generating distribution versus modelling a decision boundary or conditional prediction directly.

For every comparison, add one use case, one limitation and one diagnostic. This makes the difference operational rather than verbal.

Listen to difficult material after you have organised it

Audio is most useful for consolidation, not for replacing first exposure to code or mathematical notation. After you have built the board, use a short spoken explanation to connect sections while walking or travelling. Pause when the explanation reaches a decision point and state the answer yourself.

A useful lecture should follow one thread, such as why a model that performs well on training data can fail on unseen data. It should move from empirical risk and generalisation to overfitting, regularisation, cross-validation and the limits of the test set. Keep it short enough to replay before a practice session.

An audio lecture built from the board might look like this:

Lecture — Supervised learning: classification10 min
From a fraud dataset to a defensible classifierFollows one classification problem from target definition through validation, threshold choice and calibration.
03:4810:12
Speed1×1.25×1.5×2×

Transcript · tap any word to jump there

Start with the prediction you will actually make. If the system flags a transaction for review, the target is not simply accuracy; it is whether the review queue finds enough genuine fraud without overwhelming the investigators with false positives.

Now separate the data by the way it will arrive. A random split is unsuitable if the model will predict later transactions and patterns change over time. Fit imputation, scaling and feature selection on the training portion, then apply those fitted transformations to validation and test data.

Finally choose the threshold deliberately. Logistic regression produces a score or probability, but the threshold converts that score into a label. Lowering the threshold usually raises recall and lowers precision. Report the confusion matrix and explain which error matters most in the application; a single accuracy figure cannot do that work.

After listening, return to the board and add anything you could not explain without the audio. Then solve one new problem without replaying it. If you can repeat the lecture but cannot select a validation strategy or interpret a confusion matrix, you have practised recognition rather than retrieval.

A repeatable weekly method

Use three kinds of session across the week:

  • Build: process new notes into one board, with worked examples and assumptions.
  • Retrieve: review due cards and answer short questions without looking at notes.
  • Apply: complete a calculation, coding task or open-ended model-selection problem under a time limit.

At the end of each week, inspect errors by category: missing definition, wrong formula, poor interpretation, coding implementation, leakage, or weak justification. Repair the category rather than rereading the entire topic. If several mistakes come from one section, split it into smaller cards and add a worked example to the board.

Before an assessment, practise switching between representations. Explain a model in plain language, write its key equation, implement a small version, interpret its output and state one limitation. That combination is a better test of readiness than recognising terms in a list.

How MySummaries helps

MySummaries lets you build a machine learning revision board from your own PDFs, slides and photographed notes, then turn its sections into spaced-repetition cards, written practice and spoken explanations. For this subject, use the board for algorithms, assumptions, metrics and worked calculations; use cards for precise facts; and use audio to rehearse the reasoning that links data preparation, model choice and evaluation. Open the study platform.