How to use this machine learning flashcard deck

A useful machine learning deck tests one decision or definition at a time. It should not ask you to recite an entire algorithm from memory. Each card below is built around a fact, formula or distinction that you need to retrieve quickly when reading a paper, debugging a model or explaining a method.

Work through the deck as an actual study session. Read the question, answer before revealing the back, then grade the answer:

  • Again — you could not retrieve it or gave a materially wrong answer.
  • Hard — you remembered it, but slowly or with an important omission.
  • Good — accurate answer with normal effort.
  • Easy — immediate and precise recall.

Do not award yourself “Good” because the answer looks familiar. For a formula, say what each term means. For an algorithm, state its purpose and its main limitation. Cards graded Again or Hard should return sooner than cards graded Easy.

The following deck is one practical set for machine learning foundations. In MySummaries, a board built from your notes can be cut into a drill like this, with each card kept small enough to answer without rereading a page.

A machine learning foundations deck ends up looking like this:

Cards — Machine Learning foundations19 due

What is supervised learning?

Learning a mapping from inputs to labelled target values using example input-output pairs.

All 19 cards
What is supervised learning?Learning a mapping from inputs to labelled target values using example input-output pairs.
What is the purpose of a validation set?To choose models, hyperparameters or decision thresholds without using the final test set.
What is the difference between a model parameter and a hyperparameter?A parameter is learned from training data, such as a linear weight. A hyperparameter is set before training, such as the learning rate or tree depth.
What does overfitting mean?The model fits training data, including noise or peculiarities, too closely and performs worse on unseen data.
What does L2 regularisation add to a loss function?A penalty proportional to the sum of squared weights: λΣwⱼ². It discourages large weights but does not usually make them exactly zero.
What function converts a binary classifier's logit into a probability between 0 and 1?The sigmoid function σ(z) = 1/(1 + e⁻ᶻ).
What is precision?TP/(TP + FP): the proportion of predicted positives that are actually positive.
What is recall, also called sensitivity?TP/(TP + FN): the proportion of actual positives that the model identifies.
What is the F1 score?The harmonic mean of precision and recall: 2PR/(P + R). It is high only when both precision and recall are high.
What does a false positive represent?The model predicts the positive class when the true class is negative.
What is standardisation?Rescaling a feature as z = (x − μ)/σ, usually giving it mean 0 and standard deviation 1 based on the training data.
How does one step of gradient descent update a parameter θ?θ is replaced by θ − η∇J(θ), where η is the learning rate and ∇J(θ) is the gradient of the objective.
What does the learning rate control?The size of each optimisation step. A rate that is too large can overshoot a minimum; one that is too small makes training slow.
What is the main idea of a random forest?Train many decision trees on bootstrap samples while considering random subsets of features at splits, then aggregate their predictions.
What does a decision tree split aim to do?Reduce impurity in the child nodes, using measures such as Gini impurity or entropy for classification.
What is the kernel trick used for?It computes inner products in an implicit higher-dimensional feature space without explicitly constructing that space.
What objective does k-means minimise?The within-cluster sum of squared distances from each observation to its assigned cluster centroid.
What does the first principal component in PCA represent?The direction in feature space with the greatest variance, subject to the component being a unit vector.
What does ROC AUC measure?The area under the receiver operating characteristic curve, summarising ranking performance across classification thresholds.
An illustrative flashcard drill for machine learning foundations, with the first card open and the remaining cards queued.

That is a MySummaries deck, filled with machine learning material. Yours is written from your own notes. Start free

The order is deliberate. The first cards establish the learning setup and generalisation problem. The middle group covers classification evaluation and optimisation. The final group compares common model families and unsupervised methods. If your course uses different notation, keep the notation from your notes but preserve the single-fact structure.

The facts worth retaining

Several cards need more than word-for-word recall. They become useful only when you can apply them correctly.

Metrics are conditional on the question

Precision and recall are not interchangeable. If a positive prediction triggers an expensive investigation, precision tells you how often that prediction is correct. If missing a true positive is costly, recall tells you how many real positives are found. A model can have high recall and low precision, or the reverse.

The F1 score is useful when both matter, but it is not automatically the best metric. It ignores true negatives and depends on the chosen classification threshold. For imbalanced data, also inspect the confusion matrix and choose metrics that match the consequence of each error.

Data preparation must respect the split

The mean and standard deviation for standardisation must be calculated from the training data. Applying the full dataset's statistics before splitting allows information from the validation or test set to influence training. The same rule applies to imputation, feature selection and other learned transformations.

In practice, put preprocessing and the estimator into one pipeline. Fit that pipeline on the training portion, select settings using validation or cross-validation, and evaluate the final locked procedure on the test portion.

Regularisation changes the objective

Regularisation is not a decorative setting. It changes what the optimiser is trying to minimise. With L2 regularisation, large weights are penalised. With L1 regularisation, the penalty is proportional to the sum of absolute weights and can produce exact zeros, which may be useful for sparse feature selection.

The strength of regularisation is a hyperparameter. It should be selected without repeatedly tuning against the final test set.

Cut the deck from a must-know core

A long source document should not become a long, repetitive deck. First reduce it to a short checklist. Then make cards only for items that are independently retrievable: a definition, a formula, a contrast, an assumption or a consequence.

A good card asks one question. “Explain gradient descent” is too broad for reliable grading. “What is the parameter update in one gradient descent step?” is narrower and has a clear answer. If you need the broader explanation, create separate cards for the gradient, learning rate, stopping condition and failure modes.

The deck above was cut from this core:

Must not miss coreMachine Learning foundations
Learning setup: supervised learning uses labelled input-output examples; validation data selects settings; the test set estimates final generalisation.
Generalisation: overfitting gives low training error but poorer unseen-data performance; regularisation adds a penalty to the objective.
Classification metrics: precision = TP/(TP + FP), recall = TP/(TP + FN), and F1 = 2PR/(P + R).
Preprocessing: calculate standardisation statistics on training data only; z = (x − μ)/σ.
Optimisation and models: gradient descent uses θ ← θ − η∇J(θ); random forests aggregate decorrelated trees; k-means minimises within-cluster squared distance.
An illustrative must-not-miss core from which the machine learning flashcards were cut.

Use the core when a deck starts expanding without improving recall. If two cards have nearly the same answer, combine them or make the distinction explicit. If one card contains several unrelated facts, split it.

Remediate repeated mistakes

A missed card should produce a correction, not just another exposure. After grading Again, say the answer correctly once, then identify the cause of the miss. Was the formula confused with another metric? Did you remember the definition but omit the denominator? Did you apply a correct idea to the wrong task?

Repeated misses deserve a short remediation card. Keep it tied to the error rather than rewriting the whole topic. For example, a learner who repeatedly confuses precision and recall needs the direction of the condition made explicit: precision starts with predicted positives; recall starts with actual positives.

A remediation tray for this deck might contain:

Remediation tray

You lost this card twice: Which denominator belongs to recall, and what does the resulting fraction measure?

Add cardDismiss
An illustrative remediation card created from a repeated machine learning flashcard error.

Answer: recall is TP/(TP + FN), and it measures the proportion of actual positives identified by the model. When this card becomes easy, return it to the normal deck rather than keeping it permanently separate.

A short routine for studying the deck

Use the deck in three passes across several days:

  1. First pass: attempt every card and grade strictly. Read the correction for every Again or Hard result.
  2. Second pass: answer only the cards scheduled as due. For formulas, explain the terms; for metrics, give the operational interpretation.
  3. Application pass: take a small dataset or model output and decide which metric, preprocessing step or algorithmic distinction applies.

Do not treat a correct formula as proof that you understand the metric. Ask what happens when false positives increase, why a test split must remain untouched, or what a larger learning rate does to optimisation. These small applications expose weak recall better than rereading definitions.

How MySummaries helps

MySummaries can turn your machine learning PDFs, slides and handwritten notes into a revision board, then generate cards from the board and schedule them for spaced repetition. It can also create written practice and audio explanations from the same material, so definitions and formulas stay linked to the source you are studying.

Open MySummaries