How to use these machine learning practice questions

These machine learning practice questions are designed to test decisions, not just definitions. For each question, choose the best option, then read the explanation immediately below it. Pay attention to the reason the strongest distractor fails: that is often where an otherwise sound answer loses precision.

The set moves from core supervised-learning ideas to evaluation, preprocessing, unsupervised learning and model reliability. In MySummaries, a question set can be generated from a revision board so that missed concepts return as targeted practice rather than as another full paper.

1. Bias and variance

This question tests whether you can identify the likely cause of a large gap between training and validation performance.

Question 11 mark

A decision tree has 99% accuracy on the training set and 68% accuracy on an independent validation set. Increasing the tree depth makes training accuracy rise but validation accuracy fall. Which change is most likely to improve validation performance?

The correct answer is a shallower tree or pruning: the widening train–validation gap is consistent with high variance, or overfitting. Increasing depth is the opposite intervention. Training on the validation set would contaminate evaluation, and removing the validation set would hide the problem rather than solve it. The strongest distractor is increasing depth because a more flexible model can reduce training error, but here that flexibility is already harming generalisation.

An illustrative single-best-answer question on diagnosing overfitting.

A useful rule is to separate the diagnosis from the remedy. High training error suggests underfitting or high bias; low training error with substantially worse validation error suggests high variance. The remedy should change model complexity, regularisation or the amount and quality of training data accordingly.

2. Data leakage

This question tests whether information from outside the training split has entered the feature-generation process.

Question 21 mark

You are building a model to predict whether a patient will be readmitted. You standardise every feature using the mean and standard deviation calculated from the complete dataset before splitting it into training and test sets. What is the main problem?

The test set has influenced preprocessing. Even though the labels were not used, test-set feature statistics have entered the transformation fitted for the training data. Fit the scaler on the training split only, then apply those fitted parameters to validation and test data. The strongest distractor is that standardisation cannot be used with classification; it can be useful for many classifiers, including logistic regression and support-vector machines.

An illustrative single-best-answer question on preprocessing leakage.

The same principle applies to imputation, feature selection, dimensionality reduction and target encoding. Put the operation inside a pipeline where possible, and fit each learned preprocessing step only on the relevant training fold during cross-validation.

3. Precision and recall

This question tests metric selection when the cost of a false negative is high.

Question 31 mark

A screening model is intended to identify as many people as possible who have a serious disease. Missing a true case is much more harmful than referring a healthy person for further testing. Which metric should receive priority during initial screening?

Recall is the proportion of actual positive cases that the model identifies, so prioritising it reduces false negatives. Precision becomes more important when unnecessary positive referrals are the dominant concern. Accuracy can be misleading when the condition is uncommon. The strongest distractor is specificity: it measures the proportion of true negatives correctly rejected, but it does not directly protect against missed cases.

An illustrative single-best-answer question on classification metrics.

Do not describe a model as simply “accurate” without stating the class balance and the error that matters. A threshold change can raise recall while lowering precision, so report the trade-off and select the threshold using the deployment cost, not the default probability cut-off by habit.

4. Logistic regression

This question checks what logistic regression models and how its output is interpreted.

Question 41 mark

In binary logistic regression, what does the sigmoid function do to the linear predictor before a classification threshold is applied?

The sigmoid maps the linear predictor to a value between 0 and 1, which is interpreted as a modelled probability under the model assumptions. A threshold is then used to turn that probability into a class decision. The strongest distractor is the final option: the sigmoid produces a probability-like output, not a class label by itself.

An illustrative single-best-answer question on logistic regression output.

A coefficient is linear on the log-odds scale, not directly on the probability scale. For a one-unit increase in a feature, holding other features constant, the odds are multiplied by the exponentiated coefficient. Interpretation becomes less direct when features are correlated or have very different scales.

5. Regularisation

This question tests the difference between L1 and L2 penalties.

Question 51 mark

Which statement best describes the usual effect of adding an L1 penalty to a linear model?

L1 regularisation adds a penalty based on the absolute coefficient values and can drive some coefficients exactly to zero, producing a sparse model. It does not guarantee lower test error and does not replace validation. The strongest distractor is that it makes all coefficients equal; the penalty shrinks coefficients, but it does not impose equality between them.

An illustrative single-best-answer question on L1 regularisation.

L2 regularisation instead penalises squared coefficient values and usually shrinks coefficients towards zero without making as many exactly zero. The strength of either penalty is a hyperparameter: select it using validation or cross-validation, with the complete preprocessing procedure kept inside each training fold.

6. Decision-tree splits

This question tests impurity reduction rather than memorisation of a particular tree algorithm.

Question 61 mark

For a binary classification tree, a candidate split produces two child nodes that are each much purer than the parent node. What does this indicate?

If the child nodes are substantially purer, the split has likely produced a large reduction in impurity, such as Gini impurity or entropy. That does not prove good performance on unseen data: a deep tree can find useful-looking splits by fitting noise. The strongest distractor is that the split must be optimal on unseen data; impurity is measured on the available sample, not on future cases.

An illustrative single-best-answer question on decision-tree splitting.

Tree-based models do not make causal claims merely because a feature appears near the root. Use held-out evaluation for generalisation and a separate causal design if the question is about intervention or mechanism.

7. Scaling before k-means

This question tests why distance-based algorithms are sensitive to feature units.

Question 71 mark

You use k-means clustering with two features: annual income measured in dollars and age measured in years. What is the most important first concern?

Income values may be numerically much larger than age values, causing Euclidean distance to be driven mostly by income. Consider scaling based on the modelling goal before clustering. The strongest distractor is that scaling guarantees meaningful clusters; it can prevent one unit from dominating, but it cannot make an unsuitable number of clusters or features meaningful.

An illustrative single-best-answer question on k-means preprocessing.

Scaling is not an automatic virtue. If the magnitude of a variable is itself the intended notion of importance, changing the scale may be inappropriate. State the distance measure, scaling choice and rationale when reporting clusters.

8. Principal component analysis

This question checks what PCA optimises and what information may be lost.

Question 81 mark

What is the primary objective of the first principal component in PCA?

The first principal component is the direction that captures the greatest variance under the standard PCA formulation. PCA is unsupervised, so it does not directly maximise classification accuracy. The strongest distractor is the classification objective: a direction with high variance may contain little information about the target, so PCA can reduce useful predictive information.

An illustrative single-best-answer question on principal component analysis.

Standardise features first when their units or variances should not determine the components. Fit PCA on the training data only when it is part of a predictive workflow; fitting it before the split can leak test-set structure.

9. Probability calibration

This question tests the difference between ranking cases and producing reliable probabilities.

Question 91 mark

A classifier ranks high-risk cases correctly but its predicted probabilities are consistently too high. Which statement is most accurate?

Discrimination concerns whether higher-risk cases are ranked above lower-risk cases, whereas calibration concerns whether predicted probabilities match observed frequencies. A model can rank well while being overconfident. The strongest distractor is that the model cannot be useful: it may still support ranking decisions, but probability-based decisions require calibration assessment or recalibration.

An illustrative single-best-answer question on probability calibration.

Use a calibration plot or a suitable calibration summary on data separate from the data used to fit the calibration mapping. Also report discrimination with an appropriate metric; neither aspect replaces the other.

10. Choosing a validation design

This final question tests whether the split reflects the way predictions will be made.

Question 101 mark

You have several records from each of 200 people. You want to predict outcomes for entirely new people. Which validation approach best avoids an over-optimistic estimate?

Split by person, using grouped cross-validation or a grouped hold-out, so that the model is tested on people it has not seen. A random record-level split can let person-specific patterns leak across folds. The strongest distractor is the random split because it may produce more data in every fold, but its estimate is optimistic when repeated records are correlated within person.

An illustrative single-best-answer question on validation for grouped data.

The validation scheme should imitate deployment. Use time-based splits for future prediction, group-based splits for new groups, and stratification when class representation needs to be preserved and the grouping or time structure allows it.

A short written follow-up

After the multiple-choice set, practise explaining one decision in your own words. The marked response below is deliberately good but incomplete: it identifies leakage but does not explain how to prevent it during cross-validation.

Paper — Machine learning evaluation06:20
78%Machine learning evaluation — marked7/9 marks · 06:20 taken

Explain how you would use standardisation and PCA safely when comparing models with cross-validation.

7/9

I would split the data into cross-validation folds, standardise the features and apply PCA before fitting the model. The same transformation can then be used for the validation fold. This prevents the model from seeing the validation labels.

The answer correctly uses cross-validation and recognises that the transformation must be applied consistently. It does not explicitly say that the scaler and PCA must be fitted on the training portion of each fold, rather than on all the data before cross-validation.

Missed

Fit the scaler and PCA on each training fold only.

Apply the fitted transformations unchanged to that fold's validation data.

Model answerPlace standardisation and PCA in a pipeline with the predictive model. In each cross-validation split, fit the scaler and PCA using only the training portion of that split, then transform the validation portion using those fitted parameters. Compare the resulting validation scores, and finally refit the selected pipeline on all available training data before evaluating once on a held-out test set.

An illustrative marked short-answer response on preventing data leakage.

That is a MySummaries paper, filled with machine learning material. Yours is written from your own notes. Start free

What to revise next

Use the errors rather than the total score to choose your next session. A useful record separates conceptual errors, calculation errors and errors caused by misreading the question.

Where marks go missing
54%Validation, leakage and grouped splits5×
67%Metrics and probability calibration4×
81%Regularisation and model complexity4×
90%PCA and clustering3×
An illustrative weak-area report from the completed machine learning questions.

For each weak area, write one rule and one counterexample. For instance: “fit learned preprocessing on training folds only”; counterexample: “calculating the mean and PCA components on the complete dataset before cross-validation”. Then return to a fresh question rather than rereading the definition alone.

A repeated error can become a small remediation card for the next study session:

Remediation tray

You lost this mark twice: during cross-validation, on which data must the scaler and PCA be fitted, and how is the validation fold transformed?

Add cardDismiss
An illustrative remediation prompt created from a repeated validation mistake.

How MySummaries helps

Build a Machine Learning board from your notes, lecture slides or textbook extracts, then generate question sets from its sections. The platform can return missed concepts to your revision queue, mark written answers against the source material and show which topics are costing marks. Start at MySummaries.