What a machine learning viva is testing

A machine learning viva is not only a test of whether you can define terms. You may need to explain a model choice, defend an evaluation method, identify a source of bias, or reason through what could fail after deployment.

There is no single international format for a machine learning viva. Your university, course or professional body sets its own topics, timing and marking method, so check the official subject guide first. The method below works across most oral assessments because it trains the underlying skill: giving a technically correct answer, stating assumptions, and applying the idea to a specific situation.

A useful answer usually has five parts:

  1. Direct answer — answer the question in the first sentence.
  2. Definition or mechanism — explain the relevant concept accurately.
  3. Assumptions — state when the explanation applies.
  4. Example — connect it to a model, dataset or deployment setting.
  5. Limitation or trade-off — show what would change your decision.

For instance, do not answer “regularisation prevents overfitting” and stop. Explain that it adds a penalty or constraint that discourages complex parameter values, then distinguish L1 from L2 regularisation and say that the penalty strength must be selected without using the test set.

The most efficient practice is to turn your revision material into stations. Each station should have an opening question, two or three likely follow-ups, and a fixed time target. Answer aloud rather than silently recognising the correct paragraph in your notes. A good response must be retrievable under pressure.

Set up an examiner-style board

Start with the material you are actually expected to know: lecture slides, practical notebooks, required readings, assignment feedback and the published learning outcomes. Separate broad topics into answerable stations rather than keeping one large heading called “machine learning”.

A board might contain supervised learning, model selection, evaluation, unsupervised learning and responsible deployment. Each section should hold facts that can be said precisely, such as the difference between a validation set and a test set, or the formula for precision.

When creating a station, write the opening question in the same form an examiner might use. “What is cross-validation?” is useful for recall, but “You have 800 labelled observations and must compare three classifiers; how would you estimate generalisation performance?” tests planning and explanation as well.

A board for this topic can be organised into sections such as these:

MySummaries’ examiner view shows the topic emphasis and the areas that need an oral response:

ExaminerMachine Learning viva
Conceptual accuracyAssumptions and limitationsExperimental designClear technical communication
The examiner view lists the four marking emphases used for this machine learning oral practice.

Use the emphases as a speaking checklist, not as invented rules for your university’s assessment. If your subject guide gives different criteria, replace these with the official wording.

Four stations to prepare

A practical station set should cover both theory and applied judgement. The following four stations are deliberately different: one tests statistical intuition, one tests experimental design, one tests unsupervised learning, and one tests deployment.

The station opening questions below are the prompts to practise. Prepare a two-minute answer for each, then add follow-up questions that require a decision rather than a definition.

StationAttemptsBestAvg
Bias, variance and regularisationExplain the bias–variance trade-off and how you would diagnose overfitting in a supervised learning model.37669
Evaluation under class imbalanceA binary classifier detects a rare condition. Which metrics would you report, and how would you choose a decision threshold?27166
Clustering without labelsHow would you decide whether a clustering result is useful when no ground-truth labels are available?16868
Deployment and driftA model performs well in validation but accuracy falls after deployment. What would you investigate first?0——
The station table shows four machine learning viva prompts, with previous attempts and current oral scores.

Do not prepare a memorised essay for every station. Prepare a structure that lets you adapt to the follow-up. For example, the class-imbalance station may lead to questions about precision-recall trade-offs, calibration, resampling, threshold selection or the cost of false negatives.

Station 1: Bias, variance and regularisation

This station tests whether you can distinguish a model’s error on the training data from its ability to generalise. It also tests whether you can recommend a remedy without treating every poor result as “overfitting”.

Answer the question aloud before reading the marked response. Aim to define the terms, describe an observable diagnostic, and name a suitable intervention.

Oral — Bias, variance and regularisationMarked

Examiner

Explain the bias–variance trade-off and how you would diagnose overfitting in a supervised learning model.

2:183:00Mark answer
78%Bias, variance and regularisation — marked78/100 · Sound, with one missing diagnostic · 2:18 spoken of 3:00
Conceptual accuracy17/20

The answer correctly contrasted underfitting with overfitting and linked variance to sensitivity to the training sample.

ImproveState that bias and variance describe different components of expected generalisation error, rather than two model types.

Assumptions and limitations14/20

The answer assumed that a train–validation comparison was available but did not discuss data leakage or noisy labels.

ImproveSay that a validation gap is informative only when the splits are appropriate and preprocessing is fitted within each training split.

Experimental design16/20

The proposed learning curves and regularisation comparison were appropriate.

ImproveMention cross-validation or repeated splits when the dataset is small.

Clear technical communication14/20

The explanation was ordered and understandable, but the recommendation arrived before the diagnostic evidence.

ImproveUse the order: symptom, likely cause, check, intervention.

A strong answerBias is error from an overly restrictive set of assumptions, whereas variance is sensitivity to the particular training sample. Overfitting is suggested when training performance is much better than validation performance, although the split must be representative and free from leakage. I would inspect learning curves, compare cross-validation results with the held-out test result, and check the labels and preprocessing pipeline. If variance is the problem, I could increase the effective training data, reduce model complexity, or increase regularisation, selecting the penalty strength using validation data rather than the test set.

A marked oral response on the bias–variance trade-off shows the score, platform criteria and a stronger model answer.

That is a MySummaries station, filled with machine learning material. Yours is written from your own notes. Start free

The important correction is that regularisation is not a universal cure. A large penalty can increase bias and produce underfitting. A training–validation gap can also result from a distribution shift, a flawed split, duplicated observations, label noise or leakage. Say what evidence you would inspect before choosing the remedy.

A follow-up worth practising is: “Why must preprocessing be included inside cross-validation?” The answer is that statistics such as a feature mean, standard deviation or imputation value must be estimated from the training portion of each fold. If they are calculated using the complete dataset first, information from the validation portion enters the training process and makes the estimate optimistic.

Station 2: Evaluation with imbalanced classes

This station tests whether you can choose metrics according to the decision problem rather than reporting accuracy by habit. It also gives the examiner several ways to probe your understanding of thresholds and probabilities.

For a rare positive class, a classifier that predicts every case as negative may have high accuracy but zero recall for the positive class. Precision is TP divided by TP plus FP. Recall, also called sensitivity, is TP divided by TP plus FN. The F1 score is the harmonic mean of precision and recall, so it is not a replacement for understanding the costs of the two error types.

Oral — Evaluation under class imbalanceMarked

Examiner

A binary classifier detects a rare condition. Which metrics would you report, and how would you choose a decision threshold?

2:413:00Mark answer
84%Evaluation under class imbalance — marked84/100 · Competitive · 2:41 spoken of 3:00
Conceptual accuracy18/20

The answer correctly rejected accuracy as the sole metric and distinguished precision, recall and specificity.

ImproveState that the confusion matrix changes with the selected threshold, while a score or probability can be assessed across thresholds.

Assumptions and limitations16/20

The answer considered the cost of false negatives but did not mention calibration or prevalence shift.

ImproveExplain that a probability used for decisions should be checked for calibration in the relevant population.

Experimental design17/20

The proposed precision–recall curve, separate test set and threshold selection on validation data were appropriate.

ImproveAdd confidence intervals or repeated evaluation when the positive class is small.

Clear technical communication16/20

The answer used a concrete decision rule and clearly separated model evaluation from threshold choice.

ImproveState the operational consequence of a false positive as well as a false negative.

A strong answerI would report the confusion matrix and choose metrics that reflect the consequence of each error. Recall is important if missing a positive case is costly, while precision matters if follow-up investigations are expensive; I would also report specificity and consider a precision–recall curve because the positive class is rare. I would select the threshold on validation data using an agreed cost or target recall, then estimate final performance once on an untouched test set. If the output is interpreted as a probability, I would check calibration and confirm that the evaluation prevalence resembles the deployment population.

A marked oral response on imbalanced classification shows how metric choice and threshold selection are assessed.

A strong follow-up answer distinguishes ranking from decision-making. The area under a ROC curve summarises ranking across thresholds, but it does not tell you which threshold to deploy. A precision–recall curve can be more informative when positives are rare, yet it still does not replace a decision rule based on consequences and available resources.

You should also be prepared to discuss calibration. A model can rank cases well while its probabilities are poorly calibrated. If cases assigned a probability of 0.8 are positive only 0.5 of the time, that probability should not be treated as a reliable risk estimate without recalibration or a different modelling approach.

Station 3: Clustering without labels

This station tests whether you can avoid claiming that an algorithm has discovered “true groups” merely because it returned a plot. There is no single best number of clusters independent of purpose and representation.

Oral — Clustering without labelsMarked

Examiner

How would you decide whether a clustering result is useful when no ground-truth labels are available?

2:063:00Mark answer
72%Clustering without labels — marked72/100 · Developing · 2:06 spoken of 3:00
Conceptual accuracy15/20

The answer correctly mentioned silhouette score and the dependence on feature scaling.

ImproveExplain that a clustering metric measures a geometric property of the chosen representation, not business or scientific truth.

Assumptions and limitations13/20

The response did not discuss the cluster shape assumptions of k-means or sensitivity to distance choice.

ImproveState why the algorithm’s assumptions may be unsuitable and compare an alternative when justified.

Experimental design14/20

The answer proposed stability checks but did not specify what would be varied.

ImproveRepeat the analysis across resamples, initialisations, feature transformations and plausible values of k.

Clear technical communication15/20

The answer was concise but ended with a metric rather than a decision about usefulness.

ImproveFinish by linking the clusters to a predefined use, such as exploratory segmentation or prioritised investigation.

A strong answerI would first define what the clusters are intended to support, because a mathematically separated grouping may not be useful for the task. I would scale or transform features appropriately, choose an algorithm whose assumptions fit the data, and compare plausible settings rather than accepting one value of k. Internal measures such as silhouette score can describe separation and cohesion, but I would also test stability across resamples and inspect the clusters in the original feature space. Finally, I would check whether the groups are interpretable and useful without presenting them as ground-truth categories.

A marked oral response on unsupervised learning shows how to discuss metrics, assumptions and usefulness together.

The phrase “no ground-truth labels” is important. Silhouette score, Davies–Bouldin index and similar measures do not prove that clusters correspond to real categories. They evaluate the arrangement produced under a particular distance measure and representation.

Expect a follow-up such as “Why might k-means be inappropriate?” A precise answer includes its reliance on numeric features, distance calculations and cluster-centre representations. It can be sensitive to scaling, initialisation and outliers, and it tends to suit roughly compact, similarly shaped groups better than elongated or uneven-density structures. The best alternative depends on the data and purpose; do not name an algorithm without explaining what problem it addresses.

Station 4: Deployment and distribution shift

The fourth station has not yet been attempted. Practise it as a fresh response rather than reading the model answer first. Your first task is to separate a measurement problem from a model problem.

Start by asking whether the fall is real and whether the label definition or reporting process has changed. Then inspect the input distribution, the relationship between inputs and outcomes, data quality, missingness, preprocessing, software versions, thresholding and subgroup performance. A model may retain ranking performance while a fixed threshold becomes unsuitable because prevalence has changed.

A useful answer should also mention monitoring after deployment. Decide in advance which inputs, predictions, outcomes, calibration measures and subgroup results will be monitored, how often they will be reviewed, and what action follows an alert. Retraining is not automatically the correct response: you may need to correct a data pipeline, revise labels, recalibrate probabilities or investigate a new population.

How to answer when the examiner probes

Oral marks are often lost in the second sentence, when a candidate gives a correct headline but cannot state its boundary. Practise these follow-up patterns:

  • “What assumption are you making?” Name the assumption and say how you would check it.
  • “What would you report?” Give the metric, the population, the split and the uncertainty or limitation.
  • “Why this method?” Compare it with one plausible alternative on the relevant trade-off.
  • “What could go wrong?” Give a concrete failure mode, not just “bias”.
  • “How would you know?” Name the diagnostic, experiment or monitoring signal.

For example, if asked whether a complex model is better than a linear model, avoid answering solely with validation accuracy. Ask whether the comparison used the same preprocessing and splits, whether the metric reflects the task, whether the difference is stable, and whether the additional complexity is acceptable for interpretability, latency, maintenance or data requirements.

A recorded answer can be marked more usefully than a general feeling that it “sounded fine”. Listen for unsupported absolutes such as “always”, “proves” and “the model understands”. Replace them with bounded language: “under these assumptions”, “the result suggests”, or “the representation captures a pattern correlated with”.

The transcript view below isolates a common oral weakness: naming a method without explaining the evidence needed to trust it.

Transcript

I would use cross-validation to choose the best model. Then I would report the cross-validation score as the final performance. I would compare the models using the same folds and then evaluate the selected pipeline once on an untouched test set.

unsupported evaluation claim
Then I would report the cross-validation score as the final performance.

Cross-validation can support model selection, but using its score as the final estimate after selection can be optimistic. The response also needs to specify what happens to preprocessing and hyperparameter tuning.

Say: I would use cross-validation within the training data to compare pipelines and select hyperparameters, then report performance from one untouched test set, with preprocessing fitted within each training fold.
The wording view marks a gap in the evaluation plan and supplies a more defensible sentence.

Use this process after every station:

  1. Record the answer without restarting when you make a mistake.
  2. Mark each claim as correct, incomplete or unsupported.
  3. Add one missing threshold, formula, assumption or example to the board.
  4. Re-record only the weak section.
  5. Schedule the station again after the interval set by your revision system.

Do not judge an answer only by its length. A two-minute answer with a clear decision, assumption and limitation is stronger than five minutes of definitions. The goal is controlled reasoning: answer the question asked, justify the method, and show what evidence would change your conclusion.

The debrief can make the next practice target explicit:

Spoken feedback

You selected sensible evaluation measures, but the next mark depends on separating threshold choice from model ranking and stating where the threshold is tuned.

The spoken debrief gives one precise improvement for the next machine learning viva attempt.

How MySummaries helps

Build a Machine Learning board from your lecture slides, notes, practical work and assessment guidance. Turn each section into oral stations, record answers, and review the examiner-style feedback against conceptual accuracy, assumptions, experimental design and communication. Weak answers can return as focused cards, while the board keeps formulas such as precision, recall and F1 beside the cases where you need to apply them.

Start at MySummaries.