What this machine learning study plan covers
This machine learning study plan is designed for a learner who needs both conceptual understanding and working implementation skills. It assumes you can write basic Python and use a notebook, but it does not assume that you already understand the mathematics or terminology of machine learning.
The plan runs for six weeks. Use five study sessions per week, with about 90 minutes per session. That gives you 7.5 hours each week. If you have less time, keep the order of activities but reduce each session to 45–60 minutes. Do not remove the retrieval and coding sessions: those are where you find out whether you can use the material rather than recognise it.
Each week has four types of work:
- Learn: read or watch one carefully chosen source and summarise it in your own words.
- Implement: write or modify code, preferably first with a small dataset.
- Retrieve: answer questions without looking at your notes.
- Review: correct errors and schedule the topics you still confuse.
Keep one revision board for the whole subject, divided into sections. Put definitions, equations, assumptions, code patterns and common errors on the board. A separate board for every small algorithm usually makes revision fragmented.
Set up the first session
Before studying algorithms, collect the material you are expected to know. Include lecture notes, practical instructions, formula sheets, assessment feedback and any required reading. Label each item by topic rather than by week of teaching. For example, move a lecture on logistic regression into a section called “Classification”, not into “Week 4”.
Create a short baseline test. Without checking your notes, explain the difference between training and test data, write the expression for mean squared error, and describe what regularisation does. Add one small coding task: load a tabular dataset, split it into training and test sets, fit a simple model and report an appropriate metric. The purpose is not to score well. It is to identify the starting point for the plan.
In MySummaries, the due queue for this plan can keep retrieval work separate from new learning. A due queue on the first morning might look like this:
Start with the cards that are due before adding new cards. If a card repeatedly fails, rewrite it as a smaller question. “Explain gradient descent” is too broad for a reliable card; “What does the learning rate control in gradient descent?” is easier to retrieve and mark.
Week 1: foundations and the complete workflow
Time: 7.5 hours across five sessions.
The objective is to understand the full supervised-learning workflow before studying individual models.
- Session 1, 90 minutes: Define features, targets, samples, parameters, hyperparameters, training data, validation data and test data. Draw the flow from raw data to a reported test result.
- Session 2, 90 minutes: Study data cleaning, categorical encoding, scaling and missing values. Write down when each operation must be fitted only on the training data.
- Session 3, 90 minutes: Implement a baseline model with a train/validation/test split or cross-validation. Record the baseline score and the metric used.
- Session 4, 90 minutes: Retrieve key definitions and work through two short numerical examples involving predictions and errors.
- Session 5, 90 minutes: Rebuild the workflow from an empty notebook. Finish with a list of three mistakes you made and the correction for each.
Your first board should contain decisions, not just headings. For example, it should state why accuracy can be misleading for imbalanced classes, and why preprocessing the full dataset before cross-validation causes leakage.
A board for the first week could be organised like this:
- Feature — an input variable used to make a prediction
- Target — the value the model is trained to predict
- Parameter — learned from training data; a hyperparameter is selected outside the fitting process
Compare against a simple reference such as the majority-class classifier. Report the confusion matrix as well as a single summary metric.
- Fit preprocessing on the training fold only
- Use validation data for model or hyperparameter choices
- Keep the test set untouched until the final estimate
- Set random_state where reproducibility matters
- Inspect the target distribution before choosing a metric
- Use a pipeline when preprocessing is part of model fitting
| Metric | Definition | Caution |
|---|---|---|
| MAE | mean of |y − ŷ| | same units as target |
| MSE | mean of (y − ŷ)² | penalises large errors |
| RMSE | square root of MSE | sensitive to outliers |
Week 2: linear and logistic models
Time: 7.5 hours across five sessions.
Study linear regression and logistic regression together because both use a weighted combination of features, but their outputs and losses differ.
- Session 1: Derive the prediction form for linear regression and explain the role of the intercept and coefficients.
- Session 2: Learn mean squared error, gradient descent and the effect of the learning rate. Implement linear regression with a library, then inspect predictions rather than treating the score as the whole result.
- Session 3: Study logistic regression, the sigmoid function, log loss and the distinction between a probability estimate and a class decision.
- Session 4: Compare a classification threshold of 0.5 with another threshold on a validation set. Record the effect on precision, recall and the confusion matrix.
- Session 5: Complete a closed-book retrieval session, then explain one coefficient in context without claiming that it proves causation.
Keep the mathematical cards narrow. One card should test the sigmoid function; another should test why changing a threshold affects precision and recall. Do not put the entire derivation of logistic regression on one card.
Week 3: trees, ensembles and feature choices
Time: 7.5 hours across five sessions.
This week is about models that make decisions through splits and combinations of models.
- Session 1: Draw a decision tree for a small dataset. Define root, node, leaf, depth and impurity.
- Session 2: Compare overfitting in a deep tree with underfitting in a shallow tree. Change
max_depthand record both training and validation performance. - Session 3: Study random forests: bootstrap samples, random feature selection and aggregation across trees.
- Session 4: Study boosting at a conceptual level. Focus on the sequential correction of errors and the risks of excessive complexity.
- Session 5: Use cross-validation to compare a linear model, a decision tree and an ensemble on the same prepared data.
Make a short decision table in your notes: what each model assumes, what it handles well, what it makes difficult to explain, and which hyperparameters you changed. This is more useful than memorising a list of algorithms.
Week 4: evaluation, leakage and responsible use
Time: 7.5 hours across five sessions.
This is the week to connect model performance with the consequences of a prediction.
- Session 1: Revise accuracy, precision, recall, F1 score, specificity, ROC-AUC and PR-AUC. Write the confusion matrix before calculating any metric.
- Session 2: Work through imbalanced classification examples. Decide which error is more serious in each scenario and justify the metric choice.
- Session 3: Study cross-validation, stratification, grouped data and time-ordered data. Identify when a random split would produce an unrealistic estimate.
- Session 4: Find leakage in three deliberately flawed pipelines. Examples include scaling before a split, using a future value as a feature, and selecting features using the complete dataset.
- Session 5: Write a one-page model report covering the data, split, metric, baseline, limitations and proposed next test.
At the end of this week, stop adding algorithms temporarily. Spend the final 30 minutes checking whether you can defend an evaluation design. A model with a lower score can be the more useful model if its evaluation matches the real decision setting.
Week 5: neural networks and optimisation
Time: 7.5 hours across five sessions.
Study the components of a basic feed-forward neural network: layers, weights, biases, activation functions, loss, backpropagation and optimisation.
- Session 1: Draw a network with an input layer, one hidden layer and an output layer. Label the dimensions of each weight matrix.
- Session 2: Explain the forward pass and calculate the output of a very small network by hand.
- Session 3: Study backpropagation as repeated application of the chain rule. You do not need to memorise every derivative without understanding what is being updated.
- Session 4: Train a small network. Change the learning rate, batch size or number of epochs one at a time and record the result.
- Session 5: Diagnose underfitting and overfitting from training and validation curves. Add regularisation or early stopping where appropriate.
Use diagrams and dimensions in your notes. Many neural-network errors are shape errors or incorrect assumptions about the output layer and loss function, not failures of advanced mathematics.
When choosing the next short lecture or explanation, prioritise the topic where your recent work is weakest rather than following the original teaching order. The picker below uses that rule:
Evaluation metrics · Struggling — getting 3 of 8 cards wrong and choosing accuracy for imbalanced examples
Week 6: integration and assessment practice
Time: 7.5 hours across five sessions.
Use a small project or case study that requires the complete process.
- Session 1, 90 minutes: Define the prediction task, target, unit of observation and likely risks. Choose a baseline and an evaluation split.
- Session 2, 90 minutes: Build a reproducible preprocessing and modelling pipeline. Keep a record of every choice.
- Session 3, 90 minutes: Compare two or three models. Do not change several hyperparameters without recording what changed.
- Session 4, 90 minutes: Write answers to conceptual and numerical questions under a time limit. Mark each answer against your notes, looking for missing conditions and incorrect terminology.
- Session 5, 90 minutes: Present the project in five minutes. Explain the result, uncertainty, limitations and the next experiment you would run.
Finish by making a one-page summary from memory. Include the full workflow, the main regression and classification metrics, common leakage patterns, the purpose of regularisation, and the difference between parameters and hyperparameters.
Review your weak areas each weekend
Do not judge progress only by the number of pages read or notebooks completed. Record performance by topic. A useful review distinguishes between a knowledge gap, a calculation error, a coding error and a communication problem. Each needs a different fix.
For a learner midway through this plan, a weakness report might look like this:
If evaluation metrics are weakest, spend the next session calculating confusion-matrix measures from small tables and explaining when each metric is appropriate. If neural-network optimisation is weakest, draw the forward pass and inspect learning curves before reading another broad overview. If coding is the problem, replace a reading session with a short implementation and annotate each step.
Repeat missed cards after one day, several days and about a week, adjusting the interval when recall is unreliable. At the end of every week, remove duplicate cards and split any card that still tests several unrelated facts.
How MySummaries helps
You can build a revision board from your machine learning notes, then use it to generate cards, targeted explanations and progress views. The useful part of the method is the connection between your source material, your errors and the next study session: open MySummaries.