Develop a broad foundation in classical supervised classification across ten Jupyter notebooks. You will study how major algorithms make decisions, implement selected mechanics in Python, build scikit-learn models, visualize results, and compare classifiers with cross-validation.
The course connects mathematical ideas to code: sigmoid and log loss for logistic regression, distances and voting for KNN, probability for Naive Bayes, margins and kernels for SVMs, information gain for trees, backpropagation for neural networks, and model combination for ensembles. Focused code-completion exercises reinforce Gaussian distributions, digit recognition, and model selection.
What You Will Learn
- Derive and implement logistic regression with sigmoid, log loss, gradient descent, and scikit-learn
- Build KNN classification from distance calculations and voting, then examine K selection and kd-trees
- Apply Bayes’ theorem, parameter estimation, Multinomial Naive Bayes, and Gaussian distributions
- Compare linear and nonlinear support vector machines and common kernel functions
- Implement perceptron training and backpropagation, then classify digits with
MLPClassifier - Construct decision trees using information gain concepts, pruning, and scikit-learn
- Train Bagging, Random Forest, AdaBoost, and Gradient Boosting classifiers
- Preprocess Abalone data and compare seven classifiers with 10-fold cross-validation
Who This Course Is For
This course is for learners who already know Python and want to understand a wide range of classical classification methods beyond calling a single estimator. It suits students moving from introductory machine learning into algorithm comparison and model selection. The notebooks are marked beginner, but the mathematical derivations and hand-built components make this a poor fit for someone new to programming or basic algebra.
Prerequisites: Python, NumPy, Pandas, plotting, and basic algebra are recommended. Familiarity with probability, derivatives, train/test splits, and the idea of features and labels will help.
Learning environment: All ten activities run in a provided Ubuntu 22.04 Jupyter environment. The notebooks combine explanations, visualizations, runnable scikit-learn examples, provided data files, and code sections to complete.
Frequently Asked Questions
Does the course implement algorithms from scratch or use scikit-learn?
It does both. Several notebooks derive and code core mechanics—such as logistic gradient descent, KNN distance and voting, perceptron updates, and tree splitting—then use scikit-learn estimators for practical training and comparison.
Which classifiers are covered?
The course covers logistic regression, KNN, Naive Bayes, linear and kernel SVMs, perceptrons and multilayer neural networks, decision trees, Bagging, Random Forest, AdaBoost, and Gradient Boosting.
Does it teach model evaluation and selection?
Yes. Individual labs calculate prediction accuracy, and the final notebook preprocesses Abalone data and uses 10-fold cross-validation to compare seven classifier families under default settings.
Can I copy the examples directly into a current local scikit-learn project?
The prepared environment supports the course notebooks, but some examples use older scikit-learn-era imports or constructor patterns. When moving code to a newer local installation, check the current API and adjust deprecated names or arguments.





