Strengthen your scikit-learn problem-solving skills through seven Intermediate, code-completion challenges. Each challenge provides starter files and concrete TODO requirements, then asks you to implement functions for model training, evaluation, tuning, interpretation, or visualization.
The 20 task steps span supervised and unsupervised learning: classification and regression metrics, Naive Bayes, K-nearest neighbors, validation curves, decision trees, linear and regularized regression, and three clustering methods. Automated checks and a reference solution accompany every task.
What You Will Learn
- Evaluate classification with accuracy, precision, recall, and F1, and regression with MSE, MAE, RMSE, and R²
- Build a custom scorer and use a scaled SVR pipeline with grid search
- Train Gaussian Naive Bayes on Iris data, scale features, compare Gaussian, Multinomial, and Bernoulli variants, and report a confusion matrix
- Split Iris data, train K-nearest neighbors, evaluate predictions, and search for a better value of
k - Generate validation curves for decision-tree
max_depth, choose a depth from cross-validation scores, retrain, and test the model - Prepare Iris data in pandas, fit a
DecisionTreeClassifier, and summarize precision, recall, and F1 - Fit linear regression to supplied CSV data, calculate regression metrics, and compare Ridge and Lasso with cross-validation
- Apply K-means, agglomerative clustering, and DBSCAN to Iris data, compare silhouette scores, and visualize cluster assignments
Who This Course Is For
This course is for Python and scikit-learn learners who have completed introductory model-building exercises and want more independent implementation practice. It suits learners who can read starter code, fill function bodies, interpret requirements, and debug failed checks.
Prerequisites: Working knowledge of Python, NumPy or pandas data structures, train/test splitting, estimator fit and predict, and basic classification and regression concepts is recommended. These challenges are not designed to teach Python syntax from scratch.
Learning environment: All seven challenges run in an Ubuntu 22.04 WebIDE. They contain 20 Intermediate task steps, each with an automated check and reference solution. Six environments pin scikit-learn 1.2.2; the metrics challenge pins scikit-learn 1.1.3 and NumPy below 2 for its legacy Boston Housing exercises.
Frequently Asked Questions
Are these suitable as a first introduction to scikit-learn?
Not ideally. Despite the course-level Beginner label, all seven included challenges are individually labeled Intermediate and begin with TODO-oriented starter code. Start with a guided scikit-learn fundamentals course if fit, predict, metrics, and train/test splits are new to you.
How much guidance and feedback is provided?
Each task describes the required functions or outputs and supplies starter project files, but you implement the missing logic. One automated check and one reference solution are available for each of the 20 task steps.
Which datasets are used?
Several challenges reuse the Iris dataset. The metrics work also uses the Breast Cancer dataset and a legacy Boston Housing dataset, while the linear-regression challenge supplies its own CSV files. This is focused algorithm practice, not a survey of large production datasets.
Will the code work unchanged with the latest scikit-learn?
Not all of it. The metrics challenge intentionally installs scikit-learn 1.1.3 because it imports the Boston Housing loader removed from later releases; the other challenges pin 1.2.2. Treat that exercise as version-specific practice and use a current supported dataset when adapting the code to modern projects.




