Learn how clustering reveals structure in unlabeled data through nine Jupyter notebooks. You will move from the distinction between supervised and unsupervised learning to centroid-based, hierarchical, density-based, and graph-based methods, then compare how algorithms respond to different data shapes and sizes.
The course combines algorithm intuition, selected Python implementations, scikit-learn models, and applied exercises. You will compress an image by clustering pixels, build a wheat-seed dendrogram, locate dense shared-bike areas and outliers, and benchmark eight clustering methods.
What You Will Learn
- Distinguish unsupervised clustering from supervised prediction
- Implement key K-Means steps and examine K selection, K-Means++, and Mini-Batch K-Means
- Compress image colors by grouping pixels with Mini-Batch K-Means
- Compare linkage strategies, build agglomerative hierarchies, and explore BIRCH and PCA
- Create and truncate a hierarchical dendrogram for wheat-seed data
- Apply DBSCAN and HDBSCAN concepts to irregular clusters, density parameters, and noise
- Cluster shared-bike coordinates and mark spatial outliers with DBSCAN
- Use spectral clustering and compare eight algorithms across moons, circles, blobs, and data sizes
Who This Course Is For
This course is for Python learners with introductory machine-learning experience who want to understand when different clustering families are useful. It suits learners ready to connect distances, centroids, density, hierarchical trees, graphs, visual results, and runtime tradeoffs. The notebooks are marked beginner but move beyond a first programming course.
Prerequisites: Basic Python, NumPy, Pandas, Matplotlib, and scikit-learn workflows are recommended. Familiarity with vectors, distances, matrices, and trainable model APIs will help; prior clustering experience is not required.
Learning environment: All nine activities run in a provided Ubuntu 22.04 Jupyter environment. Prepared notebooks contain explanations, visualizations, supplied image and CSV data, runnable examples, and code-completion tasks.
Frequently Asked Questions
Does the course implement clustering algorithms or only call scikit-learn?
It does both. The notebooks walk through core K-Means, hierarchical, DBSCAN, and spectral-clustering ideas and selected implementation steps, then use scikit-learn and SciPy for practical models and dendrograms.
Which clustering methods are covered?
The curriculum includes K-Means, K-Means++, Mini-Batch K-Means, agglomerative clustering, BIRCH, DBSCAN, HDBSCAN concepts, spectral clustering, Affinity Propagation, and Mean Shift.
Are there applied exercises beyond synthetic point clouds?
Yes. You compress a Chengdu image, cluster wheat-seed measurements, and analyze shared-bike latitude and longitude data to identify dense areas and possible outliers.
Can I copy every example unchanged into the newest scikit-learn version?
The prepared course environment supports the notebooks, but some examples use older parameter names such as affinity in agglomerative clustering. Check current library documentation and update deprecated arguments when moving code to a newer environment.





